AI · LEADERBOARD

Choose the model that fits the work.

A practical ranking of the AI models teams can deploy now—compared by intelligence, speed, cost, context and operational fit.

Snapshot: 5 Aug 2026 · 13:01 TST12 models compared

Independent editorial snapshot · not live platform telemetry

The short answer

Three leaders, three different decisions.

Best overall

Closed model

Anthropic

Intelligence
60.7
Speed
61tok/s
Input / output
$5 / $25
Context
1M

Best fit

State-of-the-art reasoning and complex agentic workflows

Best open model

Open weights

Kimi

Intelligence
57.1
Speed
36tok/s
Input / output
$3 / $15
Context
1.05M

Best fit

Sovereign, open-weight deployments needing maximum intelligence

Best value

Open weights

Xiaomi

Intelligence
37.2
Speed
82tok/s
Input / output
$0.14 / $0.28
Context
1.05M

Best fit

Cost-sensitive extraction and high-volume structured work

OpenRouter usage

OpenRouter LLM usage ranking

Top models by prompt, completion, and reasoning tokens processed on OpenRouter. Showing the first 20 only.

Top 20OpenRouter data · cached hourlyUpstream · Aug 7, 2026
  1. 1.DeepSeek V4 Flash 0731by DeepSeek7.46T tokens>999%
  2. 2.Hy3by Tencent6.24T tokens30%
  3. 3.DeepSeek V4 Flash 0423by DeepSeek6.15T tokens19%
  4. 4.MiMo-V2.5by Xiaomi5.31T tokens34%
  5. 5.GPT-5.6 Luna (batch)by OpenAI4.45T tokens432%
  6. 6.GLM 5.2 (batch)by Z.ai3.18T tokens4%
  7. 7.DeepSeek V4 Proby DeepSeek2.51T tokens30%
  8. 8.Nemotron 3 Ultra (free)by Nvidia2.39T tokens7%
  9. 9.Gemini 3.6 Flash (batch)by Google2.33T tokens479%
  10. 10.Laguna S 2.1 (free)by Poolside1.85T tokens242%
  11. 11.MiniMax M3 (batch)by MiniMax1.72T tokens15%
  12. 12.Kimi K3by Moonshot AI1.37T tokens1%
  13. 13.Step 3.7 Flashby StepFun1.3T tokens24%
  14. 14.Claude Opus 5 (batch)by Anthropic1.16T tokens26%
  15. 15.Claude Sonnet 5 (batch)by Anthropic1.05T tokens3%
  16. 16.Ling-3.0-flash (free)by Inclusion AI954B tokens29%
  17. 17.Gemini 3 Flash Preview (batch)by Google912B tokens7%
  18. 18.Claude Sonnet 4.6 (batch)by Anthropic827B tokens15%
  19. 19.GPT-5.6 Terra (batch)by OpenAI776B tokens104%
  20. 20.Gemini 2.5 Flash Lite (batch)by Google611B tokens9%

From benchmark to decision

Which model for which job?

Nine recurring production scenarios, with the trade-offs that matter before a team commits to a model.

Selected use case

01 / 09

Decision view

Customer support

Why it fits

  1. 01288 tok/s keeps wait time low
  2. 02Multilingual intake
  3. 03Strong multimodal routing

How the ranking works

Transparent enough to disagree with.

The ranking combines public benchmark signals with price, measured throughput, context and deployment constraints. It is a decision aid—not a universal definition of model quality.

Dataset snapshot
5 Aug 2026 · 13:01 TST
Review cadence
Revalidated when major model data changes

Value score

(intelligence − 25)² × min(1, speed ÷ 30) ÷ (input cost + output cost × 3)

Output is weighted ×3 to reflect output-heavy production use; intelligence is squared so cheap but weak models do not dominate.

Intelligence
Public benchmark and reasoning signals form the default order.
Speed
Output tokens per second; provider and region still affect real latency.
Cost
Public input and output price per million tokens, before enterprise discounts.
Deployment
Context, weight access and hosting constraints determine production fit.

Public references

Prices and performance can change without notice. Validate the finalists against your own prompts, infrastructure and risk requirements before procurement.

FAQ

Questions worth asking before the benchmark wins the argument.

TENTEN AI · MODEL EVALUATION

The leaderboard is the start

Bring the workflow. We’ll test the shortlist.

A model only becomes a good decision after it survives your documents, permissions, latency budget and failure cases.