AI · LEADERBOARD

Choisissez le modèle adapté au travail.

Une comparaison pratique des modèles IA réellement déployables aujourd’hui.

Snapshot: 21 Jul 2026 · 10:01 TST12 models compared

Independent editorial snapshot · not live platform telemetry

The short answer

Three leaders, three different decisions.

Best overall

Closed model

Anthropic

Intelligence
59.9
Speed
70tok/s
Input / output
$10 / $50
Context
1M

Best fit

High-stakes reasoning and complex agentic workflows

Best open model

Open weights

Z AI

Intelligence
51.1
Speed
198tok/s
Input / output
$1.4 / $4.4
Context
1M

Best fit

Sovereign deployments that still need strong reasoning

Best value

Open weights

Xiaomi

Intelligence
37.2
Speed
64tok/s
Input / output
$0.14 / $0.29
Context
1.05M

Best fit

Cost-sensitive extraction and high-volume structured work

OpenRouter usage

OpenRouter LLM usage ranking

Top models by prompt, completion, and reasoning tokens processed on OpenRouter. Showing the first 20 only.

Top 20OpenRouter data · cached hourlyUpstream · Jul 21, 2026
  1. 1.Hy3 (free)by Tencent10.2T tokens24%
  2. 2.MiMo-V2.5by Xiaomi9.43T tokens29%
  3. 3.DeepSeek V4 Flashby DeepSeek5.37T tokens2%
  4. 4.GLM 5.2by Z.ai3.63T tokens16%
  5. 5.MiniMax M3by MiniMax3.23T tokens22%
  6. 6.DeepSeek V4 Proby DeepSeek2.8T tokens9%
  7. 7.Nemotron 3 Ultra (free)by Nvidia2.67T tokens2%
  8. 8.Claude Opus 4.7by Anthropic1.93T tokens8%
  9. 9.Claude Opus 4.8by Anthropic1.91T tokens13%
  10. 10.Claude Sonnet 5by Anthropic1.13T tokens21%
  11. 11.Gemini 3 Flash Previewby Google963B tokens3%
  12. 12.Step 3.7 Flashby StepFun961B tokens5%
  13. 13.Claude Sonnet 4.6by Anthropic877B tokens8%
  14. 14.Kimi K3by Moonshot AI741B tokensnew
  15. 15.Hy3by Tencent638B tokens>999%
  16. 16.GPT-5.5 (batch)by OpenAI599B tokens22%
  17. 17.Gemini 2.5 Flashby Google591B tokens3%
  18. 18.Laguna M.1 (free)by Poolside575B tokens3%
  19. 19.MiMo-V2.5-Proby Xiaomi572B tokens12%
  20. 20.Gemini 2.5 Flash Liteby Google570B tokens2%

From benchmark to decision

Which model for which job?

Nine recurring production scenarios, with the trade-offs that matter before a team commits to a model.

Selected use case

01 / 09

Decision view

Customer support

Why it fits

  1. 01294 tok/s keeps wait time low
  2. 02Multilingual intake
  3. 03Strong multimodal routing

How the ranking works

Transparent enough to disagree with.

The ranking combines public benchmark signals with price, measured throughput, context and deployment constraints. It is a decision aid—not a universal definition of model quality.

Dataset snapshot
21 Jul 2026 · 10:01 TST
Review cadence
Revalidated when major model data changes

Value score

(intelligence − 25)² × min(1, speed ÷ 30) ÷ (input cost + output cost × 3)

Output is weighted ×3 to reflect output-heavy production use; intelligence is squared so cheap but weak models do not dominate.

Intelligence
Public benchmark and reasoning signals form the default order.
Speed
Output tokens per second; provider and region still affect real latency.
Cost
Public input and output price per million tokens, before enterprise discounts.
Deployment
Context, weight access and hosting constraints determine production fit.

Public references

Prices and performance can change without notice. Validate the finalists against your own prompts, infrastructure and risk requirements before procurement.

FAQ

Questions worth asking before the benchmark wins the argument.

TENTEN AI · MODEL EVALUATION

The leaderboard is the start

Bring the workflow. We’ll test the shortlist.

A model only becomes a good decision after it survives your documents, permissions, latency budget and failure cases.