AI · LEADERBOARD

Escolha o modelo certo para o trabalho.

Uma comparação prática dos modelos de IA que equipes podem implantar agora.

Snapshot: 5 Aug 2026 · 13:01 TST12 models compared

Independent editorial snapshot · not live platform telemetry

The short answer

Three leaders, three different decisions.

Best overall

Closed model

Anthropic

Intelligence
60.7
Speed
61tok/s
Input / output
$5 / $25
Context
1M

Best fit

State-of-the-art reasoning and complex agentic workflows

Best open model

Open weights

Kimi

Intelligence
57.1
Speed
36tok/s
Input / output
$3 / $15
Context
1.05M

Best fit

Sovereign, open-weight deployments needing maximum intelligence

Best value

Open weights

Xiaomi

Intelligence
37.2
Speed
82tok/s
Input / output
$0.14 / $0.28
Context
1.05M

Best fit

Cost-sensitive extraction and high-volume structured work

OpenRouter usage

OpenRouter LLM usage ranking

Top models by prompt, completion, and reasoning tokens processed on OpenRouter. Showing the first 20 only.

Top 20OpenRouter data · cached hourlyUpstream · Aug 6, 2026
  1. 1.DeepSeek V4 Flash 0423by DeepSeek6.31T tokens17%
  2. 2.DeepSeek V4 Flash 0731by DeepSeek6.15T tokensnew
  3. 3.Hy3by Tencent5.69T tokens19%
  4. 4.MiMo-V2.5by Xiaomi5.3T tokens40%
  5. 5.GPT-5.6 Luna (batch)by OpenAI4.15T tokens900%
  6. 6.GLM 5.2 (batch)by Z.ai2.99T tokens8%
  7. 7.DeepSeek V4 Proby DeepSeek2.64T tokens27%
  8. 8.Nemotron 3 Ultra (free)by Nvidia2.33T tokens13%
  9. 9.Gemini 3.6 Flash (batch)by Google2T tokens466%
  10. 10.Laguna S 2.1 (free)by Poolside1.82T tokens466%
  11. 11.MiniMax M3 (batch)by MiniMax1.76T tokens13%
  12. 12.Kimi K3by Moonshot AI1.39T tokens3%
  13. 13.Step 3.7 Flashby StepFun1.32T tokens24%
  14. 14.Ling-3.0-flash (free)by Inclusion AI1.14T tokens12%
  15. 15.Claude Opus 5 (batch)by Anthropic1.13T tokens44%
  16. 16.Claude Sonnet 5 (batch)by Anthropic1.04T tokens2%
  17. 17.Gemini 3 Flash Preview (batch)by Google929B tokens5%
  18. 18.Claude Sonnet 4.6 (batch)by Anthropic866B tokens7%
  19. 19.GPT-5.6 Terra (batch)by OpenAI758B tokens145%
  20. 20.Gemini 2.5 Flash Lite (batch)by Google612B tokens13%

From benchmark to decision

Which model for which job?

Nine recurring production scenarios, with the trade-offs that matter before a team commits to a model.

Selected use case

01 / 09

Decision view

Customer support

Why it fits

  1. 01288 tok/s keeps wait time low
  2. 02Multilingual intake
  3. 03Strong multimodal routing

How the ranking works

Transparent enough to disagree with.

The ranking combines public benchmark signals with price, measured throughput, context and deployment constraints. It is a decision aid—not a universal definition of model quality.

Dataset snapshot
5 Aug 2026 · 13:01 TST
Review cadence
Revalidated when major model data changes

Value score

(intelligence − 25)² × min(1, speed ÷ 30) ÷ (input cost + output cost × 3)

Output is weighted ×3 to reflect output-heavy production use; intelligence is squared so cheap but weak models do not dominate.

Intelligence
Public benchmark and reasoning signals form the default order.
Speed
Output tokens per second; provider and region still affect real latency.
Cost
Public input and output price per million tokens, before enterprise discounts.
Deployment
Context, weight access and hosting constraints determine production fit.

Public references

Prices and performance can change without notice. Validate the finalists against your own prompts, infrastructure and risk requirements before procurement.

FAQ

Questions worth asking before the benchmark wins the argument.

TENTEN AI · MODEL EVALUATION

The leaderboard is the start

Bring the workflow. We’ll test the shortlist.

A model only becomes a good decision after it survives your documents, permissions, latency budget and failure cases.