AI · LEADERBOARD

Escolha o modelo certo para o trabalho.

Uma comparação prática dos modelos de IA que equipes podem implantar agora.

Snapshot: 25 Aug 2026 · 10:41 TST16 models compared

Independent editorial snapshot · not live platform telemetry

The short answer

Three leaders, three different decisions.

Best overall

Closed model

Anthropic

Intelligence
63.1
Speed
58.8tok/s
Input / output
$5 / $25
Context
1M

Best fit

State-of-the-art reasoning and complex agentic workflows

Best open model

Open weights

Kimi

Intelligence
59.7
Speed
35tok/s
Input / output
$3 / $15
Context
1.05M

Best fit

Sovereign, open-weight deployments needing maximum intelligence

Best value

Open weights

DeepSeek

Intelligence
51.8
Speed
115tok/s
Input / output
$0.14 / $0.28
Context
1.31M

Best fit

Cost-sensitive open inference and coding workloads

OpenRouter usage

OpenRouter LLM usage ranking

Top models by prompt, completion, and reasoning tokens processed on OpenRouter. Showing the first 20 only.

Top 20OpenRouter data · cached hourlyUpstream · Sep 5, 2026
  1. 1.Hy4 previewby Tencent14.1T tokens639%
  2. 2.GLM 5.3 Flash (batch)by Z.ai12.5T tokens170%
  3. 3.DeepSeek V4 Flash 0731 (batch)by DeepSeek12.3T tokens0%
  4. 4.GPT-5.6 Luna (batch)by OpenAI12.2T tokens80%
  5. 5.MiniMax M3 (free)by MiniMax5.56T tokens206%
  6. 6.DeepSeek V4 Flash 0423by DeepSeek5.24T tokens3%
  7. 7.Hy3by Tencent4.44T tokens33%
  8. 8.Nemotron 3 Ultra (free)by Nvidia3.65T tokens35%
  9. 9.GLM 5.3by Z.ai2.81T tokens127%
  10. 10.MiMo-V2.5by Xiaomi2.76T tokens72%
  11. 11.GLM 5.2 (free)by Z.ai2.3T tokens26%
  12. 12.Gemini 3.7 Flash (batch)by Google2.26T tokens44%
  13. 13.Kimi K3 (batch)by Moonshot AI2.03T tokens35%
  14. 14.GPT-5.6 Sol (batch)by OpenAI1.87T tokens8%
  15. 15.Claude Opus 5 (batch)by Anthropic1.66T tokens11%
  16. 16.MiniMax M3 (free)by MiniMax1.47T tokens2%
  17. 17.Laguna S 2.1 (free)by Poolside1.35T tokens1%
  18. 18.DeepSeek V4 Pro 0423by DeepSeek1.33T tokens28%
  19. 19.Claude Sonnet 5 (batch)by Anthropic1.31T tokens10%
  20. 20.Nemotron 3.5 Lightning (free)by Nvidia1.1T tokens22%

From benchmark to decision

Which model for which job?

Nine recurring production scenarios, with the trade-offs that matter before a team commits to a model.

Selected use case

01 / 09

Decision view

Customer support

Why it fits

  1. 01High measured throughput
  2. 02Low per-token cost
  3. 03Strong intelligence for support

How the ranking works

Transparent enough to disagree with.

The ranking combines public benchmark signals with price, measured throughput, context and deployment constraints. It is a decision aid, not a universal definition of model quality.

Dataset snapshot
25 Aug 2026 · 10:41 TST
Review cadence
Revalidated weekly and after major releases

Value score

(intelligence − 25)² × min(1, speed ÷ 30) ÷ (input cost + output cost × 3)

Output is weighted ×3 to reflect output-heavy production use; intelligence is squared so cheap but weak models do not dominate.

Intelligence
Public benchmark and reasoning signals form the default order.
Speed
Output tokens per second; provider and region still affect real latency.
Cost
Public input and output price per million tokens, before enterprise discounts.
Deployment
Context, weight access and hosting constraints determine production fit.

Public references

Prices and performance can change without notice. Validate the finalists against your own prompts, infrastructure and risk requirements before procurement.

FAQ

Questions worth asking before the benchmark wins the argument.

TENTEN AI · MODEL EVALUATION

The leaderboard is the start

Bring the workflow. We’ll test the shortlist.

A model only becomes a good decision after it survives your documents, permissions, latency budget and failure cases.