Best overall
Closed model
Anthropic
- Intelligence
- 63.1
- Speed
- 58.8tok/s
- Input / output
- $5 / $25
- Context
- 1M
Best fit
State-of-the-art reasoning and complex agentic workflows
AI · LEADERBOARD
A practical ranking of the AI models teams can deploy now, compared by intelligence, speed, cost, context and operational fit.
Independent editorial snapshot · not live platform telemetry
The short answer
Best overall
Closed model
Anthropic
Best fit
State-of-the-art reasoning and complex agentic workflows
Best open model
Open weights
Kimi
Best fit
Sovereign, open-weight deployments needing maximum intelligence
Best value
Open weights
DeepSeek
Best fit
Cost-sensitive open inference and coding workloads
OpenRouter usage
Top models by prompt, completion, and reasoning tokens processed on OpenRouter. Showing the first 20 only.
From benchmark to decision
Nine recurring production scenarios, with the trade-offs that matter before a team commits to a model.
Selected use case
01 / 09
Decision view
Why it fits
How the ranking works
The ranking combines public benchmark signals with price, measured throughput, context and deployment constraints. It is a decision aid, not a universal definition of model quality.
Value score
(intelligence − 25)² × min(1, speed ÷ 30) ÷ (input cost + output cost × 3)
Output is weighted ×3 to reflect output-heavy production use; intelligence is squared so cheap but weak models do not dominate.
Public references
Prices and performance can change without notice. Validate the finalists against your own prompts, infrastructure and risk requirements before procurement.
FAQ
The first row and Best overall card show the current intelligence-first leader. The best choice can change once price, latency, hosting and task-specific evaluation enter the decision.
The Best open model card is recalculated from models with verified downloadable weights, using the same intelligence-first ordering as the main table.
The Best value card is recalculated from the published value formula whenever verified intelligence, speed or base token pricing changes.
No. Context capacity only describes how much can be submitted. Retrieval design, instruction quality, attention behaviour and evaluation still determine whether that context produces a useful answer.
A scheduled pipeline revalidates the dataset weekly and after major releases. Failed or conflicting source checks keep the last verified snapshot online instead of publishing partial data.
The default order uses the public intelligence signal. Speed, price, context, openness and a separate value formula remain visible so teams can challenge the default order.
Tenten does not sell model tokens through this page. Public sources are linked, assumptions are shown, and we recommend a task-specific evaluation before any production decision.
TENTEN AI · MODEL EVALUATION
The leaderboard is the start
A model only becomes a good decision after it survives your documents, permissions, latency budget and failure cases.
Model explorer
No single score tells the whole story. Narrow the list, inspect the trade-offs, then compare the finalists side by side.
16models
| Compare models | Model | Best for | ||||||
|---|---|---|---|---|---|---|---|---|
| #1 | 63.1 | 58.8 tok/s | 18.1 | $5 / $25input / output | 1M | State-of-the-art reasoning and complex agentic workflows | ||
| #2 | 62.1 | 71 tok/s | 8.6 | $10 / $50input / output | 1M | High-reliability enterprise work with fallback safety | ||
| #3 | 60.9 | 73.6 tok/s | 40.3 | $2 / $10input / output | 1.1M | Enterprise reasoning and large-context problem solving | ||
| #4 | 60.9 | 61.9 tok/s | 64.5 | $2 / $6input / output | 500K | High-intelligence proprietary reasoning at competitive token pricing | ||
| #5 | 59.7 | 35 tok/s | 25.1 | $3 / $15input / output | 1.0M | Sovereign, open-weight deployments needing maximum intelligence | ||
| #6 | 58.1 | 23 tok/s | 42.0 | $2 / $6input / output | 1M | High-intelligence multilingual work and long-form translation | ||
| #7 | 56.8 | N/A | N/A | $1.25 / $4.25input / output | 1.0M | Massive-context analysis and structured data extraction | ||
| #8 | 56.6 | 121.5 tok/s | 26.2 | $2 / $12input / output | 1.1M | Fast enterprise reasoning with balanced economics | ||
| #9 | 55.8 | 58 tok/s | 47.4 | $2 / $6input / output | 500K | Fast reasoning where live information matters | ||
| #10 | 55.3 | 69 tok/s | 28.7 | $2 / $10input / output | 1M | Balanced agents, coding and long-context workflows | ||
| #11 | 53.2 | 71.7 tok/s | 70.9 | $1.12 / $3.37input / output | 1.0M | Low-cost open-weight reasoning and coding workloads | ||
| #12 | 52.6 | 152.8 tok/s | 61.4 | $1.19 / $3.74input / output | 1.0M | Sovereign deployments that still need strong reasoning | ||
| #13 | 52.3 | 140.7 tok/s | 196.4 | $0.20 / $1.2input / output | 1.1M | High-volume customer support with low latency and controlled cost | ||
| #14 | 51.8 | 115 tok/s | 732.9 | $0.14 / $0.28input / output | 1.3M | Cost-sensitive open inference and coding workloads | ||
| #15 | 51.6 | 221 tok/s | 59.0 | $0.75 / $3.75input / output | 1.0M | Real-time support, translation and multimodal intake | ||
| #16 | 38.0 | 82 tok/s | 172.4 | $0.14 / $0.28input / output | 1.1M | Cost-sensitive extraction and high-volume structured work |
Best fit
State-of-the-art reasoning and complex agentic workflows
Best fit
High-reliability enterprise work with fallback safety
Best fit
Enterprise reasoning and large-context problem solving
Best fit
High-intelligence proprietary reasoning at competitive token pricing
Best fit
Sovereign, open-weight deployments needing maximum intelligence
Best fit
High-intelligence multilingual work and long-form translation
Best fit
Massive-context analysis and structured data extraction
Best fit
Fast enterprise reasoning with balanced economics
Best fit
Fast reasoning where live information matters
Best fit
Balanced agents, coding and long-context workflows
Best fit
Low-cost open-weight reasoning and coding workloads
Best fit
Sovereign deployments that still need strong reasoning
Best fit
High-volume customer support with low latency and controlled cost
Best fit
Cost-sensitive open inference and coding workloads
Best fit
Real-time support, translation and multimodal intake
Best fit
Cost-sensitive extraction and high-volume structured work