Best overall
Closed model
Anthropic
- Intelligence
- 59.9
- Speed
- 70tok/s
- Input / output
- $10 / $50
- Context
- 1M
Best fit
High-stakes reasoning and complex agentic workflows
AI · LEADERBOARD
مقارنة عملية لنماذج الذكاء الاصطناعي التي يمكن للفرق نشرها الآن.
Independent editorial snapshot · not live platform telemetry
The short answer
Best overall
Closed model
Anthropic
Best fit
High-stakes reasoning and complex agentic workflows
Best open model
Open weights
Z AI
Best fit
Sovereign deployments that still need strong reasoning
Best value
Open weights
Xiaomi
Best fit
Cost-sensitive extraction and high-volume structured work
OpenRouter usage
Top models by prompt, completion, and reasoning tokens processed on OpenRouter. Showing the first 20 only.
From benchmark to decision
Nine recurring production scenarios, with the trade-offs that matter before a team commits to a model.
Selected use case
01 / 09
Decision view
Why it fits
How the ranking works
The ranking combines public benchmark signals with price, measured throughput, context and deployment constraints. It is a decision aid—not a universal definition of model quality.
Value score
(intelligence − 25)² × min(1, speed ÷ 30) ÷ (input cost + output cost × 3)
Output is weighted ×3 to reflect output-heavy production use; intelligence is squared so cheap but weak models do not dominate.
Public references
Prices and performance can change without notice. Validate the finalists against your own prompts, infrastructure and risk requirements before procurement.
FAQ
Claude Fable 5 leads the intelligence-first ranking in this captured dataset. The best choice can change once price, latency, hosting and task-specific evaluation enter the decision.
GLM-5.2 is the leading open-weight option in this snapshot, combining a strong intelligence signal, high throughput and sovereign deployment flexibility.
MiMo-V2.5 leads the value formula because its low token price and large context offset a lower raw intelligence score. It is particularly compelling for structured, high-volume work.
No. Context capacity only describes how much can be submitted. Retrieval design, instruction quality, attention behaviour and evaluation still determine whether that context produces a useful answer.
This page is an editorial snapshot, not live platform telemetry. We revalidate the dataset when major model releases or material pricing and performance changes occur.
The default order uses the public intelligence signal. Speed, price, context, openness and a separate value formula remain visible so teams can challenge the default order.
Tenten does not sell model tokens through this page. Public sources are linked, assumptions are shown, and we recommend a task-specific evaluation before any production decision.
TENTEN AI · MODEL EVALUATION
The leaderboard is the start
A model only becomes a good decision after it survives your documents, permissions, latency budget and failure cases.
Model explorer
No single score tells the whole story. Narrow the list, inspect the trade-offs, then compare the finalists side by side.
12models
| Compare models | Model | Best for | ||||||
|---|---|---|---|---|---|---|---|---|
| #1 | 59.9 | 70 tok/s | 5.8 | $10 / $50input / output | 1M | High-stakes reasoning and complex agentic workflows | ||
| #2 | 58.9 | 68 tok/s | 9.4 | $5 / $30input / output | 1M | Enterprise reasoning and large-context problem solving | ||
| #3 | 57.1 | 39 tok/s | 12.7 | $3 / $15input / output | 1M | Ultra-long context work with open deployment options | ||
| #4 | 55.7 | 62 tok/s | 8.8 | $5 / $25input / output | 1M | Deep analysis and complex document processing | ||
| #5 | 55.0 | 154 tok/s | 18.4 | $2.5 / $15input / output | 1M | Fast enterprise reasoning with balanced economics | ||
| #6 | 53.8 | 75 tok/s | 31.2 | $2 / $6input / output | 500K | Fast reasoning where live information matters | ||
| #7 | 53.4 | 91 tok/s | 23.1 | $2 / $10input / output | 1M | Balanced agents, coding and long-context workflows | ||
| #8 | 51.2 | 212 tok/s | 37.8 | $1 / $6input / output | 1M | High-throughput reasoning at a controlled cost | ||
| #9 | 51.1 | 198 tok/s | 45.3 | $1.4 / $4.4input / output | 1M | Sovereign deployments that still need strong reasoning | ||
| #10 | 50.6 | 122 tok/s | 43.8 | $1.25 / $4.25input / output | 1.0M | Large-scale document analysis and structured extraction | ||
| #11 | 50.2 | 294 tok/s | 24.2 | $1.5 / $9input / output | 1M | Real-time support, translation and multimodal intake | ||
| #12 | 37.2 | 64 tok/s | 171.1 | $0.14 / $0.29input / output | 1.0M | Cost-sensitive extraction and high-volume structured work |
Best fit
High-stakes reasoning and complex agentic workflows
Best fit
Enterprise reasoning and large-context problem solving
Best fit
Ultra-long context work with open deployment options
Best fit
Deep analysis and complex document processing
Best fit
Fast enterprise reasoning with balanced economics
Best fit
Fast reasoning where live information matters
Best fit
Balanced agents, coding and long-context workflows
Best fit
High-throughput reasoning at a controlled cost
Best fit
Sovereign deployments that still need strong reasoning
Best fit
Large-scale document analysis and structured extraction
Best fit
Real-time support, translation and multimodal intake
Best fit
Cost-sensitive extraction and high-volume structured work