# Model DFO Score Tok/s In $/1M Out $/1M Ctx
1GPT-5.6 SolOpenAIBEST FOR CODING87.180.9$2.00$10.001.1M
2Claude Opus 5AnthropicTOP PICKIN-HOUSE PICK87.154.3$5.00$25.001M
3Grok 4.6SpaceXAIBEST FOR CODINGIN-HOUSE PICKNEW86.664.9$2.00$6.00500K
4Qwen3.8 MaxQwenTOP PICK86.338.7$2.00$6.001M
5GLM 5.3 FlashZ.aiFRONTIERNEW86.247.4$0.0750$0.25001.3M
6GLM 5.3Z.aiFRONTIERNEW85.780.0$1.40$4.401.3M
7Claude Fable 5.1AnthropicTOP PICKNEW85.667.3$10.00$50.001M
8Claude Fable 5AnthropicTOP PICK85.270.0$10.00$50.001M
9Claude Opus 4.8Anthropic84.357.3$5.00$25.001M
10Gemini 3.7 FlashGoogleNEW84.2320$0.7500$3.751M
11Muse Spark 1.2Meta84.1245$1.25$4.251M
12Qwen3.8 2.4T A95B (batch)QwenNEW83.940.0$2.00$6.001M
13GPT-5.6 TerraOpenAI83.8108$2.00$12.001.1M
14Qwen3.8 2.4T A95BQwenNEW83.623.9$2.00$6.001M
15GPT-5.6 LunaOpenAI83.5128$0.2000$1.201.1M
16Grok 4.5xAIBEST FOR CODING83.257.6$2.00$6.00500K
17Claude Sonnet 5AnthropicIN-HOUSE PICK83.069.2$2.00$10.001M
18Kimi K3MoonshotAI82.738.7$3.00$15.001M
19Muse Spark 1.1Meta82.0178$1.25$4.251M
20Claude Opus 4.7Anthropic81.848.4$5.00$25.001M
Showing 1–20 of 490 · Data from OpenRouter, Artificial Analysis, Hugging Face & our own testing. Scores editorially curated.

We deploy these models for businesses every week. Get a recommendation for your workload.

Get Started

Models with strong tool-use and function-calling support.

Leaderboards by use case

The overall table, re-ranked for the job you're hiring a model for.

How we rank AI models

The Design for Online AI Model Leaderboard scores 751 models on a single 0–100 scale built from four weighted dimensions: intelligence (reasoning and knowledge benchmarks), technical capability (coding and tool use), content quality (writing and instruction-following) and value (capability per dollar).

Underlying data is aggregated from the OpenRouter API for pricing and availability, Artificial Analysis for intelligence, coding and agentic indices, and the Hugging Face Open LLM Leaderboard for open-model benchmarks. The fourth source is our own: we deploy these models in client agents, chatbots and automations every week, and that internal testing feeds the editorial layer, so a model that benchmarks well but is impractical to deploy will not automatically top the table.

Models are grouped into tiers (Frontier, Professional, Specialist, Efficient, Emerging and Legacy) to make like-for-like comparison easier, and newly released models are flagged so you can see what has just landed.

Leaderboard FAQ

How often is the leaderboard updated?

Pricing, availability and benchmark data are synced daily from our sources, and editorial scores are reviewed whenever a significant new model is released.

How is the overall score calculated?

Each model is graded 0–10 on intelligence, technical capability, content quality and value; those dimensions are weighted and combined into the 0–100 overall score used to rank the table.

Where does the data come from?

From four sources: the OpenRouter API, Artificial Analysis, the Hugging Face Open LLM Leaderboard, and internal testing from real deployments by the Design for Online team.

© 2026 Design for Online Ltd. Registered in England and Wales No. 10328553. VAT Registered. Design for Online® and Forerunner® are registered trademarks.