Best AI Models for Coding.

Best models for code generation and debugging.

167 Models tracked
36 Providers
Daily Data refresh
10 Editorial picks
# Model DFO Score Tok/s In $/1M Out $/1M Ctx
1Claude Fable 5AnthropicTOP PICK91.556.5$10.00$50.001M
2Claude Opus 4.8AnthropicTOP PICKIN-HOUSE PICK89.352.4$5.00$25.001M
3Gemini 3.1 Pro PreviewGoogleFRONTIER85.4126$2.00$12.001M
4GLM 5.2Z.aiBEST VALUE84.7167$0.2338$0.73481M
5Claude Sonnet 5AnthropicIN-HOUSE PICKNEW84.682.8$2.00$10.001M
6Grok 4.5xAIBEST FOR CODINGNEW84.373.8$2.00$6.00500K
7GPT-5.5OpenAIBEST FOR CODING83.368.5$5.00$30.001.1M
8Qwen3.5 397B A17BQwen82.060.7$0.3900$2.34262K
9Grok 4 FastxAI82.076.7$0.2000$0.50002M
10Gemini 2.5 Pro Preview 05-06Google82.0$1.25$10.001M
11KAT-Coder-Pro V2Kwaipilot82.0103$0.3000$1.20256K
12DeepSeek-Coder-V2DeepSeek82.0FreeFree
13DeepSeek V3.2 SpecialeDeepSeek82.0$0.2870$0.4310164K
14gpt-oss-20bOpenAI82.0220$0.0300$0.1300131K
15QwQ 32BQwen82.031.9$0.1500$0.5800131K
16GPT-5.5 (medium)OpenAI82.064.4$5.00$30.00
17KAT-Coder-Pro V1Kwaipilot82.0109$0.2070$0.8280256K
18Devstral Small 1.1Mistral82.047.5$0.1000$0.3000131K
19Qwen2.5 72B InstructQwen82.055.7$0.3600$0.4000131K
20ERNIE 5.0 Thinking PreviewBaidu82.0FreeFree
Showing 1–20 of 167 Ā· Data from OpenRouter, Artificial Analysis, Hugging Face & our own testing. Scores editorially curated.

We deploy these models for businesses every week. Get a recommendation for your workload.

Get Started

Best models for code generation and debugging.

Leaderboards by use case

The overall table, re-ranked for the job you're hiring a model for.

How we rank AI models

The Design for Online AI Model Leaderboard scores 612 models on a single 0–100 scale built from four weighted dimensions: intelligence (reasoning and knowledge benchmarks), technical capability (coding and tool use), content quality (writing and instruction-following) and value (capability per dollar).

Underlying data is aggregated from the OpenRouter API for pricing and availability, Artificial Analysis for intelligence, coding and agentic indices, and the Hugging Face Open LLM Leaderboard for open-model benchmarks. The fourth source is our own: we deploy these models in client agents, chatbots and automations every week, and that internal testing feeds the editorial layer, so a model that benchmarks well but is impractical to deploy will not automatically top the table.

Models are grouped into tiers (Frontier, Professional, Specialist, Efficient, Emerging and Legacy) to make like-for-like comparison easier, and newly released models are flagged so you can see what has just landed.

Leaderboard FAQ

How often is the leaderboard updated?

Pricing, availability and benchmark data are synced daily from our sources, and editorial scores are reviewed whenever a significant new model is released.

How is the overall score calculated?

Each model is graded 0–10 on intelligence, technical capability, content quality and value; those dimensions are weighted and combined into the 0–100 overall score used to rank the table.

Where does the data come from?

From four sources: the OpenRouter API, Artificial Analysis, the Hugging Face Open LLM Leaderboard, and internal testing from real deployments by the Design for Online team.