Best AI Models for SEO.

Best models for search engine optimisation tasks.

36 Models tracked
14 Providers
Daily Data refresh
1 Editorial picks
# Model DFO Score Tok/s In $/1M Out $/1M Ctx
1Nova 2.0 Lite (Non-reasoning)Amazon82.0204$0.3000$2.50
2DeepSeek V3.2 SpecialeDeepSeek82.0$0.2870$0.4310164K
3Qwen3.5-9BQwen82.051.5$0.1000$0.1500262K
4Llama Nemotron Super 49B v1.5 (Non-reasoning)NVIDIA82.063.7$0.1000$0.4000
5Gemini 2.5 Flash Lite Preview 09-2025Google82.0348$0.1000$0.40001M
6Mistral Small 4Mistral82.0156$0.1500$0.6000262K
7Qwen3.5 2B (Non-reasoning)Alibaba82.023.7$0.0200$0.1000
8Qwen3 30B A3B Thinking 2507Qwen82.0145$0.1300$1.5682K
9Gemma 4 31BGoogle82.035.9$0.1400$0.4000262K
10Qwen3.5 4B (Non-reasoning)Alibaba82.020.7$0.0300$0.1500
11GPT-5 NanoOpenAI82.0152$0.0500$0.4000400K
12Gemma 4 26B A4BGoogle82.065.3$0.1200$0.3500262K
13Gemini 1.5 Flash (Sep ’24)Google82.0FreeFree
14Mistral Small 3.2 24BMistral82.0145$0.1000$0.3000256K
15DeepSeek V4 FlashDeepSeekBEST VALUE82.0101$0.0980$0.19601M
16Gemini 1.5 Flash-8BGoogle82.0FreeFree
17GPT-4o-mini Search PreviewOpenAI82.053.0$0.1500$0.6000128K
18Mistral Large 3Mistral82.049.3$0.5000$1.50
19SabaMistral82.0$0.2000$0.600033K
20Magistral Small 1.2Mistral82.083.5$0.5000$1.50
Showing 1–20 of 36 Ā· Data from OpenRouter, Artificial Analysis, Hugging Face & our own testing. Scores editorially curated.

We deploy these models for businesses every week. Get a recommendation for your workload.

Get Started

Best models for search engine optimisation tasks.

Leaderboards by use case

The overall table, re-ranked for the job you're hiring a model for.

How we rank AI models

The Design for Online AI Model Leaderboard scores 620 models on a single 0–100 scale built from four weighted dimensions: intelligence (reasoning and knowledge benchmarks), technical capability (coding and tool use), content quality (writing and instruction-following) and value (capability per dollar).

Underlying data is aggregated from the OpenRouter API for pricing and availability, Artificial Analysis for intelligence, coding and agentic indices, and the Hugging Face Open LLM Leaderboard for open-model benchmarks. The fourth source is our own: we deploy these models in client agents, chatbots and automations every week, and that internal testing feeds the editorial layer, so a model that benchmarks well but is impractical to deploy will not automatically top the table.

Models are grouped into tiers (Frontier, Professional, Specialist, Efficient, Emerging and Legacy) to make like-for-like comparison easier, and newly released models are flagged so you can see what has just landed.

Leaderboard FAQ

How often is the leaderboard updated?

Pricing, availability and benchmark data are synced daily from our sources, and editorial scores are reviewed whenever a significant new model is released.

How is the overall score calculated?

Each model is graded 0–10 on intelligence, technical capability, content quality and value; those dimensions are weighted and combined into the 0–100 overall score used to rank the table.

Where does the data come from?

From four sources: the OpenRouter API, Artificial Analysis, the Hugging Face Open LLM Leaderboard, and internal testing from real deployments by the Design for Online team.