AI Leaderboards

Choosing an AI model comes down to what it costs, what it does well and whether it suits the job in hand. We track 610 models across 90 providers and refresh the data daily, so you can compare them properly. The leaderboard is free to use with nothing to sign up for.

AI Leaderboards for Business, Benchmark-Driven Rankings
Overview

We built the leaderboard because we needed it. Running AI across client projects means choosing between models constantly, and the marketing rarely answers the question you are actually asking, which is whether a given model suits your workload at a price that works.

The data comes from sources we name. Pricing and availability are pulled from the OpenRouter API, intelligence and coding indices come from Artificial Analysis, and open-model benchmarks come from the Hugging Face Open LLM Leaderboard. Those refresh daily. On top of that sits our own editorial review, which is where a model that benchmarks well but is awkward to deploy gets marked down accordingly.

The part we find most useful, and the part no other leaderboard offers, is the In-House Pick badge. It marks the models we run in production ourselves. Benchmarks tell you what a model can do under test conditions. Knowing what an agency trusts with live client work tells you something different and arguably more useful.

Two tools sit alongside the main table. The comparison tool puts up to four models next to each other on score, price, context and capability. The price calculator takes your expected token usage and estimates the monthly cost across every model, which usually reframes the shortlist considerably.

If you would rather have the answer than the data, our AI consultancy service covers model selection as part of a wider look at where AI fits in your business.

What’s included
01

610 Models, 90 Providers

610 models from 90 providers, covering Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, xAI and the open-weight releases, with pricing and context limits for each.

02

Sources You Can Verify

Pricing and availability from the OpenRouter API, intelligence and coding indices from Artificial Analysis, and open-model benchmarks from the Hugging Face Open LLM Leaderboard. We name our sources so you can check them.

03

In-House Picks

We mark the models we run in production ourselves with an In-House Pick badge. It tells you what an agency actually trusts with client work, which is a different question from what benchmarks well.

04

Side-by-Side Comparison

Select up to four models and compare scores, pricing, context limits and capabilities side by side, which is usually how the shortlist gets settled.

05

Price Calculator

Enter your expected token usage and see estimated monthly costs across every model, so the budget conversation happens before you commit to a provider.

How we deliver

Our AI Leaderboards Process

01

Step 1: Track New Releases

New releases are picked up from our data sources as they appear, with specifications, pricing and context limits standardised so models can be compared like for like.

02

Step 2: Pull the Benchmark Data

Pricing and availability sync from OpenRouter, intelligence and coding indices from Artificial Analysis, and open-model benchmarks from Hugging Face, all refreshed daily.

03

Step 3: Test the Ones That Matter

Our team tests the significant releases in real work rather than relying on published benchmarks alone, which is where the In-House Pick badges come from.

04

Step 4: Score and Publish

Models are scored, ranked and grouped into tiers from Frontier through to Legacy, with new arrivals flagged so you can see what has just landed.

Client work

How we use it ourselves

We use the leaderboard to make our own routing decisions. Most client systems we build route between several models rather than using one throughout, sending routine steps to efficient models and reserving the frontier ones for work that genuinely needs them. That approach typically cuts running costs substantially against a single premium model, and the leaderboard is how we work out which model belongs where.

Design for Online, in production
Still have questions? Use our live chat or call 01284 245 170, Mon–Fri 9am–5pm.
Questions

AI Leaderboards FAQs

Is the leaderboard free to use?

Yes, entirely. There is no sign-up, no account and no paywall. We built it for our own work and made it public because it is useful, and we would rather people found us that way.

Where does the data come from?

Pricing and availability come from the OpenRouter API, intelligence, coding and agentic indices from Artificial Analysis, and open-model benchmarks from the Hugging Face Open LLM Leaderboard. Those sync daily, and our editorial review sits on top of them.

How is the overall score worked out?

Each model is graded on intelligence, technical capability, content quality and value, and those dimensions are weighted into a single score out of 100. The top of the table is where the scoring does most of its work, and we would encourage you to look at the underlying indices and pricing alongside the headline number.

Which model should our business use?

It depends on the job and the budget, which is exactly why the comparison tool and price calculator exist. For most business workloads the answer is more than one model, routed by task. If you would like that worked out properly against your own use cases, that forms part of our AI consultancy service.

How current is the data?

Pricing, availability and benchmark data sync daily. Editorial scores and the In-House Pick badges are reviewed whenever a significant model launches, which at the moment is happening every few weeks.

Still have questions? Use our live chat or call our team on 01284 245 170, Mon–Fri 9am–5pm.
19
certifications completed
Claude Academy certified

Claude Academy training supporting our AI services.

We have completed all 19 certifications currently available, covering practical AI use, automation, development, data integrations and training. This helps us apply Claude more effectively across client projects.

Completed through Anthropic’s Claude Academy · select any badge to view its public verification page.
AI Fluency: Framework & Foundations
Covers a structured approach to planning tasks, working with AI and reviewing the results.
AI Capabilities and Limitations
Covers where AI can be useful, where its limitations apply and how its output should be checked.
Claude 101
Covers the main features of Claude and how they can be used effectively in day-to-day work.
Claude Platform 101
Provides an overview of the Claude platform and the tools available for different types of work.
Building with the Claude API
Supports our work building custom AI automations, integrations and workflows using the Claude API.
Claude Code 101
Covers the use of Claude Code for development, technical tasks and maintaining digital systems.
Claude Code in Action
Develops the practical use of Claude Code across real development workflows and projects.
Introduction to Model Context Protocol
Covers how Claude can connect with business systems and data through Model Context Protocol.
Model Context Protocol: Advanced Topics
Covers more advanced MCP integrations, including access controls and reliable system design.
Introduction to Claude Cowork
Covers the use of Claude Cowork for shared workflows, planning and collaborative tasks.
Claude with Amazon Bedrock
Covers how Claude can be deployed and managed through Amazon Bedrock within AWS.
Claude with Google Cloud's Vertex AI
Covers how Claude can be deployed and managed through Vertex AI within Google Cloud.
AI Fluency for Builders
Focuses on using AI to support the design and development of practical tools and products.
AI Fluency for Small Businesses
Looks at practical ways small businesses can introduce AI into their everyday operations.
AI Fluency for Nonprofits
Covers the practical use of AI within charities, nonprofits and resource-conscious teams.
AI Fluency for Educators
Covers how AI can be introduced and explained clearly within learning and training environments.
AI Fluency for pK–12 Educators
Covers appropriate and responsible AI use within schools and education settings.
AI Fluency for Students
Provides a foundation in using AI effectively, responsibly and with an understanding of its limitations.
Teaching AI Fluency
Supports our ability to provide structured AI training and guidance for client teams.
View our accreditations →

Latest news

ChatGPT Ads are now here in the UK — Design for Online campaign mockup showing a labelled ad in ChatGPT

ChatGPT Ads Are Here in the UK: Our Guide After Launching a Campaign

ChatGPT Ads are now available in the UK. We launched a campaign in OpenAI Ads Manager and share how they work, what they cost, and ...
Design for Online® new office at Stowmarket Innovation Gateway, Gateway 14

We’ve Got a New Home!

We’ve opened our new office and studio at Stowmarket Innovation Gateway, creating a new home for Design for Online® in Suffolk.
Celebrating 10 years of Design For Online

10 years of Design for Online®

Design for Online® celebrates 10 years in business, a new Stowmarket office and studio, and a decade of helping businesses grow across Suffolk and the ...
© 2026 Design for Online Ltd. Registered in England and Wales No. 10328553. VAT Registered. Design for Online® and Forerunner® are registered trademarks.