Home > AI Business Automation > AI Leaderboards
AI Leaderboards
Choosing an AI model comes down to what it costs, what it does well and whether it suits the job in hand. We track 610 models across 90 providers and refresh the data daily, so you can compare them properly. The leaderboard is free to use with nothing to sign up for.
Overview
We built the leaderboard because we needed it. Running AI across client projects means choosing between models constantly, and the marketing rarely answers the question you are actually asking, which is whether a given model suits your workload at a price that works.
The data comes from sources we name. Pricing and availability are pulled from the OpenRouter API, intelligence and coding indices come from Artificial Analysis, and open-model benchmarks come from the Hugging Face Open LLM Leaderboard. Those refresh daily. On top of that sits our own editorial review, which is where a model that benchmarks well but is awkward to deploy gets marked down accordingly.
The part we find most useful, and the part no other leaderboard offers, is the In-House Pick badge. It marks the models we run in production ourselves. Benchmarks tell you what a model can do under test conditions. Knowing what an agency trusts with live client work tells you something different and arguably more useful.
Two tools sit alongside the main table. The comparison tool puts up to four models next to each other on score, price, context and capability. The price calculator takes your expected token usage and estimates the monthly cost across every model, which usually reframes the shortlist considerably.
If you would rather have the answer than the data, our AI consultancy service covers model selection as part of a wider look at where AI fits in your business.
What’s included
610 Models, 90 Providers
610 models from 90 providers, covering Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, xAI and the open-weight releases, with pricing and context limits for each.
Sources You Can Verify
Pricing and availability from the OpenRouter API, intelligence and coding indices from Artificial Analysis, and open-model benchmarks from the Hugging Face Open LLM Leaderboard. We name our sources so you can check them.
In-House Picks
We mark the models we run in production ourselves with an In-House Pick badge. It tells you what an agency actually trusts with client work, which is a different question from what benchmarks well.
Side-by-Side Comparison
Select up to four models and compare scores, pricing, context limits and capabilities side by side, which is usually how the shortlist gets settled.
Price Calculator
Enter your expected token usage and see estimated monthly costs across every model, so the budget conversation happens before you commit to a provider.
How we deliver
Our AI Leaderboards Process
Step 1: Track New Releases
New releases are picked up from our data sources as they appear, with specifications, pricing and context limits standardised so models can be compared like for like.
Step 2: Pull the Benchmark Data
Pricing and availability sync from OpenRouter, intelligence and coding indices from Artificial Analysis, and open-model benchmarks from Hugging Face, all refreshed daily.
Step 3: Test the Ones That Matter
Our team tests the significant releases in real work rather than relying on published benchmarks alone, which is where the In-House Pick badges come from.
Step 4: Score and Publish
Models are scored, ranked and grouped into tiers from Frontier through to Legacy, with new arrivals flagged so you can see what has just landed.
On this page
Client work
How we use it ourselves
We use the leaderboard to make our own routing decisions. Most client systems we build route between several models rather than using one throughout, sending routine steps to efficient models and reserving the frontier ones for work that genuinely needs them. That approach typically cuts running costs substantially against a single premium model, and the leaderboard is how we work out which model belongs where.
Questions
AI Leaderboards FAQs
Is the leaderboard free to use?
Yes, entirely. There is no sign-up, no account and no paywall. We built it for our own work and made it public because it is useful, and we would rather people found us that way.
Where does the data come from?
Pricing and availability come from the OpenRouter API, intelligence, coding and agentic indices from Artificial Analysis, and open-model benchmarks from the Hugging Face Open LLM Leaderboard. Those sync daily, and our editorial review sits on top of them.
How is the overall score worked out?
Each model is graded on intelligence, technical capability, content quality and value, and those dimensions are weighted into a single score out of 100. The top of the table is where the scoring does most of its work, and we would encourage you to look at the underlying indices and pricing alongside the headline number.
Which model should our business use?
It depends on the job and the budget, which is exactly why the comparison tool and price calculator exist. For most business workloads the answer is more than one model, routed by task. If you would like that worked out properly against your own use cases, that forms part of our AI consultancy service.
How current is the data?
Pricing, availability and benchmark data sync daily. Editorial scores and the In-House Pick badges are reviewed whenever a significant model launches, which at the moment is happening every few weeks.
Claude Academy training supporting our AI services.
We have completed all 19 certifications currently available, covering practical AI use, automation, development, data integrations and training. This helps us apply Claude more effectively across client projects.





















