Here we are in August 2026, and since writing this article in February we have had to keep updating it. The leaderboard we published back then looks almost nothing like the one we are looking at today.
Anthropic put Claude Opus 5 at the top of our board in late July and it has not been shifted since. SpaceXAI released Grok 4.6 on 12 August and it climbed to second place inside a few days, undercutting the model above it by three quarters on output pricing, something we really were hoping to see!
For most business owners that pace can become a pain rather than an opportunity, especially if AI providers sunset older models setup within existing workflows. Retesting your AI stack every six weeks is not likely feasible and something which isn’t likely required unless the model is a significant improvement or offers ongoing cost savings. What matters is choosing something that will still be a sensible decision in six months, and knowing which jobs are worth paying more for.
We have ranked the five we would actually put in front of a client this month, scored for business use on our live AI Model Leaderboard rather than on raw benchmark performance. We run these on client work every day, so the notes underneath each one are what we have found in practice. Prices are per million tokens on the API and have nothing to do with what you pay for a chat subscription.
1. Claude Opus 5

Opus took the top of our board in late July on 87.1 and has stayed there, it is the model we reach for when a mistake would be expensive to put right. That means client deliverables, long documents that need reading properly rather than skimming, and coding agents that have to see a job through to the end instead of stopping three quarters of the way in. A 1M context window, vision and tool use all come as standard, so you are not paying extra to attach a file.
The catch is that Anthropic priced it for that work and no other. At $5 in and $25 out it costs more than four times what Grok charges to produce the same volume, which is perfectly reasonable on a piece of work worth a few thousand pounds and potentially overkill on a chatbot answering questions about opening hours.
2. Grok 4.6

SpaceXAI shipped 4.6 on 12 August and it reached second place on our board within days, which is quick even by this year’s standards. What made people sit up was not the 86.6 score but the price attached to it. At $2 in and $6 out it does frontier coding for roughly 40% of what Opus charges on input and less than a quarter on output, and it ranks above Opus for agentic work while it is at it.
It holds up on the jobs that usually expose a cheaper model, staying with multi-step tasks, working through real codebases, and turning a brief into something finished rather than a first draft somebody then has to rescue. Vision, files and tools are all included. This is the model which is also powering the new Grok Bot which is certainly going to be a big benefit for businesses. The one genuine compromise is a 500K context window against the 1M offered either side of it, although we have not hit that ceiling on operational work yet and most businesses will not either.
3. Qwen3.8 Max

Qwen sits third on 86.3, matches Grok at $2 and $6, and adds a 1M context window with genuinely capable handling of text, images and video. On paper it is the best value on this page. Because Chinese models are currently working very fast with their model capability and pushing cheaper and more capable models, it means there are other lower cost options available for businesses.
Governance is a caution here, for example if UK or EU client material is going through the hosted API then you need to know where it ends up, and that is a conversation worth having before the model quietly becomes part of how the business runs. Self-hosting or a regional deployment changes the answer completely.
4. GPT-5.6 Sol

Seventh place on 83.4 makes Sol look ordinary and it is anything but. It ranks #2 for coding, ahead of everything else here including Opus, and carries the largest context window on the page at 1.1M. OpenAI has priced it like a flagship at $5 and $30, and on hard engineering work it behaves like one.
Whether any of that matters to you comes down to what you are already paying for. A team running Codex, Agents or ChatGPT enterprise gets a top tier model inside the tooling and the billing it has already committed to, with Terra and Luna underneath to soak up the cheap volume, and moving away would cost more in disruption than it saved. Coming to it fresh the case is much thinner, because Grok sits two places behind it for coding and charges a fifth as much to produce the output.
5. Claude Sonnet 5

Sonnet 5 arrived in June, scores 83.0, has a 1M context window, and is still on introductory pricing of $2 and $10 until the end of this month, after which it rises to $3 and $15.
It loses to Opus on the hardest jobs, which matters far less than it sounds, because most work is not the hardest job. It writes well, it sees multi-step tasks through more reliably than any Sonnet before it, and it keeps content, research and client automation at a monthly cost you can actually forecast.
The one thing to watch is the newer tokeniser used, this can turn the same document into more tokens, where a model like Opus 5 while more expensive may achieve the task quicker.
The five at a glance
| Model | DFO score | Input / Output | Context | Best for |
|---|---|---|---|---|
| Claude Opus 5 | 87.1 | $5 / $25 | 1M | High-stakes work |
| Grok 4.6 | 86.6 | $2 / $6 | 500K | Coding and agents |
| Qwen3.8 Max | 86.3 | $2 / $6 | 1M | Multimodal agents |
| GPT-5.6 Sol | 83.4 | $5 / $30 | 1.1M | OpenAI coding stack |
| Claude Sonnet 5 | 83.0 | $2 / $10* | 1M | Everyday business use |
Head-to-head on our leaderboard
The overall score hides a lot. These are the metrics we refresh daily, and they explain why a model ranked ninth for intelligence can still be the right one for your business.
| Metric | Opus 5 | Grok 4.6 | Qwen3.8 Max | GPT-5.6 Sol | Sonnet 5 |
|---|---|---|---|---|---|
| Overall DFO score | 87.1 | 86.6 | 86.3 | 83.4 | 83.0 |
| Intelligence rank | #1 | #2 | #3 | #7 | #9 |
| Coding rank | #3 | #4 | #16 | #2 | #18 |
| Agentic rank | #7 | #6 | #7 | #9 | #16 |
| Context | 1M | 500K | 1M | 1.1M | 1M |
| Input price / 1M | $5.00 | $2.00 | $2.00 | $5.00 | $2.00 |
| Output price / 1M | $25.00 | $6.00 | $6.00 | $30.00 | $10.00 |
Our verdict
The arrangement we recommend to most businesses has changed throughout the year, with Google’s Gemini 3.1 being our initial recommendaiton, moving to Sonnet and Opus 5. Opus 5 can handle the hardest client work, Sonnet 5 takes mostly everything else, with the flagship called in for the exceptions, we are noticing another shift as newer models release from other providers and we are also keeping a close eye on Meta’s models including Muse Spark, Gemini 3.5 and Kimi.
What has changed mostly this month sits in the middle of the range. Grok 4.6 does frontier coding and agent work at $2 and $6, which simply was not on offer in July, and it is the first release in a while that has sent us back to look at where our own escalation line should sit. For teams already inside OpenAI, Sol takes the coding-led work and Terra or Luna handle the volume.
Fable 5 – this comes up in nearly every conversation we have about this, but we haven’t chose to rank it in our top five. It is still excellent on long, complex jobs, but at $10 and $50 it is a specialist rather than a daily driver, which is exactly why it sits below Opus 5 on our board. Use it for the hard cases and go back to Sonnet or Grok for everything else.
What is still due in 2026
Google will be releasing their latest models soon including Gemini 3.5 Pro, while not available yet, flash variants have shipped in the meantime, so the hold-up appears to be with the larger model rather than the family, and whenever Pro does land it has the potential to move the reasoning rankings quickly.
Grok 4.7 is already being talked about as a follow-on to 4.6, with no firm date, pricing or model card yet but it is likely only weeks away. Open-weight models keep taking on work that needed a frontier API a year ago, and the pricing is very hard to argue with, but data residency still needs checking before client material goes anywhere near them.
We will update this guide as they arrive, which on current form will be sooner than any of us expect. Day to day movement across the wider field is on the live AI Model Leaderboard.
How we can help
Choosing the model is the smallest part of this. It only pays for itself once it is wired into the systems your staff already use, with routing and approval rules that match how the work actually happens, and that is the part most businesses underestimate.
If you want a hand doing that against a real workload, have a look at our AI business automation and AI consultancy services, or compare live scores on the AI Model Leaderboard.
About Design for Online®: We are a full service digital marketing agency in Bury St Edmunds, Suffolk. Our services include web design, SEO, PPC, AI business automation, photography and videography, and we work with businesses across the UK to help them grow online. For enquiries: hello@designforonline.com