xAI: Grok 4.20 Multi-Agent Beta

x-ai · Released Mar 12, 2026 BEST FOR AGENTS
DFO score #40 / 688
54.1 leaderboard score
Speed #20 / 328
241.8 tok/s
Input Price #572 / 688
$2.00 per 1M tokens
Output Price #559 / 688
$6.00 per 1M tokens
Context #1 / 688
2M tokens

Analysis Summary

xAI's Grok 4.20 Multi-Agent Beta is a vision-capable model with a 2M context window and support for tool use. Its supplied measurements show excellent overall reasoning, particularly strong agentic performance, and strong instruction following and tool-use results. Coding capability is also strong, though it is not at the very highest coding tier represented in the broader landscape. This is a broadly capable model for complex, tool-connected tasks.

For an agency, it is a strong fit for multi-step research, long-document analysis, and agent workflows that combine visual input with external tools. The context capacity is useful for large knowledge sets, while its measured capabilities support more demanding work than routine content generation alone. Pricing is high, so using it for every low-risk request would be inefficient. Reserve it for tasks where its reasoning and agentic strengths can reduce manual effort or improve important deliverables.

Assessed September 23, 2026

Editorial notes

Grok 4.20 Multi-Agent Beta combines excellent reasoning and agentic performance with a 2M context, vision, and strong instruction following; its premium pricing favors high-value work.

The DFO score ranks models on the Artificial Analysis Intelligence Index and is refreshed as new models launch. The write-up and verdict are our own. Feedback?

DFO Verdict

Grok 4.20 Multi-Agent Beta combines excellent reasoning and agentic performance with a 2M context, vision, and strong instruction following; its premium pricing favors high-value work.

#40 of 688 overall Down 30 this week

Benchmarks

Intelligence Index 32.8 #40 of 431 · Artificial Analysis
Coding Index 28.5 #107 of 198 · Artificial Analysis
Agentic Index 46.4 #14 of 181 · Artificial Analysis
Creative Writing Not rated Arena rating
Instruction Following Not rated Arena rating
GPQA Diamond 88.5%
HLE 30%
SciCode 44.7%
TerminalBench Hard 40.9%
τ²-Bench 96.5%
IFBench 82.9%
LCR 59%

Artificial Analysis data refreshed Sep 30, 2026. Arena ratings from the Sep 25, 2026 text leaderboard, with style control. Bars: magenta = reasoning, ink = coding and agents, cyan = instructions and long context.

Artificial Analysis no longer lists this model, so its indexes are converted from their older scale to stay comparable.

How xAI: Grok 4.20 Multi-Agent Beta compares

XAI: Grok 4.20 Multi-Agent Beta ranks #40 of 431 AI models we track for overall intelligence, #107 of 198 for coding, #14 of 181 for agentic tasks. Its 2M-token context window is larger than 100% of the models we list. At $2.00 per million input tokens it is cheaper than 17% of comparable models.

Position in the field
DFO score: ahead of 94% of models #40
Speed: faster than 94% of models #20
Price: cheaper than 17% of models #572
Context: larger than 100% of models #1
worst in fieldmedianbest in field
Price vs frontier peers · $ per 1M tokens
xAI: Grok 4.20 Multi-Agent Beta $2.00 in $6.00 out
Anthropic: Claude Opus 5.5 $4.00 in $20.00 out
Anthropic: Claude Sonnet 5.5 $2.00 in $10.00 out
Anthropic: Claude Fable 5.1 $10.00 in $50.00 out

Dark bar = input · light bar = output, scaled to the priciest peer.

Context window vs peers · tokens

1M tokens ≈ 8 full-length novels or ~2,500 pages of business documents in a single request.

Pricing

Token Type Cost per 1M tokens Cost per 1K tokens
Input $2.00 $0.002000
Output $6.00 $0.006000

What would xAI: Grok 4.20 Multi-Agent Beta cost your business?

Pick the job that looks most like yours, then fine-tune with the sliders. Estimates update live.

A website chatbot handling around 100 customer conversations a day, a few short messages each.

3,000
One request is one message, email, draft or automation call.
1,200 tokens

$11.52/mo xAI: Grok 4.20 Multi-Agent Beta $0.0038 per request
$31.68/mo Anthropic: Claude Opus 5.5 $0.01 per request
$15.84/mo Anthropic: Claude Sonnet 5.5 · best value $0.0053 per request

Full calculator with 688 models → Price Calculator

DFO AI AUTOMATION

These numbers get smaller with the right architecture.

We route routine calls to cheap models and save xAI: Grok 4.20 Multi-Agent Beta for the hard ones. Most clients cut their estimate by 60-80%.

Talk to our team

About xAI: Grok 4.20 Multi-Agent Beta

Grok 4.20 Multi-Agent Beta is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information across complex tasks. Reasoning effort behavior:
- low / medium: 4 agents
- high / xhigh: 16 agents

Frequently asked questions about xAI: Grok 4.20 Multi-Agent Beta

How much does xAI: Grok 4.20 Multi-Agent Beta cost?

xAI: Grok 4.20 Multi-Agent Beta costs $2.00 per million input tokens and $6.00 per million output tokens.

What is the context window of xAI: Grok 4.20 Multi-Agent Beta?

xAI: Grok 4.20 Multi-Agent Beta has a context window of 2,000,000 tokens (2M).

Is xAI: Grok 4.20 Multi-Agent Beta good for coding?

On our coding benchmark index, xAI: Grok 4.20 Multi-Agent Beta ranks #107 of 198 models, placing it in the broader range of the field for code generation and debugging.

What can xAI: Grok 4.20 Multi-Agent Beta do?

xAI: Grok 4.20 Multi-Agent Beta supports image/vision input and tool use.

Who created xAI: Grok 4.20 Multi-Agent Beta?

xAI: Grok 4.20 Multi-Agent Beta is developed by xAI and was released on March 12, 2026.

© 2026 Design for Online Ltd. Registered in England and Wales No. 10328553. VAT Registered. Design for Online® and Forerunner® are registered trademarks.