NVIDIA: Nemotron 3 Ultra

NVIDIA: Nemotron 3 Ultra

nvidia · Released Jun 4, 2026
Intelligence #54 / 756
67.5 our score
Speed #38 / 332
195.3 tok/s
Input Price #509 / 772
$0.625 per 1M tokens
Output Price #551 / 772
$3.13 per 1M tokens
Context #226 / 772
262,144 tokens

Analysis Summary

NVIDIA's Nemotron 3 Ultra is built for tool-connected work, with the strongest agentic profile in this batch and a coding result near the excellent range. It supports function calling and tool use, while its 262K context is suitable for substantial repositories, documents, and multi-step task state.

The model fits autonomous workflow prototypes, engineering assistants, research pipelines, and structured business automation. Instruction following is strong, and the model's agentic capability is a meaningful advantage over models with similar general intelligence. Its text-only modality means teams handling image-heavy briefs will need a separate vision model. Pricing is listed at 0.625 for input and 3.125 for output.

Use Nemotron 3 Ultra as a practical agent and coding specialist. It offers a useful balance of capability, context, and operating cost for connected workflows.

Assessed September 7, 2026

Editorial notes

Nemotron 3 Ultra pairs strong agentic performance with capable coding, a 262K context, and tool calling at moderate listed pricing, making it a practical choice for connected workflows.

Rankings consider pricing, capabilities, benchmarks, and real-world applicability and are refreshed as new models launch. Feedback?

DFO Verdict

Nemotron 3 Ultra pairs strong agentic performance with capable coding, a 262K context, and tool calling at moderate listed pricing, making it a practical choice for connected workflows.

#54 of 756 overall Up 3 this week

Benchmark scores

GPQA Diamond 86.7%
HLE 26.6%
SciCode 39.9%
TerminalBench Hard 36.4%
τ²-Bench 83.3%
IFBench 81.4%
LCR 67%

Magenta = intelligence · Ink = technical/agentic · Cyan = content & long-context · Grey = community benchmarks. Data: Artificial Analysis, Hugging Face.

37.8 Intelligence Index·49.3 Coding Index·59.8 Agentic Index

How NVIDIA: Nemotron 3 Ultra compares

NVIDIA: Nemotron 3 Ultra ranks #45 of 437 AI models we track for overall intelligence, #76 of 209 for coding, #4 of 191 for agentic tasks. Its 262K-token context window is larger than 71% of the models we list. At $0.63 per million input tokens it is cheaper than 34% of comparable models.

Position in the field
Intelligence: smarter than 93% of models #54
Speed: faster than 89% of models #38
Price: cheaper than 34% of models #509
Context: larger than 71% of models #226
worst in fieldmedianbest in field
Price vs frontier peers · $ per 1M tokens
NVIDIA: Nemotron 3 Ultra $0.63 in $3.13 out
OpenAI: GPT-6 Astra $10.00 in $50.00 out
Anthropic: Claude Fable 5.1 $10.00 in $50.00 out
Claude Opus 5 $5.00 in $25.00 out

Dark bar = input · light bar = output, scaled to the priciest peer.

Context window vs peers · tokens
NVIDIA: Nemotron 3 Ultra 262K

1M tokens ≈ 8 full-length novels or ~2,500 pages of business documents in a single request.

Intelligence5.8Technical6.3Value7Content7.8
Performance profile

Strongest on business fit. The pulled-in intelligence corner is the trade-off, and if the shape matters more than the price, this is your model.

Compare shapes side-by-side →

Pricing

Token Type Cost per 1M tokens Cost per 1K tokens
Input $0.63 $0.000625
Output $3.13 $0.003125

What would NVIDIA: Nemotron 3 Ultra cost your business?

Pick the job that looks most like yours, then fine-tune with the sliders. Estimates update live.

A website chatbot handling around 100 customer conversations a day, a few short messages each.

3,000
One request is one message, email, draft or automation call.
1,200 tokens

$0/mo NVIDIA: Nemotron 3 Ultra

Full calculator with 772 models → Price Calculator

DFO AI AUTOMATION

These numbers get smaller with the right architecture.

We route routine calls to cheap models and save NVIDIA: Nemotron 3 Ultra for the hard ones. Most clients cut their estimate by 60-80%.

Talk to our team

About NVIDIA: Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it..

Frequently asked questions about NVIDIA: Nemotron 3 Ultra

How much does NVIDIA: Nemotron 3 Ultra cost?

NVIDIA: Nemotron 3 Ultra costs $0.63 per million input tokens and $3.13 per million output tokens.

What is the context window of NVIDIA: Nemotron 3 Ultra?

NVIDIA: Nemotron 3 Ultra has a context window of 262,144 tokens (262K).

Is NVIDIA: Nemotron 3 Ultra good for coding?

On our coding benchmark index, NVIDIA: Nemotron 3 Ultra ranks #76 of 209 models, placing it in the broader range of the field for code generation and debugging.

What can NVIDIA: Nemotron 3 Ultra do?

NVIDIA: Nemotron 3 Ultra supports tool use and function calling.

Who created NVIDIA: Nemotron 3 Ultra?

NVIDIA: Nemotron 3 Ultra is developed by NVIDIA and was released on June 4, 2026.

© 2026 Design for Online Ltd. Registered in England and Wales No. 10328553. VAT Registered. Design for Online® and Forerunner® are registered trademarks.