NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Specifications
| Provider | NVIDIA |
|---|---|
| Intelligence Index | 22.9 (Artificial Analysis) |
| Context window | 262,144 tokens (262K) |
| Input price | $0.5 / 1M tokens |
| Output price | $2.2 / 1M tokens |
| Input modalities | text |
| Output modalities | text |
| Model ID | nvidia/nemotron-3-ultra-550b-a55b |
Pricing is per 1M tokens, indicative pricing via OpenRouter (updated 2026-10-09); a provider's own list price may differ. Intelligence Index via Artificial Analysis.
What is Nemotron 3 Ultra?
Nemotron 3 Ultra is a large language model from NVIDIA Corporation, released on Hugging Face on June 4, 2026. NVIDIA's model card gives it 550B parameters in total with 55B active, and describes a hybrid Latent Mixture-of-Experts (LatentMoE) architecture that interleaves Mamba-2 and MoE layers with select Attention layers. The card says use of the model is governed by the OpenMDW License Agreement, version 1.1, and that the model is ready for commercial and non-commercial use.
Key specs from NVIDIA's model card
These figures are quoted as the model card prints them. They describe the checkpoint NVIDIA published, not any one hosted endpoint: the card lists a context length of "Up to 1M tokens", while a hosted endpoint may expose less. For the price and context window of the listing on this page, use the Specifications table above.
- Parameters: 550B total, 55B active.
- Architecture: LatentMoE, a Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP).
- Context length: Up to 1M tokens.
- Release date: June 4, 2026. Model dates: December 2025 - April 2026.
- Data freshness: the post-training data has a cutoff date of May 2026; the pre-training data has a cutoff date of September 2025.
- Languages: English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, and Chinese. The card's metadata block also tags Arabic and Hebrew.
- License: OpenMDW License Agreement, version 1.1.
- Runtime engine listed: NeMo 26.04.01. Deployment instructions are given for vLLM, SGLang and TensorRT-LLM.
Nemotron Ultra vs Nemotron 3 Ultra — which one do you mean?
Two different NVIDIA models carry "Ultra" in the name. The older one is Llama-3.1-Nemotron-Ultra-253B-v1, which NVIDIA's model card gives a release date of 2025-04-07. Nemotron 3 Ultra, the model on this page, is a separate model released on June 4, 2026 according to its own card. Searches for "Nemotron Ultra" can mean either; the table sets the two cards side by side, with each entry taken from the respective card.
| Attribute | Llama-3.1-Nemotron-Ultra-253B-v1 | Nemotron 3 Ultra |
|---|---|---|
| Release date | 2025-04-07 | June 4, 2026 |
| Parameters | 253B model parameters | 550B total, 55B active |
| Derived from | Meta Llama-3.1-405B-Instruct, which the card calls the reference model | The card describes NVIDIA's own LatentMoE model and does not name a Llama base |
| Architecture | Dense decoder-only Transformer, customized through Neural Architecture Search (NAS) | Mamba2-Transformer Hybrid Latent Mixture of Experts (LatentMoE) with Multi-Token Prediction (MTP) |
| Context length | 128K tokens (the card's Input section says up to 131,072 tokens) | Up to 1M tokens |
| Reasoning control | ON/OFF through the system prompt | On/off through the chat template (enable_thinking), plus a medium-effort mode |
| Hardware note | The card says the model fits on a single 8xH100 node for inference | Minimum GPU requirement: 8x GB200/B200/GB300/B300, 16x H100, 8x H200 |
| Model dates | Trained between November 2024 and April 2025 | December 2025 - April 2026 |
| License | NVIDIA Open Model License, with the Llama 3.1 Community License Agreement listed as additional information ("Built with Llama") | OpenMDW License Agreement, version 1.1 |
| Languages | English and coding languages; German, French, Italian, Portuguese, Hindi, Spanish and Thai are also listed as supported | English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, and Chinese |
In short: if you mean the 253B model derived from Llama-3.1-405B-Instruct, that is the older Llama-3.1-Nemotron-Ultra-253B-v1, and its model card is linked under Sources. If you mean the 550B-parameter model with 55B active parameters and the Mamba-2 + MoE hybrid design, that is Nemotron 3 Ultra, and the rest of this page is about it.
What it is built for
NVIDIA's model card calls Nemotron 3 Ultra a "frontier-scale general purpose reasoning and chat model intended to be used in English, Code, and supported multilingual contexts", optimized for "complex agentic workflows, long-context reasoning, and high-stakes analytical workloads". It names developers designing AI agent systems, chatbots, RAG systems and other AI-powered applications as the intended users, and says the model is suitable for complex instruction-following and long-context reasoning over very large documents and codebases.
The card says the model answers by first generating a reasoning trace and then a final response, and that reasoning is switched through the chat template: enable_thinking=True or False. It also documents a medium-effort mode (medium_effort alongside enable_thinking), which it says uses significantly fewer reasoning tokens than full thinking mode and is recommended as a starting point before tuning explicit token budgets, and an advanced reasoning_budget option that sets a hard token ceiling on the reasoning trace. For tool calling, the card's SGLang notes say a chat-completions request that passes tools must set enable_thinking and force_nonempty_content to true so reasoning and tool calls parse correctly.
Benchmarks NVIDIA reports
The card prints a benchmark table comparing Nemotron 3 Ultra (shown as N-3-Ultra) with six other models. These are NVIDIA-reported results, not independent tests; the card says they were collected with its Nemo Evaluator SDK. Seven rows, with the scores exactly as printed:
| Benchmark | N-3-Ultra 550B-A55B | MiniMax-2.7 230B-A10B | GLM-5.1 744B-A40B | Kimi-K2.6 1T-A32B | Qwen-3.5 397B-17B | DS-v4-Pro 1.6T-A49B | DS-v4-Flash 284B-A13B |
|---|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 56.4 | 55.5 | 59.3 | 67.2 | 49.9 | 49.2 | 54.2 |
| SWE-Bench Verified | 70.7 | 75.3 | 76.2 | 75.7 | 73.6 | 74.5 | 73.5 |
| LiveCodeBench (v6) | 89.0 | 77.2 | 85.7 | 90.2 | 79.3 | 92.5 | 90.9 |
| GPQA (no tools) | 87.0 | 86.6 | 86.1 | 91.0 | 87.1 | 87.8 | 88.5 |
| MMLU-Pro | 86.8 | 81.9 | 85.9 | 88.1 | 88.3 | 87.5 | 86.4 |
| IFBench (prompt loose) | 81.7 | 74.6 | 76.6 | 73.7 | 78.2 | 79.1 | 82.0 |
| RULER (1M) | 94.7 | -- | -- | -- | 90.1 | 94.2 | 87.7 |
On six of these seven rows at least one other model in the card's table has a higher printed score. On RULER (1M), 94.7 is the highest number printed in the row; three of the six comparison models have no result there ("--"). The full table in the card has many more rows across agentic, reasoning and knowledge, chat and instruction following, long-context and multilingual tests.
Can you self-host it?
NVIDIA's model card lists the minimum GPU requirement as "8x GB200/B200/GB300/B300, 16x H100, 8x H200". In its deployment notes it gives the minimum recommended hardware for the BF16 checkpoint as a single node of 8× B200 (≈1.5 TB aggregate HBM), or a multi-node setup of at least 8 GPUs across H100 / H200 / GB200 / GB300, orchestrated with Ray v2.
The card provides deployment instructions for three runtimes: vLLM (recommended container vllm/vllm-openai:v0.22.0), SGLang (container lmsysorg/sglang:v0.5.12.post1, tested on 8× B200) and TensorRT-LLM. For TensorRT-LLM the card says current support is limited to NVIDIA Blackwell, and that Hopper support is planned but not currently available. Its vLLM and SGLang examples default to a 256k context; the card says reaching 1M needs an extra environment variable and a larger max-model-len / context-length setting.
For running the model on a smaller footprint, the card points to a separate NVFP4 checkpoint, NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4. Its model card is linked below.
Using it through an API
The listing on this page comes from OpenRouter, where the model is published under the id nvidia/nemotron-3-ultra-550b-a55b. The Specifications table above shows the price and context window for that listing. The context window there can differ from the card's "Up to 1M tokens"; both figures come from OpenRouter data when the site is built.
NVIDIA also links a hosted chat page for the model at build.nvidia.com (see Sources). The card's own API examples use an OpenAI-compatible client pointed at a self-hosted server on port 8000 with the served model name nvidia/nemotron-3-ultra; a hosted provider will use its own model id, so copy the one from the Model ID row.
Sources
- NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 model card (Hugging Face)
- OpenMDW License Agreement, version 1.1 (the licence URL given in the card)
- Nemotron 3 Ultra on build.nvidia.com (NVIDIA)
- NVIDIA Nemotron 3 Ultra Technical Report (PDF, linked from the card)
- NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 model card (Hugging Face)
- Llama-3.1-Nemotron-Ultra-253B-v1 model card (Hugging Face), the older model
Editorial last reviewed 2026-10-09. Benchmark figures are reported by the model's developer, not independently verified by us.
Other Nemotron models
Frequently asked questions
How capable is Nemotron 3 Ultra?
Nemotron 3 Ultra scores 22.9 on the Artificial Analysis Intelligence Index, an aggregate of reasoning, knowledge, coding, and math benchmarks (higher is better). Compare it head-to-head with other models on the comparison pages below, or run it against them directly on Vincony.
How much does Nemotron 3 Ultra cost?
Indicative pricing (via OpenRouter, updated 2026-10-09) is $0.5 per 1M input tokens and $2.2 per 1M output tokens. A provider's own list price may differ. On Vincony, Nemotron 3 Ultra is available on credit-based pricing alongside 750+ other models.
What is Nemotron 3 Ultra's context window?
Nemotron 3 Ultra supports a context window of up to 262,144 tokens (262K), which is how much text it can consider at once.
Who makes Nemotron 3 Ultra?
Nemotron 3 Ultra is built by NVIDIA. You can access it — and compare it against other providers' models — through Vincony's multi-model platform.
How do I use Nemotron 3 Ultra?
You can use Nemotron 3 Ultra directly through NVIDIA, or via Vincony, which aggregates 750+ models (including Nemotron 3 Ultra) into one interface with side-by-side comparison and credit-based pricing.