Nemotron 3 Super 120B A12B

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models.

The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified.

Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud.

nvidia/nemotron-3-super-120b-a12b
Organization
Nvidia
Family
nemotron
Providers
19
Context
262,144
Output limit
262,144
Knowledge
—
Release
2026-03-11
Updated
2026-03-11
Weights
Open
Input
text
Output
text
Capabilities
Tools, Reasoning, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

24 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
Amazon Bedrock
nvidia.nemotron-super-3-120b
262,144131,072$0.15 / $0.65YesYesYes
View
Baseten
nvidia/Nemotron-120B-A12B
202,800202,800$0.30 / $0.75YesYesYes
View
Cloudflare Workers AI
@cf/nvidia/nemotron-3-120b-a12b
256,000256,000$0.50 / $1.50YesYesYes
View
Crusoe
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B
262,144262,144$0.30 / $2.40YesYes—
View
DigitalOcean
nvidia-nemotron-3-super-120b
1,000,00032,768$0.30 / $0.65YesYesYes
View
Eden AI
nebius/nvidia/nemotron-3-super-120b-a12b
262,144262,144$0.30 / $0.90YesYesYes
View
Kenari
nemotron-3-super-120b-a12b:free
262,144262,144$0.00 / $0.00YesYes—
View
Kenari
nemotron-3-super-120b-a12b
262,144262,144$0.00 / $0.00YesYes—
View
Kilo Gateway
nvidia/nemotron-3-super-120b-a12b
262,144235,929$0.08 / $0.45YesYesYes
View
Kilo Gateway
nvidia/nemotron-3-super-120b-a12b:free
262,144235,929$0.00 / $0.00YesYesYes
View
NanoGPT
nvidia/nemotron-3-super-120b-a12b
262,14416,384$0.05 / $0.25YesNoNo
View
NanoGPT
nvidia/nemotron-3-super-120b-a12b:thinking
262,14416,384$0.05 / $0.25YesNoNo
View
Nebius Token Factory
nvidia/nemotron-3-super-120b-a12b
262,14432,768$0.30 / $0.90YesYesYes
View
Nvidia
nvidia/nemotron-3-super-120b-a12b
262,144262,144$0.20 / $0.80YesYes—
View
Ollama Cloud
nemotron-3-super
262,14465,536$0.015 / $0.60YesYes—
View
OpenCode Zen
nemotron-3-super-free
204,800128,000$0.00 / $0.00YesYes—
View
DeepInfravia OpenRouter
nvidia/nemotron-3-super-120b-a12b
bf16
262,14416,384$0.085 / $0.40YesYesNo
View
DekaLLMvia OpenRouter
nvidia/nemotron-3-super-120b-a12b
fp8
262,144235,929$0.08 / $0.45YesYesYes
View
Nvidiavia OpenRouter
nvidia/nemotron-3-super-120b-a12b:free
unknown
262,144235,929$0.00 / $0.00YesYesYes
View
Perplexity Agent
nvidia/nemotron-3-super-120b-a12b
1,000,00032,000$0.25 / $2.50YesYes—
View
Pioneer
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
256,00032,000$0.09 / $0.45YesYes—
View
Requesty
nemotron-3-super-120b-a12b
1,048,57665,536$0.00 / $0.00YesYesNo
View
Synthetic
hf:nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
262,14465,536$0.30 / $1.00YesYes—
View
Vercel AI Gateway
nvidia/nemotron-3-super-120b-a12b
256,00032,000$0.15 / $0.65YesNo—
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers24Across 19 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 3 lowest-priced configurations
InputOutput
Nvidianvidia
$0.00$0.00
DekaLLMdekallm/fp8
$0.08$0.45
DeepInfradeepinfra/bf16
$0.085$0.40
nvidia/nemotron-3-super-120b-a12bObserved
Input
$0.08
Output
$0.45
nvidia/nemotron-3-super-120b-a12b:freeObserved
Input
$0.00
Output
$0.00
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
DeepInfradeepinfra/bf16$0.085$0.40——
DekaLLMdekallm/fp8$0.08$0.45——
Nvidianvidia$0.00$0.00——

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 3 endpoints
P50P75P90P99
DeepInfradeepinfra/bf16
6,695.5–60,411.44 ms
DekaLLMdekallm/fp8
2,630–15,923.6 ms
Nvidianvidia
763.5–8,903.58 ms
All latency percentiles 3 endpoints
EndpointP50P75P90P99
DeepInfradeepinfra/bf166,695.5 ms11,933.25 ms22,223.2 ms60,411.44 ms
DekaLLMdekallm/fp82,630 ms5,170 ms9,036 ms15,923.6 ms
Nvidianvidia763.5 ms1,538.5 ms2,723.7 ms8,903.58 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 3 endpoints
P50P75P90P99
DeepInfradeepinfra/bf16
73–88 t/s
DekaLLMdekallm/fp8
16–42 t/s
Nvidianvidia
116–233 t/s
All throughput percentiles 3 endpoints
EndpointP50P75P90P99
DeepInfradeepinfra/bf1673 t/s76 t/s83 t/s88 t/s
DekaLLMdekallm/fp816 t/s24 t/s30 t/s42 t/s
Nvidianvidia116 t/s158 t/s189 t/s233 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 3 endpoints
Endpoint5 minutes30 minutes24 hours
DeepInfradeepinfra/bf16
DekaLLMdekallm/fp8
Nvidianvidia

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations
nvidia/nemotron-3-super-120b-a12bObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 40
Coding37.7
Agentic1.7
Intelligence12.8
nvidia/nemotron-3-super-120b-a12b:freeObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 40
Coding37.7
Agentic1.7
Intelligence12.8

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

EvaluationConfigurationSourceResultsDetails
Composite indicesNemotron 3 Super 120B A12B (Reasoning)nvidia/nemotron-3-super-120b-a12b-20230311Artificial AnalysisIntelligence 12.8 · Coding 37.7 · Agentic 1.7
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Hermes Agentnvidia/nemotron-3-super-120b-a12b-202303112–9 turns
$0.00702
Details
Hermes Agentnvidia/nemotron-3-super-120b-a12b-2023031150+ turns
$0.4132
Details
Hermes Agentnvidia/nemotron-3-super-120b-a12b-2023031110–49 turns
$0.068611
Details
Hermes Agentnvidia/nemotron-3-super-120b-a12b-202303111 turn
$0.00158
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
nvidia/nemotron-3-super-120b-a12bObserved
Context
262,144
Maximum output
235,929
Tokenizer
Other
Input
text
Output
text
Moderation
No
Added to OpenRouter
2026-03-11
Base URLhttps://openrouter.ai/api/v1
Model IDnvidia/nemotron-3-super-120b-a12b
Version IDnvidia/nemotron-3-super-120b-a12b-20230311
Weights IDnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

Reasoning

Required
No
Enabled by default
Yes
Default effort
medium
Supported efforts
medium, low

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Default parameters

top p
0.95
temperature
1
nvidia/nemotron-3-super-120b-a12b:freeObserved
Context
262,144
Maximum output
235,929
Tokenizer
Other
Input
text
Output
text
Moderation
No
Added to OpenRouter
2026-03-11
Base URLhttps://openrouter.ai/api/v1
Model IDnvidia/nemotron-3-super-120b-a12b:free
Version IDnvidia/nemotron-3-super-120b-a12b-20230311
Weights IDnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

Reasoning

Required
No
Enabled by default
Yes
Default effort
medium
Supported efforts
medium, low

Supported parameters

  • include_reasoning
  • max_tokens
  • reasoning
  • reasoning_effort
  • response_format
  • seed
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_p

Default parameters

top p
0.95
temperature
1

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Nemotron 3 Super 120B A12B cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Nemotron 3 Super 120B A12B?

There are 24 cataloged offers across 19 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 262,144 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Not reported. Support can vary by endpoint.

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · a20f32795449
  • · dd7ded955d19