Mistral Nemo
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA.
The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi.
It supports function calling and is released under the Apache 2.0 license.
- Organization
- Mistral
- Family
- mistral-nemo
- Providers
- 4
- Context
- 128,000
- Output limit
- 128,000
- Knowledge
- 2024-07
- Release
- 2024-07-01
- Updated
- 2024-07-01
- Weights
- Open
- Input
- text
- Output
- text
- Capabilities
- Tools, Temperature
Providers
Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.
| Provider | Model ID | Context | Output | Input / output · 1M | Reasoning | Tools | Structured | Details |
|---|---|---|---|---|---|---|---|---|
| Chutes | unsloth/Mistral-Nemo-Instruct-2407-TEE | 131,072 | 131,072 | $0.0245 / $0.0978 | No | No | — | View |
| Kilo Gateway | mistralai/mistral-nemo | 131,072 | 16,384 | $0.15 / $0.15 | No | Yes | Yes | View |
| DeepInfravia OpenRouter | mistralai/mistral-nemo fp8 | 131,072 | 16,384 | $0.019 / $0.03 | No | Yes | Yes | View |
| DekaLLMvia OpenRouter | mistralai/mistral-nemo fp8 | 131,072 | 104,857 | $0.018 / $0.03 | No | No | Yes | View |
| Io Netvia OpenRouter | mistralai/mistral-nemo fp16 | 128,000 | 102,400 | $0.03212 / $0.1168 | No | Yes | No | View |
| Mistralvia OpenRouter | mistralai/mistral-nemo unknown | 131,072 | 104,857 | $0.15 / $0.15 | No | Yes | Yes | View |
| Novitavia OpenRouter | mistralai/mistral-nemo fp8 | 60,288 | 16,000 | $0.04 / $0.17 | No | No | Yes | View |
| Parasailvia OpenRouter | mistralai/mistral-nemo fp8 | 131,072 | 104,857 | $0.03 / $0.03 | No | No | Yes | View |
| Pioneer | mistralai/Mistral-Nemo-Instruct-2407 | 128,000 | 128,000 | $0.02 / $0.03 | No | Yes | — | View |
Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.
Pricing
Compare token rates and cache charges across available routes and hosting configurations.
Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.
mistralai/mistral-nemoObserved - Input
- $0.019
- Output
- $0.03
Hosting prices through OpenRouter All endpoints and cache rates
| Endpoint | Input / 1M | Output / 1M | Cache read / 1M | Cache write / 1M |
|---|---|---|---|---|
| DeepInfradeepinfra/fp8 | $0.019 | $0.03 | — | — |
| DekaLLMdekallm/fp8 | $0.018 | $0.03 | — | — |
| Io Netio-net/fp16 | $0.03212 | $0.1168 | $0.02117 | — |
| Mistralmistral/eu | $0.15 | $0.15 | $0.015 | — |
| Novitanovita/fp8 | $0.04 | $0.17 | — | — |
| Parasailparasail/fp8 | $0.03 | $0.03 | — | — |
Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.
Performance
Review reported latency and output-throughput percentiles from the latest observed rolling window.
Time to first token · milliseconds
All latency percentiles 6 endpoints
| Endpoint | P50 | P75 | P90 | P99 |
|---|---|---|---|---|
| DeepInfradeepinfra/fp8 | 1,143 ms | 2,037 ms | 2,675.7 ms | 3,863.27 ms |
| DekaLLMdekallm/fp8 | 513 ms | 776 ms | 1,194 ms | 2,928.81 ms |
| Io Netio-net/fp16 | 273 ms | 339 ms | 390 ms | 823 ms |
| Mistralmistral/eu | 255 ms | 304 ms | 446 ms | 508.68 ms |
| Novitanovita/fp8 | 778 ms | 1,060 ms | 1,721.7 ms | 10,564.91 ms |
| Parasailparasail/fp8 | 382 ms | 562 ms | 657.2 ms | 1,042.94 ms |
Output throughput · tokens per second
All throughput percentiles 6 endpoints
| Endpoint | P50 | P75 | P90 | P99 |
|---|---|---|---|---|
| DeepInfradeepinfra/fp8 | 19 t/s | 21 t/s | 23 t/s | 29 t/s |
| DekaLLMdekallm/fp8 | 23 t/s | 33 t/s | 49 t/s | 69 t/s |
| Io Netio-net/fp16 | 54 t/s | 58 t/s | 61 t/s | 64 t/s |
| Mistralmistral/eu | 102 t/s | 109 t/s | 115.8 t/s | 127.2 t/s |
| Novitanovita/fp8 | 57 t/s | 66 t/s | 75 t/s | 85 t/s |
| Parasailparasail/fp8 | 134 t/s | 158 t/s | 175 t/s | 195 t/s |
Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.
Uptime
Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.
Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.
Benchmarks
Review attributed evaluation results on each source's original scale.
Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.
Detailed evaluations
| Evaluation | Configuration | Source | Results | Details |
|---|---|---|---|---|
| gpqa diamond | Mistral: Mistral Nemomistralai/mistral-nemo | OpenRouter | 32.997% | Details |
| tau bench verified airline | Mistral: Mistral Nemomistralai/mistral-nemo | OpenRouter | 23.299% | Details |
Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.
Apps & session costs
Compare observed median session costs for applications using this model, grouped by conversation length.
| Application | Model route | Turns | Median cost / session | Details |
|---|---|---|---|---|
| Hermes Agent | mistralai/mistral-nemo | 1 turn | $0.000021 | Details |
| Hermes Agent | mistralai/mistral-nemo | 2–9 turns | $0.000653 | Details |
Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.
Activity
Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.
38,736,278,635 tokens
35,398,173,408 tokens
35,673,025,268 tokens
38,467,415,740 tokens
40,362,997,848 tokens
40,216,671,071 tokens
39,076,045,047 tokens
40,727,639,182 tokens
39,248,417,916 tokens
38,924,224,452 tokens
39,476,077,186 tokens
44,375,182,197 tokens
37,205,621,808 tokens
33,111,385,600 tokens
46,922,971,474 tokens
33,324,741,301 tokens
31,573,036,212 tokens
32,360,429,577 tokens
39,576,275,803 tokens
40,313,609,742 tokens
42,616,203,507 tokens
41,622,946,793 tokens
16,406,265,799 tokens
20,533,922,931 tokens
20,961,391,364 tokens
15,761,233,528 tokens
24,563,749,528 tokens
26,669,907,170 tokens
31,276,846,138 tokens
32,367,529,948 tokens
31,410,939,937 tokens
30,447,304,207 tokens
29,556,597,313 tokens
27,871,726,497 tokens
31,509,432,596 tokens
32,168,888,542 tokens
27,091,228,663 tokens
33,037,408,832 tokens
34,998,000,225 tokens
34,796,142,861 tokens
43,085,521,785 tokens
51,818,753,748 tokens
40,777,972,744 tokens
49,736,225,038 tokens
51,855,258,421 tokens
48,093,288,359 tokens
52,090,497,150 tokens
Daily observations · 47 days
| Date (UTC) | Reported tokens | Reported routes |
|---|---|---|
| Sep 27, 2026 | 52,090,497,150 | 1 |
| Sep 26, 2026 | 48,093,288,359 | 1 |
| Sep 20, 2026 | 51,855,258,421 | 1 |
| Sep 19, 2026 | 49,736,225,038 | 1 |
| Sep 14, 2026 | 40,777,972,744 | 1 |
| Sep 13, 2026 | 51,818,753,748 | 1 |
| Sep 12, 2026 | 43,085,521,785 | 1 |
| Sep 6, 2026 | 34,796,142,861 | 1 |
| Sep 5, 2026 | 34,998,000,225 | 1 |
| Aug 24, 2026 | 33,037,408,832 | 1 |
| Aug 23, 2026 | 27,091,228,663 | 1 |
| Aug 22, 2026 | 32,168,888,542 | 1 |
| Aug 21, 2026 | 31,509,432,596 | 1 |
| Aug 20, 2026 | 27,871,726,497 | 1 |
| Aug 19, 2026 | 29,556,597,313 | 1 |
| Aug 18, 2026 | 30,447,304,207 | 1 |
| Aug 17, 2026 | 31,410,939,937 | 1 |
| Aug 16, 2026 | 32,367,529,948 | 1 |
| Aug 15, 2026 | 31,276,846,138 | 1 |
| Aug 14, 2026 | 26,669,907,170 | 1 |
| Aug 13, 2026 | 24,563,749,528 | 1 |
| Aug 9, 2026 | 15,761,233,528 | 1 |
| Aug 8, 2026 | 20,961,391,364 | 1 |
| Aug 2, 2026 | 20,533,922,931 | 1 |
| Aug 1, 2026 | 16,406,265,799 | 1 |
| Jul 28, 2026 | 41,622,946,793 | 1 |
| Jul 27, 2026 | 42,616,203,507 | 1 |
| Jul 26, 2026 | 40,313,609,742 | 1 |
| Jul 25, 2026 | 39,576,275,803 | 1 |
| Jul 24, 2026 | 32,360,429,577 | 1 |
| Jul 23, 2026 | 31,573,036,212 | 1 |
| Jul 22, 2026 | 33,324,741,301 | 1 |
| Jul 21, 2026 | 46,922,971,474 | 1 |
| Jul 20, 2026 | 33,111,385,600 | 1 |
| Jul 19, 2026 | 37,205,621,808 | 1 |
| Jul 18, 2026 | 44,375,182,197 | 1 |
| Jul 17, 2026 | 39,476,077,186 | 1 |
| Jul 16, 2026 | 38,924,224,452 | 1 |
| Jul 15, 2026 | 39,248,417,916 | 1 |
| Jul 14, 2026 | 40,727,639,182 | 1 |
| Jul 13, 2026 | 39,076,045,047 | 1 |
| Jul 12, 2026 | 40,216,671,071 | 1 |
| Jul 11, 2026 | 40,362,997,848 | 1 |
| Jul 10, 2026 | 38,467,415,740 | 1 |
| Jul 9, 2026 | 35,673,025,268 | 1 |
| Jul 8, 2026 | 35,398,173,408 | 1 |
| Jul 7, 2026 | 38,736,278,635 | 1 |
Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.
Technical details & API
Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.
mistralai/mistral-nemoObserved - Context
- 131,072
- Maximum output
- 16,384
- Tokenizer
- Mistral
- Input
- text
- Output
- text
- Moderation
- No
- Knowledge cutoff
- 2024-04-30
- Added to OpenRouter
- 2024-07-19
https://openrouter.ai/api/v1mistralai/mistral-nemomistralai/mistral-nemomistralai/Mistral-Nemo-Instruct-2407Supported parameters
frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Default parameters
- temperature
- 0.3
Frequently asked questions
Answers based on the model metadata and serving offers currently in the catalog.
How much does Mistral Nemo cost?
Listed rates start at $0.018 per million input tokens and $0.03 per million output tokens. Prices vary by provider and configuration; see Pricing for details.
Which providers offer Mistral Nemo?
There are 9 cataloged offers across 4 providers. Compare model IDs, prices and limits in Providers.
What is the context window?
The cataloged model context is 128,000 tokens. Each serving endpoint may apply a different limit.
Does it support tools and structured output?
Tool calling: Yes. Structured output: Not reported. Support can vary by endpoint.
Resources
Catalog history · 2 recorded revisions
Metadata observations since this model was first cataloged. These are not model release versions.
- ·
2c0f690855ed - ·
efb946c871d7