LLM Perks
RU

Free Nemotron model via API

Updated:

Ultra, Super, and Nano from the Nemotron 3 and 4 line. The catalog mixes NVIDIA NIM and compatible aggregators without merging their quotas.

Read more

NVIDIA Build needs a developer account. A third-party :free route is a different key and a different cap.

30 models found

Power
№ ↕ Model ↕ Power ↕ Provider ↕ Parameters ↕ Context ↕ Limit ↕
1 NVIDIA: Nemotron 3.5 Lightning (free) ⚖️ OpenRouter ? 1.0M Free tier Open
2 NVIDIA: Nemotron 3.5 Content Safety (free) ⚖️ OpenRouter ? 128K Free tier Open
3 NVIDIA: Nemotron 3 Ultra (free) 🚀 OpenRouter 550B 1.0M Free tier Open
4 NVIDIA: Nemotron 3 Nano Omni (free) ⚖️ OpenRouter 30B 256K Free tier Open
5 NVIDIA: Nemotron 3 Super (free) 🚀 OpenRouter 120B 262K Free tier Open
6 nim/nvidia/llama-3.1-nemotron-70b-instruct 🚀 Together AI 70B 16K Fair Use (По мере нагрузки серверов) Open
7 nim/nvidia/llama-3.3-nemotron-super-49b-v1 ⚖️ Together AI 49B 16K Fair Use (По мере нагрузки серверов) Open
8 Nvidia Nemotron 3 Nano 30B A3b Bf16 ⚖️ Together AI 30B 262K Fair Use (По мере нагрузки серверов) Open
9 Nvidia Nemotron 3 Super 120B A12b Fp8 🚀 Together AI 120B 262K Fair Use (По мере нагрузки серверов) Open
10 Nvidia Nemotron 3 Super 120B A12b Bf16 🚀 Together AI 120B 262K Fair Use (По мере нагрузки серверов) Open
11 Nemotron 3 Nano Omni 30B A3b Reasoning Fp8 ⚖️ Together AI 30B 131K Fair Use (По мере нагрузки серверов) Open
12 LLAMA 3.1 Nemotron 51b Instruct ⚖️ NVIDIA Build 51B ? Free serverless inference for development; limits and model availability may change. Open
13 LLAMA 3.1 Nemotron 70b Instruct 🚀 NVIDIA Build 70B ? Free serverless inference for development; limits and model availability may change. Open
14 LLAMA 3.1 Nemotron Ultra 253b V1 🚀 NVIDIA Build 253B ? Free serverless inference for development; limits and model availability may change. Open
15 Mistral Nemotron ⚖️ NVIDIA Build ? ? Free serverless inference for development; limits and model availability may change. Open
16 Nemotron 3 Nano Omni 30b A3b Reasoning ⚖️ NVIDIA Build 30B ? Free serverless inference for development; limits and model availability may change. Open
17 Nemotron 3 Super 120b A12b 🚀 NVIDIA Build 120B ? Free serverless inference for development; limits and model availability may change. Open
18 Nemotron 3 Ultra 550b A55b 🚀 NVIDIA Build 550B 1.0M Free serverless inference for development; limits and model availability may change. Open
19 Nemotron 3.5 Lightning 30b A3b ⚖️ NVIDIA Build 30B ? Free serverless inference for development; limits and model availability may change. Open
20 Nemotron 4 340b Instruct 🚀 NVIDIA Build 340B ? Free serverless inference for development; limits and model availability may change. Open
21 Nemotron Nano 3 30b A3b ⚖️ NVIDIA Build 30B ? Free serverless inference for development; limits and model availability may change. Open
22 Nemotron 3 Nano Omni 30b A3b Reasoning:Free ⚖️ TokenRouter 30B ? Free models and promotional access may be time-limited; availability is checked through the Models API. Open
23 Mistral Nemotron ⚖️ UnoRouter ? 128K No-card free models use shared capacity and per-model limits; availability can change. Open
24 Nemotron 3 Nano Omni 30b A3b Reasoning ⚖️ UnoRouter 30B 262K No-card free models use shared capacity and per-model limits; availability can change. Open
25 Nemotron 3 Super 120b A12b 🚀 UnoRouter 120B 262K No-card free models use shared capacity and per-model limits; availability can change. Open
26 Nemotron 3 Ultra 550b A55b 🚀 UnoRouter 550B 1.0M No-card free models use shared capacity and per-model limits; availability can change. Open
27 Nemotron 3.5 Lightning ⚖️ UnoRouter ? 262K No-card free models use shared capacity and per-model limits; availability can change. Open
28 Nemotron 3.5 Lightning 30b A3b ⚖️ UnoRouter 30B 262K No-card free models use shared capacity and per-model limits; availability can change. Open
29 Nemotron Nano 12b V2 Vl ⚖️ UnoRouter 12B 131K No-card free models use shared capacity and per-model limits; availability can change. Open
30 Nemotron Nano 9b V2 ⚡ UnoRouter 9B 131K No-card free models use shared capacity and per-model limits; availability can change. Open