CATALOG

Models

Find the right model by capability, context and price. Compare options or try one in Chat.

Output / endpointExact metadata
16 of 647 modelsReference token prices are per million tokens.
ToolsReasoning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

by nvidiaAug 11, 2026262.14K context$0.06/M input$0.16/M outputText → Text
ToolsReasoning

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

by nvidiaJun 4, 2026262.14K context$0.6/M input$2.4/M outputText → Text
Audio input

Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across...

by nvidiaMay 27, 2026N/A contextSpecialized media pricing · open model detailsAudio → Transcription
ToolsReasoning

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

by nvidiaMar 11, 2026262.14K context$0.08/M input$0.45/M outputText → Text
FreeToolsReasoning

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

by nvidiaMar 11, 2026262.14K context$0/M input$0/M outputText → Text
ToolsReasoning

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

by nvidiaDec 14, 2025262.14K context$0.05/M input$0.2/M outputText → Text
Tools

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...

by mistralaiJul 19, 2024131.07K context$0.019/M input$0.03/M outputText → Text