NVIDIA: Llama Nemotron Rerank VL 1B V2 (free)

nvidia/llama-nemotron-rerank-vl-1b-v2:free
Compare Try it here

Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG...

Modalities
Text + Image → Rerank
Input / Output Price$0 / $0From · per 1M tokens · see provider prices
Context10.2Ktokens · maximum across providers
ReleasedJun 9, 2026Public catalog metadata
Availability & data sources Listed in CLSSAI · 1 reference endpoints
CLSSAI API · Directory listedCatalog preview · Exact metadata
ACCESS

Where this model works

Catalog preview and CLSSAI API routing are independent decisions with separate sources.

Catalog previewTestable in Playground
Testable

Exact capability metadata is available. CLSSAI can prepare the Rerank request without guessing the model contract.

Evidence
Exact ID nvidia/llama-nemotron-rerank-vl-1b-v2:free · /rerank
Mapped mode
Rerank
CLSSAI APICallable through CLSSAI API
Callable

The exact model ID is listed in the current CLSSAI public directory. This proves current route eligibility, not a completed inference for every mode.

Evidence
Exact ID nvidia/llama-nemotron-rerank-vl-1b-v2:free is directory-listed
Rerank test status

The request contract, frontend path, and a real recorded output are available for this exact model and mode.

Request contract
Build verified
/rerank · Schema tests · 2026-08-05
Frontend path
Available
Rerank form · 2026-08-05
Real model output
Verified
2026-08-05 · 3 documents ranked · 91 tokens · $0

Try this model

Rerank · NVIDIA: Llama Nemotron Rerank VL 1B V2 (free) · Runs through the CLSSAI gateway at the official price; nothing is saved.

Providers

Compare published prices and endpoint capabilities. Select a provider for context limits, caching prices and supported parameters.

1 providers
$0$0------100.00%

USD per 1M tokens. Prices and metrics describe catalog endpoints, not CLSSAI routing. “--” means no measurement was supplied. Uptime: last 30 minutes.

Pricing

Published base rates across providers, in USD per 1M tokens. Actual cost depends on the provider, caching and prompt length.

Input / 1M tokens$0
Output / 1M tokens$0
Cache read / 1M tokens--

Effective Pricing: request-weighted costs and historical prices are not supplied by the public source.

Performance

Latency and throughput history are not supplied by the public endpoint API. Check the catalog for its latest performance charts.

Uptime

Successful requests in the last 30 minutes, as reported by the catalog. A snapshot does not show historical reliability.

Nvidia100.00%

Benchmarks

No scored evaluations are included in the public model metadata. View the source for benchmarks and their methodology.

Apps

Application rankings require verified traffic data, which is not included in this source.

Activity

Historical token and request volumes are not supplied by the public API. The catalog shows activity on its own network.

Frequently asked questions

Provider availability, context limits and API capabilities.

Which providers serve NVIDIA: Llama Nemotron Rerank VL 1B V2 (free)?

Nvidia are listed in the public reference data. Check CLSSAI availability above before making a request.

What context length is available for NVIDIA: Llama Nemotron Rerank VL 1B V2 (free)?

The largest available context window is 10,240 tokens.

Which API parameters does NVIDIA: Llama Nemotron Rerank VL 1B V2 (free) support?

No source-labelled supported-parameter metadata was supplied.

Explore

Continue comparing models or start integrating.