Getting started

Models & pricing

Select callable model IDs and compare current reference metadata.

On this page

This page is for developers choosing a callable model and estimating the likely cost before integration.

OpenAI-compatible endpoints accept either a complete author/model ID or a bare model name. The gateway resolves exact matches from its alias table. For example, gpt-5.4-nano resolves to openai/gpt-5.4-nano, deepseek-v4-flash to deepseek/deepseek-v4-flash, qwen3.5-flash to qwen/qwen3.5-flash-02-23, and claude-haiku-4.5 to anthropic/claude-haiku-4.5. A bare name absent from the alias table is treated as openai/<name>.

Prefer the complete ID in production and inspect the response model field to confirm the resolved ID. Variants are not bare aliases: always write the complete author/model:free or author/model:batch value.

Anthropic and Gemini native endpoints use different namespaces. Anthropic Messages uses native hyphenated names, such as the verified claude-haiku-4-5. The gateway maps a few dotted aliases (claude-opus-4.8 resolves to claude-opus-4-8), but coverage is partial: claude-haiku-4.5 returns 404. OpenAI-compatible IDs such as anthropic/claude-haiku-4.5 are never valid there. Use public GET /anthropic/v1/models for its current list. Gemini accepts Google-native names, such as the verified gemini-3.1-flash-lite; use public GET /v1beta/models for that protocol.

Base URLs and primary endpoints#

PurposeMethod and endpoint
OpenAI-compatible model list (public)GET https://api.clssai.com/v1/models
Chat CompletionsPOST https://api.clssai.com/v1/chat/completions
Responses API (verified; data-only SSE, dispatch on JSON type)POST https://api.clssai.com/v1/responses
Legacy Completions (text models; wrong-modality models are rejected)POST https://api.clssai.com/v1/completions
EmbeddingsPOST https://api.clssai.com/v1/embeddings
RerankingPOST https://api.clssai.com/v1/rerank
Image generationPOST https://api.clssai.com/v1/images
Image edits (multipart translated to JSON; no mask; 8 MB total; always b64_json)POST https://api.clssai.com/v1/images/edits
Speech synthesisPOST https://api.clssai.com/v1/audio/speech
TranscriptionPOST https://api.clssai.com/v1/audio/transcriptions
Video pass-through (lightly tested)POST https://api.clssai.com/v1/videos
Key balanceGET https://api.clssai.com/v1/balance
Anthropic MessagesPOST https://api.clssai.com/v1/messages
Anthropic token count (not billed)POST https://api.clssai.com/v1/messages/count_tokens
Anthropic model list (public)GET https://api.clssai.com/anthropic/v1/models
Gemini generationPOST https://api.clssai.com/v1beta/models/{model}:generateContent
Gemini streamingPOST https://api.clssai.com/v1beta/models/{model}:streamGenerateContent
Gemini model list (public)GET https://api.clssai.com/v1beta/models
Request logsGET https://api.clssai.com/key/<key>/logs.json

OpenAI-compatible calls use the base URL https://api.clssai.com/v1 and accept Bearer, x-api-key, or a ?key= query parameter. Prefer a header: a key in the URL leaks through referrers, proxy logs, and browser history. The Anthropic SDK base URL is https://api.clssai.com without /v1; calls accept x-api-key or Bearer and require anthropic-version: 2023-06-01. Gemini uses https://api.clssai.com/v1beta with x-goog-api-key or ?key= and does not accept Bearer authentication.

List callable model IDs#

GET /v1/models is public, supports ETag and If-None-Match, and returns only id, object, created, and owned_by for each record.

BashSyntax highlighted
curl https://api.clssai.com/v1/models \
  -H "Content-Type: application/json" \
  -H 'If-None-Match: "your-last-etag"'
JSONSyntax highlighted
{
  "object": "list",
  "data": [
    {"id": "author/model", "object": "model", "created": 1767225600, "owned_by": "author"}
  ]
}

The API directory is the authority for callable IDs. Reference context and token prices below are server-rendered from the same catalog snapshot used by the full model browser. Non-token-priced models show an em dash instead of a token price.

Models by modality#

Reference prices; you are billed by upstream-reported usage.cost — see Usage & billing.

Chat

Model IDContextInputOutput
inclusionai/ling-3.0-flash-vl:free262.14K$0/M$0/M
nex-agi/nex-n2.5-mini:free262.14K$0/M$0/M
nex-agi/nex-n2.5-pro:free262.14K$0/M$0/M
inclusionai/ling-3.0-flash-sante:free262.14K$0/M$0/M
inclusionai/ling-3.0-flash-fin:free262.14K$0/M$0/M
dots-studio/dots-3-note-preview:free512K$0/M$0/M
liquid/lfm-2.5-2.6b:free65.54K$0/M$0/M
nvidia/nemotron-3.5-lightning:free1M$0/M$0/M
thinkingmachines/inkling-small:free1.05M$0/M$0/M
poolside/laguna-s-2.1:free262.14K$0/M$0/M
thinkingmachines/inkling:free1.05M$0/M$0/M
poolside/laguna-xs-2.1:free262.14K$0/M$0/M
cohere/north-mini-code:free256K$0/M$0/M
z-ai/glm-5.2:free32.77K$0/M$0/M
nvidia/nemotron-3.5-content-safety:free128K$0/M$0/M
nvidia/nemotron-3-ultra-550b-a55b:free1M$0/M$0/M
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free256K$0/M$0/M
google/gemma-4-26b-a4b-it:free262.14K$0/M$0/M
google/gemma-4-31b-it:free262.14K$0/M$0/M
nvidia/nemotron-3-super-120b-a12b:free262.14K$0/M$0/M

Vision

Model IDContextInputOutput
inclusionai/ling-3.0-flash-vl:free262.14K$0/M$0/M
nex-agi/nex-n2.5-mini:free262.14K$0/M$0/M
nex-agi/nex-n2.5-pro:free262.14K$0/M$0/M
dots-studio/dots-3-note-preview:free512K$0/M$0/M
thinkingmachines/inkling-small:free1.05M$0/M$0/M
thinkingmachines/inkling:free1.05M$0/M$0/M
nvidia/nemotron-3.5-content-safety:free128K$0/M$0/M
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free256K$0/M$0/M

Image generation

Model IDContextInputOutput
black-forest-labs/flux-video-editN/A
openai/gpt-image-2.5-sunburst400K
openai/gpt-image-2.5-flare400K
microsoft/mai-image-2.64.1K
microsoft/mai-image-2.6-flash4.1K
meta/muse-image65.54K
recraft/recraft-v4-styles-pro65.54K
recraft/recraft-v4-styles-vector65.54K

Embeddings

Model IDContextInputOutput
liquid/lfm-2.5-embedding-350m:free512$0/M$0/M
nvidia/nemotron-3-embed-1b:free32.77K$0/M$0/M
nvidia/llama-nemotron-embed-vl-1b-v2:free131.07K$0/M$0/M
perplexity/pplx-embed-v1-0.6b32K$0.004/M$0/M
thenlper/gte-base512$0.005/M$0/M
intfloat/e5-base-v2512$0.005/M$0/M
sentence-transformers/paraphrase-minilm-l6-v2512$0.005/M$0/M
sentence-transformers/all-minilm-l12-v2512$0.005/M$0/M

TTS

Model IDContextInputOutput
deepgram/flux-tts:freeN/A
fish-audio/s1N/A
fish-audio/s2-proN/A
fish-audio/s2.1-pro-free:freeN/A
fish-audio/s2.1-proN/A
microsoft/mai-voice-2-flashN/A
qwen/qwen-audio-3.0-tts-flashN/A
qwen/qwen-audio-3.0-tts-plusN/A

STT

Model IDContextInputOutput
thinkingmachines/inkling-small:free1.05M$0/M$0/M
thinkingmachines/inkling:free1.05M$0/M$0/M
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free256K$0/M$0/M
google/gemini-2.5-flash-lite:batch1.05M$0.05/M$0.2/M
meta/muse-spark-1.3-contributor1.05M$0.1/M$0.2/M
meta/muse-spark-1.2-contributor1.05M$0.1/M$0.2/M
mistralai/voxtral-small-24b-250732.77K$0.1/M$0.3/M
google/gemini-2.5-flash-lite1.05M$0.1/M$0.4/M

Rerank

Model IDContextInputOutput
qwen/qwen3-reranker-8b40.96K$0/M$0/M
voyageai/rerank-2.5-lite32K$0/M$0/M
voyageai/rerank-2.532K$0/M$0/M
nvidia/llama-nemotron-rerank-vl-1b-v2:free10.24K$0/M$0/M
cohere/rerank-4-pro32.77K$0/M$0/M
cohere/rerank-4-fast32.77K$0/M$0/M
cohere/rerank-v3.54.1K$0/M$0/M

Video

Model IDContextInputOutput
black-forest-labs/flux-video-editN/A
minimax/hailuo-3-maxN/A
alibaba/wan-3.0-primeN/A
alibaba/wan-3.0N/A
heygen/avatar-ivN/A
black-forest-labs/flux-video-upscaleN/A
bytedance/seedance-2.0-miniN/A
bytedance/seedance-2.5N/A
Browse all 598 models →

Common mistakes#

  • Include /v1 in the OpenAI SDK base URL.
  • Do not infer availability from a model detail URL; use GET /v1/models.
  • Prefer a complete author/model ID in production, use the response model field to audit bare-name resolution, and keep :free or :batch in the complete ID.
  • Do not send OpenAI-compatible IDs to native Anthropic or Gemini endpoints; use each protocol's public model list.
  • Do not treat a displayed reference price as the final charge; see Usage & billing.
  • Do not call inference endpoints from browser JavaScript.
  • Use the Anthropic SDK base URL without /v1.
  • Use cf-ray when tracing inference requests.
Need a hand?

Find answers to common questions or diagnose a failed request.

Frequently asked questions →Troubleshoot errors →