Exact capability metadata is available. CLSSAI can prepare the Text chat request without guessing the model contract.
- Evidence
- Exact ID deepseek/deepseek-v4-flash · /chat/completions
- Mapped mode
- Text chat
deepseek/deepseek-v4-flashDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Catalog preview and CLSSAI API routing are independent decisions with separate sources.
Exact capability metadata is available. CLSSAI can prepare the Text chat request without guessing the model contract.
The exact model ID is listed in the current CLSSAI public directory. This proves current route eligibility, not a completed inference for every mode.
The request contract and frontend path are ready. A real recorded output has not yet been captured for this exact model and mode.
Compare published prices and endpoint capabilities. Select a provider for context limits, caching prices and supported parameters.
| $0.0002 | $1.6 | $0.0002 | -- | -- | 100.00% | |
| $0.0002 | $1.28 | $0.0002 | -- | -- | 99.83% | |
| $0.025 | $0.051 | $0.0051 | -- | -- | 71.3% | |
| $0.028 | $0.056 | $0.0056 | -- | -- | 98.5% | |
| $0.09 | $0.18 | $0.018 | -- | -- | -- | |
| $0.09 | $0.18 | $0.018 | -- | -- | 99.81% | |
| $0.091 | $0.182 | $0.018 | -- | -- | 100.00% | |
| $0.097 | $0.193 | $0.02 | -- | -- | 99.51% | |
| $0.098 | $0.196 | $0.02 | -- | -- | 99.99% | |
| $0.13 | $0.28 | $0.028 | -- | -- | 99.89% | |
| $0.134 | $0.268 | $0.027 | -- | -- | 99.16% | |
| $0.14 | $0.28 | $0.028 | -- | -- | 99.93% | |
| $0.14 | $0.28 | $0.028 | -- | -- | 99.17% | |
| $0.14 | $0.28 | $0.07 | -- | -- | 99.86% | |
| $0.19 | $0.5 | -- | -- | -- | 98.3% | |
| $0.21 | $0.56 | $0.031 | -- | -- | 98.5% |
USD per 1M tokens. Prices and metrics describe catalog endpoints, not CLSSAI routing. “--” means no measurement was supplied. Uptime: last 30 minutes.
Published base rates across providers, in USD per 1M tokens. Actual cost depends on the provider, caching and prompt length.
Effective Pricing: request-weighted costs and historical prices are not supplied by the public source.
Latency and throughput history are not supplied by the public endpoint API. Check the catalog for its latest performance charts.
Successful requests in the last 30 minutes, as reported by the catalog. A snapshot does not show historical reliability.
No scored evaluations are included in the public model metadata. View the source for benchmarks and their methodology.
Application rankings require verified traffic data, which is not included in this source.
Historical token and request volumes are not supplied by the public API. The catalog shows activity on its own network.
Provider availability, context limits and API capabilities.
OpenInference Fp8, Relace Fp4, Baidu Fp8, StreamLake Fp8, Reka, DeepInfra Fp8, GMICloud Fp8, Venice, DigitalOcean, SiliconFlow Fp8, Alibaba Fp8, Novita Fp8, AtlasCloud Fp4, Parasail Fp8, Mancer 2 Fp8, Azure (US) are listed in the public reference data. Check CLSSAI availability above before making a request.
The largest available context window is 1,048,576 tokens.
The providers list 22 distinct supported parameters. Open an endpoint row to inspect them.
Continue comparing models or start integrating.