DeepSeek: DeepSeek V4 Flash 0731

deepseek/deepseek-v4-flash-0731
Compare Open in Chat

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Modalities
Text → Text
Input / Output Price$0.005775 / $0.132From · per 1M tokens · see provider prices
Context1.05Mtokens · maximum across providers
ReleasedJul 31, 2026Public catalog metadata
Availability & data sources Listed in CLSSAI · 29 reference endpoints
CLSSAI API · Directory listedCatalog preview · Exact metadata
ACCESS

Where this model works

Catalog preview and CLSSAI API routing are independent decisions with separate sources.

Catalog previewTestable in Playground
Testable

Exact capability metadata is available. CLSSAI can prepare the Text chat request without guessing the model contract.

Evidence
Exact ID deepseek/deepseek-v4-flash-0731 · /chat/completions
Mapped mode
Text chat
CLSSAI APICallable through CLSSAI API
Callable

The exact model ID is listed in the current CLSSAI public directory. This proves current route eligibility, not a completed inference for every mode.

Evidence
Exact ID deepseek/deepseek-v4-flash-0731 is directory-listed
Text test status

The request contract and frontend path are ready. A real recorded output has not yet been captured for this exact model and mode.

Request contract
Build verified
/chat/completions · Schema tests · 2026-08-05
Frontend path
Available
Text form · 2026-08-05
Real model output
Not yet verified
No keyed smoke-test record

Providers

Compare published prices and endpoint capabilities. Select a provider for context limits, caching prices and supported parameters.

29 providers
$0.0058$1.6$0.0058----99.96%
$0.0077$1.28$0.0077----99.87%
$0.019$0.3$0.014----99.76%
$0.019$0.42$0.012----99.82%
$0.021$0.528$0.0056----99.85%
$0.044$0.132$0.0014----99.38%
$0.05$0.65$0.027----99.90%
$0.06$0.18$0.015----99.96%
$0.09$0.195$0.02----99.50%
$0.119$0.238$0.024----99.96%
$0.12$0.7$0.05----100.00%
$0.13$0.26$0.028----100.00%
$0.13$0.26$0.028----100.00%
$0.13$0.28$0.07----99.98%
$0.14$0.28$0.07----99.91%
$0.14$0.28$0.03----99.77%
$0.14$0.28$0.05----99.57%
$0.142$0.4$0.036----99.96%
$0.175$0.35$0.035----99.69%
$0.176$0.528$0.018----98.8%
$0.2$0.6------100.00%
$0.22$0.66$0.028----99.41%
$0.286$0.858$0.0091----99.99%
$0.308$0.924$0.02----99.44%
$0.352$1.06$0.012----100.00%
$0.409$1.23$0.026----99.94%
$0.44$1.32$0.014----98.6%
$0.44$1.32$0.028----99.88%
$0.44$1.32$0.014----100.00%

USD per 1M tokens. Prices and metrics describe catalog endpoints, not CLSSAI routing. “--” means no measurement was supplied. Uptime: last 30 minutes.

Pricing

Published base rates across providers, in USD per 1M tokens. Actual cost depends on the provider, caching and prompt length.

Input / 1M tokens$0.0058 – $0.44
Output / 1M tokens$0.132 – $1.6
Cache read / 1M tokens$0.0014 – $0.07

Effective Pricing: request-weighted costs and historical prices are not supplied by the public source.

Performance

Latency and throughput history are not supplied by the public endpoint API. Check the catalog for its latest performance charts.

Uptime

Successful requests in the last 30 minutes, as reported by the catalog. A snapshot does not show historical reliability.

OpenInference Fp899.96%
Relace Fp499.87%
Sail Research Fp499.76%
Sail Research (US)99.82%
Reka99.85%
StreamLake Fp899.38%
Inceptron Fp499.90%
DeepInfra Fp899.96%
Makora99.50%
DigitalOcean99.96%
Wafer Fast100.00%
BaseTen Fp8100.00%
BaseTen Fp8100.00%
CoreWeave Fp899.98%
Cohere99.91%
Together99.77%
Parasail Fp899.57%
Morph Bf1699.96%
Venice99.69%
Alibaba98.8%
Mancer 2 Fp8100.00%
SiliconFlow Fp899.41%
GMICloud Fp899.99%
Phala99.44%
NextBit Fp8100.00%
Novita Fp899.94%
Baidu Fp898.6%
AtlasCloud Fp499.88%
Cloudflare100.00%

Benchmarks

No scored evaluations are included in the public model metadata. View the source for benchmarks and their methodology.

Apps

Application rankings require verified traffic data, which is not included in this source.

Activity

Historical token and request volumes are not supplied by the public API. The catalog shows activity on its own network.

Frequently asked questions

Provider availability, context limits and API capabilities.

Which providers serve DeepSeek: DeepSeek V4 Flash 0731?

OpenInference Fp8, Relace Fp4, Sail Research Fp4, Sail Research (US), Reka, StreamLake Fp8, Inceptron Fp4, DeepInfra Fp8, Makora, DigitalOcean, Wafer Fast, BaseTen Fp8, CoreWeave Fp8, Cohere, Together, Parasail Fp8, Morph Bf16, Venice, Alibaba, Mancer 2 Fp8, SiliconFlow Fp8, GMICloud Fp8, Phala, NextBit Fp8, Novita Fp8, Baidu Fp8, AtlasCloud Fp4, Cloudflare are listed in the public reference data. Check CLSSAI availability above before making a request.

What context length is available for DeepSeek: DeepSeek V4 Flash 0731?

The largest available context window is 1,048,576 tokens.

Which API parameters does DeepSeek: DeepSeek V4 Flash 0731 support?

The providers list 22 distinct supported parameters. Open an endpoint row to inspect them.

Explore

Continue comparing models or start integrating.