CLSSAI 独立产品正在逐项复刻与接入真实能力。查看进度

CATALOG

Models

按能力、模态、上下文、价格与数据策略寻找模型。默认展示官方标价,实际路由费用会在调用前透明显示。

19 of 456 modelsOutput and specialty tabs may overlap when a model supports multiple paths.
FFish Audio: S1fish-audio/s1OpenRouter preview · CLSSAI route pending

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...

by fish-audioJul 29, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
FFish Audio: S2 Profish-audio/s2-proOpenRouter preview · CLSSAI route pending

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

by fish-audioJul 29, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
FFish Audio: S2.1 Pro Free (free)fish-audio/s2.1-pro-free:freeOpenRouter preview · CLSSAI route pending
Free

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability...

by fish-audioJul 29, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
FFish Audio: S2.1 Profish-audio/s2.1-proOpenRouter preview · CLSSAI route pending

S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and...

by fish-audioJul 29, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
MMicrosoft: MAI-Voice-2-Flashmicrosoft/mai-voice-2-flashOpenRouter preview · CLSSAI route pending

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages...

by microsoftJul 23, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
QQwen: Qwen-Audio-3.0-TTS Flashqwen/qwen-audio-3.0-tts-flashOpenRouter preview · CLSSAI route pending

Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

by qwenJul 23, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
QQwen: Qwen-Audio-3.0-TTS Plusqwen/qwen-audio-3.0-tts-plusOpenRouter preview · CLSSAI route pending

Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

by qwenJul 23, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
DDeepgram: Aura-2deepgram/aura-2OpenRouter preview · CLSSAI route pending

Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multiple languages.

by deepgramJul 16, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
MMiniMax: Speech 2.8 HDminimax/speech-2.8-hdOpenRouter preview · CLSSAI route pending

MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

by minimaxJul 16, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
MMiniMax: Speech 2.8 Turbominimax/speech-2.8-turboOpenRouter preview · CLSSAI route pending

MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

by minimaxJul 16, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
MMicrosoft: MAI-Voice-2microsoft/mai-voice-2OpenRouter preview · CLSSAI route pending

MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales,...

by microsoftJun 2, 2026N/A contextSpecialized media pricing · open model detailsTextSpeech
XxAI: Grok Voice TTS 1.0x-ai/grok-voice-tts-1.0OpenRouter preview · CLSSAI route pending

Grok Voice TTS 1.0 is a text-to-speech model from xAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara,...

by x-aiMay 15, 202615K contextSpecialized media pricing · open model detailsTextSpeech

Gemini 3.1 Flash TTS Preview is a text-to-speech model from Google, and a substantial generational step up from Gemini 2.5 Flash TTS. It takes text input and produces audio output...

by googleApr 24, 202632.77K contextSpecialized media pricing · open model detailsTextSpeech
ZZyphra: Zonos v0.1 Transformerzyphra/zonos-v0.1-transformerOpenRouter preview · CLSSAI route pending

Zonos v0.1 Transformer is a text-to-speech model from Zyphra built on a pure transformer architecture. It offers the same American and British English voice coverage as the Hybrid variant, and...

by zyphraApr 23, 20264.1K contextSpecialized media pricing · open model detailsTextSpeech
ZZyphra: Zonos v0.1 Hybridzyphra/zonos-v0.1-hybridOpenRouter preview · CLSSAI route pending

Zonos v0.1 Hybrid is a text-to-speech model from Zyphra built on a hybrid architecture. It produces English speech output with coverage across American and British accents in male and female...

by zyphraApr 23, 20264.1K contextSpecialized media pricing · open model detailsTextSpeech
CCanopy Labs: Orpheus 3Bcanopylabs/orpheus-3b-0.1-ftOpenRouter preview · CLSSAI route pending

Orpheus 3B is an English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and expressive delivery. It offers 7 preset voices and is suited for narration, voice assistants, and...

by canopylabsApr 23, 20264.1K contextSpecialized media pricing · open model detailsTextSpeech
SSesame: CSM 1Bsesame/csm-1bOpenRouter preview · CLSSAI route pending

CSM 1B is a conversational speech model from Sesame. It accepts text input and produces English speech output, with voice options spanning conversational and read-speech styles. At 1B parameters, it...

by sesameApr 23, 20264.1K contextSpecialized media pricing · open model detailsTextSpeech
Hhexgrad: Kokoro 82Mhexgrad/kokoro-82mOpenRouter preview · CLSSAI route pending

Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese)...

by hexgradApr 23, 20264.1K contextSpecialized media pricing · open model detailsTextSpeech
MMistral: Voxtral Mini TTSmistralai/voxtral-mini-tts-2603OpenRouter preview · CLSSAI route pending

Voxtral Mini TTS is Mistral's text-to-speech model featuring zero-shot voice cloning and multilingual support. It converts text input into natural-sounding audio output.

by mistralaiApr 19, 20264.1K contextSpecialized media pricing · open model detailsTextSpeech