CATALOG

Models

Find the right model by capability, context and price. Compare options or try one in Chat.

Output / endpointExact metadata
12 of 647 modelsReference token prices are per million tokens.

MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is...

by microsoftOct 1, 2026N/A contextSpecialized media pricing · open model detailsText → Speech

MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is...

by microsoftOct 1, 2026N/A contextSpecialized media pricing · open model detailsText → Speech
Image outputVision

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster [MAI-Image-2.6 Flash](/microsoft/mai-image-2.6-flash). It is suited for design-ready visuals and...

by microsoftSep 4, 20264.1K contextSpecialized media pricing · open model detailsText, Image → Image
Image outputVision

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the [MAI-Image-2.6](/microsoft/mai-image-2.6) family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

by microsoftSep 4, 20264.1K contextSpecialized media pricing · open model detailsText, Image → Image
Audio input

MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech,...

by microsoftSep 3, 2026N/A contextSpecialized media pricing · open model detailsAudio → Transcription
Image outputVision

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

by microsoftJul 23, 20264.1K contextSpecialized media pricing · open model detailsText, Image → Image

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft AI for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15...

by microsoftJul 23, 2026N/A contextSpecialized media pricing · open model detailsText → Speech

MAI-Voice-2 is an expressive text-to-speech model from Microsoft AI. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18...

by microsoftJun 2, 2026N/A contextSpecialized media pricing · open model detailsText → Speech
Audio input

MAI-Transcribe 1.5 is a multilingual speech-to-text model from Microsoft AI. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, with reliable transcription across 43 languages, diverse...

by microsoftJun 2, 2026N/A contextSpecialized media pricing · open model detailsAudio → Transcription
Image outputVision

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

by microsoftJun 2, 20264.1K contextSpecialized media pricing · open model detailsText, Image → Image
Text

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...

by microsoftJan 10, 202516.38K context$0.07/M input$0.14/M outputText → Text
Text

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is...

by microsoftApr 16, 202465.54K context$0.62/M input$0.62/M outputText → Text