CATALOG

Models

Find the right model by capability, context and price. Compare options or try one in Chat.

Output / endpointExact metadata
15 of 647 modelsReference token prices are per million tokens.
VisionToolsReasoning

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

by x-aiSep 21, 2026500K context$2/M input$6/M outputText, Image, File → Text
VisionToolsReasoning

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).

by x-aiAug 12, 2026500K context$2/M input$6/M outputText, Image, File → Text
Image outputVision

Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and...

by x-aiAug 11, 202665.54K contextSpecialized media pricing · open model detailsText, Image → Image
Audio input

Grok STT is SpaceXAI's speech-to-text model, available via the REST /v1/stt endpoint. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.

by x-aiJul 23, 2026N/A contextSpecialized media pricing · open model detailsAudio → Transcription
VisionToolsReasoning

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

by x-aiMay 20, 2026256K context$1/M input$2/M outputText, Image, File → Text
Vision

Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -...

by x-aiMay 18, 2026N/A contextSpecialized media pricing · open model detailsText, Image → Video
Image outputVision

Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a...

by x-aiMay 18, 202665.54K contextSpecialized media pricing · open model detailsText, Image → Image

Grok Voice TTS 1.0 is a text-to-speech model from SpaceXAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara,...

by x-aiMay 15, 202615K contextSpecialized media pricing · open model detailsText → Speech
VisionToolsReasoning

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

by x-aiApr 30, 20261M context$1.25/M input$2.5/M outputText, Image, File → Text
VisionToolsReasoning

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

by x-aiApr 30, 20261M context$1/M input$2/M outputText, Image, File → Text
VisionToolsReasoning

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

by x-aiMar 31, 20262M context$1.25/M input$2.5/M outputText, Image, File → Text