CATALOG

Models

Find the right model by capability, context and price. Compare options or try one in Chat.

Output / endpointExact metadata
2 of 647 modelsReference token prices are per million tokens.
VisionVideo inputToolsReasoning

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

by stepfunMay 28, 2026262.14K context$0.2/M input$1.15/M outputText, Image, Video → Text
ToolsReasoning

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token....

by stepfunJan 29, 2026262.14K context$0.1/M input$0.3/M outputText → Text