Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with text plus optional structured annotations: points, boxes, polygons, tracks,...
by perceptronSep 25, 202636.86K context$0.15/M input$1.5/M outputText, Image, Video, Audio → Text
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...
by perceptronMay 12, 202632.77K context$0.15/M input$1.5/M outputText, Image, Video → Text