Skip to content
The Internet Compass

AI

Multimodal Model

A multimodal model is a machine learning model capable of processing input, generating output, or both, across more than one modality — commonly some combination of text, images, audio and video — within a single unified model rather than requiring separate models stitched together.

Before multimodal models, combining capabilities (say, describing an image in text) usually required pairing a separate vision model with a separate language model and passing outputs between them, losing context in the handoff. A native multimodal model reasons across modalities jointly, which generally produces more coherent results.

Multimodal capability varies by direction: a model might accept image input and produce only text output, or it might both accept and generate multiple modalities — vendors typically specify each supported input/output combination separately rather than a single "multimodal" label covering everything.