Multimodal
Definition
An AI model that can work with more than one type of input — not just text, but also images, audio, or video. GPT-4o is multimodal: you can show it a photo and ask what's in it, or upload a document and ask it to summarise the contents. A text-only model can't do that.
Why It Matters
Multimodal tools are becoming the standard. If your AI tool can't handle images or files yet, the next version probably will.
What does that look like in practice?
You snap a photo of your fridge and ask 'what can I cook with this?' Handling the image and the question together is multimodal.
What people actually mean when they say this
When someone says 'you can show it a picture now,' they mean the tool has gone multimodal, working with images or audio as well as text.
Last reviewed: March 2026
Knowing the word is just the start.
We help non-tech people go from looking things up to actually using AI confidently. Tutorials, prompts, plain English. All of it, for less than a cup of coffee a month.
Show Me How This Works →Start free. Cancel anytime. No judgment, ever.
