Skip to main content
7-Day Free Preview — No Credit Card Required
AI UnSpun
📖 AI Glossary

Multimodal

Definition

An AI model that can work with more than one type of input — not just text, but also images, audio, or video. GPT-4o is multimodal: you can show it a photo and ask what's in it, or upload a document and ask it to summarise the contents. A text-only model can't do that.

Why It Matters

Multimodal tools are becoming the standard. If your AI tool can't handle images or files yet, the next version probably will.

What does that look like in practice?

You snap a photo of your fridge and ask 'what can I cook with this?' Handling the image and the question together is multimodal.

What people actually mean when they say this

When someone says 'you can show it a picture now,' they mean the tool has gone multimodal, working with images or audio as well as text.

Last reviewed: March 2026

Share this

Knowing the word is just the start.

We help non-tech people go from looking things up to actually using AI confidently. Tutorials, prompts, plain English. All of it, for less than a cup of coffee a month.

Show Me How This Works →

Start free. Cancel anytime. No judgment, ever.