AJ Learning Hub logoAJ Learning Hub

Multimodal

A multimodal model can understand or generate more than one type of input or output, such as text plus images, audio, or documents.

Analogy

Like a person who can both read a letter and look at a photo, instead of only reading.

Why it matters

Multimodal models let automations process screenshots, receipts, or PDFs directly, expanding what you can build beyond plain text.

In practice

Sending Claude a photo of a receipt and asking it to extract the total as text.

Related terms:Foundation ModelComputer UseToken