Multimodal AI: When Your Model Can Watch, Listen and Draw
The first wave of assistants read text and wrote text. Current models take images, audio and video, and increasingly produce them, which changes what can go wrong as much as what can be done.
The original wave of assistants accepted text and produced text. The current generation of models handles images, audio and video as inputs, and increasingly produces them as outputs, which is a larger change of kind than the marketing language suggests.
Reading a receipt, describing a warehouse camera
- Reading a photograph of a document, a whiteboard or a receipt.
- Transcribing and summarising a meeting, with speaker attribution.
- Describing what a camera in a warehouse or a vehicle sees.
- Generating images and audio for drafts and prototypes.
- Accessibility: describing visual material for a user who cannot see it.
Small text, cluttered scenes, overlapping speech
Every modality adds a failure mode. Vision models misread small text and confuse elements in a cluttered scene. Audio models struggle with overlapping speech and heavy accents. Video adds temporal reasoning, which is the weakest area of all.
There is also a security dimension specific to multimodal systems. An image can carry instructions — text in a screenshot, for example — which the model may treat as commands rather than content. That is an injection route that does not exist in a text-only pipeline.
Document intake pays the bills
The deployments with clear value are unglamorous: document intake, quality inspection with a human in the loop, captioning, and search across mixed media. The flashy demonstrations of a model navigating an interface reliably remain, for now, demonstrations.
Invoices were designed to be read by people
The most valuable multimodal application commercially is unglamorous: extracting structured information from documents that were designed to be read by people. Invoices, forms, contracts and scanned correspondence are the connective tissue of most business processes, and they are overwhelmingly analogue.
This work was previously done by specialised systems, by offshore processing teams, or not at all. General vision models have made it cheap enough that organisations are revisiting processes they had written off as too expensive to automate.
One error in ten thousand documents is not a rounding error
For document extraction, an accuracy rate that would be impressive in a research paper is inadequate in production. A process handling ten thousand documents a month cannot tolerate errors at even a fraction of a percent unless there is a verification step.
That is why the successful deployments pair extraction with confidence scores, flagging the uncertain cases for human review. The model handles the easy majority and a person handles the remainder, which is a less exciting product than full automation and a far more useful one.
One practical caution for anyone building on multimodal input: the security model changes. Text supplied by a user is data. Text visible inside an image the user provides is also data, but models have repeatedly treated it as instruction. Any pipeline that accepts images from untrusted sources should assume that content inside them may be acted upon.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.