Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

Multimodal AI: When Your Model Can Watch, Listen and Draw

The first wave of assistants read text and wrote text. Current models take images, audio and video, and increasingly produce them, which changes what can go wrong as much as what can be done.

Audio and visual media equipment — photo by D-Kuru, licensed under CC BY-SA 3.0 at via Wikimedia Commons.

The original wave of assistants accepted text and produced text. The current generation of models handles images, audio and video as inputs, and increasingly produces them as outputs, which is a larger change of kind than the marketing language suggests.

Reading a receipt, describing a warehouse camera

  • Reading a photograph of a document, a whiteboard or a receipt.
  • Transcribing and summarising a meeting, with speaker attribution.
  • Describing what a camera in a warehouse or a vehicle sees.
  • Generating images and audio for drafts and prototypes.
  • Accessibility: describing visual material for a user who cannot see it.

Small text, cluttered scenes, overlapping speech

Every modality adds a failure mode. Vision models misread small text and confuse elements in a cluttered scene. Audio models struggle with overlapping speech and heavy accents. Video adds temporal reasoning, which is the weakest area of all.

Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.

There is also a security dimension specific to multimodal systems. An image can carry instructions — text in a screenshot, for example — which the model may treat as commands rather than content. That is an injection route that does not exist in a text-only pipeline.

Document intake pays the bills

The deployments with clear value are unglamorous: document intake, quality inspection with a human in the loop, captioning, and search across mixed media. The flashy demonstrations of a model navigating an interface reliably remain, for now, demonstrations.

Invoices were designed to be read by people

The most valuable multimodal application commercially is unglamorous: extracting structured information from documents that were designed to be read by people. Invoices, forms, contracts and scanned correspondence are the connective tissue of most business processes, and they are overwhelmingly analogue.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

This work was previously done by specialised systems, by offshore processing teams, or not at all. General vision models have made it cheap enough that organisations are revisiting processes they had written off as too expensive to automate.

One error in ten thousand documents is not a rounding error

For document extraction, an accuracy rate that would be impressive in a research paper is inadequate in production. A process handling ten thousand documents a month cannot tolerate errors at even a fraction of a percent unless there is a verification step.

That is why the successful deployments pair extraction with confidence scores, flagging the uncertain cases for human review. The model handles the easy majority and a person handles the remainder, which is a less exciting product than full automation and a far more useful one.

One practical caution for anyone building on multimodal input: the security model changes. Text supplied by a user is data. Text visible inside an image the user provides is also data, but models have repeatedly treated it as instruction. Any pipeline that accepts images from untrusted sources should assume that content inside them may be acted upon.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.