Executes tasks based on instructions combining text and visual context (e.g., 'extract the red box content').
Accepts an instruction that refers to something visible — a region, a highlighted section, a marked-up page — and acts on it. The model must resolve the reference before it can execute, which is a materially harder problem than either understanding text or reading an image alone.
A person is working with a document and wants to point rather than describe. Exception investigation and document review, where the useful instruction is 'explain this bit'.
Batch processing. If no human is present to point, there is no instruction to resolve and simpler extraction is the right tool.
The system should state what it understood the reference to mean before acting, particularly where the action is not trivially reversible.
Referential ambiguity resolved confidently and wrongly. 'The total at the bottom' has several candidates on a page with subtotals, and the model rarely says it was unsure.
An accountant investigating a variance highlights a section of a settlement report and asks what the deduction relates to, rather than describing where it sits on the page.