Generates descriptive text, alt-text, or summaries from standalone images using VLMs or LLM+Vision systems.
Produces prose describing what an image contains. The dominant production use is accessibility — alt text at scale — with search indexing and content moderation close behind.
Large image libraries needing accessible descriptions or searchable text, where writing them by hand is not viable.
Where the image content is consequential and specific. A caption is a summary and will omit whatever the model judged unimportant, which may be the thing that mattered.
Captions entering a public product need bias and safety filtering, and where a caption becomes a record it needs the same provenance marking as any other generated content.
Fluent vagueness. The caption reads well, describes the scene broadly, and misses the specific detail that made the image worth captioning.
A retailer generating alt text for several hundred thousand product images so the catalogue is navigable by screen reader.
Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.