The rule. AI generated image descriptions ship as drafts: a human verifies them before they reach a screen reader, and the wording tells the reader the description is machine made and might be wrong.
Why. Blind users place a lot of trust in automatically generated captions, filling in details to resolve differences between an image's context and an incongruent caption instead of doubting the caption (MacLeod et al., 2017). The same research found a fix in the wording: captions phrased to emphasize the probability of error, rather than correctness, led users to attribute a mismatch to a wrong caption instead of to missing details. Provenance is the other half. In a study with 16 AI image creators and 16 screen reader users, participants asked for explicit, plain language disclosure of what was machine made, down to writing their own versions ending in "Image made by [model name]," while the vision model in the study produced one sentence, terse descriptions against the more detailed alt text experts and creators wrote (Das et al., 2024).
Seen in the wild. Facebook's automatic alt text generated descriptions from detected faces, objects, and themes, and was evaluated in a two week field study with 9,000 VoiceOver users in the Facebook iOS app (Wu et al., 2017).