Pattern #27 · Accessibility

Treat AI generated alt text as a draft, not a deliverable

Terse, confident, and occasionally about a different photo.

Track Aaccessibilityalt-textscreen-readersprovenancetrustimage-description

Do

Put a human pass between generation and publication, and phrase whatever ships as an estimate: error acknowledging language plus plain provenance ("Image made by [model name]").

Don't

Let "an image of a person" close the accessibility ticket. The model met the deadline; nobody checked whether it met the photo.

The rule. AI generated image descriptions ship as drafts: a human verifies them before they reach a screen reader, and the wording tells the reader the description is machine made and might be wrong.

Why. Blind users place a lot of trust in automatically generated captions, filling in details to resolve differences between an image's context and an incongruent caption instead of doubting the caption (MacLeod et al., 2017). The same research found a fix in the wording: captions phrased to emphasize the probability of error, rather than correctness, led users to attribute a mismatch to a wrong caption instead of to missing details. Provenance is the other half. In a study with 16 AI image creators and 16 screen reader users, participants asked for explicit, plain language disclosure of what was machine made, down to writing their own versions ending in "Image made by [model name]," while the vision model in the study produced one sentence, terse descriptions against the more detailed alt text experts and creators wrote (Das et al., 2024).

Seen in the wild. Facebook's automatic alt text generated descriptions from detected faces, objects, and themes, and was evaluated in a two week field study with 9,000 VoiceOver users in the Facebook iOS app (Wu et al., 2017).

References

  1. 01

    Das, M., Fiannaca, A. J., Morris, M. R., Kane, S. K., & Bennett, C. L. (2024). From provenance to aberrations: Image creator and screen reader user perspectives on alt text for AI-generated images. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3613904.3642325

    https://doi.org/10.1145/3613904.3642325
  2. 02

    MacLeod, H., Bennett, C. L., Morris, M. R., & Cutrell, E. (2017). Understanding blind people's experiences with computer-generated captions of social media images. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3025453.3025814

    https://doi.org/10.1145/3025453.3025814
  3. 03

    Wu, S., Wieland, J., Farivar, O., & Schiller, J. (2017). Automatic alt-text: Computer-generated image descriptions for blind users on a social network service. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing (pp. 1180-1192). ACM. https://doi.org/10.1145/2998181.2998364

    https://doi.org/10.1145/2998181.2998364