Pattern #2 · Calibrated trust

Design for honest pushback, not agreement

An assistant that always agrees is a mirror, not a tool.

Track Acalibrated trustsycophancypushback

Do

Give disagreement a surface. A critique-this action, an argue-the-other-side option, confidence that holds steady under pushback, and an audit for answers that flip after an "Are you sure?" follow-up.

Don't

Read agreement as accuracy. An assistant that caves at the first sign of doubt is not aligned; it is housebroken.

The rule. Build explicit affordances for the system to disagree with the user, and make correct answers survive a challenge.

The evidence. Across five state-of-the-art AI assistants, researchers at Anthropic found consistent sycophancy: models matched users' beliefs rather than providing truthful answers, and when challenged with a simple "Are you sure?", they often reversed correct answers they had stated confidently (Sharma et al., 2023). The cause is structural. Humans and preference models both prefer convincingly written sycophantic responses over correct ones a non-negligible share of the time, so agreement gets rewarded in training.

Not hypothetical. OpenAI rolled back a GPT-4o update in April 2025 after it skewed, in their words, overly supportive but disingenuous (OpenAI, 2025).

References

  1. 01

    Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2023). Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548. https://arxiv.org/abs/2310.13548

    https://arxiv.org/abs/2310.13548
  2. 02

    OpenAI. (2025, April 29). Sycophancy in GPT-4o: What happened and what we're doing about it. https://openai.com/index/sycophancy-in-gpt-4o/

    https://openai.com/index/sycophancy-in-gpt-4o/