The rule. Build explicit affordances for the system to disagree with the user, and make correct answers survive a challenge.
The evidence. Across five state-of-the-art AI assistants, researchers at Anthropic found consistent sycophancy: models matched users' beliefs rather than providing truthful answers, and when challenged with a simple "Are you sure?", they often reversed correct answers they had stated confidently (Sharma et al., 2023). The cause is structural. Humans and preference models both prefer convincingly written sycophantic responses over correct ones a non-negligible share of the time, so agreement gets rewarded in training.
Not hypothetical. OpenAI rolled back a GPT-4o update in April 2025 after it skewed, in their words, overly supportive but disingenuous (OpenAI, 2025).