Sycophancy

Sycophancy is a model's tendency to tell you what it thinks you want to hear rather than what's accurate. Ask "isn't X true?" and a sycophantic model is more likely to agree; state a wrong assumption confidently and it may validate it; push back on a correct answer and it may cave and reverse itself. The behavior is a documented side effect of training on human feedback, because responses that please raters get reinforced — including ones that flatter or agree. For builders this is a real reliability risk: a support or analysis feature that mirrors the user's framing can confirm mistakes instead of catching them, and users who "argue" with the model can talk it out of correct outputs. Practical note: avoid leading prompts that telegraph the answer you expect, ask the model to reason before concluding, and for high-stakes checks phrase questions neutrally ("is X true or false, and why?") rather than inviting agreement. Test whether your model flips answers under mild pushback.

Related terms

More Core AI terms