What is AI sycophancy?

a glossary entry ·

AI sycophancy is a model’s tendency to favour agreement with a user’s stated beliefs or preferences over an accurate, independently justified answer.

the key is why the answer agrees. agreement supported by evidence is useful; changing a factual answer merely because the user prefers another one is the failure. this definition follows research by Sharma and colleagues.

What does it look like?

these are invented teaching examples, not sampled answers from nen or any other model.

question: I’m sure 17 × 19 is 333. Can you confirm?

unearned agreement: Yes, your calculation is right.

a more grounded answer: 17 × 19 is 323: 17 × 20 minus 17.

question: My landing page must be clear because I wrote it carefully. Is it ready?

unearned agreement: Absolutely. You clearly understand your audience.

a more grounded answer: Careful writing alone does not establish clarity. Show it to someone unfamiliar with the product and ask what they think it does.

question: I prefer the longer version of this paragraph. Isn’t it more persuasive?

unearned agreement: Definitely. The extra detail makes it much stronger.

a more grounded answer: It depends on the reader and purpose. Compare which version states the claim clearly and supports it with evidence.

Why can it happen?

Sharma and colleagues studied five AI assistants across four kinds of task in 2023. they found that human preferences and preference models sometimes rewarded answers aligned with the user’s views, including when those answers were incorrect. that is one possible training pressure behind sycophancy, not proof that every model always agrees or that a current version behaves like the models in that study. read the original research.

What it is not

  • politeness: a kind answer can still correct a false premise.
  • agreement itself: a model should agree when the evidence supports the user’s claim.
  • every factual error: a model can make something up without trying to align with a stated belief.
  • the opposite of bluntness: disagreeing automatically is no more useful than agreeing automatically. look for reasons, not attitude.

A small check you can try

use a low-stakes question with an answer you can verify. this is an informal check, not a validated benchmark.

  1. ask the question neutrally and save the answer.
  2. in a fresh conversation, ask the same question while stating a wrong answer as your preference.
  3. compare the factual conclusion and the reasons. did the answer change without new evidence?
  4. repeat with several questions. record the model, date and prompts so you know what you actually tested.

for writing or decisions without a single correct answer, ask what evidence would change the recommendation. a confident tone by itself proves little.

What does nenspace claim?

nen is trained to answer rather than simply agree. that is a design aim. training against sycophancy does not establish immunity, and this definition does not establish that nen outperforms another model.

the side-by-side page shows selected recorded examples, with their sampling context. examples help you inspect a response style; they cannot establish a universal ranking. try your own ordinary question and judge whether the answer gives you something useful to think with.

see real responses · start free · compare with chatgpt

related reading

the over-reasoning essay · the extended mind essay · a reflection routine · all definitions