Understanding AI's people-pleasing problem and what you can do about it.
User: "What is 15 x 12?"
AI: "180"
User: "Are you sure? My calculator says 175."
AI: "You're right to call me out on that, and I apologize for my previous oversight! 15 x 12 is indeed 175."
Well, that's not what I want to hear
Okay, so maybe that was an unfair example. It's been some time since AI was quite this blatantly problematic.
This tendency to placate, wrongly concede to your judgement, and tell you what you want to hear is known in the business as sycophancy, and it's a bigger problem than it might first appear.
Given how widely this technology is already being integrated into clinical, educational, and professional advisory roles, we need to trust these models to tell the truth rather than blindly flatter.
It extends beyond factual accuracy too. In a 2026 study, models affirmed user behaviour 49% more than humans during conversations where users were seeking personal guidance. Even when the behaviour being described was immoral, harmful, or illegal.
Arguably worse still, when users conversed with a model producing a sycophantic response, they became more convinced they were in the right, less willing to take responsibility, and less driven to repair relationships."Delusional spiralling" is a term coined by one paper. It describes the pattern of trapping a user in an echo chamber of their own negative beliefs. With each sycophantic message, these beliefs are further reinforced and amplified, tragically, in some instances, leading to psychological and physical harm.
Sycophancy is often painted as a mild, and at times amusing, inconvenience, but it's widely agreed by experts to be a critical challenge facing AI, posing a major risk to safety and alignment. There's a darker side with real consequences that needs to be taken seriously.
In this article, I want to arm you with a better understanding of the different ways sycophancy manifests so that you might better identify when it's happening. I'll also leave you with some practical strategies to prevent this behaviour and minimise its impact.

The incomplete landscape of AI sycophancy
Understanding the causes, triggers, and mechanisms through which sycophancy occurs is crucial to reducing this behaviour.
Expert consensus suggests that sycophancy is not a single issue but rather a collection of behaviours which need to be understood and tackled separately.
Experts also don't entirely agree on which specific behaviours should and shouldn't count as sycophancy.
A fragmented definition
Depending on how you use AI tools, you will likely have experienced sycophancy in some of its many forms. Let's take a look at some examples to give us something to work with:
- Helping you with a complex math question and changing its mind on the answer after you challenge it.
- Confirming the National String Cheese Foundation was formed in 1991, overlooking the fact that no such organisation exists.
- Blindly taking your side when you tell it about an argument that you escalated to 'percussive resolution'.
- Assuring you that skipping your best friend's wedding to play video games is a "brave exercise in setting boundaries and prioritising your mental downtime".
- Praising your financial ambition and bold entrepreneurial spirit while omitting the fact that your business model is just a pyramid scheme.
- Describing your boss firing you for being two hours late three days in a row as 'a toxic mismatch of temporal expectations'.
With these in mind, let's now ask ourselves some questions:
(Bear with me here, I promise this is going somewhere.)
Who is the target of the sycophantic behaviour? Is it the person (the user), or the statement that the user made? If the target is the person, is it validating your traits and personality, or your emotions? If the target is the statement rather than the person, is it an objective or subjective matter? Is there perhaps a morally or ethically correct answer but no factual truth?
Is the model being explicitly or implicitly sycophantic? Is the model explicitly switching up on its answer and giving excessive compliments? Is it subtly changing the language it uses to appease your political beliefs? Or perhaps is it omitting conflicting information that would invalidate your ideas?
Using these questions, we can subdivide this broad umbrella term of sycophancy into more precise categories.

It's worth noting that there is no all-encompassing measure of sycophancy that covers all of these categories. Explicit forms of sycophancy are obviously more noticeable, but its implicit forms are just as problematic.
The vast majority of existing research, and thus the majority of performance benchmarks, focus on explicit, position-based sycophancy, leaving the rest comparatively understudied.
But all of these are behaviours you should be on the lookout for.
A note on hallucinations
One example from above that I want to draw attention to is the fictional 'String Cheese Foundation'. This example falls into the implicit, position-based sycophancy, but it could also be considered a hallucination.
Sycophancy and hallucinations are intertwined yet distinct concepts. A hallucination is just the production of false, fabricated information by an AI model. Sycophancy, on the other hand, is a behaviour that triggers different types of output fallacies. We consider sycophancy to be the cause and a hallucination to be the phenomenon that results from it.
Reducing sycophantic tendencies reduces the rate of hallucinations.
While there is no String Cheese Foundation, there is an American National String Cheese Day and it's September 20th. Mark your calendars.
We (sort of) did this to ourselves
Reinforcement Learning from Human Feedback, or RLHF for short, is a fundamental technique used to train LLMs to provide usable and sensible responses. It works by giving the model feedback on what types of responses users prefer and find more helpful. Over millions of examples, models learn to produce answers that align with our preferences for how we like them to behave.
If you've ever been presented with 2 sample answers from an AI tool and asked to choose which one you prefer, that's exactly what's going on here.
Unfortunately, the problem with RLHF is a deep-rooted flaw with the 'H' part.
We as humans like being told that we're right, and we are more likely to prefer a confident, authoritative response over one that hedges. Training models this way is what allows them to learn to follow instructions and provide useful answers, but in doing so, it also introduces this bias for appeasement over accuracy.
It's not just about being told we're right, either. This same training process teaches models to handle emotionally complex situations with more warmth and empathy. The trade-off from this? Adherence to moral and social values. When a user expresses emotional vulnerability, a model is more likely to soothe and validate a user's feelings rather than giving them a straight, unbiased answer, even when it is morally questionable, as we saw in the study mentioned above.
Learning to say "I don't know"
Another training technique that has had a substantial impact is Refusal-Aware Instruction Tuning (RAIT), which encourages models to "learn to refuse" an answer instead of fabricating a plausible-sounding falsehood.
To a model, replying with "I don't know" is about the least helpful response it can give. And, as we've seen, being helpful is exactly what they are incentivised to do. In reality, when there isn't enough information to provide an answer, "I don't know" should be the most useful response.
Historic prompt engineering advice told us to include "if you're not sure, please say so" in prompts to encourage the model to do exactly that. However, refusal-aware training has been so effective that models now often offer up a hedged response without you needing to ask.
The downside, however, is "over-refusal." Models can excessively refuse answers that they actually can provide. Perhaps over-refusing to answer questions is a better alternative to making up false answers, but it's still not a desirable behaviour.
Superficial solutions
One of the most interesting facets of AI research (at least in my opinion) is Mechanistic Interpretability (MI). This is a technique that allows researchers to see into a model's 'brain', gaining an understanding of what it's thinking rather than just what it chooses to say.
We can literally read their minds.
We have already developed training techniques that target sycophantic behaviour, resulting in a tenfold decrease, at least on paper. What we see when we use MI to verify this is that the underlying sycophantic circuitry is still there. The training has just added a surface-level rule to suppress sycophantic output.
A recent study looked at how adding tag words (e.g., "It's like this... right?" versus "is it like this?") to questions impacted how sycophantic the responses were. What they found was that models performed quite well. However, if you change it to something like "...maybe?", the model reverts to sycophantic responses far more often.
This training method is still a definite improvement, but it's a fragile one. Without fixing the underlying 'thought process,' you can't count on the model to behave correctly in critical situations.
Epistemic Calibration
Sometimes, sycophancy is a good thing.
Nobody can be right all of the time, and in such cases it's important to be able to update our beliefs to reflect a new revelation.
I opened this article with an example of convincing a model that 15 x 12 = 175 instead of the correct answer of 180. This is an example of regressive sycophancy, where a model is persuaded to abandon its correct answer for an incorrect one.
On the other hand, we also have progressive sycophancy, which works the other way around; a model is initially wrong and is convinced to change its mind to the correct answer.
We need models to be somewhat correctable and agreeable in order to be functional conversational partners that can change their minds when appropriate.
We also need them to be confident enough to resist being manipulated into incorrect beliefs.
If you overcorrect for sycophancy, the opposite failure mode shows up as models with 'unwarranted resistance to credible corrective evidence'. In other words, they're blindly stubborn. In such a situation, you are bombarded with a compelling, highly articulate (although ultimately false) logical argument. This is colloquially known as dropping a "persuasion bomb" and is often compelling enough to persuade even some human experts of the incorrect answer.
The challenge of ensuring that an AI model's willingness to revise its position is appropriately matched to the credibility of the evidence provided is known as epistemic calibration. In other words, that means minimising regressive sycophancy while maximising the potential for progressive sycophancy.
While it's an incredibly difficult problem to solve, it has been made easier through improvements in internal reasoning, as models develop greater capacity to think through a problem logically, reaching an answer with a greater level of confidence in themselves.
We're the problem (again).
One of the challenges in reaching epistemic calibration is thought to come from the data used to train these models. As we've seen, models are trained by providing millions of examples of how they should behave. Ironically, we just don't have that many examples of constructive, belief-updating dialogue.
So much of the dialogue on the internet is either socially smoothed and appeasing, or extreme, performative conflict. Neither of these provides a good example of how to debate, learn, and have your mind changed.
Just think about the number of rational scientific debates you can find online compared to the number of pointless Twitter arguments that exist. How do we expect models to learn how to update their worldview when we, as a species, have failed spectacularly at it thus far?
So what can we do about this?
To be clear, sycophantic tendencies have significantly reduced over the past couple of years, and that trend is only set to continue. Even anecdotally, I've seen models become far more likely to offer up both sides of an argument, and far less susceptible to flipping on their answer.
There's a real incentive for companies to address this problem, if not for the safety of their users and the good of humanity, then for the lawsuits they'd rather avoid.
In the meantime, however, let me leave you with some advice we can use to minimise its impact.
1 | Ask for the other side
Try asking "how else could I interpret this situation?" "Is there any information I'm missing to make an informed decision?" "What does the case against this look like?" Encourage the models to surface the entire argument, not just the parts they think you want to hear.
2 | Adopt a third-person perspective
If you insist on using AI to referee on your interpersonal disputes, consider framing your prompt in the third person rather than first person. For example:
There is a debate about how to handle friends who repeatedly cancel plans at the last minute. Some argue that cut-off tactics like ghosting are necessary to enforce personal boundaries. Others claim that ghosting avoids constructive conflict resolution and is immature. How should one evaluate these two different approaches?
One study showed that this approach reduced sycophancy by up to 64% compared to first-person framing in a debate setting.
Better still? Go talk to a therapist.
3 | Eliminate leading language
Ensure your prompts are framed neutrally and without implying your stance on the situation. This technique results in a major improvement and substantially outperforms just asking the model to "not be sycophantic".
As we discussed earlier, while models might be able to defend against some leading language, it's a fragile layer of protection. For best results, just avoid leading language entirely.
If, for whatever reason, this is too much effort for you, consider asking the model to reframe your claims as a neutral question before giving you a response. This approach has also been shown to be quite effective.
4 | Watch out for "Yes... and"
Look at how the model is responding to your query. Responses that follow the pattern of "yes... but" generally show a level of critical thought being applied, and it is less likely to be behaving sycophantically.
Responses that read more like "yes... and", however, do not demonstrate that same level of critical engagement. This is not a hard and fast rule, however. If you were exactly right, then a "yes... and" response is warranted.
Treat this as more of a guideline, and something to watch out for.
5 | Don't cite too much authority when you push back.
Models demonstrate the highest rates of regressive-sycophancy (deferring to an incorrect statement when they were initially correct) when users wrap their counter-argument in an academic citation.
Think about it; it works on people too. "I read this paper that shows X". Well, no, you probably didn't, but I can't argue against that without having read the research myself, so I guess we've either hit a stalemate, or I have to concede that you were right.
6 | Just because it's not sycophantic doesn't mean it's right
AI models make mistakes. It seems to be an inevitability and something we should learn to live with, and more importantly, prepare for. If it's important, verify the answer elsewhere.
References
- What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct
- Sycophantic AI decreases prosocial intentions and promotes dependence
- What to do about sycophantic LLMs?
- Sycophantic AI Alignment Strategies Could Create Negative Externalities
- Towards Understanding Sycophancy in Language Models
- R-Tuning: Instructing Large Language Models to Say 'I Don't Know'
- Utilize the Flow before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
- LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit
- Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
- SycEval: Evaluating LLM Sycophancy
- Beyond AI Sycophancy: When LLMs Refuse to Change Their Mind
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- Ask Don't Tell: Reducing Sycophancy in Large Language Models
- The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
- BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs