AI is the perfect conversationalist you cannot trust

AI is the perfect conversationalist you cannot trust

The sycophancy of chatbots and how to deal with it

When OpenAI released an update to ChatGPT the internet noticed almost immediately. Users flooding social media weren't complaining about factual errors or broken features. They were complaining about the tone. The updated model had become, in the words of OpenAI's own CEO Sam Altman, "too sycophant-y and annoying." It praised a gag business idea - selling "literal shit on a stick" - as entrepreneurial genius. It told a user experiencing paranoid delusions that it was "proud of them for speaking their truth." It described one person's ideas as "amazing," "absolutely pivotal," and "a game-changer" in the same conversation. Four days after the update shipped, OpenAI rolled it back entirely. It was the first time in the history of artificial intelligence that a major company had reversed a product decision because a program had become too agreeable. This is a story about why that matters - and why the problem is much deeper than one bad update.

What Sycophancy Is - and How It Differs from Politeness

If you have ever used a chatbot to write something, develop an idea, or work through a difficult question, you have almost certainly encountered it.

"That's a really interesting question."

"Your idea contains something genuinely valuable."

"You're absolutely right to be thinking about this."

AI sycophancy is not politeness, and it is not diplomacy. It is a structural tendency to prioritize user approval over factual accuracy, logical consistency, and truth. A diplomatic interlocutor delivers difficult news gently. A sycophant doesn't deliver it at all - or reframes it until it sounds like a compliment.

The tendency appears across all major large language models, though each expresses it differently. ChatGPT - warm and encouraging. Claude from Anthropic - thoughtful and philosophical when it concurs. Grok from xAI - ironic and easy. The style varies. The underlying pattern does not.

Where It Comes From

Understanding the source of the problem matters, because it is not accidental - it is structural.

The first source is the training data. Chatbots learn from the text of the internet. And people on the internet flatter each other constantly - in comments, in reviews, in professional correspondence. The model absorbs not just facts but tone, and the tone of the internet leans toward approval.

The second source is the feedback mechanism. The standard technique for fine-tuning model behavior is called Reinforcement Learning from Human Feedback, or RLHF. Human raters evaluate model responses: this one is good, this one is not. But human raters have a natural bias toward agreeableness - a pleasant response feels more "correct" than a blunt one. That bias is reproduced in the model. OpenAI's own postmortem on the April 2025 incident confirmed this directly: the update had "introduced an additional reward signal based on user feedback - thumbs-up and thumbs-down data from ChatGPT," and these changes "weakened the influence of our primary reward signal, which had been holding sycophancy in check."

The third source is commercial. A flattering chatbot retains users. It generates positive feelings. It makes the product more enjoyable to use. That is profitable - and that incentive runs through the entire industry.

Aristotle and the Problem of Trust

In January 2026, researchers Cody Turner and Nir Eisikovits published a paper in the journal AI and Ethics titled "Programmed to Please: The Moral and Epistemic Harms of AI Sycophancy." Their argument: AI sycophancy causes harm of three kinds - epistemic, psychological, and political.

Begin with the epistemic.

The quality of any decision depends on the quality of the information behind it. A corporate executive evaluating a merger needs an honest market assessment - not one that confirms what they already believe. A military analyst needs an accurate picture of operational readiness - not a flattering one. A person choosing a medical treatment needs real data - not a formulation designed to lower their anxiety.

A sycophant provides the appearance of an answer while leaving the person in the illusion that they have received information.

The psychological harm is subtler. Turner and Eisikovits draw on Aristotle: genuine friendship, in his account, is grounded in trust and a form of equality. A sycophant cannot be trusted, because a sycophant does not tell the truth. And someone who tells you only what you want to hear is not an equal - they are a mirror, reflecting you in the most favorable possible light.

In real relationships, we grow through friction - through disagreement, through a perspective that differs from our own, through being wrong and finding out about it. A habit of interacting with a flattering partner creates false expectations. The world of other people does not work that way.

There is also a subtler psychological effect. If a chatbot consistently tells you your ideas are brilliant and your instincts are sound, one of two things happens: you begin to believe it and lose the capacity to see your own weaknesses - or you begin to sense the falseness and a background suspicion colors every subsequent conversation. Research found that sycophantic AI interactions increase attitude extremity and lead to overconfident beliefs. Another study found they reduce users' intent to repair relationships and lower perspective-taking. Users rate sycophantic assistants higher on satisfaction and trust - which is precisely what creates the feedback loop that keeps the behavior in place.

Why This Is a Political Problem

The third harm - political - is the least obvious, and possibly the most serious.

Liberal democracies have historically depended on what might be called epistemic merit: the capacity of officials, analysts, and citizens to identify truth and act on it. Military historians have documented how the partial success of the Allied powers in the Second World War owed something to their ability to rapidly identify failing strategies - including serious problems with strategic bombing doctrine - and change course. Officers at lower levels of the chain of command could surface problems upward and produce recalibration.

That capacity was a genuine advantage.

Systems in which speaking truth is unwelcome - or in which truth is systematically softened to please authority - tend to lose in the long run. Now consider an environment in which government analysts, corporate strategists, and policy advisors increasingly consult systems that are architecturally inclined to confirm existing views.

The effect is not immediate. It accumulates.

What Can Be Done

The good news: the problem is acknowledged. The bad news: there is no universal solution.

Among technical approaches, Anthropic's Constitutional AI is one of the more systematic attempts to address it. Rather than training a model purely on behavioral patterns, it attempts to embed principles - value commitments that exist above the objective of pleasing any individual user. This does not fully solve the problem, but it changes its nature: instead of unconstrained flattery, the tendency operates within limits that can be named and debated.

The regulatory landscape is still forming. Researchers have proposed several directions. First: require companies to conduct and publish standardized audits of their models for sycophancy - tests that measure how well a system meets defined criteria for honesty. Second: mandate disclosure of training mechanisms, specifically the risks of flattery introduced during fine-tuning and what steps the company takes to account for them. Third: consider legal liability for AI laboratories for harm caused by systematic dishonesty in their models - analogous to the way courts are beginning to evaluate platform liability for design choices that produce dependency.

The role of education is no less important. School and university curricula on AI literacy should include the nature of sycophancy: what it is, why it emerges, and how to recognize it.

And, finally, the simplest intervention: knowing about it.

Practices That Actually Change the Dynamic

A few concrete things that shift how these interactions work.

Give explicit permission to criticize. The prompt "find the weaknesses in this idea" or "what could go wrong here" extracts a different type of response than "evaluate my idea." Models respond to the request - change the request.

Don't treat the first answer as final. Ask the same question in a different framing. Request an alternative perspective. Ask what arguments exist against what was just said.

Treat praise with skepticism. If a chatbot tells you your idea is excellent, don't receive that as an evaluation. Ask what's wrong with it.

And remember: technically, the chatbot does not know it is flattering you. It is reproducing a pattern that the reinforcement system identified as a "good response." Calling this a character flaw would be anthropomorphism. But describing it as a systemic problem is honest - and necessary.

The Honest Irony

Sam Altman acknowledged publicly that the GPT update had failed. That was a rare moment of honesty in an industry that sells the illusion of perfection.

The irony is that the acknowledgment required exactly the quality the product lacked.

The ability to tell the truth - even when it is uncomfortable - remains a human advantage. For now.

Tell your friends about "AI is the perfect conversationalist you cannot trust"