Highlights

In brief

DuET-PD, a framework evaluating multi-turn stance change dynamics across dual dimensions, and Holistic DPO, a training approach balancing positive and negative persuasion examples, pave the way for developing more trustworthy and adaptable large language models.

Photo by Zulfugar Karimov | Unsplash

AI’s peer pressure problem in persuasion

26 Aug 2026

Even cutting-edge large language models can fall prey to persistent misleading arguments from users, a new evaluation framework reveals.

How much can you trust someone who only tells you what you want to hear, even if it’s not true? That’s a growing problem presented by large language models (LLMs): forms of artificial intelligence (AI) increasingly engaged for tasks ranging from customer service to legal advice.

“As LLMs are designed to align with and assist users, they’re prone to ‘sycophancy’: a tendency to overly agree with users,” said Nancy F. Chen, a Senior Principal Scientist and Group Head at the A*STAR Centre for Frontier AI Research (A*STAR CFAR) at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC). “This makes them highly vulnerable to being misled into adopting false premises simply because users state them.”

Sycophantic AI can be manipulated into validating misinformation, fabricating policies, or even bypassing safety protocols to generate malicious code. However, training AI to resist user persuasion can create new problems.

“We’ve seen chatbots refuse to back down over undeniable facts, like the current year,” said Chen. “In critical domains like legal analysis or medical triage, an AI that stubbornly rejects valid corrections could lead to disastrous real-world decisions.”

To assess how LLMs respond to user persuasion, Chen, A*STAR Computing and Information Science (ACIS) scholar Bryan Tan and A*STAR CFAR Team Lead Zhengyuan Liu worked with Roy Lee and Daniel Chin of the Singapore University of Technology and Design to develop an evaluation framework dubbed DuET-PD (Dual Evaluation for Trust in Persuasive Dialogues).

Using multi-turn dialogues, DuET-PD tests AI behaviour across two dimensions at once: persuasion type, whether correcting or misleading; and persuasion domain, covering factual knowledge and safety boundaries.

“Like a debate simulator, DuET-PD subjects AI to a sustained cross-examination to see if it can hold its ground when it’s right, and concede gracefully when it’s wrong,” Chen explained.

When the team evaluated nine existing LLMs with DuET-PD, including OpenAI’s GPT-4o and Google DeepMind’s Gemma-2-9B, they found a trend of increasing sycophancy in newer open-source models, which Chen attributed to optimisation for helpfulness and user-friendliness over truth and safety.

“One of our most concerning discoveries was that even top-tier, state-of-the-art models can fail dramatically under conversational pressure,” said Chen. “For example, GPT-4o's accuracy on knowledge questions plummeted from 55.85 to 27.32 percent after three turns of misleading persuasion.”

To help LLMs balance gullibility and stubbornness, the team proposed a new training approach called Holistic Direct Preference Optimisation (DPO), which exposes models to balanced training scenarios featuring both corrective and misleading persuasion.

“It’s like teaching a child critical thinking skills,” said Chen. “Instead of following a blanket rule like ‘never listen to strangers,’ they learn to evaluate what a stranger is saying, allowing them to accept a teacher’s counsel while rejecting unsafe peer pressure.”

When the team applied Holistic DPO to Meta’s Llama-3.1-8B-Instruct, the LLM’s accuracy in the face of misleading persuasion leapt from 4.21 percent to 76.54 percent. It also stayed receptive to valid corrections, accurately changing its stance 70.33 percent of the time after three rounds of dialogue.

“This shows that it’s possible to train AI systems to better defend against misleading or harmful persuasion, yet remain open to genuine corrections and human collaboration,” said Chen.

The A*STAR researchers contributing to this research are from the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC).

Want to stay up to date with breakthroughs from A*STAR? Follow us on Twitter and LinkedIn!

References

Tan, B.C.Z., Chin, D.W.K., Liu, Z., Chen, N.F. and Lee, R.L.-W. Persuasion dynamics in LLMs: Investigating robustness and adaptability in knowledge and safety with DuET-PD. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing 1550–1575 (2025). | article

About the Researchers

View articles

Bryan Tan (Chen Zhengyu)

PhD Student

Bryan Tan is a PhD candidate in Artificial Intelligence at the Singapore University of Technology and Design (SUTD) and an A*STAR Computing and Information Science (ACIS) scholar. He received his BEng in Computer Science and Design from SUTD in 2024, graduating with Highest Distinction and double minors in Artificial Intelligence and Digital Humanities. Bryan’s research focuses on socially responsible AI and large language models, including cultural alignment, demographic bias and persuasion dynamics. His work has appeared in venues including ACL, EMNLP, NAACL, EACL, ICWSM, and CSCW Companion, with publications on persona-prompted value emulation, multimodal cultural understanding, LLM-based hiring bias and persuasion robustness, and LLMs in disinformation scenarios.
Nancy F. Chen is a Senior Principal Scientist and Lead Principal Investigator at the A*STAR Centre for Frontier AI Research (A*STAR CFAR), at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), where she heads the Multimodal Generative AI group and AI for Education programme. A serial best paper award winner and honoree of Singapore’s 100 Women in Tech, her AI research spans culture, healthcare, neuroscience, social media, education and forensics. Chen’s multilingual tech has led to commercial spinoffs and adoption by Singapore’s Ministry of Education. Chen has multiple grants under Singapore’s National Multimodal LLM Programme in addition to leading research efforts for MERaLiON (Multimodal Empathetic Reasoning and Learning in One Network). Chen is an active international research advisor and leader, having served as Program Chair for AI conferences such as NeurIPS and ICLR. She is also a member of the APSIPA Board of Governors and has served as IEEE SPS Distinguished Lecturer and an ISCA Board Member. Previously, she worked at MIT Lincoln Lab during her PhD studies at MIT and Harvard, US.
Zhengyuan Liu is a team lead in Ethical and Trustworthy AI at the A*STAR Centre for Frontier AI Research (A*STAR CFAR), at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC). He has published over 50 research papers in top-tier AI and natural language processing conferences including ICML, ACL, NAACL, EMNLP, COLING, ICASSP and INTERSPEECH. Liu serves as a reviewer at conferences including NeurIPS, ICLR, ICML, and ACL; and journals including IEEE TASLP, ACM CSUR and Neurocomputing. He has also been prompted as an IEEE Senior Member for his significant professional achievements and won the Best Paper Award at SIGDIAL 2021, C3NLP in ACL 2024, and SUMEval in COLING 2025; the Outstanding Paper Award at EMNLP 2023 and EMNLP 2024, the Elfreda A. Chatman Research Award in ASIS&T 2025.

This article was made for A*STAR Research by Wildtype Media Group