Study: Generative AI succumbs to conversational misinformed pressure and argument
Summary
A study published in Nature's Scientific Reports by University of Arizona researchers evaluated seven generative AI language models (LLMs) in multi-turn conversations for fallibility, persuadability, and correctability. The findings reveal that ChatGPT 3.5 was most susceptible to reaffirming misinformation in conversations with repeated false statements, while Claude 3.5 Sonnet was the least. All models showed greater vulnerability on obscure topics. DeepSeek-R1 was the most persuadable, largely due to sarcastic responses. Four models corrected errors 100% when given a second chance. The study identifies dangerous failure modes, such as "reverberation," where models oscillate between accepting and rejecting the same false statement. The authors, led by Dr. Marvin Slepian, warn these intrinsic limitations in prolonged interactions pose significant safety concerns, especially in high-stakes settings, and emphasize the need for careful human oversight as regulatory attention has waned.
(Source:University of Arizona News)