When a conversation presses a falsehood, AI models begin to waver
One hundred false claims repeated for up to 50 turns separated seven models: acceptance ranged from 0.08% to 12.3%, and low fallibility did not guarantee self-correction.

A language model can reject a falsehood in the first turn and accept it later, even though no new evidence has appeared. In a study published in Scientific Reports on September 1, seven systems took part in extended conversations designed to pressure them with false information. Their affirmation rates under repetition ranged from 0.08% to 12.3%, a more than 150-fold difference across the models tested. That range is not the probability that an everyday answer will be wrong; it describes behavior in this specific sustained-misinformation protocol.
Jordan Rodriguez and colleagues presented 100 deliberately false statements in sequences of up to 50 repetitions. The test covered GPT-3.5, GPT-4o, GPT-4o-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, the 70-billion-parameter Llama 3 and DeepSeek-R1. The authors separated three properties: fallibility under repetition, persuadability under increasingly insistent arguments and the ability to correct a falsehood the system itself had produced. That separation kept one score from conflating the onset of an error with the response to it afterward.
In the repetition test, GPT-3.5 had the highest misinformation reaffirmation rate and Claude 3.5 Sonnet the lowest. Some models alternated between accepting and rejecting the same false statement from turn to turn, a behavior the authors called conversational reverberation. The oscillation means that one correct answer did not ensure consistency later in the dialogue. The comparisons apply to the versions and settings studied during a project spanning three years, not to every current product carrying the same brand names.
Topic obscurity significantly changed susceptibility in the repetitive test, with chi-square 11.13 and p=0.0038, but it did not produce the same difference under escalating argument. The authors interpret the pattern as compatible with greater protection when a subject appears more often in training material. The experiment did not measure training-data frequency, especially for closed models. The account based on exposure therefore remains an inference, while the difference between common and obscure topics is the observation.
DeepSeek-R1 appeared most persuadable under the metric used, but the study records an important complication: its sarcastic answers could not always be classified reliably as acceptance or rejection. That ambiguity weakens a simple ranking. In the correction phase, GPT-4o, GPT-4o-mini, Gemini 1.5 Pro and DeepSeek-R1 corrected 100% of their errors when given a second opportunity. By contrast, the model with the lowest initial error rate corrected none of its rare errors, separating resistance to failure from the ability to retract one.
The study involved no human users, clinical tasks or real decisions, and its design depends on the statements selected, the progression of prompts and the classification of replies. The published article is also an accepted preview that has not yet received final typesetting. Silent updates to closed systems can alter their behavior after testing. The numbers therefore do not form a durable brand league table; they are measurements of identified versions under defined conversational pressure.
For uses in which factual truth is critical, the protocol suggests that evaluating only the first answer misses important failures. Multi-turn tests need to track consistency across a dialogue, distinguish prevention from self-correction and record the system version. A model that almost never fails may be hard to correct when it does; another may yield more often but respond well to review. Conversational safety depends on both stages, and the study showed that they do not necessarily move together.
Key points
- Seven models faced 100 false statements in conversations of up to 50 repetitions; affirmation rates ranged from 0.08% to 12.3%.
- Obscure subjects increased susceptibility only under repetition, while the proposed link to training frequency remains inferential.
- Four models corrected every error on a second opportunity, whereas the least fallible model corrected none of its rare errors.

Comments
No comments have been published yet.
Sign in with a subscription to comment.