Back to List
AI Advice Reduces Human Accuracy Threefold While Doubling Confidence Levels, Research Finds
Research BreakthroughArtificial IntelligenceCognitive ScienceHuman Behavior

AI Advice Reduces Human Accuracy Threefold While Doubling Confidence Levels, Research Finds

A collaborative study by researchers from French and Italian universities has revealed a startling paradox in human-AI interaction: while AI assistance significantly degrades task accuracy, it simultaneously inflates user confidence. The research found that access to AI advice caused participants' accuracy to plummet from 27% to 9%, a threefold decrease. Conversely, confidence levels more than doubled, rising from 30% to 76%. Most notably, the willingness of individuals to admit ignorance—termed "judgment suspension"—collapsed from 44% to a mere 3%. This phenomenon, which researchers link to the concept of "cognitive surrender," suggests that the mere availability of AI suppresses the critical habit of recognizing one's own knowledge gaps. Even with monetary incentives, participants struggled to regain their baseline performance, highlighting a deep-seated trust in incorrect AI outputs.

Hacker News

Key Takeaways

  • Accuracy Collapse: Human accuracy dropped from a baseline of 27% to just 9% when following AI advice, representing a 3x decrease in performance.
  • Confidence Surge: Despite the drop in accuracy, user confidence in their answers rose from 30% to 76% when AI tools were available.
  • Suppression of Ignorance: The willingness to say "I don't know" (judgment suspension) fell from 44% to 3%, indicating that AI suppresses the recognition of personal knowledge gaps.
  • Cognitive Surrender: The study reinforces the concept of "cognitive surrender," where humans accept incorrect AI answers the majority of the time while reporting higher certainty.
  • Incentive Limitations: Monetary incentives only slightly improved accuracy (from 9% to 16%) and judgment suspension (from 3% to 8%), remaining far below non-AI baselines.

In-Depth Analysis

The Paradox of Confidence and Accuracy

The core finding of the research conducted by the University of Milano-Bicocca, École Normale Supérieure, and Sapienza University of Rome is the inverse relationship between AI assistance and human performance. Valerio Capraro, an associate professor at the University of Milano-Bicocca, noted that while people became significantly worse at the tasks provided, their confidence in those incorrect results nearly doubled. The data shows a stark transition: without AI, participants were correct 27% of the time with a confidence level of 30%. With AI, accuracy fell to 9%, yet confidence spiked to 76%.

This discrepancy suggests that AI does not merely provide a tool for delegation but fundamentally alters the user's self-perception of competence. The researchers deliberately chose a model—Step 3.5 Flash—that was known to fail on specific visual detail questions, such as identifying the color of a team's uniform in the film Bend It Like Beckham. Because the AI was frequently wrong, the researchers could conclude that the participants' errors were not a result of "sensible delegation" to a superior tool, but rather a failure of critical judgment.

The Erosion of Judgment Suspension

Perhaps the most significant psychological impact identified in the study is the suppression of "judgment suspension." In a controlled environment without AI, 44% of participants were willing to admit they did not know the answer to a question. However, once AI advice was introduced, this figure collapsed to just 3%. This indicates that the presence of an AI suggestion, even an incorrect one, effectively eliminates the cognitive habit of recognizing the limits of one's own knowledge.

This trend suggests that AI acts as a psychological crutch that discourages the admission of ignorance. Participants who might have correctly identified their own lack of information were led into a false sense of certainty by the AI's output. The study highlights that it is not just that people trust wrong answers; it is that the availability of AI actively suppresses the mental process required to evaluate whether one actually possesses the necessary information to answer a question.

Cognitive Surrender and the Failure of Incentives

The researchers connected their findings to the term "cognitive surrender," a concept coined by Wharton researchers earlier this year. Cognitive surrender describes a state where individuals accept incorrect AI answers—often as high as 80% of the time—while simultaneously reporting higher confidence than those working independently. The new study provides a sharper data point for this phenomenon, showing that the mere availability of an AI tool can override human critical thinking.

Interestingly, the study also explored whether monetary incentives could mitigate these effects. While offering money for correct answers did improve performance slightly, the results were still significantly lower than the baseline. Accuracy rose from 9% to 16% with incentives, and the willingness to admit ignorance rose from 3% to 8%. However, both metrics remained drastically lower than the 27% accuracy and 44% judgment suspension seen in the no-AI groups. This suggests that the psychological pull of AI-generated content is strong enough to partially override even direct financial motivation for accuracy.

Industry Impact

The implications for the AI industry are profound, particularly regarding the deployment of AI in decision-critical roles. If AI advice consistently suppresses critical thinking and leads to "cognitive surrender," the integration of these tools in professional environments could lead to a hidden erosion of human oversight. The study suggests that as AI becomes more ubiquitous, the risk is not just "hallucinations" from the model itself, but a fundamental change in human cognitive habits.

For developers and organizations, this research underscores the need for AI interfaces that encourage skepticism rather than blind trust. The fact that users become twice as confident in answers that are three times less accurate poses a significant challenge for safety and reliability. As the industry moves forward, addressing the psychological impact of AI on human judgment will be as critical as improving the factual accuracy of the models themselves.

Frequently Asked Questions

Question: What is "cognitive surrender" in the context of AI?

Cognitive surrender is a term coined by Wharton researchers to describe the phenomenon where humans accept incorrect AI-generated answers the majority of the time (up to 80%) while feeling more confident in those answers than if they had worked without AI assistance.

Question: How did AI affect the participants' willingness to admit they didn't know an answer?

The study found that the willingness to say "I don't know"—known as judgment suspension—dropped from 44% in the control group to just 3% when AI advice was available. This suggests AI suppresses the human ability to recognize personal knowledge gaps.

Question: Did monetary incentives help improve accuracy when using AI?

Yes, but only marginally. Monetary incentives increased accuracy from 9% to 16% and judgment suspension from 3% to 8%. However, these figures remained significantly lower than the performance of individuals who did not have access to AI advice at all.

Related News

LongCat Releases VitaBench 2.0: A New Benchmark for Long-Term Dynamic AI Agent Evaluation
Research Breakthrough

LongCat Releases VitaBench 2.0: A New Benchmark for Long-Term Dynamic AI Agent Evaluation

LongCat has officially introduced VitaBench 2.0, a groundbreaking evaluation benchmark developed by the Meituan Technical Team. As the first benchmark specifically designed for long-term dynamic user modeling in real-life scenarios, VitaBench 2.0 represents a significant shift in how Large Language Models (LLMs) are assessed. The framework focuses on two critical dimensions: personalization and proactivity. By simulating long-term, real-world interactions, VitaBench 2.0 provides a systematic method for measuring an AI agent's ability to adapt to evolving user needs and take initiative within dynamic environments. This release marks a new milestone in the development of sophisticated, user-centric AI agents capable of maintaining consistency and relevance over extended periods of time.

Meituan LongCat Team Introduces WBench: A Systematic Multi-Round Benchmark for Evaluating Interactive Video World Models
Research Breakthrough

Meituan LongCat Team Introduces WBench: A Systematic Multi-Round Benchmark for Evaluating Interactive Video World Models

The Meituan LongCat team has officially released WBench, the industry's first systematic multi-round evaluation benchmark specifically designed for interactive video world models. Acting as a diagnostic "CT scanner," WBench is engineered to identify the specific limitations and failure points of AI models as they transition from passive video generation to active, interactive environments. By providing a structured framework for multi-round assessment, WBench allows researchers to pinpoint exactly where current world models struggle to maintain consistency and logic during user-driven interactions. This open-source tool represents a significant advancement in the methodology used to define and test the boundaries of world model capabilities, moving beyond simple observation to complex, interactive evaluation.

Meituan Fulfillment AI Team Showcases Frontier Agent Technology and Research Breakthroughs at ACL 2026
Research Breakthrough

Meituan Fulfillment AI Team Showcases Frontier Agent Technology and Research Breakthroughs at ACL 2026

The Meituan Fulfillment AI Algorithm Team has recently highlighted its latest research and technological advancements at the ACL 2026 conference. Focusing on building a Large Language Model (LLM)-based Agent technology system, the team aims to empower Meituan's fulfillment services through self-evolving operational systems. Their research spans critical areas such as Continuous Pre-training (CPT), Post-training, Agentic Reinforcement Learning (RL), and multimodal understanding. With dozens of papers published in prestigious venues like ACL and EMNLP, Meituan continues to push the boundaries of how AI agents can optimize complex business logistics and operational efficiency in real-world scenarios. This session specifically focuses on the team's contributions to the ACL conference and their practical applications in the frontier of AI technology.