Back to List
DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency
Research BreakthroughDeepSeekARC-AGIAI Benchmarks

DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency

DeepSeek has unveiled the latest benchmark results for its V4 Flash 0731 model, demonstrating exceptional performance on the ARC-AGI (Abstraction and Reasoning Corpus) benchmarks. The model achieved a peak score of 89.0% on the ARC-AGI-1 Semi-Private benchmark and 61.4% on the ARC-AGI-2 Semi-Private benchmark under 'Max effort' conditions. Notably, DeepSeek has optimized these reasoning tasks for extreme cost-efficiency, with costs ranging from $0.02 to $0.04 per task. The results highlight the model's ability to handle complex logical reasoning through three distinct variants—Max, High, and Low—each offering a different balance of accuracy and computational intensity. These findings, verified across 120 tasks in the ARC-AGI-2 Public Eval, position DeepSeek V4 Flash 0731 as a significant contender in the pursuit of advanced machine reasoning.

Hacker News

Key Takeaways

  • Record-Breaking Accuracy: DeepSeek V4 Flash 0731 reached an 89.0% accuracy rate on the ARC-AGI-1 Semi-Private benchmark.
  • Tiered Reasoning Variants: The model utilizes three reasoning levels—Max, High, and Low—allowing for flexible performance based on task complexity.
  • Disruptive Cost-Efficiency: Reasoning tasks are performed at a remarkably low cost, specifically $0.02 per task for ARC-AGI-1 and $0.04 for ARC-AGI-2.
  • Robust ARC-AGI-2 Performance: Even on the more challenging ARC-AGI-2 Semi-Private benchmark, the model maintained a 61.4% success rate at maximum effort.
  • Granular Task Success: Detailed public evaluation data shows consistent performance across a wide array of specific logical tasks, though some high-difficulty tasks remain unsolved.

In-Depth Analysis

Benchmarking Reasoning Variants and Performance Scaling

The performance of DeepSeek V4 Flash 0731 is categorized into three distinct reasoning variants: Max, High, and Low. This tiered approach reveals how the model's accuracy scales with the amount of 'effort' or computational reasoning applied to a problem. On the ARC-AGI-1 Semi-Private benchmark, the 'Max' variant achieved the headline score of 89.0%, while the 'High' and 'Low' variants followed closely at 87.0% and 84.0%, respectively. This narrow 5% spread suggests that the model possesses a strong baseline reasoning capability even at lower effort levels for ARC-AGI-1 tasks.

However, the ARC-AGI-2 Semi-Private benchmark presents a more significant challenge, where the performance gap between variants becomes more pronounced. The 'Max' effort variant achieved 61.4%, but this dropped to 56.0% for 'High' and further down to 46.0% for 'Low'. This 15.4% performance delta between Max and Low variants indicates that the more complex logic required for ARC-AGI-2 benefits significantly from increased reasoning depth. The data suggests that for advanced AGI-style reasoning, the 'Max' effort configuration is essential for maintaining competitive accuracy.

Cost-Efficiency and Task-Level Granularity

One of the most striking aspects of the DeepSeek V4 Flash 0731 results is the economic feasibility of its reasoning. Achieving an 89% score on ARC-AGI-1 at a cost of only $0.02 per task represents a significant milestone in making high-level AI reasoning accessible. Even the more demanding ARC-AGI-2 tasks, which require more computational resources, only cost $0.04 per task. This pricing structure suggests that DeepSeek has optimized the 'Flash' architecture to deliver high-order cognitive processing without the traditional overhead associated with large-scale reasoning models.

The ARC-AGI-2 Public Eval data, which includes 120 specific tasks, provides a transparent look at where the model excels and where it struggles. For instance, tasks such as '135a2760', '136b0064', and '1818057f' saw successful passes across all three reasoning variants (Max, High, and Low). Conversely, several tasks like '13e47133', '16b78196', and '20a9e565' resulted in failures across all levels, indicating specific types of abstract reasoning that the current V4 Flash architecture has yet to master. Interestingly, some tasks like '195c6913' and '20270e3b' only passed at the 'Max' or 'High' levels, validating the necessity of the multi-tiered reasoning approach for edge-case logical puzzles.

Industry Impact

The results from DeepSeek V4 Flash 0731 have substantial implications for the AI industry, particularly in the evaluation of Artificial General Intelligence (AGI). The ARC-AGI benchmark is widely regarded as a 'hard' test for AI because it measures the ability to learn new skills and solve novel problems rather than relying on memorized training data. DeepSeek's high scores suggest that the industry is moving closer to models that can perform genuine abstract reasoning.

Furthermore, the low cost-per-task model challenges the current industry standard where high-reasoning capabilities are often synonymous with high operational costs. By providing a 'Flash' version that maintains high accuracy on ARC-AGI, DeepSeek is setting a new benchmark for 'performance-per-dollar.' This could accelerate the adoption of AI in fields requiring complex logical verification, such as software engineering, mathematical proofing, and automated scientific discovery, where thousands of reasoning tasks may need to be performed at scale.

Frequently Asked Questions

Question: What is the difference between ARC-AGI-1 and ARC-AGI-2 in these results?

ARC-AGI-1 and ARC-AGI-2 represent different sets of tasks within the Abstraction and Reasoning Corpus. Based on the results, ARC-AGI-2 appears to be significantly more difficult, as DeepSeek V4 Flash 0731 scored 89.0% on ARC-AGI-1 but dropped to 61.4% on ARC-AGI-2 under the same 'Max effort' conditions.

Question: How much does it cost to run reasoning tasks with DeepSeek V4 Flash 0731?

The model is highly cost-optimized. According to the published results, it costs $0.02 per task for ARC-AGI-1 and $0.04 per task for ARC-AGI-2 when running at maximum effort.

Question: What are the three reasoning variants mentioned in the report?

The model utilizes 'Max', 'High', and 'Low' reasoning variants. These represent different levels of computational effort applied to solve a task. The 'Max' variant consistently yields the highest accuracy, particularly on the more complex ARC-AGI-2 benchmark.

Related News

AI Tutoring and the 'TutorMoments' Challenge: When Should AI Help or Hold Back?
Research Breakthrough

AI Tutoring and the 'TutorMoments' Challenge: When Should AI Help or Hold Back?

The emergence of 'TutorMoments,' a project by the Allen Institute for AI (AllenAI) hosted on Hugging Face, highlights a critical frontier in educational technology: the timing of AI intervention. While modern Large Language Models (LLMs) are optimized for immediate helpfulness, effective pedagogy often requires 'holding back' to allow for productive struggle. This analysis explores the core question posed by the TutorMoments initiative: whether AI tutors can discern the optimal moments to provide assistance versus when to remain silent to foster independent problem-solving. By examining the tension between being a 'helpful assistant' and a 'transformative educator,' we delve into the technical and pedagogical implications of this research for the future of personalized, AI-driven learning environments and the shift toward more sophisticated, Socratic digital tutoring systems.

Microsoft Research Unveils Orchard: A New Open Framework for Scalable Agentic AI Systems
Research Breakthrough

Microsoft Research Unveils Orchard: A New Open Framework for Scalable Agentic AI Systems

Microsoft Research has announced the development of Orchard, an open framework specifically designed to address the challenges of scalable agentic AI. Authored by a prominent research team including Baolin Peng and Jianfeng Gao, the project focuses on providing a robust infrastructure for autonomous AI agents. As the industry shifts from simple conversational models to complex, multi-agent systems, Orchard aims to provide the necessary scalability and openness required for broad implementation. The framework represents a strategic move by Microsoft to standardize the development of agent-based architectures, ensuring that AI systems can operate efficiently at scale while remaining accessible to the global research and development community through an open-source approach.

Research Breakthrough

The Computational Theory of Mind: Exploring the Foundations of Cognitive Science and Artificial Intelligence

The Computational Theory of Mind (CTM) posits that the human mind functions as a sophisticated computational system, a concept that gained significant traction during the computer revolution. Originally achieving orthodox status within cognitive science during the 1960s and 1970s, CTM suggests that mental processes—including reasoning, perception, and linguistic comprehension—can be understood as computational operations. However, the theory currently faces pressure from alternative paradigms. To sustain the validity of CTM, researchers must address three critical challenges: defining the nature of mental computation, proving its existence within the human mind, and reconciling computational models with both neurophysiological data and intentional representational states. This analysis explores the historical dominance of CTM, its reliance on Turing machine concepts, and the ongoing philosophical efforts to bridge the gap between biological brains and thinking machines.