Back to list
DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency
Research BreakthroughDeepSeekARC-AGIAI Benchmarks

DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency

DeepSeek has unveiled the latest benchmark results for its V4 Flash 0731 model, demonstrating exceptional performance on the ARC-AGI (Abstraction and Reasoning Corpus) benchmarks. The model achieved a peak score of 89.0% on the ARC-AGI-1 Semi-Private benchmark and 61.4% on the ARC-AGI-2 Semi-Private benchmark under 'Max effort' conditions. Notably, DeepSeek has optimized these reasoning tasks for extreme cost-efficiency, with costs ranging from $0.02 to $0.04 per task. The results highlight the model's ability to handle complex logical reasoning through three distinct variants—Max, High, and Low—each offering a different balance of accuracy and computational intensity. These findings, verified across 120 tasks in the ARC-AGI-2 Public Eval, position DeepSeek V4 Flash 0731 as a significant contender in the pursuit of advanced machine reasoning.

Hacker News

Key Takeaways

  • Record-Breaking Accuracy: DeepSeek V4 Flash 0731 reached an 89.0% accuracy rate on the ARC-AGI-1 Semi-Private benchmark.
  • Tiered Reasoning Variants: The model utilizes three reasoning levels—Max, High, and Low—allowing for flexible performance based on task complexity.
  • Disruptive Cost-Efficiency: Reasoning tasks are performed at a remarkably low cost, specifically $0.02 per task for ARC-AGI-1 and $0.04 for ARC-AGI-2.
  • Robust ARC-AGI-2 Performance: Even on the more challenging ARC-AGI-2 Semi-Private benchmark, the model maintained a 61.4% success rate at maximum effort.
  • Granular Task Success: Detailed public evaluation data shows consistent performance across a wide array of specific logical tasks, though some high-difficulty tasks remain unsolved.

In-Depth Analysis

Benchmarking Reasoning Variants and Performance Scaling

The performance of DeepSeek V4 Flash 0731 is categorized into three distinct reasoning variants: Max, High, and Low. This tiered approach reveals how the model's accuracy scales with the amount of 'effort' or computational reasoning applied to a problem. On the ARC-AGI-1 Semi-Private benchmark, the 'Max' variant achieved the headline score of 89.0%, while the 'High' and 'Low' variants followed closely at 87.0% and 84.0%, respectively. This narrow 5% spread suggests that the model possesses a strong baseline reasoning capability even at lower effort levels for ARC-AGI-1 tasks.

However, the ARC-AGI-2 Semi-Private benchmark presents a more significant challenge, where the performance gap between variants becomes more pronounced. The 'Max' effort variant achieved 61.4%, but this dropped to 56.0% for 'High' and further down to 46.0% for 'Low'. This 15.4% performance delta between Max and Low variants indicates that the more complex logic required for ARC-AGI-2 benefits significantly from increased reasoning depth. The data suggests that for advanced AGI-style reasoning, the 'Max' effort configuration is essential for maintaining competitive accuracy.

Cost-Efficiency and Task-Level Granularity

One of the most striking aspects of the DeepSeek V4 Flash 0731 results is the economic feasibility of its reasoning. Achieving an 89% score on ARC-AGI-1 at a cost of only $0.02 per task represents a significant milestone in making high-level AI reasoning accessible. Even the more demanding ARC-AGI-2 tasks, which require more computational resources, only cost $0.04 per task. This pricing structure suggests that DeepSeek has optimized the 'Flash' architecture to deliver high-order cognitive processing without the traditional overhead associated with large-scale reasoning models.

The ARC-AGI-2 Public Eval data, which includes 120 specific tasks, provides a transparent look at where the model excels and where it struggles. For instance, tasks such as '135a2760', '136b0064', and '1818057f' saw successful passes across all three reasoning variants (Max, High, and Low). Conversely, several tasks like '13e47133', '16b78196', and '20a9e565' resulted in failures across all levels, indicating specific types of abstract reasoning that the current V4 Flash architecture has yet to master. Interestingly, some tasks like '195c6913' and '20270e3b' only passed at the 'Max' or 'High' levels, validating the necessity of the multi-tiered reasoning approach for edge-case logical puzzles.

Industry Impact

The results from DeepSeek V4 Flash 0731 have substantial implications for the AI industry, particularly in the evaluation of Artificial General Intelligence (AGI). The ARC-AGI benchmark is widely regarded as a 'hard' test for AI because it measures the ability to learn new skills and solve novel problems rather than relying on memorized training data. DeepSeek's high scores suggest that the industry is moving closer to models that can perform genuine abstract reasoning.

Furthermore, the low cost-per-task model challenges the current industry standard where high-reasoning capabilities are often synonymous with high operational costs. By providing a 'Flash' version that maintains high accuracy on ARC-AGI, DeepSeek is setting a new benchmark for 'performance-per-dollar.' This could accelerate the adoption of AI in fields requiring complex logical verification, such as software engineering, mathematical proofing, and automated scientific discovery, where thousands of reasoning tasks may need to be performed at scale.

Frequently Asked Questions

Question: What is the difference between ARC-AGI-1 and ARC-AGI-2 in these results?

ARC-AGI-1 and ARC-AGI-2 represent different sets of tasks within the Abstraction and Reasoning Corpus. Based on the results, ARC-AGI-2 appears to be significantly more difficult, as DeepSeek V4 Flash 0731 scored 89.0% on ARC-AGI-1 but dropped to 61.4% on ARC-AGI-2 under the same 'Max effort' conditions.

Question: How much does it cost to run reasoning tasks with DeepSeek V4 Flash 0731?

The model is highly cost-optimized. According to the published results, it costs $0.02 per task for ARC-AGI-1 and $0.04 per task for ARC-AGI-2 when running at maximum effort.

Question: What are the three reasoning variants mentioned in the report?

The model utilizes 'Max', 'High', and 'Low' reasoning variants. These represent different levels of computational effort applied to solve a task. The 'Max' variant consistently yields the highest accuracy, particularly on the more complex ARC-AGI-2 benchmark.

Related News

OpenAI Unveils 722 Mathematics Manuscripts Solving Long-Standing Problems with Unreleased Frontier Model
Research Breakthrough

OpenAI Unveils 722 Mathematics Manuscripts Solving Long-Standing Problems with Unreleased Frontier Model

OpenAI has revealed solutions to a collection of long-standing mathematics problems produced by an unreleased frontier model, presenting the findings in a massive batch of 722 manuscripts categorized into 372 result families that group related papers. The disclosure extends an ongoing series of breakthroughs that have concurrently impressed and unsettled members of the mathematical community. While the results demonstrate advanced computational problem-solving, the publication has simultaneously prompted critical questions regarding research ethics and academic norms. Because the underlying frontier model remains unreleased, researchers are left to examine the vast volume of paper families while navigating the complex implications of proprietary AI-driven scientific discovery.

Research Breakthrough

OpenAI Shares New Mathematical Research and Lean Proof Formalizations from Internal Frontier Model

OpenAI has published new research results addressing open problems in mathematics achieved by an internal frontier model. Alongside these findings, the organization has made Lean proof formalizations and comprehensive research details publicly accessible on GitHub. This release highlights the application of frontier artificial intelligence systems to advanced mathematical problem-solving and formal verification. By releasing formal proofs in the Lean interactive theorem prover, OpenAI allows the mathematical and machine learning communities to inspect, verify, and build upon the frontier model's technical outputs. The update represents a significant step in documenting mathematical reasoning capabilities within advanced AI architectures.

AI Labs Shake Pure Mathematics: Breakthroughs, Millennium Prize Drama, and the Push for Formal Proofs
Research Breakthrough

AI Labs Shake Pure Mathematics: Breakthroughs, Millennium Prize Drama, and the Push for Formal Proofs

Leading artificial intelligence research labs, including OpenAI and Anthropic, have initiated a dramatic shift in pure mathematics over the past year by claiming solutions to long-standing mathematical problems, including one of the prestigious Millennium Prize challenges. These achievements have pushed artificial intelligence systems far beyond what researchers previously anticipated. However, the aggressive Silicon Valley ethos of moving fast and breaking things has generated substantial friction with the traditional mathematical community. As researchers confront black-box outputs that lack formal verification and transparent step-by-step logic, intense debate has erupted over academic rigor versus rapid technological deployment. This analysis examines the technical implications, cultural clashes, and systemic challenges reshaping the frontier where advanced machine learning meets fundamental mathematical discovery.