Back to list
DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency
Research BreakthroughDeepSeekARC-AGIAI Benchmarks

DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency

DeepSeek has unveiled the latest benchmark results for its V4 Flash 0731 model, demonstrating exceptional performance on the ARC-AGI (Abstraction and Reasoning Corpus) benchmarks. The model achieved a peak score of 89.0% on the ARC-AGI-1 Semi-Private benchmark and 61.4% on the ARC-AGI-2 Semi-Private benchmark under 'Max effort' conditions. Notably, DeepSeek has optimized these reasoning tasks for extreme cost-efficiency, with costs ranging from $0.02 to $0.04 per task. The results highlight the model's ability to handle complex logical reasoning through three distinct variants—Max, High, and Low—each offering a different balance of accuracy and computational intensity. These findings, verified across 120 tasks in the ARC-AGI-2 Public Eval, position DeepSeek V4 Flash 0731 as a significant contender in the pursuit of advanced machine reasoning.

Hacker News

Key Takeaways

  • Record-Breaking Accuracy: DeepSeek V4 Flash 0731 reached an 89.0% accuracy rate on the ARC-AGI-1 Semi-Private benchmark.
  • Tiered Reasoning Variants: The model utilizes three reasoning levels—Max, High, and Low—allowing for flexible performance based on task complexity.
  • Disruptive Cost-Efficiency: Reasoning tasks are performed at a remarkably low cost, specifically $0.02 per task for ARC-AGI-1 and $0.04 for ARC-AGI-2.
  • Robust ARC-AGI-2 Performance: Even on the more challenging ARC-AGI-2 Semi-Private benchmark, the model maintained a 61.4% success rate at maximum effort.
  • Granular Task Success: Detailed public evaluation data shows consistent performance across a wide array of specific logical tasks, though some high-difficulty tasks remain unsolved.

In-Depth Analysis

Benchmarking Reasoning Variants and Performance Scaling

The performance of DeepSeek V4 Flash 0731 is categorized into three distinct reasoning variants: Max, High, and Low. This tiered approach reveals how the model's accuracy scales with the amount of 'effort' or computational reasoning applied to a problem. On the ARC-AGI-1 Semi-Private benchmark, the 'Max' variant achieved the headline score of 89.0%, while the 'High' and 'Low' variants followed closely at 87.0% and 84.0%, respectively. This narrow 5% spread suggests that the model possesses a strong baseline reasoning capability even at lower effort levels for ARC-AGI-1 tasks.

However, the ARC-AGI-2 Semi-Private benchmark presents a more significant challenge, where the performance gap between variants becomes more pronounced. The 'Max' effort variant achieved 61.4%, but this dropped to 56.0% for 'High' and further down to 46.0% for 'Low'. This 15.4% performance delta between Max and Low variants indicates that the more complex logic required for ARC-AGI-2 benefits significantly from increased reasoning depth. The data suggests that for advanced AGI-style reasoning, the 'Max' effort configuration is essential for maintaining competitive accuracy.

Cost-Efficiency and Task-Level Granularity

One of the most striking aspects of the DeepSeek V4 Flash 0731 results is the economic feasibility of its reasoning. Achieving an 89% score on ARC-AGI-1 at a cost of only $0.02 per task represents a significant milestone in making high-level AI reasoning accessible. Even the more demanding ARC-AGI-2 tasks, which require more computational resources, only cost $0.04 per task. This pricing structure suggests that DeepSeek has optimized the 'Flash' architecture to deliver high-order cognitive processing without the traditional overhead associated with large-scale reasoning models.

The ARC-AGI-2 Public Eval data, which includes 120 specific tasks, provides a transparent look at where the model excels and where it struggles. For instance, tasks such as '135a2760', '136b0064', and '1818057f' saw successful passes across all three reasoning variants (Max, High, and Low). Conversely, several tasks like '13e47133', '16b78196', and '20a9e565' resulted in failures across all levels, indicating specific types of abstract reasoning that the current V4 Flash architecture has yet to master. Interestingly, some tasks like '195c6913' and '20270e3b' only passed at the 'Max' or 'High' levels, validating the necessity of the multi-tiered reasoning approach for edge-case logical puzzles.

Industry Impact

The results from DeepSeek V4 Flash 0731 have substantial implications for the AI industry, particularly in the evaluation of Artificial General Intelligence (AGI). The ARC-AGI benchmark is widely regarded as a 'hard' test for AI because it measures the ability to learn new skills and solve novel problems rather than relying on memorized training data. DeepSeek's high scores suggest that the industry is moving closer to models that can perform genuine abstract reasoning.

Furthermore, the low cost-per-task model challenges the current industry standard where high-reasoning capabilities are often synonymous with high operational costs. By providing a 'Flash' version that maintains high accuracy on ARC-AGI, DeepSeek is setting a new benchmark for 'performance-per-dollar.' This could accelerate the adoption of AI in fields requiring complex logical verification, such as software engineering, mathematical proofing, and automated scientific discovery, where thousands of reasoning tasks may need to be performed at scale.

Frequently Asked Questions

Question: What is the difference between ARC-AGI-1 and ARC-AGI-2 in these results?

ARC-AGI-1 and ARC-AGI-2 represent different sets of tasks within the Abstraction and Reasoning Corpus. Based on the results, ARC-AGI-2 appears to be significantly more difficult, as DeepSeek V4 Flash 0731 scored 89.0% on ARC-AGI-1 but dropped to 61.4% on ARC-AGI-2 under the same 'Max effort' conditions.

Question: How much does it cost to run reasoning tasks with DeepSeek V4 Flash 0731?

The model is highly cost-optimized. According to the published results, it costs $0.02 per task for ARC-AGI-1 and $0.04 per task for ARC-AGI-2 when running at maximum effort.

Question: What are the three reasoning variants mentioned in the report?

The model utilizes 'Max', 'High', and 'Low' reasoning variants. These represent different levels of computational effort applied to solve a task. The 'Max' variant consistently yields the highest accuracy, particularly on the more complex ARC-AGI-2 benchmark.

Related News

Google Research Introduces Planetary Prediction Engine: Automating Global Models via Earth AI
Research Breakthrough

Google Research Introduces Planetary Prediction Engine: Automating Global Models via Earth AI

Google Research has announced the development of the Planetary Prediction Engine, a sophisticated framework designed to automate global modeling through the application of Earth AI. This initiative represents a significant advancement in the field of planetary science, focusing on the transition from manual modeling processes to automated, AI-driven systems. By leveraging Earth AI, the engine aims to streamline the creation and deployment of models that analyze and predict phenomena on a global scale. This development highlights Google's ongoing commitment to utilizing artificial intelligence for environmental and planetary-scale insights, potentially transforming how researchers interact with complex global datasets and improving the efficiency of planetary predictions.

Research Breakthrough

Impact of ChatGPT and Critical-Thinking Training on University Student Performance and Originality: A Comprehensive Study Analysis

A significant randomized study involving over 1,000 university students has been conducted to evaluate the intersection of ChatGPT usage and critical-thinking training. The research focuses on how these elements influence student performance, the quality of their answers, and the breadth of their thinking during real-world academic assignments. By examining key metrics such as originality and overall academic achievement, the study aims to provide data-driven insights into the role of generative AI in higher education. This analysis explores the scope of the research and its focus on determining whether AI tools, when paired with specific cognitive training, can enhance the learning process without compromising the integrity and uniqueness of student work.

Google Research Announces GlucoFM: A Specialized Foundation Model for Continuous Glucose Monitoring
Research Breakthrough

Google Research Announces GlucoFM: A Specialized Foundation Model for Continuous Glucose Monitoring

Google Research has unveiled GlucoFM, a groundbreaking foundation model specifically engineered for continuous glucose monitoring (CGM). Categorized under Health & Bioscience, this development signifies a major leap in applying large-scale artificial intelligence to physiological data. GlucoFM represents the adaptation of foundation model architectures—which have revolutionized natural language processing—to the specialized field of metabolic health. By focusing on the continuous streams of data generated by CGM devices, Google Research aims to enhance the precision and utility of glucose tracking. This initiative underscores the increasing role of specialized AI in chronic disease management and the broader evolution of personalized healthcare technology.