DeepSeek V4 Flash 0731 Achieves Breakthrough ARC-AGI Scores with High Cost-Efficiency
DeepSeek has unveiled the latest benchmark results for its V4 Flash 0731 model, demonstrating exceptional performance on the ARC-AGI (Abstraction and Reasoning Corpus) benchmarks. The model achieved a peak score of 89.0% on the ARC-AGI-1 Semi-Private benchmark and 61.4% on the ARC-AGI-2 Semi-Private benchmark under 'Max effort' conditions. Notably, DeepSeek has optimized these reasoning tasks for extreme cost-efficiency, with costs ranging from $0.02 to $0.04 per task. The results highlight the model's ability to handle complex logical reasoning through three distinct variants—Max, High, and Low—each offering a different balance of accuracy and computational intensity. These findings, verified across 120 tasks in the ARC-AGI-2 Public Eval, position DeepSeek V4 Flash 0731 as a significant contender in the pursuit of advanced machine reasoning.
Key Takeaways
- Record-Breaking Accuracy: DeepSeek V4 Flash 0731 reached an 89.0% accuracy rate on the ARC-AGI-1 Semi-Private benchmark.
- Tiered Reasoning Variants: The model utilizes three reasoning levels—Max, High, and Low—allowing for flexible performance based on task complexity.
- Disruptive Cost-Efficiency: Reasoning tasks are performed at a remarkably low cost, specifically $0.02 per task for ARC-AGI-1 and $0.04 for ARC-AGI-2.
- Robust ARC-AGI-2 Performance: Even on the more challenging ARC-AGI-2 Semi-Private benchmark, the model maintained a 61.4% success rate at maximum effort.
- Granular Task Success: Detailed public evaluation data shows consistent performance across a wide array of specific logical tasks, though some high-difficulty tasks remain unsolved.
In-Depth Analysis
Benchmarking Reasoning Variants and Performance Scaling
The performance of DeepSeek V4 Flash 0731 is categorized into three distinct reasoning variants: Max, High, and Low. This tiered approach reveals how the model's accuracy scales with the amount of 'effort' or computational reasoning applied to a problem. On the ARC-AGI-1 Semi-Private benchmark, the 'Max' variant achieved the headline score of 89.0%, while the 'High' and 'Low' variants followed closely at 87.0% and 84.0%, respectively. This narrow 5% spread suggests that the model possesses a strong baseline reasoning capability even at lower effort levels for ARC-AGI-1 tasks.
However, the ARC-AGI-2 Semi-Private benchmark presents a more significant challenge, where the performance gap between variants becomes more pronounced. The 'Max' effort variant achieved 61.4%, but this dropped to 56.0% for 'High' and further down to 46.0% for 'Low'. This 15.4% performance delta between Max and Low variants indicates that the more complex logic required for ARC-AGI-2 benefits significantly from increased reasoning depth. The data suggests that for advanced AGI-style reasoning, the 'Max' effort configuration is essential for maintaining competitive accuracy.
Cost-Efficiency and Task-Level Granularity
One of the most striking aspects of the DeepSeek V4 Flash 0731 results is the economic feasibility of its reasoning. Achieving an 89% score on ARC-AGI-1 at a cost of only $0.02 per task represents a significant milestone in making high-level AI reasoning accessible. Even the more demanding ARC-AGI-2 tasks, which require more computational resources, only cost $0.04 per task. This pricing structure suggests that DeepSeek has optimized the 'Flash' architecture to deliver high-order cognitive processing without the traditional overhead associated with large-scale reasoning models.
The ARC-AGI-2 Public Eval data, which includes 120 specific tasks, provides a transparent look at where the model excels and where it struggles. For instance, tasks such as '135a2760', '136b0064', and '1818057f' saw successful passes across all three reasoning variants (Max, High, and Low). Conversely, several tasks like '13e47133', '16b78196', and '20a9e565' resulted in failures across all levels, indicating specific types of abstract reasoning that the current V4 Flash architecture has yet to master. Interestingly, some tasks like '195c6913' and '20270e3b' only passed at the 'Max' or 'High' levels, validating the necessity of the multi-tiered reasoning approach for edge-case logical puzzles.
Industry Impact
The results from DeepSeek V4 Flash 0731 have substantial implications for the AI industry, particularly in the evaluation of Artificial General Intelligence (AGI). The ARC-AGI benchmark is widely regarded as a 'hard' test for AI because it measures the ability to learn new skills and solve novel problems rather than relying on memorized training data. DeepSeek's high scores suggest that the industry is moving closer to models that can perform genuine abstract reasoning.
Furthermore, the low cost-per-task model challenges the current industry standard where high-reasoning capabilities are often synonymous with high operational costs. By providing a 'Flash' version that maintains high accuracy on ARC-AGI, DeepSeek is setting a new benchmark for 'performance-per-dollar.' This could accelerate the adoption of AI in fields requiring complex logical verification, such as software engineering, mathematical proofing, and automated scientific discovery, where thousands of reasoning tasks may need to be performed at scale.
Frequently Asked Questions
Question: What is the difference between ARC-AGI-1 and ARC-AGI-2 in these results?
ARC-AGI-1 and ARC-AGI-2 represent different sets of tasks within the Abstraction and Reasoning Corpus. Based on the results, ARC-AGI-2 appears to be significantly more difficult, as DeepSeek V4 Flash 0731 scored 89.0% on ARC-AGI-1 but dropped to 61.4% on ARC-AGI-2 under the same 'Max effort' conditions.
Question: How much does it cost to run reasoning tasks with DeepSeek V4 Flash 0731?
The model is highly cost-optimized. According to the published results, it costs $0.02 per task for ARC-AGI-1 and $0.04 per task for ARC-AGI-2 when running at maximum effort.
Question: What are the three reasoning variants mentioned in the report?
The model utilizes 'Max', 'High', and 'Low' reasoning variants. These represent different levels of computational effort applied to solve a task. The 'Max' variant consistently yields the highest accuracy, particularly on the more complex ARC-AGI-2 benchmark.

