Back to List
Meituan LongCat Releases General 365 Reasoning Benchmark: Top Models Struggle to Surpass 63% Accuracy
Research BreakthroughMeituanLongCatAI Benchmark

Meituan LongCat Releases General 365 Reasoning Benchmark: Top Models Struggle to Surpass 63% Accuracy

The Meituan LongCat team has officially open-sourced General 365, a new benchmark designed to evaluate the reasoning capabilities of large language models. In a comprehensive assessment involving 26 mainstream AI models, the results highlight a significant performance gap in complex reasoning. Gemini 3 Pro, currently the top-performing model in this evaluation, achieved an accuracy rate of only 62.8%. Notably, the vast majority of the models tested failed to reach the 60% accuracy threshold, which is considered the passing mark for this benchmark. This release aims to establish a more rigorous standard for AI reasoning, exposing the current limitations of even the most advanced models in the industry.

美团技术团队

Key Takeaways

  • New Reasoning Benchmark: Meituan's LongCat team has officially released and open-sourced "General 365," a specialized tool for evaluating AI reasoning.
  • Comprehensive Testing: The benchmark was used to assess 26 mainstream large language models to determine their logical and reasoning proficiency.
  • Performance Ceiling: Gemini 3 Pro emerged as the leader in the test, yet it only managed an accuracy rate of 62.8%.
  • Widespread Underperformance: Most models involved in the study were unable to reach the 60% passing threshold, indicating a significant challenge in current AI reasoning capabilities.

In-Depth Analysis

The Emergence of General 365: A New Standard in Reasoning

The Meituan LongCat team has introduced General 365 at a critical juncture in the evolution of artificial intelligence. As large language models (LLMs) become increasingly integrated into complex workflows, the need for a rigorous, specialized evaluation of their reasoning capabilities has become paramount. By open-sourcing General 365, the LongCat team is providing the global AI community with a new "yardstick" to measure progress. This benchmark is specifically designed to move beyond simple knowledge retrieval and focus on the intricate logical processes that define true reasoning.

The decision to test 26 different mainstream models provides a broad and representative cross-section of the current AI landscape. This comprehensive approach ensures that the benchmark's findings are not limited to a specific architecture or provider but instead reflect the general state of the industry. The results suggest that General 365 is a high-bar evaluation tool, designed to challenge models in ways that existing benchmarks might not, thereby revealing the true depth—or lack thereof—of their reasoning faculties.

Analyzing the Performance Gap and the 60% Threshold

The data released by the LongCat team reveals a stark reality: there is a significant performance gap in the realm of AI reasoning. The fact that Gemini 3 Pro, a model recognized for its advanced capabilities, achieved only a 62.8% accuracy rate is highly telling. This score represents the current "ceiling" of performance on the General 365 benchmark, suggesting that even the industry's most sophisticated models have a long way to go before mastering complex reasoning tasks.

Perhaps more concerning is the observation that the vast majority of the 26 models tested could not even reach the 60% mark. In many academic and professional contexts, 60% is considered the minimum passing grade. The failure of most mainstream models to hit this target on General 365 indicates that the benchmark has successfully identified a widespread limitation in current LLM development. This "60% barrier" serves as a clear indicator that while models are becoming more fluent and knowledgeable, their ability to consistently apply logic and reason through complex problems remains a significant hurdle.

Industry Impact

The introduction of General 365 is poised to have a lasting impact on the AI industry by shifting the focus of model evaluation. For a long time, the industry has prioritized scale and general knowledge, but Meituan's new benchmark highlights that reasoning is the next major frontier. By making General 365 open-source, the LongCat team is encouraging transparency and healthy competition among AI developers.

This benchmark provides a clear target for research teams worldwide. The specific data points—such as the 62.8% peak and the sub-60% average—provide a baseline that will likely drive future innovations in model architecture and training methodologies. As developers strive to surpass the benchmarks set by General 365, we can expect a renewed focus on logical consistency and multi-step reasoning, which are essential for the next generation of AI applications.

Frequently Asked Questions

Question: What is General 365 and who developed it?

General 365 is a reasoning evaluation benchmark developed and open-sourced by the Meituan LongCat team. It is designed to provide a rigorous standard for testing the logical reasoning capabilities of large language models.

Question: How did the top AI models perform on this benchmark?

According to the test results of 26 mainstream models, Gemini 3 Pro was the top performer with an accuracy of 62.8%. However, the majority of the other models tested failed to reach the 60% accuracy threshold, highlighting a general struggle with the reasoning tasks presented in the benchmark.

Question: Why is General 365 considered a "new yardstick" for the industry?

It is considered a new yardstick because it sets a high difficulty level that current mainstream models struggle to meet. By focusing specifically on reasoning and revealing that most models score below 60%, it establishes a more challenging and precise standard for evaluating the true intelligence of AI systems.

Related News

Understanding AI Catastrophic Risks: A New Taxonomy of Omnicidal Futures by Andrew Critch and Jacob Tsimerman
Research Breakthrough

Understanding AI Catastrophic Risks: A New Taxonomy of Omnicidal Futures by Andrew Critch and Jacob Tsimerman

A significant research paper titled 'A Taxonomy of Omnicidal Futures Involving Artificial Intelligence' has been released by authors Andrew Critch and Jacob Tsimerman. The report provides a structured classification of potential 'omnicidal' events—scenarios where artificial intelligence could lead to the death of all or nearly all human beings. Rather than presenting these outcomes as unavoidable, the authors emphasize that these are possibilities intended to be studied and avoided. The primary goal of the taxonomy is to increase public awareness and generate the necessary support for large institutions to implement preventive measures. By documenting these catastrophic risks, the research seeks to provide a framework for global safety efforts and institutional policy-making to mitigate the most extreme threats posed by advanced AI systems.

Google Research Introduces SymptomAI: Advancing Conversational AI for Everyday Symptom Assessment
Research Breakthrough

Google Research Introduces SymptomAI: Advancing Conversational AI for Everyday Symptom Assessment

Google Research has announced the development of SymptomAI, a novel conversational AI agent specifically designed for everyday symptom assessment. This initiative represents a significant intersection of general science and artificial intelligence, aiming to provide users with a structured, dialogue-based approach to understanding their health concerns. By focusing on conversational interfaces, SymptomAI seeks to bridge the gap between complex medical information and user-friendly health evaluations. The research highlights the potential for AI agents to assist in the preliminary stages of health monitoring, offering a more interactive and accessible method for individuals to track and describe their symptoms. This development underscores Google's ongoing commitment to applying advanced AI research to practical, everyday health challenges, potentially transforming how the public interacts with digital health tools.

Towards a Quantum Computer That Learns From Its Errors: Google Research and Machine Intelligence
Research Breakthrough

Towards a Quantum Computer That Learns From Its Errors: Google Research and Machine Intelligence

Google Research has announced a significant step in the evolution of quantum computing, focusing on systems that can learn from their own errors. This development, categorized under Machine Intelligence, represents a shift from traditional error correction methods toward more autonomous, intelligent quantum systems. By enabling quantum hardware to identify and adapt to errors, this research aims to overcome one of the most persistent challenges in the field: the high sensitivity of qubits to environmental noise. The integration of machine intelligence suggests a future where quantum processors are not only faster but also inherently more reliable through self-learning mechanisms. This approach could potentially accelerate the timeline for practical, large-scale quantum applications by addressing the stability issues that currently limit the technology's scalability.