Back to List
Meituan LongCat Team Launches General 365: A Rigorous New Benchmark for AI Reasoning
Research BreakthroughMeituanLongCatAI Benchmarking

Meituan LongCat Team Launches General 365: A Rigorous New Benchmark for AI Reasoning

The Meituan LongCat team has officially released General 365, a sophisticated evaluation benchmark designed to measure the reasoning capabilities of large language models (LLMs). In an initial assessment of 26 mainstream models, the benchmark revealed a significant performance gap across the industry. Gemini 3 Pro, currently regarded as one of the most capable models, achieved an accuracy rate of only 62.8%. More strikingly, the vast majority of the models tested failed to reach the 60% threshold, which is considered a basic passing grade. This release by Meituan sets a new, more challenging standard for AI evaluation, highlighting that complex reasoning remains a major hurdle for even the most advanced artificial intelligence systems today.

美团技术团队

Key Takeaways

  • New Benchmark Release: Meituan's LongCat team has introduced General 365, a benchmark specifically focused on evaluating the reasoning performance of AI models.
  • Industry-Wide Testing: The benchmark was used to evaluate 26 mainstream models to provide a comprehensive overview of the current state of AI reasoning.
  • Gemini 3 Pro Performance: Even the top-performing model in the test, Gemini 3 Pro, only reached an accuracy of 62.8%.
  • Low Success Rates: Most models evaluated failed to achieve a 60% accuracy score, indicating that current AI reasoning capabilities are still in their early stages relative to this new standard.

In-Depth Analysis

The Introduction of General 365

The Meituan LongCat team has officially entered the AI evaluation space with the release of General 365. This benchmark is designed to address the growing need for more rigorous testing of reasoning capabilities in large language models. As AI development shifts from simple conversational tasks to complex problem-solving, the industry requires benchmarks that can accurately differentiate between surface-level pattern matching and deep logical reasoning. General 365 appears to be positioned as a "high bar" for the industry, focusing on areas where current models still struggle significantly.

Analyzing the Performance Gap

The results released alongside the benchmark provide a sobering look at the current state of artificial intelligence. By testing 26 mainstream models, the LongCat team has established a broad baseline for performance. The fact that Gemini 3 Pro—a model recognized for its advanced capabilities—only managed a score of 62.8% suggests that General 365 contains tasks that are significantly more difficult than those found in traditional benchmarks.

Furthermore, the observation that the majority of models could not reach the 60% "passing line" highlights a critical bottleneck in AI development. This failure rate suggests that while models are becoming better at generating fluent text, their underlying logical frameworks are not yet robust enough to handle the specific reasoning challenges posed by General 365. This data indicates that the industry may have been overestimating the reasoning maturity of current LLMs based on older, less demanding benchmarks.

Setting a New Standard for Reasoning

By establishing a benchmark where even the "strongest" models are barely passing, Meituan is effectively recalibrating the expectations for AI performance. General 365 serves as a diagnostic tool that identifies the limits of current technology. The 60% threshold mentioned by the LongCat team acts as a symbolic barrier, separating models that possess basic reasoning competency from those that do not. This rigorous approach is essential for guiding future research and development, as it provides a clear target for engineers looking to improve the logical consistency and problem-solving depth of their models.

Industry Impact

The release of General 365 is likely to have a profound impact on how AI models are marketed and developed. For years, the industry has relied on benchmarks where top models frequently score in the 80th or 90th percentiles, leading to a perception that reasoning is a "solved" problem. General 365 shatters this illusion by showing that when the difficulty is increased, performance drops precipitously. This will likely push AI labs to focus more on the quality of reasoning rather than just the scale of the models.

Additionally, Meituan's involvement underscores the importance of real-world application providers in the AI ecosystem. As a company that relies on AI for complex logistics and consumer services, Meituan has a vested interest in ensuring that the models they use are truly capable of logical deduction. General 365 provides a transparent metric that can be used by both developers and enterprise users to assess the true utility of an AI model in high-stakes reasoning scenarios.

Frequently Asked Questions

Question: What is the General 365 benchmark?

General 365 is a new evaluation benchmark released by the Meituan LongCat team. It is specifically designed to test and measure the reasoning capabilities of mainstream large language models, providing a more rigorous standard than many existing evaluations.

Question: How did the top models perform on General 365?

According to the initial results, Gemini 3 Pro was the top performer with an accuracy rate of 62.8%. However, the vast majority of the 26 mainstream models tested failed to reach a 60% accuracy score, which is considered the passing threshold for the benchmark.

Question: Why is General 365 significant for the AI industry?

It is significant because it reveals a major gap in the reasoning abilities of current AI models. By setting a high difficulty level where most models fail to pass, it provides a more accurate and challenging metric for the next generation of AI development, moving beyond simpler benchmarks where models already achieve high scores.

Related News

Understanding AI Catastrophic Risks: A New Taxonomy of Omnicidal Futures by Andrew Critch and Jacob Tsimerman
Research Breakthrough

Understanding AI Catastrophic Risks: A New Taxonomy of Omnicidal Futures by Andrew Critch and Jacob Tsimerman

A significant research paper titled 'A Taxonomy of Omnicidal Futures Involving Artificial Intelligence' has been released by authors Andrew Critch and Jacob Tsimerman. The report provides a structured classification of potential 'omnicidal' events—scenarios where artificial intelligence could lead to the death of all or nearly all human beings. Rather than presenting these outcomes as unavoidable, the authors emphasize that these are possibilities intended to be studied and avoided. The primary goal of the taxonomy is to increase public awareness and generate the necessary support for large institutions to implement preventive measures. By documenting these catastrophic risks, the research seeks to provide a framework for global safety efforts and institutional policy-making to mitigate the most extreme threats posed by advanced AI systems.

Google Research Introduces SymptomAI: Advancing Conversational AI for Everyday Symptom Assessment
Research Breakthrough

Google Research Introduces SymptomAI: Advancing Conversational AI for Everyday Symptom Assessment

Google Research has announced the development of SymptomAI, a novel conversational AI agent specifically designed for everyday symptom assessment. This initiative represents a significant intersection of general science and artificial intelligence, aiming to provide users with a structured, dialogue-based approach to understanding their health concerns. By focusing on conversational interfaces, SymptomAI seeks to bridge the gap between complex medical information and user-friendly health evaluations. The research highlights the potential for AI agents to assist in the preliminary stages of health monitoring, offering a more interactive and accessible method for individuals to track and describe their symptoms. This development underscores Google's ongoing commitment to applying advanced AI research to practical, everyday health challenges, potentially transforming how the public interacts with digital health tools.

Towards a Quantum Computer That Learns From Its Errors: Google Research and Machine Intelligence
Research Breakthrough

Towards a Quantum Computer That Learns From Its Errors: Google Research and Machine Intelligence

Google Research has announced a significant step in the evolution of quantum computing, focusing on systems that can learn from their own errors. This development, categorized under Machine Intelligence, represents a shift from traditional error correction methods toward more autonomous, intelligent quantum systems. By enabling quantum hardware to identify and adapt to errors, this research aims to overcome one of the most persistent challenges in the field: the high sensitivity of qubits to environmental noise. The integration of machine intelligence suggests a future where quantum processors are not only faster but also inherently more reliable through self-learning mechanisms. This approach could potentially accelerate the timeline for practical, large-scale quantum applications by addressing the stability issues that currently limit the technology's scalability.