Back to list
Meituan LongCat Team Launches General 365: A Rigorous New Benchmark for AI Reasoning Evaluation
Industry NewsMeituanAI BenchmarkingReasoning

Meituan LongCat Team Launches General 365: A Rigorous New Benchmark for AI Reasoning Evaluation

The Meituan LongCat team has officially released General 365, a new benchmark designed to evaluate the reasoning capabilities of large language models (LLMs). In an initial assessment of 26 mainstream models, the benchmark revealed a significant performance gap in the industry. Gemini 3 Pro, currently regarded as one of the most advanced models, achieved a top accuracy rate of only 62.8%. More strikingly, the vast majority of the models tested failed to reach the 60% accuracy threshold, which is traditionally considered a passing grade. This release by Meituan's technical team establishes a more demanding standard for measuring AI reasoning, highlighting that current models still face substantial challenges in complex logical tasks.

美团技术团队

Key Takeaways

  • New Evaluation Standard: Meituan's LongCat team has introduced General 365, specifically designed to test the reasoning limits of AI models.
  • Industry-Wide Testing: The benchmark was applied to 26 mainstream models to provide a comprehensive overview of current AI capabilities.
  • Performance Ceiling: Gemini 3 Pro emerged as the top performer but only managed an accuracy rate of 62.8%.
  • Reasoning Deficit: Most tested models failed to achieve a 60% score, indicating a widespread struggle with the reasoning tasks presented in General 365.

In-Depth Analysis

The Introduction of General 365

The Meituan LongCat team has officially open-sourced General 365, positioning it as a new yardstick for the evaluation of artificial intelligence. Unlike traditional benchmarks that may focus on general knowledge or linguistic fluency, General 365 appears to target the core cognitive function of reasoning. By releasing this tool, the LongCat team provides the developer community with a rigorous framework to identify the strengths and weaknesses of various large language models (LLMs) in logical processing.

The decision to open-source this benchmark suggests a move toward greater transparency and standardization in how AI progress is measured. As models become more sophisticated, the industry requires more difficult and nuanced testing environments to differentiate between superficial pattern matching and genuine logical reasoning.

Benchmarking the Leaders: Gemini 3 Pro and Beyond

In the initial testing phase conducted by the LongCat team, 26 mainstream models were put to the test. The results offer a sobering look at the current state of AI development. Gemini 3 Pro, which is currently identified as the strongest model in the field, reached an accuracy of 62.8%. While this represents the leading edge of current technology, it also highlights a significant margin for improvement.

The data reveals a steep drop-off in performance beyond the top-tier models. The fact that the majority of the 26 models could not reach a 60% accuracy level—often considered the minimum standard for competency—suggests that General 365 is a highly challenging benchmark. This performance gap underscores the difficulty of the reasoning tasks included in the set and indicates that many current LLMs may still struggle when faced with complex, multi-step logical requirements.

Industry Impact

The release of General 365 is significant for the AI industry as it shifts the focus from simple performance metrics to deep reasoning capabilities. By setting a benchmark where even the most advanced models score near the 60% mark, Meituan is effectively raising the bar for what constitutes a "high-performing" model. This encourages AI researchers and developers to move beyond optimizing for existing, potentially saturated benchmarks and instead focus on the fundamental challenges of machine reasoning.

Furthermore, the benchmark serves as a reality check for the industry. While marketing for AI models often emphasizes human-like capabilities, the General 365 results demonstrate that there is still a long way to go before AI can consistently master complex reasoning tasks. This new standard will likely drive a new wave of innovation focused on cognitive depth rather than just model size or data volume.

Frequently Asked Questions

Question: What is General 365?

General 365 is a new reasoning evaluation benchmark released by Meituan's LongCat team. It is designed to provide a rigorous standard for testing the logical reasoning capabilities of large language models.

Question: How did mainstream models perform on this benchmark?

In a test of 26 mainstream models, the performance was generally low. Gemini 3 Pro led the group with a 62.8% accuracy rate, but the majority of models failed to reach a 60% score.

Question: Why is the 60% score significant in this context?

The 60% mark is often viewed as a basic passing grade or a threshold for competency. The fact that most models fell below this line indicates that General 365 is a particularly difficult test that exposes the reasoning limitations of current AI technology.

Related News

Stripe Agrees to Acquire AI Startup OpenRouter Following $1.3 Billion Valuation Milestone
Industry News

Stripe Agrees to Acquire AI Startup OpenRouter Following $1.3 Billion Valuation Milestone

Financial infrastructure giant Stripe has entered into an agreement to acquire OpenRouter, a prominent US-based artificial intelligence startup. This strategic acquisition follows a period of significant financial growth for OpenRouter, which recently concluded a US$113 million Series B funding round. The funding round had propelled the startup to a reported valuation of approximately US$1.3 billion prior to the acquisition announcement. The deal marks a major consolidation in the AI sector, as Stripe integrates a high-value AI platform into its existing ecosystem. The transition from a newly minted unicorn to a subsidiary of Stripe highlights the rapid pace of investment and acquisition within the current artificial intelligence landscape, emphasizing the strategic value placed on established AI infrastructure and talent.

OpenAI Reportedly Disbands Preparedness Team Responsible for Assessing and Mitigating Serious AI Model Risks
Industry News

OpenAI Reportedly Disbands Preparedness Team Responsible for Assessing and Mitigating Serious AI Model Risks

OpenAI has reportedly dissolved its internal preparedness team, a specialized group formerly tasked with identifying and mitigating catastrophic risks associated with advanced AI models. According to reports from the Financial Times and The Verge, the team’s primary mandate was to evaluate whether AI models could pose serious threats, such as the potential for a model to "go rogue" or engage in unauthorized hacking activities against other organizations. The responsibility for these critical safety assessments is reportedly being redistributed within the company following the team's disbandment at the end of last month. This organizational shift marks a significant change in OpenAI's approach to internal risk management and preparedness as it continues to develop increasingly powerful artificial intelligence technologies.

Stripe Reportedly Set to Acquire AI Gateway Startup OpenRouter in Landmark $7 Billion Strategic Deal
Industry News

Stripe Reportedly Set to Acquire AI Gateway Startup OpenRouter in Landmark $7 Billion Strategic Deal

Financial technology leader Stripe is reportedly in the process of acquiring OpenRouter, a prominent startup specializing in AI gateway infrastructure. The deal, valued at over $7 billion, marks a significant consolidation between the fintech and artificial intelligence sectors. OpenRouter has gained attention for its role as a unified interface for AI model access, a position emphasized by its CEO’s description of the company as the "Stripe for AI." This acquisition highlights Stripe's aggressive expansion into the AI ecosystem, aiming to provide the underlying infrastructure for AI integration. The reported $7 billion price tag underscores the immense value placed on middleware that simplifies the deployment and management of diverse artificial intelligence models for developers and enterprises globally.