Back to List
Meituan LongCat Releases General 365: A New Reasoning Benchmark Where Most AI Models Fail to Pass
Industry NewsArtificial IntelligenceBenchmarkingMeituan

Meituan LongCat Releases General 365: A New Reasoning Benchmark Where Most AI Models Fail to Pass

The Meituan LongCat team has officially open-sourced 'General 365,' a rigorous new benchmark designed to evaluate the reasoning capabilities of large language models. In an initial assessment of 26 mainstream AI models, the results highlight a significant gap in current cognitive performance. Even Gemini 3 Pro, identified as the top performer in the test, achieved an accuracy rate of only 62.8%. Furthermore, the vast majority of the models tested were unable to reach the 60% passing threshold. This release by Meituan's technology team provides a new standard for the industry, revealing that complex reasoning remains a substantial challenge for even the most advanced artificial intelligence systems currently available.

美团技术团队

Key Takeaways

  • Meituan's LongCat team has officially released and open-sourced the General 365 reasoning benchmark.
  • Evaluation of 26 mainstream models reveals that most current AI systems struggle with complex reasoning tasks.
  • Gemini 3 Pro emerged as the top performer with an accuracy of 62.8%, yet this remains relatively low for a leading model.
  • The majority of tested models failed to reach a 60% accuracy score, establishing a high difficulty ceiling for the benchmark.

In-Depth Analysis

The Launch of General 365

The Meituan LongCat team has introduced General 365 as a specialized tool for evaluating the reasoning depth of artificial intelligence. By open-sourcing this benchmark, the team provides the global AI community with a new metric to measure progress beyond simple linguistic fluency. The focus of General 365 is specifically on 'General' reasoning, suggesting a broad application across various logical domains. The release comes at a time when the industry is shifting its focus from model size to the quality of logical output and problem-solving efficiency.

Performance Gap in Mainstream Models

The results released alongside the benchmark provide a sobering look at the current state of AI. Out of 26 mainstream models tested, the performance was notably lower than what is typically seen on standard benchmarks. The fact that Gemini 3 Pro, currently regarded as one of the most capable models globally, only secured a 62.8% accuracy rate indicates that General 365 contains highly challenging reasoning problems. This data point serves as a critical indicator that even the 'strongest' models have significant room for improvement when faced with the specific criteria set by the LongCat team.

The 60% Passing Threshold

A striking finding from the LongCat team's report is that the vast majority of the 26 models failed to reach the 60% mark. In many academic and professional contexts, 60% is considered the minimum threshold for a 'passing' grade. The failure of most mainstream models to meet this baseline suggests that current LLM (Large Language Model) architectures may still lack the robust logical frameworks required for consistent reasoning. This gap between current performance and the 'passing' line highlights the rigorous nature of General 365 as a evaluative standard.

Industry Impact

The introduction of General 365 by Meituan is significant for the AI industry as it establishes a more demanding yardstick for reasoning. By making the benchmark open-source, Meituan allows other developers to stress-test their models against the same 26-model baseline. This could lead to a shift in development priorities, moving away from general knowledge retrieval and toward the enhancement of internal logic and multi-step reasoning. As models strive to exceed the 62.8% mark set by Gemini 3 Pro, General 365 will likely become a key reference point for future iterations of large language models.

Frequently Asked Questions

Question: What is the significance of the 62.8% score achieved by Gemini 3 Pro?

Within the context of the General 365 benchmark, 62.8% represents the highest accuracy among 26 mainstream models. While it leads the field, the score suggests that even top-tier AI models face difficulty with the reasoning tasks included in this specific evaluation.

Question: Why did most models fail to reach the 60% mark on General 365?

The failure of the majority of models to reach 60% indicates that the General 365 benchmark is designed with a high level of difficulty that targets the weaknesses in current AI reasoning capabilities, setting a new and more difficult standard for the industry.

Question: Is General 365 available for public use?

Yes, the Meituan LongCat team has officially open-sourced General 365, allowing the broader technology community to use it for evaluating and improving AI model reasoning.

Related News

Industry News

The Tragedy of the Commons in the AI Era: A Deep Dive into Resource Depletion

This analysis explores the application of the 'Tragedy of the Commons' economic theory to the current landscape of artificial intelligence development. As AI companies compete for a finite pool of high-quality, human-generated data, the shared digital ecosystem faces significant risks of depletion and degradation. The article examines how the rapid consumption of public data for model training creates a paradox where the very resources that enable AI progress are being exhausted or 'polluted' by synthetic content. By viewing the internet as a digital commons, we can better understand the emerging challenges of data scarcity, the threat of model collapse, and the potential shift toward a more enclosed and proprietary data economy. This conceptual framework highlights the urgent need for sustainable resource management within the AI industry.

Amazon's Planned Texas Data Center Power Plant Could Become the Largest Climate Polluter in the United States
Industry News

Amazon's Planned Texas Data Center Power Plant Could Become the Largest Climate Polluter in the United States

Amazon is currently investing in a major data center project in Texas that includes the construction of an on-site power plant. According to reports, this facility has the potential to become the single largest source of climate pollution in the United States. The project highlights a significant shift in how tech giants manage their energy needs, moving toward dedicated on-site generation to support massive data infrastructure. However, the scale of the projected emissions from this specific Texas site has raised alarms regarding its environmental footprint. This development places Amazon's infrastructure expansion at the center of national climate discussions, as the facility's impact could surpass all other individual pollution sources in the country.

OpenAI Strategically Acquires Presentation Startup NextSlide to Enhance ChatGPT's Productivity and Visual Capabilities
Industry News

OpenAI Strategically Acquires Presentation Startup NextSlide to Enhance ChatGPT's Productivity and Visual Capabilities

OpenAI has officially acquired NextSlide, a startup specializing in presentation technology, marking a significant expansion of its development team. Following the acquisition, the NextSlide team has transitioned to working directly on ChatGPT. This move highlights OpenAI's commitment to integrating specialized expertise in structured content and visual storytelling into its flagship AI model. While specific financial details of the deal have not been disclosed, the integration of the NextSlide team suggests a strategic focus on evolving ChatGPT from a conversational interface into a more robust productivity tool capable of handling complex presentation-related tasks. This acquisition underscores the ongoing trend of major AI companies absorbing niche startups to bolster their internal capabilities and accelerate the development of multi-modal features within the competitive artificial intelligence landscape.