Back to List
LARYBench Released: A New Benchmark Defining the ImageNet for Embodied Action Representation and Generalization
Research BreakthroughEmbodied AIComputer VisionRobotics

LARYBench Released: A New Benchmark Defining the ImageNet for Embodied Action Representation and Generalization

The Meituan Technical Team has officially introduced LARYBench (Latent Action Representation Yielding Benchmark), a systematic evaluation framework designed to guide the learning of general latent action representations from large-scale visual data. Positioned as the 'ImageNet' for the embodied AI field, LARYBench provides a standardized way to measure how well models can understand and execute actions. The benchmark's initial experimental results reveal a significant shift in AI development: general-purpose vision models consistently outperform specialized embodied AI expert models in both action generalization and control precision. Furthermore, the research confirms that sophisticated embodied action representations can naturally emerge from training on extensive human video datasets, offering a scalable path for future robotic intelligence and autonomous systems.

美团技术团队

Key Takeaways

  • Introduction of LARYBench: A systematic benchmark designed to evaluate and guide the development of general latent action representations from visual data.
  • Superiority of General Models: Experimental data indicates that general vision models outperform specialized embodied AI expert models in generalization and precision.
  • Emergent Intelligence from Human Videos: The study proves that embodied action representations can emerge from large-scale human video data without specialized robotic training.
  • New Industry Standard: LARYBench is being recognized as the 'ImageNet' for embodied action, providing a critical metric for the industry.

In-Depth Analysis

Establishing a Systematic Standard for Embodied AI

The release of LARYBench (Latent Action Representation Yielding Benchmark) marks a significant milestone in the evolution of embodied AI. Much like how ImageNet revolutionized computer vision by providing a massive, standardized dataset for object recognition, LARYBench aims to do the same for action representation. By focusing on "latent action representations," the benchmark moves beyond simple command-following and looks at the underlying structures of how an AI perceives and prepares to execute physical movements. This systematic approach allows researchers to evaluate how effectively a model can translate visual information into actionable intelligence, providing a clear roadmap for developing more versatile and capable autonomous agents.

General Vision Models vs. Specialized Action Experts

One of the most striking findings presented by the Meituan Technical Team is the performance gap between general vision models and specialized embodied action expert models. Traditionally, the industry has leaned toward creating "expert" models—AI systems specifically trained on robotic data to perform specific tasks. However, LARYBench's experimental results show that general vision models, which are trained on a much broader array of visual data, exhibit significantly better action generalization and control precision. This suggests that the breadth of information contained in general vision models provides a more robust foundation for physical interaction than the narrow, task-specific training of expert models. This finding could lead to a paradigm shift in how robotic controllers are designed, favoring large-scale general pre-training over niche specialization.

The Power of Large-Scale Human Video Data

The research highlights a critical breakthrough in data sourcing for embodied AI: the emergence of action representations from human video data. Previously, it was often assumed that to teach a robot how to move, one needed data specifically from robots (teleoperation or simulation). LARYBench demonstrates that by analyzing large-scale human videos, AI models can learn the nuances of movement, spatial relationships, and physical interaction. This "emergence" of embodied intelligence from non-robotic data sources is a game-changer for the industry. It suggests that the vast libraries of human video content available today can serve as a primary training ground for the next generation of embodied AI, drastically reducing the reliance on expensive and hard-to-collect robotic execution data.

Industry Impact

The introduction of LARYBench is expected to have a profound impact on the AI and robotics industries. By providing a standardized metric for action representation, it allows for more transparent comparisons between different AI architectures. The discovery that general vision models are superior for action generalization suggests that the future of robotics lies in the integration of Large Vision Models (LVMs) rather than isolated robotic controllers. Furthermore, the ability to leverage human video data for training opens the door for rapid scaling in embodied AI, potentially accelerating the deployment of autonomous systems in complex, real-world environments such as logistics, manufacturing, and domestic assistance.

Frequently Asked Questions

Question: What is the primary purpose of LARYBench?

LARYBench is a systematic evaluation benchmark designed to measure and guide the learning of general latent action representations from large-scale visual data, serving as a foundational tool for embodied AI development.

Question: Why are general vision models performing better than specialized models in this benchmark?

According to the research, general vision models demonstrate superior action generalization and control precision because they benefit from a broader understanding of visual contexts, which proves more effective for complex embodied tasks than the narrow training of specialized expert models.

Question: Can AI learn to control robots just by watching human videos?

Yes, the findings from LARYBench show that embodied action representations can emerge from large-scale human video data, suggesting that models can learn the fundamental principles of action and movement by observing human behavior at scale.

Related News

Meituan LongCat Team Open-Sources WBench: The First Systematic Multi-Round Benchmark for Interactive Video World Models
Research Breakthrough

Meituan LongCat Team Open-Sources WBench: The First Systematic Multi-Round Benchmark for Interactive Video World Models

The Meituan LongCat team has officially released and open-sourced WBench, a pioneering systematic benchmark designed for the evaluation of interactive video world models. WBench represents a significant shift in AI assessment, moving beyond traditional single-instance testing to a multi-round evaluation framework. Described by the developers as a "CT scanner" for AI, the tool is designed to pinpoint the exact limitations and boundaries of current world models as they attempt to transition from passive video generation to active, user-driven interaction. By testing scenarios ranging from lunar walks to complex cybernetic urban environments, WBench provides a rigorous diagnostic environment to identify where models fail in maintaining consistency and logic during interactive sequences.

Meituan Fulfillment AI Team Showcases LLM-Based Agent Technology and Research Breakthroughs at ACL 2026
Research Breakthrough

Meituan Fulfillment AI Team Showcases LLM-Based Agent Technology and Research Breakthroughs at ACL 2026

Meituan's Fulfillment AI Algorithm Team has highlighted its latest research and technological advancements at the ACL 2026 conference. The team is dedicated to developing a sophisticated Agent technology system powered by Large Language Models (LLMs) to enhance Meituan's fulfillment operations. Their core research focuses on several frontier areas, including Continual Pre-Training (CPT), Post-training, Agentic Reinforcement Learning (RL), and multimodal understanding. By building self-evolving Agent operating systems, the team aims to integrate AI deeply into business processes. Having published numerous papers in top-tier international conferences like ACL and EMNLP, Meituan continues to demonstrate its leadership in applying cutting-edge AI to real-world logistics and fulfillment challenges through this featured technical session.

Meituan Technical Team Announces Six Research Papers Accepted at ACL 2026 for AI Innovation
Research Breakthrough

Meituan Technical Team Announces Six Research Papers Accepted at ACL 2026 for AI Innovation

The Meituan technical team has reached a significant milestone in artificial intelligence research, with six of its papers being accepted for the ACL 2026 conference. ACL, a premier international event for computational linguistics and natural language processing (NLP), will feature Meituan's latest findings across several high-impact domains. The research spans large language model (LLM) evaluation, complex process reasoning, and the optimization of competition-level mathematical thinking. Additionally, the papers delve into reinforcement learning and generative recommendation systems. This collection of research highlights Meituan's strategic focus on building a new paradigm for generative AI, emphasizing both the theoretical evaluation of model capabilities and the practical optimization of reasoning and performance in real-world applications.