Back to List
LARYBench Launch: Defining the ImageNet for Embodied Action Representations and Measuring Generalization from Human Video Data
Research BreakthroughEmbodied IntelligenceComputer VisionRobotics

LARYBench Launch: Defining the ImageNet for Embodied Action Representations and Measuring Generalization from Human Video Data

The Meituan Technical Team has introduced LARYBench (Latent Action Representation Yielding Benchmark), a systematic evaluation framework designed to guide the learning of general latent action representations from large-scale visual data. This benchmark serves as a foundational tool, akin to ImageNet for computer vision, but specifically tailored for embodied intelligence. Experimental results from the benchmark reveal a significant discovery: general vision models demonstrate superior performance in action generalization and control precision compared to specialized action expert models designed specifically for embodied AI. This indicates that sophisticated embodied action representations can emerge naturally from training on extensive human video datasets, suggesting a new pathway for developing robotic control systems through general-purpose visual learning.

美团技术团队

Key Takeaways

  • Introduction of LARYBench: A systematic benchmark designed to evaluate and guide the development of general latent action representations from large-scale visual datasets.
  • Superiority of General Models: Experimental data shows that general vision models significantly outperform specialized embodied AI expert models in both generalization and control precision.
  • Emergence from Human Video: The benchmark proves that embodied action representations can emerge from large-scale human video data, rather than requiring exclusively robot-specific data.
  • New Standard for Embodied AI: LARYBench aims to define the "ImageNet moment" for embodied action, providing a standardized metric for measuring how well models understand and execute physical actions.

In-Depth Analysis

The Paradigm Shift: General Vision Models vs. Action Experts

The release of LARYBench (Latent Action Representation Yielding Benchmark) marks a critical turning point in the field of embodied intelligence. For years, the industry has focused on developing "action expert models"—specialized AI systems trained specifically on robotic trajectories and narrow physical tasks. However, the findings presented by the Meituan Technical Team challenge this specialized approach.

According to the benchmark results, general vision models—those trained on broad, diverse visual data—exhibit a higher degree of action generalization and control precision than their specialized counterparts. This suggests that the underlying features required for physical interaction are not necessarily unique to robotic data but are instead embedded within the broader context of visual understanding. By outperforming expert models, general vision systems demonstrate a more robust ability to adapt to new environments and tasks, which is a primary hurdle in the quest for universal embodied AI.

The Emergence of Action from Human Video Data

One of the most significant insights provided by LARYBench is the validation of human video data as a primary source for learning embodied actions. The benchmark demonstrates that latent action representations—the internal mappings an AI uses to translate visual input into physical movement—can "emerge" from large-scale human video datasets.

This finding is transformative because human video data is far more abundant and diverse than specialized robotic data. If embodied action can be learned by observing humans, the bottleneck of data collection for robotics could be significantly alleviated. LARYBench provides the first systematic measurement of this phenomenon, proving that the visual patterns of human movement contain sufficient information to inform the control precision and generalization capabilities of AI models in embodied contexts. This effectively bridges the gap between passive observation and active physical execution.

Industry Impact

The introduction of LARYBench is poised to redefine the development pipeline for robotics and embodied AI. By establishing a systematic evaluation standard, it allows researchers to measure progress in a way that was previously fragmented. The revelation that general vision models are more effective than specialized ones may lead to a shift in investment and research focus, moving away from narrow task-specific training toward the development of large-scale general visual learners for physical tasks.

Furthermore, the ability to leverage human video data for action representation means that the scaling laws observed in Large Language Models (LLMs) may soon be fully realized in robotics. As models are exposed to more diverse human activities through video, their ability to perform complex, precise, and generalized actions in the physical world is expected to improve, accelerating the deployment of autonomous systems in domestic and industrial environments.

Frequently Asked Questions

Question: What is the primary purpose of LARYBench?

LARYBench (Latent Action Representation Yielding Benchmark) is a systematic evaluation framework designed to measure how well AI models learn general latent action representations from large-scale visual data. It aims to provide a standardized metric for embodied intelligence, similar to what ImageNet provided for general computer vision.

Question: Why are general vision models performing better than specialized expert models?

Based on the experimental results from LARYBench, general vision models show superior action generalization and control precision. This suggests that the broad visual features learned from diverse datasets provide a more robust foundation for understanding physical actions than the narrow, task-specific data used to train specialized embodied AI expert models.

Question: Can robots really learn to move by watching human videos?

Yes, the LARYBench findings indicate that embodied action representations can emerge from large-scale human video data. This means that by analyzing how humans interact with the world in videos, AI models can develop the necessary latent representations to perform actions with high precision and generalization in robotic or embodied contexts.

Related News

LongCat Releases VitaBench 2.0: A New Benchmark for Long-Term Dynamic AI Agent Evaluation
Research Breakthrough

LongCat Releases VitaBench 2.0: A New Benchmark for Long-Term Dynamic AI Agent Evaluation

LongCat has officially introduced VitaBench 2.0, a groundbreaking evaluation benchmark developed by the Meituan Technical Team. As the first benchmark specifically designed for long-term dynamic user modeling in real-life scenarios, VitaBench 2.0 represents a significant shift in how Large Language Models (LLMs) are assessed. The framework focuses on two critical dimensions: personalization and proactivity. By simulating long-term, real-world interactions, VitaBench 2.0 provides a systematic method for measuring an AI agent's ability to adapt to evolving user needs and take initiative within dynamic environments. This release marks a new milestone in the development of sophisticated, user-centric AI agents capable of maintaining consistency and relevance over extended periods of time.

Meituan LongCat Team Introduces WBench: A Systematic Multi-Round Benchmark for Evaluating Interactive Video World Models
Research Breakthrough

Meituan LongCat Team Introduces WBench: A Systematic Multi-Round Benchmark for Evaluating Interactive Video World Models

The Meituan LongCat team has officially released WBench, the industry's first systematic multi-round evaluation benchmark specifically designed for interactive video world models. Acting as a diagnostic "CT scanner," WBench is engineered to identify the specific limitations and failure points of AI models as they transition from passive video generation to active, interactive environments. By providing a structured framework for multi-round assessment, WBench allows researchers to pinpoint exactly where current world models struggle to maintain consistency and logic during user-driven interactions. This open-source tool represents a significant advancement in the methodology used to define and test the boundaries of world model capabilities, moving beyond simple observation to complex, interactive evaluation.

Meituan Fulfillment AI Team Showcases Frontier Agent Technology and Research Breakthroughs at ACL 2026
Research Breakthrough

Meituan Fulfillment AI Team Showcases Frontier Agent Technology and Research Breakthroughs at ACL 2026

The Meituan Fulfillment AI Algorithm Team has recently highlighted its latest research and technological advancements at the ACL 2026 conference. Focusing on building a Large Language Model (LLM)-based Agent technology system, the team aims to empower Meituan's fulfillment services through self-evolving operational systems. Their research spans critical areas such as Continuous Pre-training (CPT), Post-training, Agentic Reinforcement Learning (RL), and multimodal understanding. With dozens of papers published in prestigious venues like ACL and EMNLP, Meituan continues to push the boundaries of how AI agents can optimize complex business logistics and operational efficiency in real-world scenarios. This session specifically focuses on the team's contributions to the ACL conference and their practical applications in the frontier of AI technology.