Back to List
Meituan LongCat Team Launches WBench: The First Systematic Multi-Turn Benchmark for Interactive Video World Models
Research BreakthroughWorld ModelsAI BenchmarkingMeituan

Meituan LongCat Team Launches WBench: The First Systematic Multi-Turn Benchmark for Interactive Video World Models

The Meituan LongCat team has introduced and open-sourced WBench, a pioneering systematic multi-turn evaluation benchmark specifically designed for interactive video world models. As the AI industry shifts from generating static or passive video content toward creating dynamic, interactive environments, WBench serves as a critical diagnostic tool. Likened to a "CT scanner," the benchmark is engineered to precisely identify the technical bottlenecks that occur when world models attempt to transition from "passive viewing" to "active interaction." By providing a structured framework for multi-turn assessment, WBench offers a standardized method for researchers to evaluate how well AI models maintain consistency and responsiveness during complex, interactive sequences, marking a significant advancement in the field of world model research.

美团技术团队

Key Takeaways

  • Pioneering Benchmark: WBench is the first systematic multi-turn evaluation benchmark specifically targeting interactive video world models.
  • Open-Source Contribution: Developed and open-sourced by Meituan's LongCat team to foster industry-wide standardization.
  • Diagnostic Precision: The tool acts as a "CT scanner" for AI, pinpointing exactly where models fail during the transition from passive observation to active interaction.
  • Focus on Interaction: It addresses the critical gap in current AI capabilities, moving beyond simple video generation to complex, multi-turn interactive scenarios.

In-Depth Analysis

The Transition from Passive Viewing to Active Interaction

The emergence of world models represents a significant leap in artificial intelligence, moving beyond text and image generation into the realm of physical and spatial understanding. However, as the Meituan LongCat team identifies, a major hurdle remains: the transition from "passive viewing" to "active interaction." Most current models excel at generating a continuous stream of video based on a prompt, but they often struggle when a user or an external agent attempts to interact with that environment in real-time. WBench is designed to map these boundaries, exploring scenarios ranging from "moonwalks" to complex "cyber cities" to see how well the model maintains its internal logic when subjected to interactive changes.

WBench as a Diagnostic "CT Scanner"

One of the most compelling aspects of WBench is its role as a diagnostic tool. The LongCat team describes it as a "CT scanner" for world models. This metaphor suggests that WBench does not merely provide a pass/fail grade but offers a deep, structural look into the model's performance. By utilizing a systematic multi-turn evaluation process, WBench can isolate specific moments where a model's understanding of the "world" breaks down. Whether it is a loss of visual consistency, a failure in physics simulation, or a breakdown in multi-turn logic, WBench provides the granular data necessary for developers to understand the limitations of their current architectures.

Systematic Multi-Turn Evaluation Framework

Unlike traditional benchmarks that might evaluate a single output, WBench emphasizes "multi-turn" evaluation. In an interactive world, a single action is rarely isolated; it is part of a sequence of cause-and-effect relationships. A world model must be able to process an initial state, respond to an interaction, and then continue to evolve that world consistently in subsequent turns. WBench provides the first systematic framework to measure this specific capability. This is essential for applications in robotics, autonomous driving, and immersive virtual environments, where the AI must constantly update its understanding of the world based on continuous feedback and interaction.

Industry Impact

Standardizing World Model Evaluation

The release of WBench provides the AI research community with a much-needed standardized metric for world models. As more companies and research institutions develop their own versions of world models (similar to Sora or other video generation tools), having a common benchmark like WBench allows for objective comparison and faster iteration. By open-sourcing the tool, Meituan is positioning itself at the forefront of this foundational research, encouraging a collaborative approach to solving the "interaction" problem in AI.

Accelerating the Path to Embodied AI

Interactive world models are a cornerstone of Embodied AI—systems that can perceive, reason, and act in the physical or simulated world. By identifying the "boundaries" of current models, WBench helps researchers focus their efforts on the most critical weaknesses. This could accelerate the development of more sophisticated simulations for training robots or creating more realistic and responsive digital twins. The ability to move from "passive" to "active" is the key to making AI truly useful in real-world, dynamic environments.

Frequently Asked Questions

Question: What is the primary purpose of WBench?

WBench is a systematic multi-turn evaluation benchmark designed to test the interactive capabilities of video world models. It helps researchers identify where models struggle when moving from passive video generation to active, user-driven interaction.

Question: Who developed WBench and is it available to the public?

WBench was developed by the LongCat team at Meituan (Meituan Technology Team). It has been open-sourced to allow the broader AI research community to use it for evaluating and improving world models.

Question: Why is "multi-turn" evaluation important for world models?

Multi-turn evaluation is crucial because interactive environments require the AI to maintain consistency over a sequence of actions and reactions. Traditional single-turn benchmarks cannot adequately measure how a model handles the long-term consequences of interactions within a simulated world.

Related News

LongCat Releases VitaBench 2.0: A New Benchmark for Long-Term Dynamic AI Agent Evaluation
Research Breakthrough

LongCat Releases VitaBench 2.0: A New Benchmark for Long-Term Dynamic AI Agent Evaluation

LongCat has officially introduced VitaBench 2.0, a groundbreaking evaluation benchmark developed by the Meituan Technical Team. As the first benchmark specifically designed for long-term dynamic user modeling in real-life scenarios, VitaBench 2.0 represents a significant shift in how Large Language Models (LLMs) are assessed. The framework focuses on two critical dimensions: personalization and proactivity. By simulating long-term, real-world interactions, VitaBench 2.0 provides a systematic method for measuring an AI agent's ability to adapt to evolving user needs and take initiative within dynamic environments. This release marks a new milestone in the development of sophisticated, user-centric AI agents capable of maintaining consistency and relevance over extended periods of time.

Meituan LongCat Team Introduces WBench: A Systematic Multi-Round Benchmark for Evaluating Interactive Video World Models
Research Breakthrough

Meituan LongCat Team Introduces WBench: A Systematic Multi-Round Benchmark for Evaluating Interactive Video World Models

The Meituan LongCat team has officially released WBench, the industry's first systematic multi-round evaluation benchmark specifically designed for interactive video world models. Acting as a diagnostic "CT scanner," WBench is engineered to identify the specific limitations and failure points of AI models as they transition from passive video generation to active, interactive environments. By providing a structured framework for multi-round assessment, WBench allows researchers to pinpoint exactly where current world models struggle to maintain consistency and logic during user-driven interactions. This open-source tool represents a significant advancement in the methodology used to define and test the boundaries of world model capabilities, moving beyond simple observation to complex, interactive evaluation.

Meituan Fulfillment AI Team Showcases Frontier Agent Technology and Research Breakthroughs at ACL 2026
Research Breakthrough

Meituan Fulfillment AI Team Showcases Frontier Agent Technology and Research Breakthroughs at ACL 2026

The Meituan Fulfillment AI Algorithm Team has recently highlighted its latest research and technological advancements at the ACL 2026 conference. Focusing on building a Large Language Model (LLM)-based Agent technology system, the team aims to empower Meituan's fulfillment services through self-evolving operational systems. Their research spans critical areas such as Continuous Pre-training (CPT), Post-training, Agentic Reinforcement Learning (RL), and multimodal understanding. With dozens of papers published in prestigious venues like ACL and EMNLP, Meituan continues to push the boundaries of how AI agents can optimize complex business logistics and operational efficiency in real-world scenarios. This session specifically focuses on the team's contributions to the ACL conference and their practical applications in the frontier of AI technology.