Odyssey 3
Odyssey 3 is an autoregressive diffusion transformer world model designed to simulate real-time physical dynamics and power autonomous systems.
Odyssey 3 is an autoregressive diffusion transformer world model designed to simulate real-time physical dynamics and power autonomous systems.
What the product does and how it is positioned
Odyssey 3 is a foundation world model implemented as an autoregressive diffusion transformer that simulates how objects, dynamics, and physical interactions evolve across time and space.
Trained on video observations, gameplay inputs, and simulated rigid-body physics, the model provides an interactive environment generator and visual representation foundation for robotics and physical AI policies.
Source-supported ways to use the product
The company presents this as a customer use case where Flexion built humanoid control policies on Odyssey 3 that maintained task execution under environmental and lighting changes.
Developers adapted Odyssey 3 to predict waypoints for closed-loop driving demonstrations using a frozen model backbone.
Developers trained manipulation policies on robot demonstration data to execute tasks and exhibit recovery behaviors such as regrasping.
Odyssey 3 operates as a multi-step video diffusion transformer that incorporates temporally resolved prompts and action inputs to generate future visual states.
Through causal masking and teacher forcing, the system functions autoregressively by conditioning future predictions on prior visual history, while an adversarial and distribution-matching distillation pipeline reduces step count to enable real-time response.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
Odyssey 3 is an autoregressive diffusion transformer world model that predicts physical dynamics, object motion, and environmental changes over time.
The documentation reports that Odyssey 3 generates at 832x480 resolution, whereas Odyssey 3 Pro operates at 1280x720 resolution.
Developers adapt the foundation model to specific robots or vehicles by training an action decoder or policy on paired observation and action data.
The preview supports first-person navigation, third-person navigation, and independent camera positioning during environment generation.
The model was trained on annotated internet video, gameplay recordings aligned with keyboard and mouse inputs, and captioned rigid-body simulations.