OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling

[Paper] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling

OmniWorld, a groundbreaking dataset for 4D world modeling, addresses critical data limitations in visual intelligence systems. Developed by researchers from Shanghai AI Laboratory and Zhejiang University, it combines the OmniWorld-Game dataset with public datasets from various domains. This comprehensive resource offers unprecedented scale, diversity, and modal richness, surpassing existing synthetic datasets in both size and modality coverage. OmniWorld establishes a new benchmark for 3D geometric foundation models and camera-controlled video generation, revealing limitations in current approaches while providing a pathway for improvement through fine-tuning. The dataset's rich multi-modal annotations and diverse scenarios enable machines to better understand, simulate, and interact with the physical world. Fine-tuning experiments demonstrate significant performance enhancements across multiple datasets and metrics, accelerating progress in visual intelligence systems.

The rapid advancement of world modeling technologies represents a significant milestone in the evolution of visual intelligence systems. Recently, researchers from Shanghai AI Laboratory and Zhejiang University have introduced OmniWorld, a groundbreaking large-scale, multi-domain, multi-modal dataset specifically designed to address the critical data bottleneck in 4D world modeling. This comprehensive dataset enables machines to better understand, simulate, and interact with the physical world by combining both spatial geometry and temporal dynamics.

OmniWorld goes beyond conventional datasets by offering unprecedented scale, diversity, and modal richness. At its core is the newly collected OmniWorld-Game dataset, which surpasses existing synthetic datasets in both scale and modality coverage. When combined with several public datasets from various domains, OmniWorld establishes a powerful benchmark that reveals the limitations of current state-of-the-art approaches while providing a pathway for significant performance improvements through fine-tuning.

The development of world models has become central to advancing visual intelligence systems that can reason about the physical world. These systems need to go beyond static perception to simulate dynamic environments, predict object motion, infer causality, and generate content adhering to physical laws. Two fundamental tasks that reflect a model’s world modeling capability have garnered significant attention: 3D geometric foundation models and camera-controlled video generation models. The former extracts comprehensive 3D geometric information from 2D image inputs, while the latter focuses on generating dynamic video content following precise spatio-temporal instructions.

However, existing benchmarks and datasets face significant limitations. For 3D geometric foundation models, current benchmarks often feature short sequence lengths that constrain evaluation of a model’s long-term robustness. For example, the widely used Sintel dataset consists of videos averaging only 50 frames. Additionally, the limited motion amplitude and single-action types within these datasets fail to comprehensively evaluate model performance in complex, dynamic environments. Similarly, in camera-controlled video generation, mainstream datasets like RealEstate10K primarily consist of static scenes with smooth camera trajectories, creating a noticeable gap between dataset content and real-world scenarios.

To address these shortcomings, the researchers created OmniWorld, which comprises two main components. First is OmniWorld-Game, a massive synthetic video dataset featuring over 96,000 clips and more than 18 million frames, with a total duration exceeding 214 hours. This dataset is captured from diverse game environments with 720P RGB images, dense ground truth depth maps, accurate camera poses, text captions, optical flow, and foreground masks. The second component involves the integration of datasets from four key domains—simulator, robot, human, and internet—providing a wide range of real-world and virtual scenarios that enhance data diversity.

What makes OmniWorld particularly valuable is its rich suite of multi-modal annotations. The dataset provides comprehensive annotated information across multiple modalities, including RGB images, depth maps, camera poses, text captions, optical flow, and foreground masks. These annotations are crucial for detailed world modeling tasks. To ensure high quality, the researchers implemented sophisticated data acquisition and annotation pipelines, employing both automated and semi-automated approaches tailored to different data domains.

The dataset’s statistics are impressive. OmniWorld collectively contains over 600,000 video sequences and more than 300 million frames, with a significant portion having a resolution of 720P or higher. The human domain constitutes the largest share, underscoring the dataset’s richness in reflecting real-world human activities and interactions. Furthermore, OmniWorld-Game exhibits exceptional diversity across multiple dimensions, including scene types (outdoor-urban, outdoor-natural, indoor, and mixed), camera perspectives (first-person and third-person), historical eras (ancient, modern, and futuristic), and dominant object types (terrain, architecture, vehicles, and mixed).

Based on OmniWorld-Game, the researchers established a new benchmark for both 3D geometric foundation models and camera-controlled video generation models. This benchmark provides challenging, complex scenarios that accurately reflect a model’s true world capabilities. When evaluating current state-of-the-art models on this benchmark, the researchers found significant room for improvement, particularly in handling complex dynamics and camera control.

The most compelling evidence of OmniWorld’s value comes from fine-tuning experiments. The researchers selected several state-of-the-art models for two core tasks—3D geometric foundation models and camera-controlled video generation models—and fine-tuned them using OmniWorld. The results consistently showed significant performance improvements over the original published versions across multiple datasets and metrics. For example, models fine-tuned with OmniWorld demonstrated enhanced depth estimation accuracy, improved camera pose estimation, and better adherence to camera-controlled instructions in video generation.

The introduction of OmniWorld represents a significant contribution to the field of 4D world modeling. By providing a large-scale, multi-domain, and multi-modal dataset with rich annotations, it addresses a critical data bottleneck that has limited progress in this area. Moreover, by establishing a challenging benchmark and demonstrating the efficacy of fine-tuning with OmniWorld, the researchers have provided a clear pathway for advancing the next generation of models with stronger spatio-temporal consistency.

As world modeling continues to evolve as a central pursuit in visual intelligence systems, datasets like OmniWorld will play an increasingly crucial role. They not only provide the necessary training data for developing more sophisticated models but also establish benchmarks that guide research directions and measure progress. With its unprecedented scale, diversity, and modal richness, OmniWorld represents a significant step forward in enabling machines to better understand, simulate, and interact with the physical world.

The research community now has access to a powerful resource that can accelerate the development of more general and robust models for understanding and interacting with the real physical world. The OmniWorld dataset and associated resources are available through the project’s official GitHub repository, allowing researchers to build upon this foundation for future innovations in visual intelligence systems. As artificial intelligence continues to advance toward more sophisticated spatial-temporal reasoning capabilities, resources like OmniWorld will be instrumental in bridging the gap between virtual simulations and physical reality, ultimately leading to more capable and adaptable AI systems.

🚀 Unlock Ads-Free Experience At $5/year

14 days free trial Cancel anytime

Leave a Comment

Your email address will not be published. Required fields are marked *