Humanoid Robot Model Omega-0 Achieves High Multitasking Success
- tech360.tv

- 4 minutes ago
- 3 min read
Researchers from Nanyang Technological University, Peking University, HKUST (Guangzhou), and Beijing Academy of Artificial Intelligence (BAAI) have released Omega-0. This latent prediction world action model allows a humanoid robot to simultaneously walk, observe its surroundings, and perform tasks. It achieved an 81.8 percent success rate on 11 real world domestic tasks, surpassing previous models such as pi 0.5, EgoVLA, GR00T N1.7, and psi 0.

The introduction of Omega-0 provides a new approach for training humanoid robots. The model's success rate on specific household duties significantly exceeds that of its predecessors. Its architecture, not just performance, is notable. The model forecasts future visual features rather than future pixels, then employs a diffusion transformer to generate full body actions.
Omega-0 operates via a three stage training architecture. Stage one involves whole body action vision language model pretraining. The team developed a Whole Body FAST tokenizer, converting continuous, high dimensional full body movement trajectories into discrete action tokens. And Qwen3 VL 2B Instruct was fine tuned into a specialised whole body action VLM, incorporating learnable viewpoint tokens for first person and third person perspectives.
Stage two focuses on human robot action latent pretraining. The team drew upon the V JEPA concept, integrating vision, language, and robot proprioception. This process involves predicting future visual features through video queries, then injecting these predicted environmental features directly into the action representation. A diffusion transformer generates the complete full body action latent.
The final stage, stage three, involves real home fine tuning. This stage incorporates a real time chunking strategy, which helps ensure action continuity for the robot's movements. So, to facilitate the development of this architecture, the team established the Omega HOME dataset. This comprehensive dataset comprises 40.3 hours of real environment data, encompassing 4,827 distinct trajectories, and covers 24 specific types of tasks.
Each trajectory within the Omega HOME dataset includes several synchronised data points. These consist of language instructions, first person and third person video footage, proprioception data, and full body motion information. The dataset addresses eight distinct categories of common household tasks, including the retrieval of objects, the cleaning of surfaces, the operation of appliances, and general tidying activities. Trajectories were generated through teleoperation and converted into supervision signals for end to end learning systems.
The methodology employed for benchmarking represented a more rigorous aspect of the research paper, according to Pandaily. To prevent any contamination of the evaluation by pretraining data, the researchers removed the 11 fine tuning tasks from the Omega HOME dataset. The remaining data was then combined with publicly available human demonstration data for the second stage of pretraining. But this measure ensures the reported gains are genuine, not a result of test set leakage. The model's name, Omega 0, suggests it is the initial entry in what may become a sequence of developments.
The broader implications extend to embodied artificial intelligence development. Domestic environments pose the most significant deployment challenge for humanoid robots. This is due to the continuous environmental sensing and full body coordination required for such settings. Omega 0 connects future environment prediction with the generation of whole body actions. This represents one of the more structured attempts to manage this combined challenge within a singular model. The prospect of a robot reliably completing tasks like laundry remains distant. However, the architectural approach appears clearer than it was six months prior.
Outstanding challenges continue to exist within this field. These include long horizon autonomy, ensuring safe interaction with surroundings, generalising tasks across varied scenarios, and adapting to complex environments. The 81.8 percent success rate, while useful for measurement, does not signify the completion of research. It simply marks progress within an ongoing programme of work.
Omega 0 achieved 81.8 percent success on 11 real world domestic tasks.
The model's architecture predicts future visual features and uses a diffusion transformer for full body actions.
Training involved three stages and utilised the Omega HOME dataset, comprising 40.3 hours of real environment data across 24 task types.
The benchmark methodology specifically avoided pretraining data contamination for evaluation integrity.
Omega 0 represents a structured attempt to combine future environment prediction with whole body action generation for humanoid robots.
Source: Pandaily


