top of page

BYD AI Introduces Hybrid World Model, Exceeds Autonomous Driving Benchmark

  • Writer: tech360.tv
    tech360.tv
  • 22 hours ago
  • 3 min read

BYD's AI division has presented its inaugural research paper, HyWorldVLA, introducing a hybrid world model that achieved a 90.59 PDMS score on the NAVSIM v1 benchmark. This result reportedly surpasses all prior leading methods, indicating the organisation's entry into autonomous driving foundation models. The paper reveals a previously unpublicised AI research group, according to Pandaily, which includes individuals with backgrounds in robotics research from the Harbin Institute of Technology.


Close-up of a sleek car taillight with BYD TECH text, glossy black and silver details, outdoors in soft sunlight.
Credits: UNSPLASH

The HyWorldVLA model combines pixel level and latent space world modelling with a Vision Language Action architecture, specifically designed for autonomous driving applications. This system recorded 90.59 PDMS on the NAVSIM v1 public benchmark. This benchmark assesses real driving quality across several metrics including collision safety, drivable area compliance, maintenance of safety distance, goal completion, and motion comfort. This achievement represents a notable shift for BYD, an organisation previously perceived as concentrating on supply chain integration and manufacturing scale, rather than extensive code level AI research.


The architecture of this hybrid world model addresses a fundamental problem inherent in autonomous driving world models. Pixel level models project future video frames to capture environmental detail, but they often face challenges such as high computational costs and susceptibility to visual interference. Conversely, latent space models operate within compressed, hidden representations to gain efficiency, yet they risk losing real world detail crucial for making safe driving decisions. But HyWorldVLA incorporates both approaches.


It uses pixel level supervision during its training phase, functioning as a safeguard to ensure the latent space preserves sufficient information for accurate frame reconstruction. The system then conducts inference in the compressed space for operational efficiency. An internal ablation study confirmed the necessity of both components. Removing pixel level prediction caused the PDMS score to fall from 90.59 to 87.50. Removing latent space processing resulted in a score drop to 89.91, validating the contribution of each element.


The training programme for HyWorldVLA operates in three distinct phases. First, a video Variational Autoencoder compresses continuous video frames into compact latent features. This process is guided by text descriptions providing semantic context. Second, features relating to image, action, language, and latent representations are converted into unified, discrete tokens. These tokens are used for autoregressive next token prediction, benefiting from dual supervision across both pixel and latent prediction. The pixel token prediction specifically acts as a quality assurance measure, preventing any collapse of the latent representation.


Third, an action generation module integrates the latent features with historical data to produce trajectory outputs. This output undergoes verification through two methods: scoring candidate trajectories, and direct continuous trajectory generation. The weight applied to latent feature supervision proved to be a critical factor in performance. Fine tuning without any constraint yielded a score of 90.17. An optimal weight set at 0.1 resulted in the peak score of 90.59. Conversely, an excessive weight caused a decline in performance, with the score dropping to 89.75.


So, this data suggests that a moderate level of constraint helps maintain previously trained semantic information without unduly restricting the model's ability to discern representations relevant to driving. The paper also indicates BYD's team is developing long term autonomous driving foundation models, moving beyond merely production specific systems. This publication positions BYD as a serious contributor to autonomous driving AI research, potentially drawing in leading tier AI talent and signalling a strategic intention to develop core autonomous driving technology internally, rather than relying solely on external suppliers. This places BYD in competition with organisations such as NIO, XPeng, and Li Auto in the pursuit of autonomous driving AI leadership. However, BYD benefits from possessing the largest production base, which allows for the deployment of developed technologies at a substantial scale.


  • BYD's AI team has achieved a 90.59 PDMS score on the NAVSIM v1 benchmark with its HyWorldVLA model.

  • The HyWorldVLA model integrates both pixel level and latent space world modelling within a Vision Language Action architecture.

  • The development signifies BYD's strategic entry into core autonomous driving AI research.

  • The organisation plans to develop long term autonomous driving foundation models internally.


Source: Pandaily

Technology increasingly permeates every facet of our lives, making informed decision making an essential pursuit. We bridge this gap by combining the precision of AI with the irreplaceable discernment of human expertise. Our team produces rigorous product reviews that offer unique insights, honest critiques, and trustworthy recommendations. We also leverage AI to synthesise complex news from reliable sources into clear, actionable updates, ensuring that every story is carefully fact checked by our editorial staff before publication. Accuracy remains our priority. Should you identify any discrepancies, please contact us at editorial@tech360.tv. Your feedback is a vital part of our process in maintaining the high standards our readers deserve.

Tech360tv is Singapore's Tech News and Gadget Reviews platform. Join us for our in depth PC reviews, Smartphone reviews, Audio reviews, Camera reviews and other gadget reviews.

  • YouTube
  • Facebook
  • TikTok
  • Instagram
  • Twitter
  • LinkedIn

© 2021 tech360.tv. All rights reserved.

bottom of page