Physics-Coherent Autonomous Driving Scene Forecasting through Joint Image-to-Video Generation and 3D World Alignment
Keywords:
autonomous driving, scene forecasting, image-to-video generation, 3D world alignment, physics coherence, generative models, neural rendering, safety assurance, socio-technical systemsAbstract
The reliable forecasting of dynamic traffic scenes remains a foundational challenge for the safe deployment of autonomous driving systems. Existing prediction paradigms frequently operate either in the abstract latent space of bird’s-eye-view representations or through purely photorealistic video generation without explicit grounding in three-dimensional physical constraints, leading to predictions that violate fundamental laws of motion, occlusion geometry, and object permanence. This paper presents a comprehensive system-level investigation into a novel architectural paradigm that jointly optimizes image-to-video generative models and explicit 3D world alignment modules to produce physics-coherent future scene forecasts. By integrating differentiable volumetric rendering, neural radiance fields, and diffusion-based video synthesis within a shared representation framework, the approach enforces consistency across appearance, geometry, and dynamics. We examine the structural trade-offs involved in balancing generative expressiveness with deterministic physical reasoning, analyze the governance and safety implications of deploying such forecasting models in high-stakes environments, and explore the infrastructure and sustainability challenges inherent in training and serving large-scale generative world models. The discussion extends to robustness under distributional shift, fairness across diverse road user categories, and the evolving regulatory landscape surrounding learned forecasting components. Through deep interdisciplinary analysis, this paper articulates the systemic requirements, risks, and opportunities for joint generation-alignment architectures as a pathway toward trustworthy and physically grounded autonomous scene intelligence.
References
1. Badue, C., Guidolini, R., Carneiro, R. V., Azevedo, P., Cardoso, V. B., Forechi, A., ... & De Souza, A. F. (2021). Self-driving cars: A survey. Expert Systems with Applications, 165, 113816.
2. Finn, C., Goodfellow, I., & Levine, S. (2016). Unsupervised learning for physical interaction through video prediction. In Advances in neural information processing systems (pp. 64-72).
3. Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., & Ng, R. (2020). NeRF: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision (pp. 405-421). Springer.
4. Hu, A., Natan, A., Salvador, J., Ramanan, D., & Gutfreund, D. (2023). GAIA-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080.
5. Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., ... & Salimans, T. (2022). Imagen Video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303.
6. Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., ... & Dai, J. (2022). BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In European conference on computer vision (pp. 1-18). Springer.
7. Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., ... & Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 11621-11631).
8. Raissi, M., Perdikaris, P., & Karniadakis, G. E. (2019). Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378, 686-707.
9. Salzmann, T., Ivanovic, B., Chakravarty, P., & Pavone, M. (2020). Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data. In European Conference on Computer Vision (pp. 683-700). Springer.
10. Vondrick, C., Pirsiavash, H., & Torralba, A. (2016). Generating videos with scene dynamics. In Advances in neural information processing systems (pp. 613-621).
11. Xiong, Z., Song, Y., He, L., Xiong, W., Yuan, Y., Qiao, F., & Jacobs, N. (2026). PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment. arXiv preprint arXiv:2603.13770.
12. Seo, H., Kim, S., Shin, J., & Kim, J. (2023). Dense depth-guided generalizable NeRF. IEEE Robotics and Automation Letters, 8(5), 2896-2903.
13. Scholkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., & Bengio, Y. (2021). Toward causal representation learning. Proceedings of the IEEE, 109(5), 612-634.
14. Wilson, B., Hoffman, J., & Morgenstern, J. (2019). Predictive inequity in object detection. arXiv preprint arXiv:1902.11097.
15. Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., & Abbeel, P. (2017). Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS) (pp. 23-30). IEEE.
16. Kendall, A., & Gal, Y. (2017). What uncertainties do we need in Bayesian deep learning for computer vision? In Advances in neural information processing systems (pp. 5574-5584).
17. Ryan, M. (2020). The future of transportation: ethical, legal, social and economic impacts of self-driving vehicles in the year 2025. Science and Engineering Ethics, 26, 1185-1208.
18. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th annual meeting of the Association for Computational Linguistics (pp. 3645-3650).
19. Li, T., Sahu, A. K., Talwalkar, A., & Smith, V. (2020). Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3), 50-60.
20. Prakash, A., Chitta, K., & Geiger, A. (2021). Multi-modal fusion transformer for end-to-end autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7077-7087).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Engineering Systems and Digital Innovation

This work is licensed under a Creative Commons Attribution 4.0 International License.