Multimodal Large Language Models for Autonomous Driving: A Comprehensive Survey of Perception, Reasoning, Planning, and Safety Assurance
DOI:
https://doi.org/10.70882/josrar.2026.v3i4.245Keywords:
Autonomous driving, Multimodal large language models, Vision-language models, Driving benchmarks, ISO 26262, ISO 21448 (SOTIF), UL 4600 safety case, Sim-to-real transferAbstract
Autonomous driving has progressed from rule-based subsystems and modular perception, prediction, and planning stacks toward unified data-driven architectures, and multimodal large language models (MLLMs) are increasingly proposed as the cognitive substrate of the next generation of highly automated road vehicles. This 2026 survey synthesises 39 primary sources selected from an initial corpus of 274 candidate records screened over 2020-2026, organises the field around a five-role pipeline taxonomy (perception, prediction, planning, control, and human-machine interaction), and compares six representative driving MLLMs (DriveGPT-4, LMDrive, Senna, DriveLM, GPT-4V-AD, and Cosmos-1) on accuracy, latency, and parameter footprint. A benchmark coverage matrix over LingoQA, BDD-X, DriveLM, nuScenes-QA, AutoHallu, and CODA-LM exposes evaluation gaps in prediction and planning. Model behaviour is translated into safety-assurance terms by mapping four MLLM failure-mode families to the functional-safety standard ISO 26262, the Safety of the Intended Functionality standard ISO 21448 (SOTIF), and the autonomous-systems safety-case standard UL 4600. A three-tier vehicle, edge, and cloud deployment topology is described together with the digital-twin and over-the-air update infrastructure that surrounds it. The strongest empirical finding is that Cosmos-1 delivers the best accuracy among models with sub-150 ms latency (76.6 percent mean reasoning accuracy at 480 ms), leaving verifiable safety certification as the single most important open problem for closed-loop deployment. The survey closes with a six-item research agenda spanning sub-100 ms real-time inference, out-of-distribution generalisation, multi-agent intent reasoning, verifiable safety certification, long-tail corner-case coverage, and closed-loop sim-to-real transfer. The article is intended as a reference for automotive system architects, safety engineers, regulators, and machine-learning researchers preparing the next generation of automated driving systems.
References
Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L. D., Monfort, M., Muller, U., Zhang, J., Zhang, X., Zhao, J., & Zieba, K. (2016). End to end learning for self-driving cars. arXiv Preprint arXiv:1604.07316. https://arxiv.org/abs/1604.07316
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., & Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11621–11631). IEEE.
Casas, S., Sadat, A., & Urtasun, R. (2021). MP3: A unified model to map, perceive, predict, and plan. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 14403–14412). IEEE.
Chen, H., Liu, X., & Wang, T. (2025). Trustworthiness of autonomous-driving foundation models. (Preprint; details available on request).
Chen, L., Sinavski, O., Hunermann, J., Karnsund, A., Willmott, A. J., Birch, D., Maund, D., & Shotton, J. (2023). Driving with LLMs: Fusing object-level vector modality for explainable autonomous driving. arXiv Preprint arXiv:2310.01957. https://arxiv.org/abs/2310.01957
Choi, J., Park, J., & Lee, S. (2026). Diffusion Models for End-to-End Autonomous Driving: A Survey of Perception, Prediction, Planning, and Control. IEEE Access.
Cui, C., Ma, Y., Cao, X., Ye, W., Zhou, Y., Liang, K., Chen, J., Lu, J., Yang, Z., Liao, K.-D., Gao, T., Li, E., Tang, K., Cao, Z., Zhou, T., Liu, A., Yan, X., Mei, S., Cao, J., Wang, Z., & Zheng, C. (2024). A survey on multimodal large language models for autonomous driving. In Proceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW) (pp. 958–979). IEEE. https://doi.org/10.1109/WACVW60836.2024.00106
European Commission. (2024). Regulation (EU) 2024/1689 — The AI Act. Official Journal of the European Union.
Fu, D., Li, X., Wen, L., Dou, M., Cai, P., Shi, B., & Qiao, Y. (2023). Drive like a human: Rethinking autonomous driving with large language models. arXiv Preprint arXiv:2307.07162. https://arxiv.org/abs/2307.07162
Grigorescu, S., Trasnea, B., Cocias, T., & Macesanu, G. (2020). A survey of deep learning techniques for autonomous driving. Journal of Field Robotics, 37(3), 362–386. https://doi.org/10.1002/rob.21918
Hawke, A., Shen, R., Gurau, C., Sharma, S., Reda, D., Nikolov, N., Mazur, P., Micklethwaite, S., Griffiths, N., Shah, A., & Kendall, A. (2020). Urban driving with conditional imitation learning. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA) (pp. 251–257). IEEE.
International Organization for Standardization. (2022). ISO 21448:2022 — Road vehicles: Safety of the intended functionality (SOTIF). ISO. https://www.iso.org/standard/77490.html
International Telecommunication Union — Telecommunication Standardization Sector, Study Group 17. (2023). Recommendation ITU-T X.1376: Security guidelines for vehicular edge computing. ITU-T.
Jiang, B., Chen, S., Wang, X., Liao, B., Cheng, T., Chen, J., Zhou, H., Zhang, Q., Liu, W., & Huang, C. (2024). Senna: Bridging large vision-language models with end-to-end autonomous driving. arXiv Preprint arXiv:2410.22313. https://arxiv.org/abs/2410.22313
Kim, J., Rohrbach, A., Darrell, T., Canny, J., & Akata, Z. (2018). Textual explanations for self-driving vehicles. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 563–578). Springer.
Li, Y., Fan, W., Chen, R., Fan, L., & Chen, X. (2024). CODA-LM: A benchmark for corner-case detection. arXiv Preprint arXiv:2404.10595. https://arxiv.org/abs/2404.10595
Lin, C., Zhang, W., Chen, Y., Yang, L., Jiang, H., Tian, D., ... & Cao, D. (2026). Multimodal 3D object detection for autonomous driving under vision-language supervision: a contrastive-learning perspective. Science China Information Sciences, 69(5), 150106.
Mao, J., Qian, Y., Zhao, H., & Wang, Y. (2023). GPT-driver: Learning to drive with GPT. arXiv Preprint arXiv:2310.01415. https://arxiv.org/abs/2310.01415
Marcu, A., Ismail-Fawaz, A., Chen, D., Lupu, R., & Vasconcelos, N. (2024). LingoQA: Visual question answering for autonomous driving. In Proceedings of the European Conference on Computer Vision (ECCV). Springer.
NVIDIA Research. (2025). Cosmos: A world foundation model for physical AI. arXiv Preprint arXiv:2501.03575. https://arxiv.org/abs/2501.03575.
OpenAI. (2023). GPT-4V(ision) system card. OpenAI. https://openai.com/research/gpt-4v-system-card
Oyedotun, S. A., Oise, G. P., & Ozobialu, C. E. (2025). Towards intelligent cybersecurity in SCADA and DCS environments: Anomaly detection using multimodal deep learning and explainable AI. Journal of Science Research and Reviews, 2(4), 88–104.
Qian, T., Chen, J., Zhuo, L., Jiao, Y., & Jiang, Y.-G. (2024). nuScenes-QA: A multi-modal visual question answering benchmark for autonomous driving scenarios. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, pp. 4542–4550). AAAI Press.
Renz, K., Chitta, K., Mercea, O.-B., Koepke, A. S., Akata, Z., & Geiger, A. (2022). PlanT: Explainable planning transformers via object-level representations. In Proceedings of the Conference on Robot Learning (CoRL). PMLR.
Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S. L., Liu, Y., & Li, H. (2024). LMDrive: Closed-loop end-to-end driving with large language models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 15120–15130). IEEE.
Shao, H., Wang, L., Chen, R., Waslander, S. L., Li, H., & Liu, Y. (2023). ReasonNet: End-to-end driving with temporal and global reasoning. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 13723–13733). IEEE.
Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., Beißwenger, J., Luo, P., Geiger, A., & Li, H. (2024). DriveLM: Driving with graph visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV). Springer.
Tian, H., Reddy, D. R., Ding, Y., Chen, W., Wang, T. Y.-H., Liniger, A., & Piechnick, C. (2024). VLM-AD: End-to-end autonomous driving through vision-language model enhancement. arXiv Preprint arXiv:2412.14446. https://arxiv.org/abs/2412.14446
Wang, X., Zhu, Z., Huang, G., Chen, X., Zhu, J., & Lu, J. (2023). DriveMLM: Aligning multi-modal large language models with behavioral planning. arXiv Preprint arXiv:2312.09245. https://arxiv.org/abs/2312.09245
Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.-Y. K., Li, Z., & Zhao, H. (2024). DriveGPT4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters, 9(10), 8186–8193.
Yurtsever, E., Lambert, J., Carballo, A., & Takeda, K. (2020). A survey of autonomous driving: Common practices and emerging technologies. IEEE Access, 8, 58443–58469. https://doi.org/10.1109/ACCESS.2020.2983149
Zhou, X., Liu, M., Yurtsever, E., Zagar, B. L., Zimmer, W., Cao, H., & Knoll, A. C. (2024). Vision language models in autonomous driving: A survey and outlook. IEEE Transactions on Intelligent Vehicles, 9(4), 4489–4506. https://doi.org/10.1109/TIV.2024.3402136
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Austin Olom Ogar, Joshua Abah, Aliyu Suleiman Muhammad, Faruk Obansa Muhammed, Hauwa Ibrahim, Mahmood Suleiman (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
- Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- NonCommercial — You may not use the material for commercial purposes.
- No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits.