Multimodal Large Language Models for Autonomous Driving: A Comprehensive Survey of Perception, Reasoning, Planning, and Safety Assurance

Authors

DOI:

https://doi.org/10.70882/josrar.2026.v3i4.245

Keywords:

Autonomous driving, Multimodal large language models, Vision-language models, Driving benchmarks, ISO 26262, ISO 21448 (SOTIF), UL 4600 safety case, Sim-to-real transfer

Abstract

Autonomous driving has progressed from rule-based subsystems and modular perception, prediction, and planning stacks toward unified data-driven architectures, and multimodal large language models (MLLMs) are increasingly proposed as the cognitive substrate of the next generation of highly automated road vehicles. This 2026 survey synthesises 39 primary sources selected from an initial corpus of 274 candidate records screened over 2020-2026, organises the field around a five-role pipeline taxonomy (perception, prediction, planning, control, and human-machine interaction), and compares six representative driving MLLMs (DriveGPT-4, LMDrive, Senna, DriveLM, GPT-4V-AD, and Cosmos-1) on accuracy, latency, and parameter footprint. A benchmark coverage matrix over LingoQA, BDD-X, DriveLM, nuScenes-QA, AutoHallu, and CODA-LM exposes evaluation gaps in prediction and planning. Model behaviour is translated into safety-assurance terms by mapping four MLLM failure-mode families to the functional-safety standard ISO 26262, the Safety of the Intended Functionality standard ISO 21448 (SOTIF), and the autonomous-systems safety-case standard UL 4600. A three-tier vehicle, edge, and cloud deployment topology is described together with the digital-twin and over-the-air update infrastructure that surrounds it. The strongest empirical finding is that Cosmos-1 delivers the best accuracy among models with sub-150 ms latency (76.6 percent mean reasoning accuracy at 480 ms), leaving verifiable safety certification as the single most important open problem for closed-loop deployment. The survey closes with a six-item research agenda spanning sub-100 ms real-time inference, out-of-distribution generalisation, multi-agent intent reasoning, verifiable safety certification, long-tail corner-case coverage, and closed-loop sim-to-real transfer. The article is intended as a reference for automotive system architects, safety engineers, regulators, and machine-learning researchers preparing the next generation of automated driving systems.

References

Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L. D., Monfort, M., Muller, U., Zhang, J., Zhang, X., Zhao, J., & Zieba, K. (2016). End to end learning for self-driving cars. arXiv Preprint arXiv:1604.07316. https://arxiv.org/abs/1604.07316

Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., & Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11621–11631). IEEE.

Casas, S., Sadat, A., & Urtasun, R. (2021). MP3: A unified model to map, perceive, predict, and plan. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 14403–14412). IEEE.

Chen, H., Liu, X., & Wang, T. (2025). Trustworthiness of autonomous-driving foundation models. (Preprint; details available on request).

Chen, L., Sinavski, O., Hunermann, J., Karnsund, A., Willmott, A. J., Birch, D., Maund, D., & Shotton, J. (2023). Driving with LLMs: Fusing object-level vector modality for explainable autonomous driving. arXiv Preprint arXiv:2310.01957. https://arxiv.org/abs/2310.01957

Choi, J., Park, J., & Lee, S. (2026). Diffusion Models for End-to-End Autonomous Driving: A Survey of Perception, Prediction, Planning, and Control. IEEE Access.

Cui, C., Ma, Y., Cao, X., Ye, W., Zhou, Y., Liang, K., Chen, J., Lu, J., Yang, Z., Liao, K.-D., Gao, T., Li, E., Tang, K., Cao, Z., Zhou, T., Liu, A., Yan, X., Mei, S., Cao, J., Wang, Z., & Zheng, C. (2024). A survey on multimodal large language models for autonomous driving. In Proceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW) (pp. 958–979). IEEE. https://doi.org/10.1109/WACVW60836.2024.00106

European Commission. (2024). Regulation (EU) 2024/1689 — The AI Act. Official Journal of the European Union.

Fu, D., Li, X., Wen, L., Dou, M., Cai, P., Shi, B., & Qiao, Y. (2023). Drive like a human: Rethinking autonomous driving with large language models. arXiv Preprint arXiv:2307.07162. https://arxiv.org/abs/2307.07162

Grigorescu, S., Trasnea, B., Cocias, T., & Macesanu, G. (2020). A survey of deep learning techniques for autonomous driving. Journal of Field Robotics, 37(3), 362–386. https://doi.org/10.1002/rob.21918

Hawke, A., Shen, R., Gurau, C., Sharma, S., Reda, D., Nikolov, N., Mazur, P., Micklethwaite, S., Griffiths, N., Shah, A., & Kendall, A. (2020). Urban driving with conditional imitation learning. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA) (pp. 251–257). IEEE.

International Organization for Standardization. (2022). ISO 21448:2022 — Road vehicles: Safety of the intended functionality (SOTIF). ISO. https://www.iso.org/standard/77490.html

International Telecommunication Union — Telecommunication Standardization Sector, Study Group 17. (2023). Recommendation ITU-T X.1376: Security guidelines for vehicular edge computing. ITU-T.

Jiang, B., Chen, S., Wang, X., Liao, B., Cheng, T., Chen, J., Zhou, H., Zhang, Q., Liu, W., & Huang, C. (2024). Senna: Bridging large vision-language models with end-to-end autonomous driving. arXiv Preprint arXiv:2410.22313. https://arxiv.org/abs/2410.22313

Kim, J., Rohrbach, A., Darrell, T., Canny, J., & Akata, Z. (2018). Textual explanations for self-driving vehicles. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 563–578). Springer.

Li, Y., Fan, W., Chen, R., Fan, L., & Chen, X. (2024). CODA-LM: A benchmark for corner-case detection. arXiv Preprint arXiv:2404.10595. https://arxiv.org/abs/2404.10595

Lin, C., Zhang, W., Chen, Y., Yang, L., Jiang, H., Tian, D., ... & Cao, D. (2026). Multimodal 3D object detection for autonomous driving under vision-language supervision: a contrastive-learning perspective. Science China Information Sciences, 69(5), 150106.

Mao, J., Qian, Y., Zhao, H., & Wang, Y. (2023). GPT-driver: Learning to drive with GPT. arXiv Preprint arXiv:2310.01415. https://arxiv.org/abs/2310.01415

Marcu, A., Ismail-Fawaz, A., Chen, D., Lupu, R., & Vasconcelos, N. (2024). LingoQA: Visual question answering for autonomous driving. In Proceedings of the European Conference on Computer Vision (ECCV). Springer.

NVIDIA Research. (2025). Cosmos: A world foundation model for physical AI. arXiv Preprint arXiv:2501.03575. https://arxiv.org/abs/2501.03575.

OpenAI. (2023). GPT-4V(ision) system card. OpenAI. https://openai.com/research/gpt-4v-system-card

Oyedotun, S. A., Oise, G. P., & Ozobialu, C. E. (2025). Towards intelligent cybersecurity in SCADA and DCS environments: Anomaly detection using multimodal deep learning and explainable AI. Journal of Science Research and Reviews, 2(4), 88–104.

Qian, T., Chen, J., Zhuo, L., Jiao, Y., & Jiang, Y.-G. (2024). nuScenes-QA: A multi-modal visual question answering benchmark for autonomous driving scenarios. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, pp. 4542–4550). AAAI Press.

Renz, K., Chitta, K., Mercea, O.-B., Koepke, A. S., Akata, Z., & Geiger, A. (2022). PlanT: Explainable planning transformers via object-level representations. In Proceedings of the Conference on Robot Learning (CoRL). PMLR.

Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S. L., Liu, Y., & Li, H. (2024). LMDrive: Closed-loop end-to-end driving with large language models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 15120–15130). IEEE.

Shao, H., Wang, L., Chen, R., Waslander, S. L., Li, H., & Liu, Y. (2023). ReasonNet: End-to-end driving with temporal and global reasoning. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 13723–13733). IEEE.

Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., Beißwenger, J., Luo, P., Geiger, A., & Li, H. (2024). DriveLM: Driving with graph visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV). Springer.

Tian, H., Reddy, D. R., Ding, Y., Chen, W., Wang, T. Y.-H., Liniger, A., & Piechnick, C. (2024). VLM-AD: End-to-end autonomous driving through vision-language model enhancement. arXiv Preprint arXiv:2412.14446. https://arxiv.org/abs/2412.14446

Wang, X., Zhu, Z., Huang, G., Chen, X., Zhu, J., & Lu, J. (2023). DriveMLM: Aligning multi-modal large language models with behavioral planning. arXiv Preprint arXiv:2312.09245. https://arxiv.org/abs/2312.09245

Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.-Y. K., Li, Z., & Zhao, H. (2024). DriveGPT4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters, 9(10), 8186–8193.

Yurtsever, E., Lambert, J., Carballo, A., & Takeda, K. (2020). A survey of autonomous driving: Common practices and emerging technologies. IEEE Access, 8, 58443–58469. https://doi.org/10.1109/ACCESS.2020.2983149

Zhou, X., Liu, M., Yurtsever, E., Zagar, B. L., Zimmer, W., Cao, H., & Knoll, A. C. (2024). Vision language models in autonomous driving: A survey and outlook. IEEE Transactions on Intelligent Vehicles, 9(4), 4489–4506. https://doi.org/10.1109/TIV.2024.3402136

Five-role pipeline taxonomy for autonomous-driving MLLMs

Downloads

Published

2026-08-04

How to Cite

Ogar, A. O., Abah, J., Muhammad, A. S., Muhammed, F. O., Ibrahim, H., & Suleiman, M. (2026). Multimodal Large Language Models for Autonomous Driving: A Comprehensive Survey of Perception, Reasoning, Planning, and Safety Assurance. Journal of Science Research and Reviews, 3(4), 130-139. https://doi.org/10.70882/josrar.2026.v3i4.245