Research arXiv cs.AI

DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

autonomous drivingVLM benchmarkclosed-loop simulationDriveHierarchy

Evaluating VLM-based autonomous driving is difficult because driving competence is composite: a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving, but they provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement.

DriveHierarchy addresses this by organizing VLM-based autonomous driving into four ranks: perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate the hierarchy, the authors integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark containing 76,798 question-answer pairs over 84,279 frames. They also develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation.

Experiments on 15 VLMs show that DriveHierarchy captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. The benchmark is therefore presented as a unified framework for evaluating and improving VLM-based autonomous driving systems, and an anonymized project has been released at https://github.com/PerfectXu88/DriveHierarchy.

Read original →

← Back to home