DigestAI news desk

Cut through the AI noise.

Research

Researchers release DriveHierarchy benchmark for VLM driving tests

DriveHierarchy is a new benchmark for evaluating vision-language models (VLMs) in autonomous driving. The framework breaks down driving competence into four ranks: perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. It combines 76,798 question-answer pairs from 84,279 open-source driving frames with 100 simulated closed-loop scenarios on real-world road networks.

1 source primary source

Key points

  • DriveHierarchy ranks VLM driving skills in four stages: perception, memory, reasoning, and execution
  • Uses 76,798 Q&A pairs from 84,279 open-source driving frames and 100 closed-loop scenarios
  • Tests 15 VLMs, showing structured capability gaps and open-loop-to-closed-loop connections

The project integrates multiple open-source datasets into a unified platform. Tests on 15 VLMs show the benchmark captures distinct capability variations and links open-loop understanding to closed-loop performance. The goal is to help diagnose and improve VLM-based autonomous driving systems. An anonymized version is available on GitHub.

Read the original at arXiv cs.AI · by Chengkai Xu, Jiaqi Liu, Yicheng Guo, Peng Hang, Jian Sun primary sourceOpen source ↗
Topics · follow one to build your own front page
DriveHierarchyPerfectXu88

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories