DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
Summary
Introduces DriveDNA, a large-scale multimodal naturalistic driving dataset with 4,121 drives from 465 drivers across 115 vehicle models, and a benchmark for driving style identification through tasks like few-shot driver re-identification and personalized behavior prediction.
View Cached Full Text
Cached at: 07/28/26, 06:34 AM
Paper page - DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
Source: https://huggingface.co/papers/2607.23822 Published on Jul 26
·
Submitted byhttps://huggingface.co/HenryYHW
Henryon Jul 28
Abstract
Drivingstylecapturesstable,driver-specificpatternsinhowavehicleisdriven.Innaturalisticdata,however,thissignalishardtoisolatebecausedriversareobservedindifferentvehicles,ondifferentroads,andunderdifferentconditions,somodelsmaymistakevehicle-orsituation-specificregularitiesfordriver-specificstyle.WeintroduceDriveDNA,alarge-scalenaturalisticdatasetandbenchmarkforpersonalizeddriving-stylemodeling,comprising4,121drivesfrom465driversacross115vehiclemodelsandtotaling975hoursofhuman-controlleddrivingat10Hzwithforwardvideo,collectedfromcommunitydriversineverydayuse.DriveDNAdefinesdrivingstyleasaconsistent,driver-specificbehavioralpatterninhowavehiclemovesundersimilarconditions.Thebenchmarkevaluatesthissignalthroughthreecoretasks:few-shotdriverre-identification,personalizedbehaviorprediction,andcondition-matchedcomparison,andprovidesbehavioralannotationsplus276,248rule-generatedmaneuvereventsacrosssixclasseswithlarge-scalehumanauditing.Weevaluatebaselinesspanningclassicaldescriptors,supervisedandself-supervisedtime-seriesencoders,multimodalfusion,probabilisticprediction,andzero-shotfoundationmodelsunderafixedmulti-seedprotocol.Learnedrepresentationssubstantiallyoutperformclassicaldescriptorsonunseendrivers(AUROC.935vs..707)andretaindriver-specificinformationundermatcheddrivingconditions,whiledescriptorperformanceapproacheschance.Video-onlymodelsachievecomparablere-identificationaccuracybutexhibitsevererouteleakage,showingthatstrongrecognitionmayarisefromcontextualshortcutsratherthandrivingbehavior.Thesefindingsshowthatreliabledriving-styleevaluationmustassessboththebehavioralvalueoflearnedrepresentationsandtheirrobustnesstovehicle,drive,andconditionconfounds.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2607\.23822
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.23822 in a model README.md to link it from this page.
Datasets citing this paper1
#### HenryYHW/DriveDNA Updatedabout 4 hours ago • 19
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.23822 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
PersonaDrive: Human-Style Retrieval-Augmented VLA Agents for Closed-Loop Driving Simulation
This paper introduces PersonaDrive, a pipeline that conditions a vision-language-action (VLA) driving agent on retrieved demonstrations from a style-instructed human driving dataset, enabling style-diverse non-ego agents for closed-loop simulation and improving driving scores on Bench2Drive.
The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset
KITScenes Multimodal is a high-fidelity European autonomous driving dataset with synchronized sensors, complete 3D HD maps, and four benchmarks for spatial learning and embodied AI research.
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
Introduces CMU-Drive, a closed-loop benchmark for cooperative multi-agent autonomous driving, and V2V-VLA, a vision-language-action model that jointly generates driving actions, waypoints, reasoning, and communication policies. This provides the first benchmark and baseline for cooperative VLA driving.
DriveZero: End-to-End Driving Beyond Human Demonstrations
DriveZero is an end-to-end autonomous driving system that combines a vision foundation model for perception with reinforcement learning for action, achieving driving behaviors beyond human demonstrations and state-of-the-art performance on benchmarks like nuPlan and NAVSIM.
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Introduces DF3DV-1K, a large-scale real-world dataset with 1,048 scenes and 89,924 images for distractor-free novel view synthesis, along with a benchmark of nine methods and an application improving radiance field methods via fine-tuning a diffusion-based 2D enhancer.