OpenLongTail: Generative Scaling of Long-Tail Driving Data
Summary
OpenLongTail is an open-source generative data engine that transforms heterogeneous long-tail driving data into view-aligned multi-view assets for training robust autonomous driving policies, improving closed-loop driving robustness.
View Cached Full Text
Cached at: 07/21/26, 06:35 AM
Paper page - OpenLongTail: Generative Scaling of Long-Tail Driving Data
Source: https://huggingface.co/papers/2607.09655 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Scalingrobustdrivingpoliciesisfundamentallybottleneckedbythescarcityofedgecasesincurateddatasets.Whiletherealworldcontinuouslycapturesthesecriticalevents,suchlong-taileventsremainunderutilizedwhencollectedfromheterogeneoussources.Specifically,diversebutvaluablein-the-wildlong-tailvideoslackthefullviewcoveragerequiredfortrainingpolicymodels,oftenmissingmulti-viewposesororiginatingsolelyfrommonoculardashcameras.Thismodalitygappreventstheseubiquitousobservationsfrombeingconvertedintoscalabletrainingdataforlong-tailgeneralization.WeintroduceOpenLongTail,anopen-sourcegenerativedataengineforscalingautonomousdrivingpoliciesunderlong-tailevents.Totransformheterogeneousdatasourcesintoview-alignedandtemporallycoherentmulti-viewassetsthatareusefulforpolicylearning,wedevelopapose-informedextrapolativeviewsynthesispipelinethatgeneratesthemissingviews.Wefurtherenhancecross-viewconsistencyandthetemporalalignmentforthenewlygeneratedviewsbyinjectingPlückerraygeometryintothescalablegenerationengine.Bysynthesizingheterogeneouslong-taildata,weobserveasignificantimprovementinclosed-loopdrivingrobustnessinhandlinglong-tailevents.Bymeasuringtheextrapolativeviewsynthesisandposemetrics,wevalidatetheeffectivenessofOpenLongTailinvisualfidelity,cross-viewconsistency,andego-trajectoryrecovery.
View arXiv pageView PDFProject pageGitHub24Add to collection
Get this paper in your agent:
hf papers read 2607\.09655
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.09655 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.09655 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.09655 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
TailBooster is a dual-layer generative framework that synthesizes operationally valid extreme air-transport events using statistical tail extraction and autoencoder-based cleaning, significantly improving extreme-event prediction accuracy.
REARL: A Closed-loop Autonomous Driving Simulation Enhancement Framework with Real Traffic Data and Large Language Models
The paper proposes REARL, a closed-loop framework using real traffic data and large language models to enhance autonomous driving simulation by continuously monitoring and adjusting vehicle behavior for greater realism.
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek LLM is an open-source language model project that develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
OpenResearcher presents a reproducible pipeline for training deep research agents using offline search environments and synthesized trajectories, achieving significant accuracy improvements on benchmark tasks like BrowseComp-Plus.
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Introduces OctoLong, a context engineering pipeline for curating dependency-rich cross-repository code contexts, and OctoLong-Instruct, a suite of long-context open LMs trained on this data. Experiments show that replacing 12% of traditional long-context corpora with OctoLong data yields substantial gains in long-range retrieval, state tracking, repository-level code understanding, and agentic tasks.