Tag
This paper benchmarks six frontier LLMs on coding crash attributes from police crash narratives against an official fatal-crash database, finding that while GPT-5.5 High leads among LLMs, simple baselines rival or beat LLM performance and attribute-specific differences outweigh model differences.