Introduction: Road traffic crashes pose a significant public health challenge, with traditional crash data often relying on broad classifications that obscure critical details. This study addresses the knowledge gap created by ambiguous categories, “Other Improper Action” and “Not Discernible,” within the Ohio crash dataset. Method: Employing a multi-faceted analytical framework that combines descriptive statistics, N-gram analysis, and Latent Dirichlet Allocation (LDA) topic modeling on over 67,000 free-text crash narratives from 2020 to 2024, the study uncovers the latent contributing circumstances previously masked by these labels. Results: The analysis reveals that “Other Improper Action” incidents are disproportionately linked to adverse environmental conditions and a lack of formal traffic control. Text mining further extracted hidden behavioral and environmental factors, predominantly severe spatial awareness deficits, striking legally parked vehicles in urban environments and environmentally induced loss of vehicle control resulting in infrastructure strikes. In contrast, “Not Discernible” crashes are more prevalent in daylight and at signalized intersections. Rather than a lack of physical information, the NLP models revealed that investigative ambiguity primarily stems from conflicting driver accounts, specifically mutual lane change encroachment and right-of-way disputes. Conclusions: These findings not only validate the presence of recognized crash mechanisms, but transform vague data classifications into actionable intelligence. Practical Applications: This methodology provides transportation safety professionals with the granular evidence needed to refine statewide data collection protocols and develop targeted, engineering-backed safety interventions.
更多
查看译文
关键词
Traffic safety,Crash narratives,Contributing circumstances,LDA,Text mining,Crash data quality