TraffCOCO: A Scalable Framework and Dataset for Traffic Scene Object Detection
Alam, M. S., Bazilinskyy, P.
In preparation.
ABSTRACT Detection of traffic objects is critical for automated driving and Intelligent Transportation Systems (ITS). While recent advancements in the field of object detection have demonstrated remarkable performance on object recognition tasks, the ability to generalise object recognition models to diverse geographical areas is hampered due to the lack of geographical diversity of traffic datasets. Differences in road infrastructure, traffic laws, traffic control devices, types of vehicles, and weather conditions are obstacles that hinder the development of globally deployable perception models. This paper describes a framework for the development of a geographically representative YOLO-based traffic object detection model. The proposed framework overcomes the shortcomings of existing traffic object detection datasets by using an automated semantic-first annotation pipeline that uses Vision-Language Models (VLMs), ontology-guided semantic normalisation, open-vocabulary object localisation, and automatic generation of COCO-style annotations to create a geographically diverse traffic dataset. The created traffic dataset is used for training a geographically diverse YOLO-based object detection model capable of detecting a wide variety of traffic-related objects in different driving environments.