Publications

Working documents

  • Understanding global pedestrian behaviour in 565 cities with dashcam videos on YouTube
    Alam, M. S., Martens, M., Bazilinska, O., Bazilinskyy, P.
    In preparation.
    The interactions between future cars and pedestrians should be designed to be understandable and safe worldwide. Although previous research has studied vehicle-pedestrian interactions within specific cities or countries, this study offers a more scalable and robust approach by examining pedestrian behaviour worldwide. We present a dataset, Pedestrians in YouTube (PYT), which includes 1562.80 hours of YouTube day and night dashcam footage from 565 cities in 104 countries. The included videos feature continuous urban driving, are at least 10 min long, feature no atypical events and represent everyday conditions, and are from cities with a minimum of 20,000 population. We detected pedestrian movements, focussing on the speed and the pedestrian crossing decision time during road crossings based on the bounding boxes given by YOLOv11x. The results revealed statistically significant variations in pedestrian behaviour influenced by socioeconomic and environmental factors such as Gross Metropolitan Product, traffic-related mortality, Gini coefficient, traffic index, and literacy. The dataset is publicly available to encourage further research on global pedestrian behaviour.
  • Deep Reinforcement Learning based Eye Tracking for Unity Environment
    Alam, M. S., Bazilinskyy, P.
    Working project.
  • Nineteen years of ASMR on YouTube: A multilingual, theme-level analysis of 89,241 videos
    Alam, M. S., Bazilinskyy, P.
    In preparation.
    Autonomous sensory meridian response (ASMR) videos are commonly framed around whispered speech, gentle sounds, close-up attention, relaxation, sleep, and comfort, yet longitudinal, large-scale descriptions of this labelled YouTube genre are scarce. We analyse 89,241 ASMR YouTube videos uploaded between 1 January 2008 and 31 July 2026, with metadata processed on 1 August 2026, from 18,024 channels across 81 language categories. Using title and description text mining with duration, engagement, growth, theme-label indicators, and exploratory clustering metrics, we provide a multilingual longitudinal map of explicitly labelled ASMR on YouTube. Mean growth was 1,125.41 views/day (SD 11,196.23). Among videos with known duration, those under 10 min grew fastest on average (2,075.55 views/day), videos between 10 and 30 min averaged 585.11 views/day, and videos over 180 min averaged 1,293.15 views/day. The most prevalent themes were sleep-related content (17.62%), whisper-focused content (15.49\%), role-play (14.09\%), mukbang or eating (7.08\%), and visual triggers (6.85\%); driving-related content accounted for 3.67\%. An exploratory k-means solution grouped videos into 11 clusters, and a full-dataset UMAP projection visualised the resulting structure for all 89,241 videos without treating those groups as definitive ASMR subgenres.
  • A global dataset of continuous urban dashcam driving
    Alam, M. S., Bazilinska, O., Bazilinskyy, P.
    In preparation.
    We introduce CROWD (City Road Observations With Dashcams), a manually curated dataset of ordinary, variable duration, temporally contiguous, unedited, front facing urban dashcam segments screened and segmented from publicly available YouTube videos. CROWD is designed to support cross-domain robustness and interaction analysis by prioritising routine driving and explicitly excluding crashes, crash aftermath, and other edited or incident-focused content. The release contains 75,147 segment records spanning 26,273.75 hours (58,467 unique uploads), covering 8,049 named inhabited places in 238 countries and territories across all six inhabited continents (Africa, Asia, Europe, North America, South America and Oceania), with segment level manual labels for time of day (day or night) and vehicle type. To lower the barrier for reproducible baseline analysis, we provide per-segment CSV files of automatic detections for all 80 MS-COCO classes produced with YOLOv11x, together with segment-local multi-object tracks generated with BoT-SORT; e.g. person, bicycle, motorcycle, car, bus, truck, traffic light, stop sign, etc. CROWD is distributed as video identifiers with segment boundaries and derived detection and tracking pseudo-labels, enabling reproducible research without redistributing the underlying videos.
  • Collision Patterns and Reporting Blind Spots in 971 California Autonomous Vehicle Crash Reports
    Alam, M. S., Zhang, L., Li, Z., Dou, F., Bazilinskyy, P.
    In preparation.
    Autonomous vehicle collision reports offer a rare view of how autonomous driving systems perform in mixed traffic, but they are difficult to analyse at scale because they combine structured form fields, visually marked elements, and free text narratives. We analysed 970 publicly available California Department of Motor Vehicles collision reports, with dated reports spanning October 2014 to March 2026, using a ChatGPT 5.4 Thinking extraction pipeline to derive a structured dataset for empirical analysis. The largest scenario classes were rear end crashes stopped by AV (266, 27.4%), intersection lateral conflicts (180, 18.6%) and lane change or merge conflicts (156, 16.1%). Reports captured coarse scenario structure much more reliably than fine interaction detail, with a mean coarse context score of 0.97 versus 0.48 for fine context. The results suggest that many reported collisions occur in mixed traffic situations where prediction, coordination, and road user expectations are difficult to align.
  • Trial ordering in repeated exposure virtual reality pedestrian research: comparing fixed and participant specific randomised sequences

    In preparation.
    When every participant receives the same trial sequence, each condition is tied to one trial position and preceding experience. We compared two non concurrent runs of a virtual reality pedestrian interaction experiment. Participants were not assigned at random to study run. Fifty participants received participant specific randomised sequences and 50 received one fixed sequence of the same 40 trials. The main measure was the percentage of 100 ms prepassage intervals in which a trigger indicated perceived crossing unsafety. The estimate was 40.48% in the randomised run and 42.82% in the fixed run, a difference of 2.34 percentage points (95% CI [-5.16, 9.83], p=0.54). The confidence interval did not establish equivalence. No reliable overall differences were detected for additional trigger or head movement measures, although condition response profiles differed for several outcomes. The strongest condition specific contrast concerned the perceived influence of pedestrian spacing in a yielding, eHMI off, participant second, 2 m condition. The fixed run was 23.34 rating points lower (globally adjusted p=0.011), and this condition always occurred first in the fixed sequence. Its effect could not be separated from early session experience or sample differences. In the randomised run, the estimated eHMI contribution to reported vehicle intention understanding increased with trial position and prior relevant exposure, but these changes did not differ reliably between runs. Confounding of scheduling with study run precludes causal attribution, but the comparison demonstrates why fixed sequences weaken the interpretation of condition specific and experience related effects.
  • Perceived Crossing Risk in Multipedestrian Encounters with Automated Vehicles: Effects of Yielding, eHMIs, and Pedestrian Position
    Alam, M. S., Dey, D., Martens, M., Bazilinskyy, P.
    In preparation.
    This virtual reality study examined how pedestrian spacing and relative position influence perceived crossing risk during encounters with an automated vehicle. Fifty participants completed a within participant experiment varying vehicle yielding, conditional external Human-Machine Interface (eHMI) logic, relative pedestrian order, and five spacings from 2 to 10~m. Participants pressed a controller trigger whenever they judged that initiating a crossing would be unsafe. The primary outcome was the percentage of a fixed five second interval before vehicle passage classified as unsafe. A participant clustered marginal binomial model showed that yielding reduced predicted perceived unsafety by 38.71--52.66 percentage points across the four eHMI and order contexts, with Holm adjusted p<0.001 throughout. Conditional eHMI logic reduced perceived unsafety by 9.30 points in yielding trials where the participant was first in the vehicle's path (95% CI [-14.85,-3.75], adjusted p=0.004). A 2,000 sample participant bootstrap supported this contrast (95% percentile interval [-15.31,-4.25]) and all four yielding contrasts. No statistically reliable eHMI effect was found in either non-yielding order condition or in the yielding trials where the avatar was first. Complete model repetitions at trigger thresholds of 0.05, 0.10, and 0.50 produced the same substantive conclusions. Analyses that held participant position constant or aligned data to vehicle events did not support a consistent independent spacing effect. For pedestrian safety assessment, yielding kinematics remain the dominant cue, and any eHMI benefit should be evaluated across multipedestrian geometries rather than assumed to generalise across crossing contexts.
  • TraffCOCO: A Scalable Framework and Dataset for Traffic Scene Object Detection
    Alam, M. S., Bazilinskyy, P.
    In preparation.
    Detection of traffic objects is critical for automated driving and Intelligent Transportation Systems (ITS). While recent advancements in the field of object detection have demonstrated remarkable performance on object recognition tasks, the ability to generalise object recognition models to diverse geographical areas is hampered due to the lack of geographical diversity of traffic datasets. Differences in road infrastructure, traffic laws, traffic control devices, types of vehicles, and weather conditions are obstacles that hinder the development of globally deployable perception models. This paper describes a framework for the development of a geographically representative YOLO-based traffic object detection model. The proposed framework overcomes the shortcomings of existing traffic object detection datasets by using an automated semantic-first annotation pipeline that uses Vision-Language Models (VLMs), ontology-guided semantic normalisation, open-vocabulary object localisation, and automatic generation of COCO-style annotations to create a geographically diverse traffic dataset. The created traffic dataset is used for training a geographically diverse YOLO-based object detection model capable of detecting a wide variety of traffic-related objects in different driving environments.
  • Enhancing driver experience in SAE level 3 automated vehicles through multimodal and emotion-aware in-vehicle agents
    Zeng, X., Alam, M. S., Bazilinskyy, P.
    In preparation.
    Multimodal in-vehicle agents (IVAs) combine voice, gesture, facial expression, and affect-responsive behaviour, but the appropriate role and limits of each channel remain unclear. We report two exploratory studies of a physically embodied robot-like IVA using pre-recorded highway-driving videos in a controlled Wizard-of-Oz style setup. Study~1 compared voice alone, voice with facial expressions, and voice with gestures (N=12). Descriptive workload averages were lower in the two non-verbal conditions (baseline: 33%, facial expressions: 23%, gestures: 18%), but the fixed baseline-first order prevents attributing this pattern to modality alone. Participant accounts instead revealed a clearer functional distinction: gestures were valued for noticeable, anticipatory, and action-oriented signalling, whereas facial expressions were valued for social, affective, and aesthetic communication. Study~2 used a revised prototype to explore affect-aware feedback under different user-availability conditions (N=12). Variable feedback exposure, perceived facial-expression misclassification, interruption timing, and mismatch with users' self-perceived states constrained the appropriateness of the interaction. Quantitative comparisons in both studies are treated as descriptive context for the qualitative findings, while secondary statistical checks are retained only for transparency. The contribution is a set of multimodal design boundaries rather than evidence of improved driver experience: voice for actionable information, gestures for interpretable anticipation, facial expressions for low-urgency social communication, and affect-aware feedback only when sensing confidence, user availability, and timing support its use.

2026

  • A Survey of Day-Night Illumination Domain Translation for Outdoor Vision: Methods, Datasets, and Evaluation Protocols
    Alam, M. S., Singh, P., Bazilinskyy, P.
    Machine Vision and Application (2026)

    Day-night appearance shift degrades vision for driving and surveillance. Low illumination, mixed lighting, glare, and sensor noise weaken cues for detection, segmentation, localisation, and tracking. We survey illumination domain translation for images and video, focusing on day to night and night to day mapping that changes illumination while preserving geometry, semantics, and temporal coherence. We relate illumination modelling and colour transfer to learning based methods, and develop an IDT-specific constraint centric taxonomy linking supervision, five domain gap factors, and five families of constraints and priors to typical failure modes. Using this taxonomy, we organise 30 representative methods and summarise 23 datasets. We also report an artefact availability audit of 34 published methods: 29 release code, 22 provide pretrained weights, 21 specify licences, and 19 provide reproducibility packages. Finally, we recommend evaluation spanning perceptual quality, semantic preservation, downstream utility, and temporal stability, and we synthesise the literature using an evidence-aligned P/S/D/T protocol that highlights recurring failure modes and evaluation gaps.
  • What Can Public Traffic Cameras Reveal? A Short Horizon Privacy Audit Using Open Source Vehicle Tracking
    Alam, M. S., Bazilinskyy, P.
    Companion of the 2026 ACM International Joint Conference on Pervasive and Ubiquitous Computing (2026)

    Public traffic cameras are often treated as low risk because they rarely expose clear faces or licence plates. We study a different leakage channel: the behavioural structure extracted from ordinary public traffic footage using open source computer vision. Our preliminary pipeline converts timestamped public YouTube livestream footage from one traffic camera, spanning a 319.64 hour wall clock period, into route based vehicle events and measures how coarse signatures affect confusability. In the current run, 877,704 raw tracks were filtered into 135,488 vehicle events with full wall clock alignment. Using only class, route, and half hour time bin produced 231 baseline signatures and 178 low confusability events. Adding approximate size, speed, duration, and coarse colour increased the signature space to 11,466 signatures and 14,382 low confusability events, with 5,160 rare recurrence candidate signatures. The analysis does not perform person tracking, licence plate recognition, or exact vehicle re identification. Instead, weak, non explicit cues can make public traffic events more distinctive than aggregate counts suggest.
  • Exploring Veo 3's Capabilities for Generating Urban Traffic Scenes in 76 Cities Worldwide
    Alam, M. S., Wang, Z., Zhang, L., Bazilinskyy, P.
    18th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (AutoUI). Gothenburg, Sweden (2026)

    This study explores the potential of Google Veo 3, a generative video model, to synthesise 8-second dashcam-style urban traffic scenes solely based on text prompts in 76 cities across six continents. YOLOv11x was used to count facts like the number of road users, traffic lights, and stop signs, revealing variations across cities: Karachi had the most objects detected (79), while Muscat had only four cars. Audio analysis using dBFS showed that Montevideo was the loudest, while Copenhagen was the loudest. Through a qualitative visual analysis, the authors assessed and confirmed the perceived authenticity of most traffic scenes and highlighted AI errors, including the inability to handle non-English languages in these videos. Moreover, we compared 10 synthetic videos of New York City and Kampala, each, and verified that Veo 3 is consistent. To summarise, Veo 3 is capable of synthesising authentic, logical traffic scenes worldwide; nevertheless, it still poses non-negligible errors.
  • Personalised Electric Vehicle Acoustics with Generative AI: Dynamic Sonification and User Acceptance Study
    Verstelle, W., Alam, M. S., Bazilinskyy, P.
    18th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (AutoUI). Gothenburg, Sweden (2026)

    Electric vehicles (EVs) reduce powertrain noise, creating safety challenges and opportunities for sound design. This paper examines whether generative audio supports personalised dynamic acoustics for future EVs. We report a research through design study in which AI-generated sound samples were selected, edited into seamless loops and embedded in a sonification (the process of translating non-auditory data into sound) prototype. The prototype connects a Processing vehicle interface to a Pure Data audio engine using Open Sound Control. Vehicle speed and throttle input modulate pitch, amplitude and load related parameters. A user study with 20 participants combined the Acceptance Scale with open questions. The results show moderately positive acceptance, with an average usefulness of 0.88 and a satisfaction of 0.62 on a scale from -2 to +2. Participants valued personalisation and responsiveness, but requested stronger recognisability and better throttle mapping. Generative AI is useful for early EV acoustic prototyping, while the final design requires expert refinement and safety evaluation.

2025

  • Generating realistic traffic scenarios: A deep learning approach using generative adversarial networks (GANs)
    Alam, M. S., Martens, M., Bazilinskyy, P.
    Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications. Málaga, Spain (2025)

    Traffic simulations are crucial for testing systems and human behaviour in transportation research. This study investigates the potential efficacy of Unsupervised Recycle Generative Adversarial Networks (Recycle–GANs) in generating realistic traffic videos by transforming daytime scenes into nighttime environments and vice-versa. By leveraging Unsupervised Recycle-GANs, we bridge the gap between data availability during day and night traffic scenarios, enhancing the robustness and applicability of deep learning algorithms for real - world applications. GPT-4V was provided with two sets of six different frames from each day and night time from the generated videos and queried whether the scenes were artificially created based on lightning, shadow behaviour, perspective, scale, texture, detail and presence of edge artefacts. The analysis of GPT-4V output did not reveal evidence of artificial manipulation, which supports the credibility and authenticity of the generated scenes. Furthermore, the generated transition videos were evaluated by 15 participants who rated their realism on a scale of 1 to 10, achieving a mean score of 7.21. Two persons identified the videos as deep-fake generated without pointing out what was fake in the video; they did mention that the traffic was generated.
  • Cross or Nah? LLMs Get in the Mindset of a Pedestrian in front of Automated Car with an eHMI
    Alam, M. S., Bazilinskyy, P.
    17th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (AutoUI). Brisbane, QLD, Australia (2025)

    This study evaluates the effectiveness of large language model-based personas for assessing external Human-Machine Interfaces (eHMIs) in automated vehicles. 13 different models namely BakLLaVA, ChatGPT-4o, DeepSeek-VL2-Tiny, Gemma3:12B, Gemma3:27B, Granite Vision 3.2, LLaMA 3.2 Vision, LLaVA-13B, LLaVA-34B, LLaVA-LLaMA-3, LLaVA-Phi3, MiniCPM-V and Moondream were tasked with simulating pedestrian decision making for 227 vehicle images equipped with eHMI. Confidence scores (0-100) were collected under two conditions: no memory (images independently assessed) and memory-enabled (conversation history preserved), each in 15 independent trials. The model outputs were compared with the ratings of 1,438 human participants. Gemma3:27B achieved the highest correlation with humans without memory (r = 0.85), while ChatGPT-4o performed best with memory (r = 0.81). DeepSeek-VL2-Tiny and BakLLaVA showed little sensitivity to context, and LLaVA-LLaMA-3, LLaVA-Phi3, LLaVA-13B and Moondream consistently produced limited-range output.
  • Deep learning approach for realistic traffic video changes across lighting and weather conditions
    Alam, M. S., Parmar, S. H., Martens, M. H., Bazilinskyy, P.
    8th International Conference on Information and Computer Technologies (ICICT). Hilo, HI, USA (2025)

    Recent advances in GAN-based architectures have led to innovative methods for image transformation. The lack of diversity of environmental factors, such as lighting conditions and seasons in public data, prevents researchers from effectively studying the differences in the behaviour of road users under varying conditions. This study introduces a deep learning pipeline that combines CycleGAN-turbo and Real-ESRGAN to improve video transformations of traffic scenes. Evaluated using dashcam videos from Los Angeles, London, and Hong Kong, our pipeline demonstrates a notable improvement in T-SIMM for temporal consistency during night-to-day transformations, achieving a 7.97% increase for Hong Kong, 7.35% for Los Angeles, and 3.41% for London compared to CycleGAN-turbo. PSNR and VPQ scores are comparable, but the pipeline performs better in DINO structure similarity and KL divergence, with up to 153.49% better structural fidelity in Hong Kong compared to Pix2Pix and 107.32% better compared to ToDayGAN. This approach demonstrates better realism and temporal coherence in day-to-night, night-to-day, and clear-to-rainy transitions.
  • Pedestrian planet: What YouTube driving from 233 countries and territories teaches us about the world
    Alam, M. S., Martens, M. H., Bazilinskyy, P.
    17th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (AutoUI). Brisbane, QLD, Australia (2025)

    Pedestrian crossing behaviour varies globally. This study analyses dashcam footage from the CROWD dataset, covering 233 countries and territories, to examine crossing initiation time, crossing speed, and contextual variables, including detected vehicles, traffic mortality, GDP, and Gini coefficient. Qatar had the longest mean crossing initiation time (6.44 s), while China exhibited the fastest crossing speed (1.69 m/s). On average, worldwide, pedestrians exhibited a crossing initiation time of 3.18 s and crossing speed 1.20 m/s. Crossing speed and crossing initiation time are negatively correlated (r = -0.18), indicating slower crossings after longer hesitation. Crossing speed is negatively correlated with Gini coefficient (r = -0.19) and positively correlated with traffic mortality (r = 0.18). Similar crossing times in countries with different infrastructures, such as Bangladesh (3.42 s) and the Netherlands (3.40 s), underscore the complex interaction between infrastructure and behavioural adaptation. These findings emphasise the importance of culturally aware road design and the development of adaptive interfaces for vehicles.
  • Pedestrian crossing behaviour in front of electric vehicles emitting synthetic sounds: A virtual reality experiment
    Bazilinskyy, P., Alam, M. S., Merino-Martınez, R.
    54th International Institute of Noise Control Engineering (Internoise). São Paulo, Brazil (2025)

    The increasing adoption of electric vehicles (EVs), which operate more quietly than internal combustion engine vehicles, raises concerns about their detectability, particularly for visually impaired road users. Regulations mandate exterior sound signals for EVs, ensuring minimum sound pressure levels at low speeds. However, these signals are often used in already noisy urban environments, creating a challenge: enhancing detectability without adding excessive noise pollution. This study explores the use of synthetic exterior sounds that balance high noticeability with low annoyance. An audiovisual experiment was conducted with 20 participants in 15 virtual reality scenarios featuring an EV passing in front of them. Different sound signals, including pure, intermittent, and complex tones at varying frequencies, were tested alongside two baseline cases (a diesel engine and tyre noise alone, i.e., no synthetic sound added). Participants rated sounds for annoyance, noticeability, and informativeness using 11-point ICBEN scales. Trigger measurements provided additional insights into their willingness to cross in front of the EV. The results highlight optimal sound characteristics for EVs, offering guidance on improving pedestrian safety while minimising noise pollution. By refining exterior sound design, this research contributes to the development of effective and user-friendly EV sound standards, ensuring safer and more inclusive urban environments.
  • Psychoacoustic assessment of synthetic sounds for electric vehicles in a virtual reality experiment
    Bazilinskyy, P., Alam, M. S., Merino-Martınez, R.
    11th Convention of the European Acoustics Association (Euronoise). Málaga, Spain (2025)

    The growing adoption of electric vehicles, known for their quieter operation compared to internal combustion engine vehicles, raises concerns about their detectability, particularly for vulnerable road users. To address this, regulations mandate the inclusion of exterior sound signals for electric vehicles, specifying minimum sound pressure levels at low speeds. These synthetic exterior sounds are often used in noisy urban environments, creating the challenge of enhancing detectability without introducing excessive noise annoyance. This study investigates the design of synthetic exterior sound signals that balance high noticeability with low annoyance. An audiovisual experiment with 14 participants was conducted using 15 virtual reality scenarios featuring a passing car. The scenarios included various sound signals, such as pure, intermittent, and complex tones at different frequencies. Two baseline cases, a diesel engine and only tyre noise, were also tested. Participants rated sounds for annoyance, noticeability, and informativeness using 11-point ICBEN scales. The findings highlight how psychoacoustic sound quality metrics predict annoyance ratings better than conventional sound metrics, providing insight into optimising sound design for electric vehicles. By improving pedestrian safety while minimising noise pollution, this research supports the development of effective and user-friendly exterior sound standards for electric vehicles.
  • Vibe Coding in Practice: Building a Driving Simulator Without Expert Programming Skills
    Fortes-Ferreira, M., & Alam, M. S., Bazilinskyy, P.
    17th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (AutoUI). Brisbane, QLD, Australia (2025)

    The emergence of large language models has introduced new opportunities in software development, particularly through a revolutionary paradigm known as vibe coding or ``coding by vibes'', in which developers express their software ideas in natural language and where the LLM generates the code. This paper investigates the potential of vibe coding to support novice programmers. The first author, without coding experience, attempted to create a 3D driving simulator using the Cursor platform and Three.js. The iterative prompting process improved the simulation's functionality and visual quality. The results indicated that LLM can reduce barriers to creative development and expand access to computational tools. However, challenges remain: prompts often required refinements, output code can be logically flawed, and debugging demanded a foundational understanding of programming concepts. These findings highlight that while vibe coding increases accessibility, it does not completely eliminate the need for technical reasoning and understanding prompt engineering.

2024

  • Harnessing traditional controllers for fast-track training of deep reinforcement learning control strategies
    Alam, M. S., Carlucho, I.
    Journal of Marine Engineering & Technology (2024)

    In recent years, Autonomous Ships have become a focal point for research, with a specific emphasis on improving ship autonomy. Machine Learning Controllers, especially those based on Reinforcement Learning, have seen significant progress. However, addressing the substantial computational demands and intricate reward structures required for their training remains critical. This paper introduces a novel approach, “Leveraging Traditional Controllers for Accelerated Deep Reinforcement Learning (DRL) Training,” aimed at bridging conventional maritime control methods with cutting-edge DRL techniques for vessels. This innovative approach explores the synergies between stable traditional controllers and adaptive DRL methodologies, known for their complexity handling capabilities. To tackle the time-intensive nature of DRL training, we propose a solution: utilizing existing traditional controllers to expedite DRL training by transferring knowledge from these controllers to guide DRL exploration. We rigorously assess the effectiveness of this approach through various ship maneuvering scenarios, including different trajectories and external disturbances like winds. The results unequivocally demonstrate accelerated DRL training while maintaining stringent safety standards. This groundbreaking approach has the potential to bridge the gap between traditional maritime practices and contemporary DRL advancements, facilitating the seamless integration of autonomous systems into maritime operations, with promising implications for enhanced vessel efficiency, cost-effectiveness, and overall safety.
  • From A to B with ease: User-centric interfaces for shuttle buses
    Alam, M. S., Subramanian, T., Martens, M., Remlinger, W., Bazilinskyy, P.
    16th International Conference on Automotive User Interfaces and Interactive Vehicular Applications (AutoUI). Brisbane, QLD, Australia (2024)

    User interfaces are crucial for easy travel. To understand user preferences for travel information during automated shuttle rides, we conducted an online survey with 51 participants from 8 countries. The survey focused on the information passengers wish to access and their preferences for using mobile, private, and public screens during boarding and travelling on the bus. It also gathered opinions on the usage of Near-Field Communication (NFC) for shuttle bus confirmation and viewing assistance to help passengers stand precisely where the shuttle will arrive, overcoming navigation and language barriers. Results showed that 72.6% of participants indicated a need for NFC and 82.4% for viewing assistance. There was a strong correlation between preferences for shuttle bus schedules, route details (r=0.55), and next-stop information (r=0.57) on mobile screens, suggesting that passengers who value one type of information are likely to value related kinds too.

2023

  • AI on the water: Applying drl to autonomous vessel navigation
    Alam, M. S.,Sanjeev Kumar, R.S., Somayajula, A.
    Proceedings of the Sixth International Conference in Ocean Engineering (ICOE2023). Aachen, Germany (2023)

    Human decision-making errors cause a majority of globally reported marine accidents. As a result, automation in the marine industry has been gaining more attention in recent years. Obstacle avoidance becomes very challenging for an autonomous surface vehicle in an unknown environment. We explore the feasibility of using Deep Q-Learning (DQN), a deep reinforcement learning approach, for controlling an underactuated autonomous surface vehicle to follow a known path while avoiding collisions with static and dynamic obstacles. The ship's motion is described using a three-degree-of-freedom (3-DOF) dynamic model. The KRISO container ship (KCS) is chosen for this study because it is a benchmark hull used in several studies, and its hydrodynamic coefficients are readily available for numerical modelling. This study shows that Deep Reinforcement Learning (DRL) can achieve path following and collision avoidance successfully and can be a potential candidate that may be investigated further to achieve human-level or even better decision-making for autonomous marine vehicles.
  • Navigating the Ocean with DRL: Path following for marine vessels
    Jose, J., & Alam, M. S., Somayajula, A.S.
    Proceedings of the Sixth International Conference in Ocean Engineering (ICOE2023). Aachen, Germany (2023)

    Human error is a substantial factor in marine accidents, accounting for 85% of all reported incidents. By reducing the need for human intervention in vessel navigation, AI-based methods can potentially reduce the risk of accidents. AI techniques, such as Deep Reinforcement Learning (DRL), have the potential to improve vessel navigation in challenging conditions, such as in restricted waterways and in the presence of obstacles. This is because DRL algorithms can optimize multiple objectives, such as path following and collision avoidance, while being more efficient to implement compared to traditional methods. In this study, a DRL agent is trained using the Deep Deterministic Policy Gradient (DDPG) algorithm for path following and waypoint tracking. Furthermore, the trained agent is evaluated against a traditional PD controller with an Integral Line of Sight (ILOS) guidance system for the same. This study uses the Kriso Container Ship (KCS) as a test case for evaluating the performance of different controllers. The ship's dynamics are modeled using the maneuvering Modelling Group (MMG) model. This mathematical simulation is used to train a DRL-based controller and to tune the gains of a traditional PD controller. The simulation environment is also used to assess the controller's effectiveness in the presence of wind.
  • Data Driven Control for marine vehicle maneuvering
    Alam, M. S.
    Indian Institute of Technology Madras. Chennai, TN, India (2023)

    The majority of global marine accidents are caused by human decision-making errors, which has resulted in increased interest in automation within the marine industry. However, obstacle avoidance for autonomous surface vehicles in unknown environments is particularly difficult. This study investigates the possibility of utilizing a deep reinforcement learning (DRL) approach to control an underactuated autonomous surface vehicle following a predetermined path while avoiding collisions with static and dynamic obstacles. The ship’s movement is modelled using a three-degree-of-freedom (3-DOF) dynamic model, with the KRISO container ship (KCS) being selected for the study due to its extensive use in previous research and readily available hydrodynamic coefficients for numerical modelling. The study evaluates the performance of various DRL algorithms, such as Deep Q-Network (DQN), Deep Deterministic Policy Gradient (DDPG), and Proximal Policy Optimization (PPO) algorithms, for path following and their effectiveness in the presence of wind, as well as comparing them to the traditional PD controller. The study also explores DQN and DDPG algorithms for both static and dynamic obstacle avoidance and proposes a hybrid network that uses two networks for improved path following and obstacle avoidance capabilities.
  • Deep reinforcement learning based controller for ship navigation
    Deraj, R., Sanjeev Kumar, R. S., Alam, M. S., Somayajula, A.
    Ocean Engineering (2023)

    A majority of marine accidents that occur can be attributed to errors in human decisions. Through automation, the occurrence of such incidents can be minimized. Therefore, automation in the marine industry has been receiving increased attention in the recent years. This paper investigates the automation of the path following action of a ship. A deep Q-learning approach is proposed to solve the path-following problem of a ship. This method comes under the broader area of deep reinforcement learning (DRL) and is well suited for such tasks, as it can learn to take optimal decisions through sufficient experience. This algorithm also balances the exploration and the exploitation schemes of an agent operating in an environment. A three-degree-of-freedom (3-DOF) dynamic model is adopted to describe the ship’s motion. The Krisco container ship (KCS) is chosen for this study as it is a benchmark hull that is used in several studies and its hydrodynamic coefficients are readily available for numerical modeling. Numerical simulations for the turning circle and zig-zag maneuver tests are performed to verify the accuracy of the proposed dynamic model. A reinforcement learning (RL) agent is trained to interact with this numerical model to achieve waypoint tracking. Finally, the proposed approach is investigated not only by numerical simulations but also by model experiments using 1:75.5 scaled model.
  • Comparison of path following in ships using modern and traditional controllers
    Sanjeev Kumar, R. S., Alam, M. S., Reddy, B., Somayajula, A.S.
    Proceedings of the Sixth International Conference in Ocean Engineering (ICOE2023). Aachen, Germany (2023)

    Vessel navigation is difficult in restricted waterways and in the presence of static and dynamic obstacles. This difficulty can be attributed to the high-level decisions taken by humans during these maneuvers, which is evident from the fact that 85% of the reported marine accidents are traced back to human errors. Artificial intelligence-based methods offer us a way to eliminate human intervention in vessel navigation. Newer methods like Deep Reinforcement Learning (DRL) can optimize multiple objectives like path following and collision avoidance at the same time while being computationally cheaper to implement in comparison to traditional approaches. Before addressing the challenge of collision avoidance along with path following, the performance of DRL-based controllers on the path following task alone must be established. Therefore, this study trains a DRL agent using Proximal Policy Optimization (PPO) algorithm and tests it against a traditional PD controller guided by an Integral Line of Sight (ILOS) guidance system. The Krisco Container Ship (KCS) is chosen to test the different controllers. The ship dynamics are mathematically simulated using the Maneuvering Modelling Group (MMG) model developed by the Japanese. The simulation environment is used to train the deep reinforcement learning-based controller and is also used to tune the gains of the traditional PD controller. The effectiveness of the controllers in the presence of wind is also investigated.

Download all papers in bib file here.

* Joint first author.