Enhancing driver experience in SAE level 3 automated vehicles through multimodal and emotion-aware in-vehicle agents
Zeng, X., Alam, M. S., Bazilinskyy, P.
In preparation.
ABSTRACT Multimodal in-vehicle agents (IVAs) combine voice, gesture, facial expression, and affect-responsive behaviour, but the appropriate role and limits of each channel remain unclear. We report two exploratory studies of a physically embodied robot-like IVA using pre-recorded highway-driving videos in a controlled Wizard-of-Oz style setup. Study~1 compared voice alone, voice with facial expressions, and voice with gestures (N=12). Descriptive workload averages were lower in the two non-verbal conditions (baseline: 33%, facial expressions: 23%, gestures: 18%), but the fixed baseline-first order prevents attributing this pattern to modality alone. Participant accounts instead revealed a clearer functional distinction: gestures were valued for noticeable, anticipatory, and action-oriented signalling, whereas facial expressions were valued for social, affective, and aesthetic communication. Study~2 used a revised prototype to explore affect-aware feedback under different user-availability conditions (N=12). Variable feedback exposure, perceived facial-expression misclassification, interruption timing, and mismatch with users' self-perceived states constrained the appropriateness of the interaction. Quantitative comparisons in both studies are treated as descriptive context for the qualitative findings, while secondary statistical checks are retained only for transparency. The contribution is a set of multimodal design boundaries rather than evidence of improved driver experience: voice for actionable information, gestures for interpretable anticipation, facial expressions for low-urgency social communication, and affect-aware feedback only when sensing confidence, user availability, and timing support its use.