
Research Article
EmoFedProto: Privacy-Preserving Vietnamese Speech Emotion Recognition via Prototype-Based Federated Learning
@ARTICLE{10.4108/airo.11595, author={Quang-Anh Nguyen-Duc and Duc Minh Pham and Thai Dinh Kim and Thao Phuong Pham and Minh-Anh Nguyen and Xuan-Hai Le and Van-Ninh Nguyen}, title={EmoFedProto: Privacy-Preserving Vietnamese Speech Emotion Recognition via Prototype-Based Federated Learning}, journal={EAI Endorsed Transactions on AI and Robotics}, volume={5}, number={1}, publisher={EAI}, journal_a={AIRO}, year={2026}, month={4}, keywords={Speech Emotion Recognition, Federated Learning, Prototype-Based Learning, Non-IID Data, Low-Resource Languages, Vietnamese Speech}, doi={10.4108/airo.11595} }- Quang-Anh Nguyen-Duc
Duc Minh Pham
Thai Dinh Kim
Thao Phuong Pham
Minh-Anh Nguyen
Xuan-Hai Le
Van-Ninh Nguyen
Year: 2026
EmoFedProto: Privacy-Preserving Vietnamese Speech Emotion Recognition via Prototype-Based Federated Learning
AIRO
EAI
DOI: 10.4108/airo.11595
Abstract
Speech Emotion Recognition (SER) plays a fundamental role in affective computing by enabling machines to infer human emotional states from vocal expressions. However, most existing SER systems rely on centralized training paradigms, which raise serious privacy concerns due to the sensitive nature of speech data. Federated Learning (FL) offers a privacy-preserving alternative by allowing collaborative model training without sharing raw data, yet its performance often degrades significantly under non-IID data distributions, a common characteristic of speech emotion datasets caused by speaker variability and emotion imbalance. To address these challenges, we propose EmoFedProto, a prototype-based federated learning framework with clustering-enhanced prototype aggregation tailored for Vietnamese speech emotion recognition in low-resource settings. Instead of exchanging full model parameters, EmoFedProto communicates class-level feature prototypes, enabling more robust alignment across heterogeneous clients. Experiments conducted on the VNEMOS dataset under realistic non-IID and few-shot conditions demonstrate that EmoFedProto achieves an accuracy of 0.875, outperforming the baseline FedProto (0.825), while reducing performance variability by 44%. These results indicate that clustering-based prototype federated learning is an effective and communication-efficient solution for privacy-preserving speech emotion recognition, particularly in low-resource languages and realworld federated environments.
Copyright © 2026 Quang-Anh Nguyen-Duc et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.


