Research Papers research paper arxiv nlp language-models

Human-Guided Reasoning with Large Language Models for Vietnamese Speech Emotion Recognition

arXivApril 2, 20262 min read2 views

Vietnamese Speech Emotion Recognition (SER) remains challenging due to ambiguous acoustic patterns and the lack of reliable annotated data, especially in real-world conditions where emotional boundaries are not clearly separable. To address this problem, this paper proposes a human-machine collaborative framework that integrates human knowledge into the learning process rather than relying solely on data-driven models. The proposed framework is centered around LLM-based reasoning, where acoustic feature-based models are used to provide auxiliary signals such as confidence and feature-level evi — Truc Nguyen, Then Tran, Binh Truong

View PDF HTML (experimental)

Abstract:Vietnamese Speech Emotion Recognition (SER) remains challenging due to ambiguous acoustic patterns and the lack of reliable annotated data, especially in real-world conditions where emotional boundaries are not clearly separable. To address this problem, this paper proposes a human-machine collaborative framework that integrates human knowledge into the learning process rather than relying solely on data-driven models. The proposed framework is centered around LLM-based reasoning, where acoustic feature-based models are used to provide auxiliary signals such as confidence and feature-level evidence. A confidence-based routing mechanism is introduced to distinguish between easy and ambiguous samples, allowing uncertain cases to be delegated to LLMs for deeper reasoning guided by structured rules derived from human annotation behavior. In addition, an iterative refinement strategy is employed to continuously improve system performance through error analysis and rule updates. Experiments are conducted on a Vietnamese speech dataset of 2,764 samples across three emotion classes (calm, angry, panic), with high inter-annotator agreement (Fleiss Kappa = 0.8574), ensuring reliable ground truth. The proposed method achieves strong performance, reaching up to 86.59% accuracy and Macro F1 around 0.85-0.86, demonstrating its effectiveness in handling ambiguous and hard-to-classify cases. Overall, this work highlights the importance of combining data-driven models with human reasoning, providing a robust and model-agnostic approach for speech emotion recognition in low-resource settings.

Comments: 6 pages, 2 figures. Dataset of 2,764 Vietnamese speech samples across three emotion classes

Subjects:

Computation and Language (cs.CL)

Cite as: arXiv:2604.01711 [cs.CL]

(or arXiv:2604.01711v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2604.01711

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Phuoc Nguyen T. H. [view email] [v1] Thu, 2 Apr 2026 07:24:14 UTC (257 KB)

Original source

arXiv

https://arxiv.org/abs/2604.01711v1

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Market NewsFresh

[D] The memory chip market lost tens of billions over a paper this community would have understood in 10 minutes

TurboQuant was teased recently and tens of billions gone from memory chip market in 48 hours but anyone in this community who read the paper would have seen the problem with the panic immediately. TurboQuant compresses the KV cache down to 3 bits per value from the standard 16 using polar coordinate quantization. But the KV cache is inference memory. Training memory, activations, gradients, optimizer states, is a completely different thing and completely untouched. And majority of HBM demand comes from training. An inference compression paper doesn't move that number. And the commercial inference baseline already runs at 4 to 8 bit precision. The 6x headline is benchmarked against 16 bit full precision. The real marginal gain over what's actually deployed is considerably smaller than that

Reddit r/MachineLearning

2mabout 4 hours ago

Research PapersFresh

[D] Is research in semantic segmentation saturated?

Nowadays I dont see a lot of papers addressing 2D semantic segmentation problem statements be it supervised, semi-supervised, domain adaptation. Is the problem statement saturated? Are there any promising research directions in segmentation except open-set segmentation? submitted by /u/Hot_Version_6403 [link] [comments]

Reddit r/MachineLearning

1mabout 6 hours ago

Research Papers

AI, quantum computing, fusion energy remain Energy’s top research priorities - Nextgov/FCW

AI, quantum computing, fusion energy remain Energy’s top research priorities Nextgov/FCW

GNews AI quantum

1m4 months ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 150 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

Human-Guided Reasoning with Large Language Models for Vietnamese Speech Emotion Recognition

Submission history

Daily AI Digest

More about

[D] The memory chip market lost tens of billions over a paper this community would have understood in 10 minutes

[D] Is research in semantic segmentation saturated?

AI, quantum computing, fusion energy remain Energy’s top research priorities - Nextgov/FCW

Knowledge Map

Connected Articles — Knowledge Graph

Discussion

More in Research Papers

[D] Is research in semantic segmentation saturated?

AI, quantum computing, fusion energy remain Energy’s top research priorities - Nextgov/FCW

This Ancient Roman Game Board Was a Mystery. Researchers Used A.I. to Figure Out How to Play - Smithsonian Magazine

URI Day Highlights Student Research and the Future of AI Education in Rhode Island - uri.edu