{"ID":23475076,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.19826","arxiv_id":"2609.19826","title":"Consensus-Guided Shared-Specific Tri-View Learning for Speech Emotion Recognition","abstract":"Speech emotion recognition (SER) benefits from heterogeneous acoustic representations, but views derived from the same utterance contain both overlapping emotional evidence and representation-dependent cues. Direct fusion may therefore propagate redundant information or obscure complementary details. To address this issue, we propose Tri-view Consensus-Guided Fusion (TriCGF) for jointly modeling spectrogram, Mel-frequency cepstral coefficients, and HuBERT representations. TriCGF organizes each view into common and view-specific components before fusion. Cross-view Consensus Learning aggregates the common components into a global reference, while View-wise Gated Integration adaptively combines this reference with each view-specific component. A soft difference regularizer further discourages excessive information overlap. Under speaker-independent evaluation, TriCGF achieves 74.19% weighted accuracy (WA) and 75.17% unweighted accuracy (UA) on IEMOCAP, and 94.36% WA and 94.28% UA on EmoDB, outperforming representative SER methods on both datasets.","short_abstract":"Speech emotion recognition (SER) benefits from heterogeneous acoustic representations, but views derived from the same utterance contain both overlapping emotional evidence and representation-dependent cues. Direct fusion may therefore propagate redundant information or obscure complementary details. To address this is...","url_abs":"https://arxiv.org/abs/2609.19826","url_pdf":"https://arxiv.org/pdf/2609.19826v1","authors":"[\"Bing Huang\",\"Yujian Ma\",\"Xikun Lu\",\"Xianquan Jiang\",\"Jinqiu Sang\"]","published":"2026-09-17T07:32:38Z","proceeding":"eess.AS","tasks":"[\"eess.AS\"]","methods":"[\"Generative Adversarial Network\"]","has_code":false}
