CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion Recognition. Hou, B., Li, R., Zhu, Y., Yu, L., & Wang, Y. August, 2026.
abstract   bibtex   
Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal conflict. We propose CONFER, a graph-based conflict-aware evidence negotiation framework for weakly supervised multimodal emotion recognition. CONFER represents each modality expert as a node with a predictive belief, boundary-based uncertainty, and runtime reliability estimated from historical out-of-fold performance and current-sample uncertainty. Uncertaintyaware compatibility and reliability-directed asymmetric edge weights govern iterative message-passing negotiation, followed by peer-supported prediction readout. Conflict reduction, residual disagreement, and mean modality uncertainty further characterize three regimes—Consensus, Dissent, and Ambiguity—for sample-specific weak-label calibration. We evaluate CONFER on AMIGOS, MAHNOB-HCI, and DEAP under subject-dependent 10-fold and strict leave-one-subjectout (LOSO) protocols. CONFER achieves competitive performance, reaching 0.873 accuracy on AMIGOS-V and 0.854 accuracy on MAHNOB-V under strict LOSO evaluation. Further analyses show larger negotiation gains on high-conflict samples and improved robustness to weak-label corruption, indicating that cross-modal conflict provides useful information for both directional modality coordination and supervisionreliability estimation.
@article{hou_confer:_2026,
	title = {{CONFER}: {Conflict}-{Aware} {Evidence} {Negotiation} for {Regime}-{Calibrated} {Weak} {Supervision} in {Multimodal} {Emotion} {Recognition}},
	abstract = {Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal conflict. We propose CONFER, a graph-based conflict-aware evidence negotiation framework for weakly supervised multimodal emotion recognition. CONFER represents each modality expert as a node with a predictive belief, boundary-based uncertainty, and runtime reliability estimated from historical out-of-fold performance and current-sample uncertainty. Uncertaintyaware compatibility and reliability-directed asymmetric edge weights govern iterative message-passing negotiation, followed by peer-supported prediction readout. Conflict reduction, residual disagreement, and mean modality uncertainty further characterize three regimes—Consensus, Dissent, and Ambiguity—for sample-specific weak-label calibration. We evaluate CONFER on AMIGOS, MAHNOB-HCI, and DEAP under subject-dependent 10-fold and strict leave-one-subjectout (LOSO) protocols. CONFER achieves competitive performance, reaching 0.873 accuracy on AMIGOS-V and 0.854 accuracy on MAHNOB-V under strict LOSO evaluation. Further analyses show larger negotiation gains on high-conflict samples and improved robustness to weak-label corruption, indicating that cross-modal conflict provides useful information for both directional modality coordination and supervisionreliability estimation.},
	language = {en},
	author = {Hou, Bojing and Li, Ruohao and Zhu, Yitong and Yu, Luwen and Wang, Yuyang},
	month = aug,
	year = {2026},
}

Downloads: 0