Modality Discrepancy Transformer reaches 0.7408 Macro F1 on BAH dataset
A new paper on arXiv proposes the Modality Discrepancy Transformer (MDT) to detect ambivalence and hesitancy in clinical videos, where facial, vocal, and linguistic cues conflict. MDT expands a prior 6‑token design to nine tokens: three modality embeddings, three absolute‑difference features, and three Hadamard‑product discrepancy features, processed by Transformer self‑attention. The…
Key points
- MDT uses a 9-token multimodal representation: three modality embeddings, three absolute-difference, three Hadamard-product features.
- Achieves 0.7408 Macro F1 on BAH test split, 0.7368 on private leaderboard, over 10 points above prior best.
- Trains in under 20 minutes on a single GPU using FiLM‑based text modulation and LoRA fine‑tuning.
On the BAH dataset from the 3rd ABAW Challenge, MDT attains a Macro F1 of 0.7408 on the labelled test split and 0.7368 on the private leaderboard, surpassing the strongest published baseline by more than ten points. Training completes in under 20 minutes on a single GPU, demonstrating both high performance and efficiency for cross‑modal affective state recognition.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Generative AI & Models
All →- OpenAI's GPT-6 Astra: Revolutionizing Work and Research · 22 src
- DeepSeek releases V4.1-Flash with 1M context and MIT license · 6 src
- FreedomIntelligence releases HuatuoGPT-3-9B medical LLM with open usage guides · 1 src
- Google moves Chrome to a 2‑week update cycle to narrow patch gap · 1 src
- Google launches Gemini 3.8 Live and Extended Thinking voice models · 7 src
Comments
via GitHub Discussions