Study shows large language models can classify Fine-Grained emotions in app reviews
The paper evaluates how large language models (LLMs) can perform multi‑label emotion classification on mobile app reviews, a task useful for requirements engineering and feature prioritisation. Researchers compared encoder‑only fine‑tuning (both multi‑label and binary‑ensemble formulations) with decoder‑only zero‑ and few‑shot prompting across open‑source and proprietary models, and tested…
Key points
- Decoder‑only few‑shot prompting reached macro‑F1 0.642, beating fine‑tuned encoders
- Encoder with generative augmentation and weighted loss added +0.204 macro‑F1, closing most of the gap
- Encoder inference latency up to three orders of magnitude lower; rare emotion F1 gains up to +0.501
Decoder‑only few‑shot prompting achieved the highest macro‑F1 of 0.642, far above the baseline encoder results (multi‑label 0.387, binary ensemble 0.450). Adding a synthetic‑review generator, positive‑weighted loss and data augmentation to the best encoder lifted its macro‑F1 by +0.204, narrowing the gap, and produced the largest gains on the rarest emotions, with improvements up to +0.501 F1. The encoder‑based approach also ran up to three orders of magnitude faster at inference time. The authors release the full experimental pipeline, synthetic corpora and fine‑tuned checkpoints for replication.
The study demonstrates that LLMs make fine‑grained, multi‑label emotion detection feasible for app‑review analysis, offering a modest but practical performance boost while keeping latency low, which can be integrated into existing engineering workflows.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic study: task understanding beats job title for AI success · 1 src
- Researchers test fixed token codes for language models at 100B-token scale · 1 src
- Baibaichuchu at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text? · 1 src
- Behavioral history outperforms descriptions for LLM synthetic personas · 1 src
- Researchers introduce ROAR to unify AI-driven research system runs · 1 src
Comments
via GitHub Discussions