{"version":1,"type":"story","url":"https://digestai.news/story/study-shows-large-language-models-can-classify-fine-grained-emotions-i","json":"https://digestai.news/story/study-shows-large-language-models-can-classify-fine-grained-emotions-i.json","markdown":"https://digestai.news/story/study-shows-large-language-models-can-classify-fine-grained-emotions-i.md","slug":"study-shows-large-language-models-can-classify-fine-grained-emotions-i","headline":"Study shows large language models can classify Fine-Grained emotions in app reviews","summary":"The paper evaluates how large language models (LLMs) can perform multi‑label emotion classification on mobile app reviews, a task useful for requirements engineering and feature prioritisation. Researchers compared encoder‑only fine‑tuning (both multi‑label and binary‑ensemble formulations) with decoder‑only zero‑ and few‑shot prompting across open‑source and proprietary models, and tested several class‑imbalance mitigation techniques such as loss reweighting, resampling and generative data augmentation.\n\nDecoder‑only few‑shot prompting achieved the highest macro‑F1 of **0.642**, far above the baseline encoder results (multi‑label 0.387, binary ensemble 0.450). Adding a synthetic‑review generator, positive‑weighted loss and data augmentation to the best encoder lifted its macro‑F1 by **+0.204**, narrowing the gap, and produced the largest gains on the rarest emotions, with improvements up to **+0.501 F1**. The encoder‑based approach also ran up to three orders of magnitude faster at inference time. The authors release the full experimental pipeline, synthetic corpora and fine‑tuned checkpoints for replication.\n\nThe study demonstrates that LLMs make fine‑grained, multi‑label emotion detection feasible for app‑review analysis, offering a modest but practical performance boost while keeping latency low, which can be integrated into existing engineering workflows.","keyPoints":["Decoder‑only few‑shot prompting reached macro‑F1 0.642, beating fine‑tuned encoders","Encoder with generative augmentation and weighted loss added +0.204 macro‑F1, closing most of the gap","Encoder inference latency up to three orders of magnitude lower; rare emotion F1 gains up to +0.501"],"whyItMatters":"Enables app developers to automatically detect nuanced user emotions, improving issue prioritisation and feature planning without heavy compute costs.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-10-06T04:00:00Z","updatedAt":"2026-10-06T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models","url":"https://arxiv.org/abs/2610.03802","publishedAt":"2026-10-06T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Study shows large language models can classify Fine-Grained emotions in app reviews\", 6 October 2026, https://digestai.news/story/study-shows-large-language-models-can-classify-fine-grained-emotions-i","publisher":"Digest AI","title":"Study shows large language models can classify Fine-Grained emotions in app reviews","datePublished":"2026-10-06T04:00:00Z","url":"https://digestai.news/story/study-shows-large-language-models-can-classify-fine-grained-emotions-i"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}