Synthetic personas underperform plain LLMs in predicting click-through, study finds
A new arXiv paper evaluates whether large language models (LLMs) used as synthetic personas can forecast real audience reactions to copy. The authors compare a ten‑persona panel, built from demographic data, against a zero‑shot baseline that simply asks the model how likely a typical reader is to click. Using the Upworthy Research Archive – thousands of headline A/B tests with measured…
Key points
- No‑persona baseline reached Kendall τ 0.361 and 49.2% top‑1 accuracy on 399 reliable A/B tests.
- Persona‑conditioned panel scored τ 0.084 and 34.6% top‑1 accuracy, worse than baseline.
- Findings replicated across three Gemini tiers, OpenAI gpt‑4.1, and a separate news dataset.
On this subset, the no‑persona baseline achieves a Kendall τ of 0.361 and top‑1 accuracy of 49.2%, while the persona‑conditioned panel records τ of 0.084 and top‑1 accuracy of 34.6%, with non‑overlapping confidence intervals. The advantage of the plain LLM holds across three Gemini tiers, OpenAI’s gpt‑4.1, and a separate news‑domain dataset, and is robust to seed, prompt phrasing, and model choice. The authors conclude that for aggregate engagement prediction, a simple LLM ranker outperforms persona simulation.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Semantic Routing Calibration mitigates LLM over-refusal · 1 src
- AIBuildAI-2.5 ranks first on MLE-Bench with 73.3% medal rate · 1 src
- Researchers fine-tune 406M model for meeting summaries with retrieved text spans · 1 src
- Researchers test how language models handle numerical formats in word problems · 1 src
- Author pretrains language model End-to-End in Rust for $164 · 1 src
Comments
via GitHub Discussions