Researchers test Profession-Specific prompts on science tasks with Gemini 3.8
A new study on arXiv evaluates whether detailed, profession-specific system prompts improve AI performance on scientific tasks. The authors tested 503 open-source AGENTS.md profiles—designed for various scientific roles—against four controls: a minimal baseline, the profile’s opening role sentence, a generic scientific rigor guide, and an unrelated domain profile. They ran tests using Gemini 3.8…
Key points
- 503 open-source AGENTS.md profiles tested against four controls in 9 science benchmarks with 4,531 questions
- matched profiles cost 2.2–4.5x more per call but showed no clear accuracy gain over baseline
- longer prompts reduced API failures on SuperGPQA but did not improve task-solving in bioinformatics
The results show no consistent accuracy gain: matched profiles performed 0.6 percentage points worse on average than the baseline (95% bootstrap interval: [-1.5, +0.2]). However, they generated 1.5–2.3 times more tokens and cost 2.2–4.5 times more per successful call. On 60 tool-using BioMysteryBench bioinformatics problems, the profile-based approach solved only 46.7% of tasks, compared to 56.7% with the baseline—a 10.0 percentage-point drop (95% interval: [-16.7, -3.3]). One exception emerged: longer prompts reduced API drops on SuperGPQA, improving first-pass accuracy from 54.0% to 71.6%, likely due to prompt length rather than domain expertise. The study concludes that loading full profession profiles by default does not improve accuracy and may increase costs.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- arXiv study finds reading LLM judges from first token overstates position bias · 1 src
- Study compares On-Device NER models for speed, cost and accuracy · 1 src
Comments
via GitHub Discussions