DigestAI news desk

Cut through the AI noise.

Research

Researchers test Profession-Specific prompts on science tasks with Gemini 3.8

A new study on arXiv evaluates whether detailed, profession-specific system prompts improve AI performance on scientific tasks. The authors tested 503 open-source AGENTS.md profiles—designed for various scientific roles—against four controls: a minimal baseline, the profile’s opening role sentence, a generic scientific rigor guide, and an unrelated domain profile. They ran tests using Gemini 3.8…

1 source primary source

Key points

  • 503 open-source AGENTS.md profiles tested against four controls in 9 science benchmarks with 4,531 questions
  • matched profiles cost 2.2–4.5x more per call but showed no clear accuracy gain over baseline
  • longer prompts reduced API failures on SuperGPQA but did not improve task-solving in bioinformatics

The results show no consistent accuracy gain: matched profiles performed 0.6 percentage points worse on average than the baseline (95% bootstrap interval: [-1.5, +0.2]). However, they generated 1.5–2.3 times more tokens and cost 2.2–4.5 times more per successful call. On 60 tool-using BioMysteryBench bioinformatics problems, the profile-based approach solved only 46.7% of tasks, compared to 56.7% with the baseline—a 10.0 percentage-point drop (95% interval: [-16.7, -3.3]). One exception emerged: longer prompts reduced API drops on SuperGPQA, improving first-pass accuracy from 54.0% to 71.6%, likely due to prompt length rather than domain expertise. The study concludes that loading full profession profiles by default does not improve accuracy and may increase costs.

Read the original at arXiv cs.AI · by Timothy Kassis primary sourceOpen source ↗
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories