{"version":1,"type":"story","url":"https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem","json":"https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem.json","markdown":"https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem.md","slug":"researchers-test-profession-specific-prompts-on-science-tasks-with-gem","headline":"Researchers test Profession-Specific prompts on science tasks with Gemini 3.8","summary":"A new study on arXiv evaluates whether detailed, profession-specific system prompts improve AI performance on scientific tasks. The authors tested **503 open-source AGENTS.md profiles**—designed for various scientific roles—against four controls: a minimal baseline, the profile’s opening role sentence, a generic scientific rigor guide, and an unrelated domain profile. They ran tests using **Gemini 3.8 Flash** via OpenRouter in the Pi agent harness across **9 text-based science benchmarks** with **4,531 sampled questions** and **100 matched profiles**. After API-error retries, **4,488 items** completed all five conditions, graded automatically.\n\nThe results show no consistent accuracy gain: matched profiles performed **0.6 percentage points worse** on average than the baseline (95% bootstrap interval: [-1.5, +0.2]). However, they generated **1.5–2.3 times more tokens** and cost **2.2–4.5 times more per successful call**. On **60 tool-using BioMysteryBench bioinformatics problems**, the profile-based approach solved only **46.7%** of tasks, compared to **56.7%** with the baseline—a **10.0 percentage-point drop** (95% interval: [-16.7, -3.3]). One exception emerged: longer prompts reduced API drops on **SuperGPQA**, improving first-pass accuracy from **54.0%** to **71.6%**, likely due to prompt length rather than domain expertise. The study concludes that loading full profession profiles by default does not improve accuracy and may increase costs.","keyPoints":["503 open-source AGENTS.md profiles tested against four controls in 9 science benchmarks with 4,531 questions","matched profiles cost 2.2–4.5x more per call but showed no clear accuracy gain over baseline","longer prompts reduced API failures on SuperGPQA but did not improve task-solving in bioinformatics"],"whyItMatters":"The findings challenge assumptions about whether detailed, role-specific prompts enhance AI performance in scientific workflows, with potential cost and reliability trade-offs.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["Gemini 3.8 Flash"],"people":[]},"firstPublishedAt":"2026-10-02T04:00:00Z","updatedAt":"2026-10-02T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks","url":"https://arxiv.org/abs/2610.00084","publishedAt":"2026-10-02T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers test Profession-Specific prompts on science tasks with Gemini 3.8\", 2 October 2026, https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem","publisher":"Digest AI","title":"Researchers test Profession-Specific prompts on science tasks with Gemini 3.8","datePublished":"2026-10-02T04:00:00Z","url":"https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}