# Researchers test Profession-Specific prompts on science tasks with Gemini 3.8

Digest AI · Research · published 2026-10-02T04:00:00Z

Canonical: https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem

## Summary

A new study on arXiv evaluates whether detailed, profession-specific system prompts improve AI performance on scientific tasks. The authors tested **503 open-source AGENTS.md profiles**—designed for various scientific roles—against four controls: a minimal baseline, the profile’s opening role sentence, a generic scientific rigor guide, and an unrelated domain profile. They ran tests using **Gemini 3.8 Flash** via OpenRouter in the Pi agent harness across **9 text-based science benchmarks** with **4,531 sampled questions** and **100 matched profiles**. After API-error retries, **4,488 items** completed all five conditions, graded automatically.

The results show no consistent accuracy gain: matched profiles performed **0.6 percentage points worse** on average than the baseline (95% bootstrap interval: [-1.5, +0.2]). However, they generated **1.5–2.3 times more tokens** and cost **2.2–4.5 times more per successful call**. On **60 tool-using BioMysteryBench bioinformatics problems**, the profile-based approach solved only **46.7%** of tasks, compared to **56.7%** with the baseline—a **10.0 percentage-point drop** (95% interval: [-16.7, -3.3]). One exception emerged: longer prompts reduced API drops on **SuperGPQA**, improving first-pass accuracy from **54.0%** to **71.6%**, likely due to prompt length rather than domain expertise. The study concludes that loading full profession profiles by default does not improve accuracy and may increase costs.

## Key points

- 503 open-source AGENTS.md profiles tested against four controls in 9 science benchmarks with 4,531 questions
- matched profiles cost 2.2–4.5x more per call but showed no clear accuracy gain over baseline
- longer prompts reduced API failures on SuperGPQA but did not improve task-solving in bioinformatics

## Why it matters

The findings challenge assumptions about whether detailed, role-specific prompts enhance AI performance in scientific workflows, with potential cost and reliability trade-offs.

## Sources

1. [Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks](https://arxiv.org/abs/2610.00084) (arXiv cs.AI, 2026-10-02, primary source)

## Cite

Digest AI, "Researchers test Profession-Specific prompts on science tasks with Gemini 3.8", 2 October 2026, https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/researchers-test-profession-specific-prompts-on-science-tasks-with-gem.json
