{"version":1,"type":"story","url":"https://digestai.news/story/new-framework-splits-llm-agent-roles-in-factor-research","json":"https://digestai.news/story/new-framework-splits-llm-agent-roles-in-factor-research.json","markdown":"https://digestai.news/story/new-framework-splits-llm-agent-roles-in-factor-research.md","slug":"new-framework-splits-llm-agent-roles-in-factor-research","headline":"New framework splits LLM agent roles in factor research","summary":"A new arXiv paper proposes a \"governed self-evolution\" framework for language model agents in quantitative finance. The core idea is to separate the roles of proposing investment factors and judging their validity. The agent is allowed to propose factors and write diagnostic probes, but a frozen statistical referee, which the agent cannot modify, must judge them. This referee uses betting on market outcomes to ensure false-discovery guarantees hold at any stopping time.\n\nThe authors tested this setup with three proposers (a script, a bandit, and a language model) against both the frozen referee and three \"leaky\" alternatives. They used a synthetic environment with known truths and a ten-year walk-forward test on the CSI 500 index. The results show that the frozen referee admitted 5-11 times fewer sub-threshold factors than the leaky referees when using a scripted proposer. No proposer type was able to close this gap in false admissions.\n\nRegarding performance, the language model proposer outperformed the script and matched the bandit, with the added benefit of creating its own diagnostic probes. However, the strict judging process comes with a time cost: true factors wait about 500 trading days for admission, causing the certified portfolio's Sharpe ratio to trail that of an ungated one. The study concludes that judging should remain a procedural function, while proposing and instrumentation belong to the agent.","keyPoints":["Frozen statistical referee admits 5-11 times fewer false factors than leaky alternatives.","Language model proposers match bandit performance and add capability to write diagnostic probes.","Strict validation causes true factors to wait about 500 trading days for admission."],"whyItMatters":"This research offers a rigorous method to prevent AI agents from overfitting or hallucinating valid financial signals, addressing a key trust barrier in autonomous quantitative trading systems.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-09-24T04:00:00Z","updatedAt":"2026-09-24T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors","url":"https://arxiv.org/abs/2609.27051","publishedAt":"2026-09-24T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"New framework splits LLM agent roles in factor research\", 24 September 2026, https://digestai.news/story/new-framework-splits-llm-agent-roles-in-factor-research","publisher":"Digest AI","title":"New framework splits LLM agent roles in factor research","datePublished":"2026-09-24T04:00:00Z","url":"https://digestai.news/story/new-framework-splits-llm-agent-roles-in-factor-research"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}