{"version":1,"type":"story","url":"https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks","json":"https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks.json","markdown":"https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks.md","slug":"researchers-release-emailbench-to-test-ai-agents-on-email-tasks","headline":"Researchers release EmailBench to test AI agents on email tasks","summary":"A team of researchers has introduced **EmailBench**, a new benchmark designed to evaluate AI agents on enterprise email and productivity tasks. The framework includes 206 scenarios across 16 task categories, such as scheduling, expense tracking, and project coordination. It uses a synthetic, deterministic email corpus inspired by the Enron dataset and a hybrid evaluation protocol combining executable assertions and LLM-based rubrics.\n\nThe benchmark tests eight AI model configurations on a single-user corpus, revealing that even the best-performing agent passed only **33.5%** of scenarios despite **99.7%** of its tool calls succeeding. The researchers note that valid tool execution does not guarantee task completion, highlighting gaps in current AI agent capabilities. Future work includes expanding tool coverage, multi-persona testing, and repeated-run evaluations.","keyPoints":["EmailBench tests 206 email/productivity scenarios across 16 task categories, using a synthetic Enron-like corpus","Best AI agent passed only 33.5% of scenarios despite 99.7% of tool calls succeeding","Benchmark includes hybrid evaluation with executable assertions and LLM rubrics for task completion"],"whyItMatters":"EmailBench provides a rigorous, standardized way to measure AI agents’ real-world email productivity, exposing flaws in current tool execution vs. task success metrics.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-09-29T04:00:00Z","updatedAt":"2026-09-29T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks","url":"https://arxiv.org/abs/2609.31906","publishedAt":"2026-09-29T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers release EmailBench to test AI agents on email tasks\", 29 September 2026, https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks","publisher":"Digest AI","title":"Researchers release EmailBench to test AI agents on email tasks","datePublished":"2026-09-29T04:00:00Z","url":"https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}