# Researchers release EmailBench to test AI agents on email tasks

Digest AI · Agents & Tools · published 2026-09-29T04:00:00Z

Canonical: https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks

## Summary

A team of researchers has introduced **EmailBench**, a new benchmark designed to evaluate AI agents on enterprise email and productivity tasks. The framework includes 206 scenarios across 16 task categories, such as scheduling, expense tracking, and project coordination. It uses a synthetic, deterministic email corpus inspired by the Enron dataset and a hybrid evaluation protocol combining executable assertions and LLM-based rubrics.

The benchmark tests eight AI model configurations on a single-user corpus, revealing that even the best-performing agent passed only **33.5%** of scenarios despite **99.7%** of its tool calls succeeding. The researchers note that valid tool execution does not guarantee task completion, highlighting gaps in current AI agent capabilities. Future work includes expanding tool coverage, multi-persona testing, and repeated-run evaluations.

## Key points

- EmailBench tests 206 email/productivity scenarios across 16 task categories, using a synthetic Enron-like corpus
- Best AI agent passed only 33.5% of scenarios despite 99.7% of tool calls succeeding
- Benchmark includes hybrid evaluation with executable assertions and LLM rubrics for task completion

## Why it matters

EmailBench provides a rigorous, standardized way to measure AI agents’ real-world email productivity, exposing flaws in current tool execution vs. task success metrics.

## Sources

1. [EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks](https://arxiv.org/abs/2609.31906) (arXiv cs.AI, 2026-09-29, primary source)

## Cite

Digest AI, "Researchers release EmailBench to test AI agents on email tasks", 29 September 2026, https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/researchers-release-emailbench-to-test-ai-agents-on-email-tasks.json
