Researchers release EmailBench to test AI agents on email tasks
A team of researchers has introduced EmailBench, a new benchmark designed to evaluate AI agents on enterprise email and productivity tasks. The framework includes 206 scenarios across 16 task categories, such as scheduling, expense tracking, and project coordination. It uses a synthetic, deterministic email corpus inspired by the Enron dataset and a hybrid evaluation protocol combining…
Key points
- EmailBench tests 206 email/productivity scenarios across 16 task categories, using a synthetic Enron-like corpus
- Best AI agent passed only 33.5% of scenarios despite 99.7% of tool calls succeeding
- Benchmark includes hybrid evaluation with executable assertions and LLM rubrics for task completion
The benchmark tests eight AI model configurations on a single-user corpus, revealing that even the best-performing agent passed only 33.5% of scenarios despite 99.7% of its tool calls succeeding. The researchers note that valid tool execution does not guarantee task completion, highlighting gaps in current AI agent capabilities. Future work includes expanding tool coverage, multi-persona testing, and repeated-run evaluations.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- OpenAI may unveil Aeon agent at DevDay to compete with Meta’s Muse · 2 src
- Manus launches Manus 2.0 and Cue with standalone email, phone, wallet · 1 src
- Cloudflare launches cf CLI to let agents use its full API · 1 src
- H company launches Holo4 agentic models for desktop and API tasks · 2 src
- Google to replace Gemini Gems with skills by November 17 · 2 src
Comments
via GitHub Discussions