Artificial Analysis releases open-source tool to benchmark local AI agents on laptops and workstations
Artificial Analysis has launched AA-AgentPerf-Local, an open-source benchmarking tool designed to measure how fast AI agents perform on laptops and workstations. The tool replays real agent trajectories—eight recorded tasks spanning 168 model turns—with a context window growing to ~56K tokens. It isolates inference speed by default but can also simulate tool delays or live CPU calls.
Key points
- AA-AgentPerf-Local benchmarks local AI agent performance on laptops/workstations with 8 recorded tasks (~56K tokens)
- RTX 5090 fastest for models fitting in 32GB; DGX Spark outperforms Ryzen AI Halo due to compute efficiency
- Tool supports any OpenAI-compatible server and will expand to multi-agent scenarios and live CPU tool-calling
The initial results cover four hardware setups: NVIDIA DGX Spark (128GB), RTX 5090 (32GB), AMD Ryzen AI Halo (128GB), and MacBook Pro M5 Pro (64GB). Benchmarks focus on models like Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash (124B/5B active), all tested at 4-bit quantization. The RTX 5090 emerged as the fastest system for models fitting in its memory, while the DGX Spark outperformed the Ryzen AI Halo due to higher compute efficiency. The MacBook Pro (M5 Pro) showed competitive results, finishing within 2–21% of the Ryzen AI Halo for some models. The tool also highlights how prefill speeds—reading new input tokens—can dominate latency, especially on systems with lower compute-to-memory bandwidth ratios. Future updates will expand hardware coverage, add multi-agent scenarios, and include user-submitted leaderboards.
The story so far
2 episodes →- Artificial Analysis releases open-source tool to benchmark local AI agents on laptops and workstationsthis story
AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations
artificialanalysis.ai · 30 September 2026
Loading the full article…
This text was published by artificialanalysis.ai and written by Artificial Analysis. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Hardware & Compute
All →- Ollama cheat sheet guides local AI model management · 1 src
- OpenAI VP details Jalapeño ASIC’s AI-assisted design and efficiency focus · 4 src
- DeepSeek says it has partnered with Huawei to develop Ascend chip programming tools · 1 src
- AMD says PerfOpt in Linux 7.4 may boost AI/LLM performance on Radeon iGPUs by up to 23% · 1 src
- Netlist sues for US import ban on Micron chips in Google, Nvidia AI hardware · 1 src
Comments
via GitHub Discussions