🛠 AA-AgentPerf-Local: Testing agents on local hardware
Artificial Analysis released an open benchmark for agentic inference speed on laptops and workstations: 8 real trajectories, 168 model steps, context up to ~56k tokens. Works with any OpenAI-compatible server, code is published.
🌍 First public measurement of such workloads: RTX 5090 is more than 3.5x faster than all others (1792 GB/s), DGX Spark outperformed Ryzen AI Halo by 1.4–1.7x at a price of $4000. For MoE, speed predicts the number of active parameters, not the total size.
👤 Run your own "hardware + llama.cpp/vLLM + model" setup and compare with the leaderboard. Best configs — with speculative decoding: +30–120%. Mac M5 Pro ($3700) is close to Ryzen AI Halo. No independent verification yet.
Source 1: https://artificialanalysis.ai/articles/aa-agentperf-local Source 2: https://github.com/ArtificialAnalysis/aa-agentperf-local
