Artificial Analysis has released AA-AgentPerf-Local — an open benchmark for the inference speed of agentic AI models on local hardware: laptops and workstations. This is the first public measurement of specifically agentic workloads on local machines: instead of a synthetic token stream, the test replays recorded real agent trajectories. The code, data, and ready-made serving configurations are published in an open repository, so you can already compare your own hardware, server, and model combination against a public leaderboard.


What happened
The benchmark replays eight recorded real agentic trajectories totaling 168 model steps, with the context growing to approximately 56,000 tokens per request during the session. Generation is deterministic: each step outputs exactly the recorded number of tokens, and tool execution is skipped by default to isolate inference speed. The test works with any OpenAI API-compatible server, including llama.cpp and vLLM. The code, data, and 14 serving configurations are fully open, all with speculative decoding: they are published in the ArtificialAnalysis/aa-agentperf-local repository. The first measurement covered four platforms: CUDA, ROCm, Vulkan-compatible Halo, and Metal, and four models in 4-bit quantization: Qwen3.5-9B, dense Qwen3.8-27B, the MoE model Qwen3.6-35B-A3B with three billion active parameters, and Ling 3.0 Flash with 124 billion parameters, of which five billion are active.
Context
Until now, the choice of local stack for agents relied on measurements of abstract tokens per second, which were weakly correlated with real agent work. An agentic session is structured differently: the context constantly grows, and in a single session, about 196,000 tokens are fed to the model as input, compared to approximately 31,000 generated, so the decisive stage becomes prefill, the processing of new input tokens, rather than decoding. Deterministic generation is fundamentally important in such a methodology: the variance in results between systems is attributed to the hardware and serving stack, not to sampling randomness, and the same test can be included in continuous integration as a performance regression to track speed degradation after updates.
Why this matters for the industry
For the AI industry, the first public benchmarks for agentic workloads on local hardware have appeared, and they revealed non-trivial mechanisms. The number of active parameters in a MoE model predicts speed better than the total size: Qwen3.6-35B-A3B with three billion active parameters is 2.5–3.3x faster than dense Qwen3.8-27B, but the rule is not absolute, since Ling 3.0 Flash with five active billion is still slower than dense Qwen3.5-9B, meaning architecture, implementation, and quantization are also important. Prefill accounted for 22–41% of end-to-end time, meaning local agentic sessions are bottlenecked by memory bandwidth when processing input. The gap between DGX Spark and Halo with identical specifications is explained by the maturity of the CUDA stack, and the measured 1.4–1.7x factor at equal price gives ROCm vendors a specific measurable goal. It has become cheaper for startups to test the "agent on user hardware" scenario: competition is shifting from the question of "is it possible locally" to the question of "what latency for what dollar," and on-prem hardware procurement can be checked against a public leaderboard instead of marketing figures.
Why this matters for users
The benchmark is open, and you can already run your own "hardware plus server plus model" combination today and compare the result with the published leaderboard, and the ready-made configurations from the repository lower the barrier to entry. From applicable conclusions: speculative decoding is present in all the best configurations and gives +30–120% to decoding speed beyond the bandwidth ceiling, so it should be enabled immediately. In agentic sessions, pauses during prefill are noticeable: on Ryzen AI Halo, an agent can "think" for 24 seconds, on a MacBook for 39 seconds, compared to about 2 seconds on an RTX 5090 with 1792 GB/s memory bandwidth. The Mac M5 Pro at $3700 is close to the Ryzen AI Halo on some models, while the market prices for DGX Spark and Halo are noticeably higher. The bottom line for the reader is simple: instead of guessing whether a local agent will work, you can calculate the latency budget for a specific combination and choose a configuration for your scenario.
What is still unknown / limitations
So far, the benchmark is described by a single source: the Artificial Analysis article and its repository, and the related Hacker News thread had zero comments and one point at the time of preparation, so there are no independent re-runs or external criticism of the methodology yet. The scope of the first measurement is also narrow: a limited set of platforms, four models, and a single quantization precision, so conclusions cannot be automatically transferred to other models and formats. In addition, the test isolates inference speed and does not execute tools, meaning the published numbers characterize the model, not the full latency of a real agentic cycle with actual tool calls.
Sources
- AA-AgentPerf-Local: Benchmarking Local AI Agents on Laptops & Workstations — Artificial Analysis article
- GitHub: ArtificialAnalysis/aa-agentperf-local — code, data, and 14 benchmark configurations
- Artificial Analysis — Laptops & Workstations Inference Leaderboard (measurement results)
- Hacker News — discussion of AA-AgentPerf-Local
Author
Look at AI, editorial team
