Earlier this year, leaders at Amazon Web Services delivered a new mandate to their engineers: They need to conserve CPU cycles at all costs. AWS has reportedly experienced an explosion in wait times for CPU server capacity as AI workloads strain the company’s cloud infrastructure.
The issue seemingly took AWS off guard, and for good reason. The AI boom led to a surge in demand for GPUs and, later, memory. CPUs were mostly left out of the story, as their relative lack of parallelization made them a poor fit for AI model inference, the process of running and serving large language models (LLM) to users.
But the rise of agentic AI systems, which allow AI models to operate autonomously and call on sub-agents, is changing the narrative.
Matt Kimball, vice president and principal data center analyst at Moor Insights & Strategy in Austin, Texas, says 2026 has brought a spike in CPU demand, much of it due to agentic AI. “It’s one thing to have this agentic workload, and let’s say it spawns 100 agents. If I’m going to roll this out across my enterprise, those 100 become tens of thousands, hundreds of thousands, or millions of agents,” Kimball says. “You have agents spawning sub-agents, making API [application programming interface] calls and talking to more agents through [Anthropic’s] model context protocol.”
Kimball’s comments refer in part to “tool use,” which is shorthand for an LLM’s ability to access the internet, open files on a desktop, and generally use a variety of software to accomplish its task.
Source link







