MikhbarMIKHBAR
Artificial Intelligence

AI Agents Outpace Humans in Token Usage, Threatening RAM

Data from OpenRouter shows AI agents consuming significantly more tokens than human users, a trend expected to accelerate toward a tenfold increase.

AI Agents Outpace Humans in Token Usage, Threatening RAM

AI Agents Surpass Human Token Consumption

Artificial intelligence agents are driving a massive surge in token utilization, consuming five times more tokens than human users according to data highlighted by Futurum Group CEO Daniel Newman in an an X post published on September 30, 2026. Newman's insights draw from an Andreessen Horowitz (a16z) chart utilizing OpenRouter metrics, indicating that agent token usage has rapidly outpaced human interaction.

According to the underlying statistics featured in reporting by Tom's Hardware, agent activity officially surpassed human utilization in February. By August, agent token volume climbed to 7.3 trillion tokens, while human usage stood at 1.4 trillion. OpenRouter metrics indicate that since that crossover point, agent token consumption has multiplied by 14, whereas human usage has grown at a more modest 2.8x rate.

The OpenRouter logo in lime green and white on a black background
(Image credit: OpenRouter) · Source: Tom's Hardware

The Mechanics Behind the Surge: Cached Prompts

Despite the immense volume of tokens handled by autonomous agents, the vast majority of this activity does not represent new computation. Data from a16z and OpenRouter reveals that more than 85% of agent tokens originate from cached prompts. Essentially, automated agents are repeatedly rereading information they have already processed rather than generating entirely novel inputs.

This repetitive consumption pattern is reflected across various operational logs. For instance, a call center consultancy testing DeepSeek on rented Nvidia hardware observed that 96% of all input during September usage on Claude Code consisted of rereading previous conversation history. While cached tokens incur significantly lower processing costs compared to generating prompts from scratch, they introduce unique hardware and memory overhead constraints.

Broader Industry Adoption and Platform Insights

The rapid proliferation of automated agents is not isolated to a single gateway platform. According to McKinsey’s 2026 State of AI survey, 40% of respondents hailing from large organizations reported actively scaling AI agents, marking a notable jump from 27% recorded just one year prior.

OpenRouter classifies API keys into agentic, mixed, and human categories utilizing a complex seven-signal weighted composite score. This evaluation criteria incorporates metrics such as turn count, tool call rate, and gap timing. While mixed traffic categories have also experienced notable growth, the dominant trajectory points toward increasingly autonomous agent workflows scaling well beyond standard human interaction benchmarks.

OpenRouter chart of seven-day average token usage by agents, humans, and mixed traffic from September 2025 to August 2026
(Image credit: OpenRouter) · Source: Tom's Hardware

Implications for Global Memory and RAM Shortages

The heavy reliance on cached prompts means that models must continuously retain stored context within a KV cache. According to industry analyses, the demands placed by the KV cache are rapidly outgrowing available GPU high-bandwidth memory (HBM) capacity. Because memory manufacturers are prioritizing HBM production for AI data centers, traditional hardware markets face compounding supply constraints.

Micron has projected that RAM and storage shortages will worsen throughout 2027 and 2028, leading to increased component pricing for regular consumers. As predictions suggest agent token consumption could eventually escalate to ten times human levels or higher, regular PC buyers will increasingly find themselves competing against massive artificial intelligence workflows for limited global memory resources.

Sources

  • Tom's HardwareAI agents use 5x more tokens than humans as cached prompts explode, headed for 10x

Continue chronologically

Related entity coverage