AI agents consuming five times more tokens than humans as prompt caching drives exponential growth
According to OpenRouter platform data analyzed by Futurum Group's CEO, machine-driven AI agents have dramatically outpaced human usage, with projections suggesting the gap will widen further as cached prompts dominate token consumption.

Machine learning agents are now processing substantially more tokens than human users on the OpenRouter platform, marking a significant shift in how artificial intelligence systems are being deployed. Futurum Group Chief Executive Daniel Newman highlighted this trend in a recent social media post, noting that the disparity between agent and human token consumption continues to accelerate.
Newman stated: "AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher." His analysis draws from data compiled by Andreessen Horowitz using OpenRouter figures, which recorded agents consuming 7.3 trillion tokens against 1.4 trillion tokens for human users as of August. This represents a significant milestone reached just six months after agent usage first exceeded human usage on the platform.
The surge in agent token consumption is largely driven by cached prompts, which account for more than 85 percent of all agent tokens according to Andreessen Horowitz's analysis of OpenRouter data. OpenRouter, described as "a leading AI model gateway and routing platform," categorizes API keys into three distinct groups—agentic, mixed, or human—using what it calls a "7-signal weighted composite score that includes inputs such as tool call rate, turn count, gap timing, and others."
Peter Walker, OpenRouter's head of insights, released data showing that since agents surpassed humans in February, agent token usage has grown 14 times while human usage increased 2.8 times. The mixed category, which may represent behavior spanning both agent and human characteristics, expanded 4.7 times during the same period. The growth trajectory has not been entirely linear, with usage dips recorded in April and July, though the overall trend remains upward.

Newman added context to the implications: "We keep speaking to human adoption when trying to determine ROI, but the utilization and scale is exponentially larger than that." The observation underscores how enterprise focus on human user metrics may be missing the broader picture of AI system deployment and resource consumption.
Evidence of this pattern extends beyond OpenRouter. A call center consultancy that evaluated DeepSeek using rented Nvidia H200 hardware found that in its own agents' September usage on Claude Code, "96% of all input was re-reading old conversation." This alignment with OpenRouter's findings suggests that while token counts may overstate actual computational costs, the underlying memory requirements remain substantial.
Cached tokens are less expensive to process than fresh prompts, but they must still be maintained in memory. Andreessen Horowitz, which holds an investment stake in OpenRouter, has connected this memory demand to growing pressure on high-bandwidth memory supplies. AI models store this cached context in structures known as KV caches, and according to industry reporting, "the KV cache is outgrowing GPU HBM capacity."
McKinsey's 2026 State of AI survey provides additional confirmation that agent deployment is expanding across the industry. The survey found that 40 percent of respondents from large organizations reported scaling AI agents, compared to 27 percent in the previous year, indicating that the OpenRouter trend reflects broader enterprise adoption patterns.
Memory constraints are already emerging as a bottleneck. Micron has projected that RAM and storage shortages will intensify during 2027 and 2028, with customers facing higher prices while memory manufacturers prioritize high-bandwidth memory allocations for AI data centers. If Newman's projection holds true and the agent-to-human token ratio reaches "10X, 20X, 30X," competition for memory resources between AI systems and traditional computing will intensify, potentially affecting availability and pricing for personal computer buyers.