AI agents now consume more tokens than humans on OpenRouter
AI agents have consumed more tokens than human users on OpenRouter since February 2025, with agentic usage growing 14x since then versus 2.8x growth in human usage. Nearly 70% of agent token consumption comes from cheap cached prompts, so actual infrastructure costs are rising more slowly than the raw token volume suggests.
Why it matters: This is early hard evidence that agentic workloads, not chat, are becoming the dominant driver of inference demand — which has direct implications for how model providers price caching, and for the memory and compute capacity planning behind the current AI infrastructure buildout. It also complicates simple 'AI usage is exploding' narratives, since heavy caching means costs aren't scaling as fast as raw token counts imply.