For a while, enterprise AI seemed to have a simple rule: if a model could process more information, it would produce better answers. As context windows expanded from thousands of tokens to millions, bigger numbers became synonymous with progress, and in response, organisations fed AI assistants entire document repositories and conversation histories.
At Elastic, this shift was evident as customers moved from pilot programs into production. The challenge was never whether AI could handle more data, but whether it was reasoning over the right data.
Tokenmaxxing began as a Silicon Valley workplace trend in early 2026: maximising AI token consumption as a signal of productivity. Some companies built internal leaderboards ranking employees by tokens burned, and high usage became a status symbol regardless of output quality. The trend drew sharp criticism because it measures inputs, not outcomes. “It is the AI-era equivalent of judging developers by lines of code," as Ravindra Ramnani, Head of Field Engineering, India, Elastic puts it.
The same fallacy shows up at the architecture level, where teams include irrelevant documents and conversation histories in prompts, assuming more information improves output. Models don’t filter unnecessary material like skilled analysts, resulting in more noise, slower responses, and outputs needing significant human review.
The gap became clear once organisations moved into production, where an agent acting on incomplete context doesn't just cost more; it produces consequential errors. ' That question lies at the heart of context engineering ," says Ramnani.
Source link







