Neil Movva, founder of Sale Research, is building a 'token factory' optimized for long-running AI agents (hours/days) rather than real-time chatbots. He argues that by focusing on cost per token over latency, using a scavenger strategy for cheap chips and power, AI can become 10x cheaper, enabling new use cases like deep research and autonomous cybersecurity. The conversation also covers GPU trade-offs, the KV cache bottleneck, and why the premium for frontier AI labs may not last.
Summarized by Podsumo
The shift from real-time chatbots to long-running background agents (hours/days) as the future of AI inference, prioritizing cost over latency.
The 'scavenger strategy': buying cheap chips (AMD, etc.) and intermittent power (solar/wind) at deep discounts, tolerating 95% uptime to build a distributed token factory.
Explains the fundamental GPU trade-off between latency and throughput, and how most inference stacks are built for the wrong use case.
KV cache (dynamic memory for conversation context) is a major hidden bottleneck; compression here could yield massive efficiency gains.
The premium for frontier AI labs (3-6 months ahead) may not endure due to open-source distillation and enterprise adoption lags.
"The best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. You didn't even have to ask for it."
"I want abundant tokens and diverse harnesses. Every company, every user even, should own their intelligence. Make the agent your own."
"The premium for being three to six months ahead may not last. Enterprise deployments don't move at that speed. A lot of enterprises are probably still on Opus 4.6 or 4.7."