Your Agentic Workflow's Cache Keepalive Costs 8x Too Much
Summary
Maxim Khailo empirically benchmarks KV-cache keepalive strategies across four major LLM providers (Anthropic, OpenAI, Gemini, DeepSeek) and finds the industry-standard 30-second ping interval costs 8× more than necessary. The optimal interval is ~4 minutes, and whether a keepalive is financially worthwhile at all depends entirely on the provider's cache eviction behavior and pricing structure. At a 10-minute pause, only Anthropic's pricing model makes the keepalive a money-saving decision.
Key Insight
The widely copied 30-second cache keepalive interval is an expensive folk convention — the mathematically optimal interval is ~4 minutes, keepalives only save money on Anthropic at typical pause lengths, and past a provider-specific break-even horizon the disciplined move is to stop pinging entirely and pay the re-prefill.
Spicy Quotes (click to share)
- 7
The convention is to ping every 30 seconds. That convention costs 8× more than necessary, and the surprise is bigger: at the ten-minute pause I measured, only one of the four major providers saved money with a keepalive.
- 7
The convention isn't cautious. It's just expensive.
- 4
Pay one premium every 4 minutes and you can afford twelve of them before the premiums exceed the claim. Twelve premiums at 4 minutes apart is 46 minutes of pause. That is the line.
- 6
The arbitrage is real, and it has an expiry date.
- 3
The prescription is per-provider, not universal.
Tone
analytical
