Notes on running AI infrastructure, and on the cost of getting it wrong.
Cache multipliers vary 20x across models, and most cost tools treat them as a constant. The number stays plausible, which is why nobody catches it.