FinOps for the LLM era: why your cloud playbook doesn't translate
FinOps principles still apply, but token-based pricing breaks almost every traditional cost allocation pattern. Here's exactly what changes.
Cloud FinOps grew up on a simple unit: an instance-hour. You could tag it, allocate it, rightsize it, and reserve capacity against it. LLM spend has none of those properties.
A single feature ships across 6 providers, 14 models, and a routing layer that changes the mix every hour. Tags don't survive the proxy. Reservations don't exist. The unit of work — a token — is invisible to your existing FOCUS exports.
The teams getting this right are doing three things: instrumenting at the prompt level (not the API key level), allocating cost to features rather than services, and treating model selection as a continuous optimization rather than a quarterly procurement decision.