SpendTensor Journal

Field notes on AI cost optimization

Engineering deep dives, FinOps research, customer stories, and product updates from the team running $2.4B of LLM spend.

AI spend management: the 2026 playbook to control LLM costs without slowing your team
Guide

AI spend management: the 2026 playbook to control LLM costs without slowing your team

Why token-based pricing breaks traditional FinOps, and the 7-step framework top engineering teams use to cut LLM bills 30–60% while shipping faster.

Helena Voss 14 min
Grok
Research

Grok 3 vs GPT-4o: cost and performance benchmark for FinOps teams

Grok 3 vs GPT-4o on cost per million tokens, latency, and quality scores — with routing recommendations for engineering and FinOps leaders.

Priya Natarajan 10 min
OpenAI
Engineering

OpenAI vs. Anthropic prompt caching: which delivers better AI cost reduction?

OpenAI vs Anthropic prompt caching for high-volume workloads: pricing math, break-even thresholds, and which provider wins for your traffic pattern.

Marcus Lee 11 min
How
Customer Story

How Parallax cut their Claude bill 47% without a single customer complaint

The routing rules, prompt caching, and quality guardrails that took $1.2M/yr in Anthropic spend down to $640k — with faster responses.

Anika Roy 9 min
The
Engineering

The real cost of prompt caching (and when it actually pays off)

The break-even math on prompt caching is non-obvious. We benchmarked 14 production workloads to find when caching saves money — and when it costs more.

Marcus Lee 12 min
FinOps
FinOps

FinOps for the LLM era: why your cloud playbook doesn't translate

FinOps principles still apply, but token-based pricing breaks almost every traditional cost allocation pattern. Here's exactly what changes.

Helena Voss 8 min
Routing
Engineering

Routing without regret: a quality-aware fallback design

The hardest part of a router isn't picking the cheapest model — it's deciding when to escalate. We open-sourced the eval harness we use internally.

Daniel Okonkwo 11 min
Forecasting
Research

Forecasting AI spend with cohort modeling

Linear extrapolation is why your end-of-month bill keeps surprising you. Cohort-based forecasts predict the next invoice within ±3%.

Priya Nair 7 min
Introducing
Product

Introducing Semantic Cache v2 — now with embedding-aware invalidation

Our second-generation semantic cache lands today: embedding-aware invalidation, per-tenant namespacing, and 38% fewer input tokens on average.

Theo Bauer 5 min

Get the weekly drop

One email every Tuesday with the best new writing on LLM cost, routing, and FinOps. No spam.