Stop runaway AI agent spend before the invoice arrives.
llm-guard sits in front of your model API calls, attributes every dollar to the key that spent it, and detects a runaway loop while it is still running. Python standard library only — nothing to install, and nothing to audit but our code.
Runaway-spend detection
window: last 15 min · io ratio > 30:1 · velocity > 25x baseline
CRITICAL io_ratio [key: prod-agent]
prod-agent is running a 74:1 input-to-output ratio across 67 calls
in the last 15 min (12,006,400 in / 162,216 out, $38.45). Normal
traffic sits at 5:1-15:1.
→ what to do
The prompt is being re-sent in full on every step. Cap the context
(summarise or truncate history), and enable prompt caching so the
repeated prefix is billed at the cache rate. OWASP attributes ~62%
of agent bills to re-sent context.
WARN retry_storm [key: prod-batch]
prod-batch had 42 failed calls out of 56 (75%) in the last 15 min.Where the money actually goes
A budget dashboard tells you what you already spent. The failure that produces five-figure bills is a single agent chain looping, retrying or fanning out — and a daily cap fires hours after the money is gone.
Why it is built this way
Three decisions that shape everything else, including what the tool deliberately does not do.
No dependencies
A gateway sits on the request path and holds your API keys. No FastAPI, no httpx, no uvicorn — about 7,900 lines you can read in an afternoon, and a surface small enough to get past a security review.
Unknown over wrong
Cache reads cost between 2.5% and 50% of input depending on the model. Treating that as one constant is a silent four-to-five-times error. If a model has no price on file, the cost is recorded as NULL and reported as unpriced. A plausibly wrong number is worse than an obvious gap.
Advisory, not destructive
Detectors never block traffic. A detector that silently kills production is worse than the bill it prevents. The one exception is the optional per-stream cap, and the number it acts on is a local estimate, never used for billing.
What it catches
The failure patterns are not our invention. They come from OWASP AISVS C9.1 and a published catalogue of 63 confirmed production incidents across 21 orchestration frameworks. Each detection cites the pattern it matched, the number of recorded incidents, and the largest reported loss.
| Pattern | Observable signal | Recorded incidents |
|---|---|---|
| Unbounded context loop | Input/output ratio above 30:1 | 11 |
| Delegation fan-out race | Concurrent burst across many keys | 11 |
| Retry storm | Failure rate above 5% | 9 |
| Premium model as default | One model at 8x the average cost per call | — |
| Uncached repeated prefix | Cache hit rate below 35% | — |
| Velocity spike without a ceiling | Peak day at 8x the median day | — |
| No cost attribution | All spend on a single key | — |
| Unpriced traffic | Successful calls with no rate on file | — |
Run llm-guard patterns to see the signals, thresholds and citations in full.
Start in one command
No API key, no network, no signup. The demo data ships with a real 74:1 loop so the detector has something to find on the first run.
$ pip install git+https://github.com/leyao-daily/llm-guard.git
$ python3 -m llmguard seed --reset --compare-days 30
Seeded 13,254 requests spanning 60 days ($227.10 of simulated spend).
Last 30 days is the reporting window; the preceding 30 days gives the
period-over-period delta.
$ python3 -m llmguard anomalies
$ python3 -m llmguard diagnose --client "Acme Corp" --out report.htmlIf your AI bill is unpredictable
We find out why. A fixed-price diagnosis delivered as a written report: what is driving it, what to change first, and what each change is worth per month. No infrastructure to deploy, read-only access to usage data, and no prompts or completions ever leave your account.
If you would rather run it yourself
llm-guard is MIT licensed and complete. Zero runtime dependencies, standard library only, budgets and runaway-loop detection included. There is no crippled edition and no feature held back for paying customers.
We are a small team in Hong Kong and would rather talk to people running this in production than ship a newsletter. [email protected]