llm-guard

A self-hosted gateway that sits in front of your OpenAI, Anthropic or Gemini calls. It records what each request actually cost, attributes that cost to the API key, project and end user behind it, detects runaway agent spend while it is still happening, and enforces hard budgets.

It runs on the Python standard library. No pip install, no virtualenv, no third-party package anywhere in the request path.

Source on GitHub → · Read the notes →


The problem it solves

A budget that checks between requests cannot stop the request that is crossing the line. For a short chat call that hardly matters. For a 200-step agent loop it is the entire problem: OWASP records that a 200-step loop costs more than 100x a single call, and that 62% of agent bills come from context being re-sent on every step.

llm-guard watches for three shapes of runaway spend:

Detection only ever reports. Enforcement is a separate, explicit decision: a per-key budget with action=block returns HTTP 429, and optional per-stream caps abort a single streaming response mid-flight.

What makes it different

Zero runtime dependenciesStandard library only. Auditable in one sitting, which matters because a gateway on the request path is a high-value target.
Cost from billing truthPrices come from the usage block the provider returned, never from a local token estimate.
Per-model cache economicsCache reads are billed at 2.5% to 50% of input depending on the model. Treating that as a constant is a silent 4-5x error.
No data leaves your networkNo telemetry, no CDN, no phone-home. The dashboard is server-rendered SVG.
Unknown is a valid answerA model with no price on file is reported as unpriced, not estimated.

Try it in thirty seconds

No API key, no network access, no signup:

git clone https://github.com/leyao-daily/llm-guard.git
cd llm-guard
python3 -m llmguard seed --reset --compare-days 30
python3 -m llmguard anomalies
python3 -m llmguard dashboard --out dash.html

The demo dataset contains a real unbounded-context loop, so the detector has something to find on the first run.

Measured overhead

Benchmarked against a zero-latency local upstream, which is the worst case for a ratio — real LLM calls take 400 to 4000ms:

Scenarioreq/smeanp95
Direct to upstream10,1362.87 ms3.20 ms
Through llm-guard3,6888.12 ms11.81 ms
Through llm-guard (SSE streaming)3,8517.70 ms8.33 ms

About 5 ms per request, and streaming responses are metered correctly, including usage that only arrives in the final frame.

What it does not do

It does not terminate inbound TLS — put a load balancer in front of it. It does not cache responses, retry requests, or store prompts and completions. Those are deliberate omissions rather than a roadmap.

Licensed MIT. Copyright LYE LABS LIMITED.