LLM Cost Diagnosis
We find out where a company's AI spend is going, and what to change. Fixed price, fixed scope, no infrastructure to deploy.
The problem we see most often
An agent workload starts re-sending its entire conversation history on every step. Input tokens grow while output stays flat. Nothing crashes, nothing alerts, and the first signal is a bill.
OWASP's research into agent execution budgets catalogues 63 confirmed production overruns across 21 orchestration frameworks, and puts Fortune 500 leakage at roughly $400M in unbudgeted spend. It also records that runtime visibility exists in only about a fifth of organisations.
The recorded failure patterns are not mysterious. They are just not visible until somebody looks at the token counts.
How a diagnosis works
No deployment. No agents in your VPC. No code changes.
1. You give us read access, not infrastructure
Two ways, whichever you prefer:
- Your provider's admin API. Both OpenAI and Anthropic expose organisation-level usage and cost endpoints. One read-only key, and we pull the data directly.
- A file. Export it yourself and send it. We will tell you exactly which columns we need.
Either way it is read-only. We do not ask for a write key, and we do not need access to prompts or completions.
2. We analyse metadata, never content
Token counts, model ids, timestamps, attribution labels, and what the provider actually billed. That is the whole input.
We do not read your prompts. We do not read your completions. We could not tell you what your application does if you asked.
3. You get a written report
A short document, A4, the kind of thing you can forward to a CTO without rewriting it. It contains:
- A verdict in plain language: what is driving the bill.
- Every finding ranked by what it is worth per month, as a range, not a single confident number.
- The evidence behind each one, so your engineers can check our arithmetic rather than trust it.
- A confidence level per finding. Some of these are arithmetic on your data. Some are a judgement about your architecture. The report says which is which.
- What the analysis cannot tell you. There is always something, and it belongs on the page rather than in our heads.
- Where a finding matches a catalogued production failure, the citation, the number of recorded incidents, and the largest reported loss.
What it costs
| Price | What you get | |
|---|---|---|
| Scoping call | Free, 15 minutes | We confirm the data is available and tell you honestly whether a diagnosis is worth doing yet. Sometimes the answer is no. |
| Diagnosis | $1,500 | The written report, delivered within 5 working days of receiving data. |
| Deployment and monitoring | $2,000 setup, then $1,000/month | We install guardrails in your infrastructure, wire up alerting, keep the price table current, and answer when something looks wrong. |
| Ongoing cost governance | From $2,500/month | We keep watching: routing changes, cache policy, context budgets, a quarterly review. The alternative is hiring for it. |
If the diagnosis finds nothing worth acting on, we refund it. We would rather lose the fee than defend a report that says "your spend is fine but here are some charts".
That is not a marketing promise. It is a consequence of the data model: we can see before we quote whether there is anything in there, because the scoping call already showed us the shape of the usage.
What we do not do
- We do not need your prompts. If a vendor asks for them, ask why.
- We do not sell a dashboard you have to babysit. We sell the finding and the fix. The tooling is free and open source if you want to run it yourself: llm-guard.
- We do not take a cut of your spend. No percentage-of-spend pricing. We are paid the same whether your bill goes up or down, which is the only arrangement where the advice is worth anything.
Why the tooling is free
llm-guard is MIT licensed and complete. Zero runtime dependencies, runs on the standard library, enforces budgets and detects runaway loops. If your team wants to run it themselves, do that. It is genuinely the whole thing, not a crippled edition.
We charge for the part a tool cannot do: knowing which patterns matter, calibrated against what we have seen, and being accountable for the answer.
Get started
Email [email protected] with:
- Roughly what you spend per month on model APIs.
- Which providers.
- Whether you have had a cost surprise in the last year. If yes, roughly how much.
That is enough for a scoping call. If it looks like there is something to find, we will say so. If it does not, we will say that too.
LYE LABS LIMITED (灵野科技有限公司) · Hong Kong