从零到一份成本报告,大约两分钟。
不需要账号、不需要 API key、不需要联网。演示数据集里带了一个真实的失控循环,所以下面每条命令第一次跑就有东西可看。
安装,然后确认它能跑
只有一个包,没有别的要装,因为运行期只用 Python 标准库。doctor 会在你依赖它之前,先检查解释器、SQLite、数据库和价格表。
$ pip install git+https://github.com/leyao-daily/llm-guard.git
$ llm-guard doctor
llm-guard doctor
python 3.12.14 (Darwin)
sqlite 3.53.1
database ./llmguard.db
rows 0
priced models 8
budgets 0
all checks passedpip install llm-guardPyPI 上这个名字属于 Protect AI 的 prompt 注入防护工具,是另一个无关的项目(26 个版本)。装它会得到完全不同的软件。本项目的分发包名是 llm-cost-guard,命令行仍然是 llm-guard。
载入演示数据
六十天的模拟流量,跨六把 API key,其中包括一个一直在以 74:1 的输入输出比循环的客服 agent。最近三十天是报告窗口,再往前三十天用来做环比。
$ llm-guard seed --reset --compare-days 30
Seeded 13,254 requests spanning 60 days ($227.10 of simulated spend).
Last 30 days is the reporting window; the preceding 30 days gives the
period-over-period delta.
Added 4 demo budgets.
Next: llm-guard report在循环还在跑的时候抓到它
这就是这个产品存在的理由。它看最近十五分钟的流量,报告任何看起来像失控的东西,并给出支撑这个结论的具体数字。注意它同时在那把 API key 和那个 project 上触发 —— 同一批流量在两个归因标签下都可见,这就是你查清「这是谁的工作负载」的方式。
Runaway-spend detection
window: last 15 min · io ratio > 30:1 · velocity > 25x baseline
─────────────────────────────────────────────────────────────
CRITICAL io_ratio [key: prod-agent]
prod-agent is running a 74:1 input-to-output ratio across 67 calls
in the last 15 min (12,006,400 in / 162,216 out, $38.45). Normal
traffic sits at 5:1-15:1.
→ what to do
The prompt is being re-sent in full on every step. Cap the context
(summarise or truncate history), and enable prompt caching so the
repeated prefix is billed at the cache rate. OWASP attributes ~62%
of agent bills to re-sent context.
WARN retry_storm [key: prod-batch]
prod-batch had 42 failed calls out of 56 (75%) in the last 15 min.这里全部只是告警。如果你想让某个东西真的拦住流量,那是一个 action=block 的预算,需要你自己明确设置 —— 见第 6 步。
读报告
anomalies 是实时的、窄的;report 是三十天的、宽的。它先讲该改什么,而不是发生了什么,并且给每条建议标上金额。
LLM Spend Report
window: last 30 days · 13,254 requests recorded in total
────────────────────────────────────────────────────────────
Total spend $149.80
vs previous period ▲ +88.2% ($79.60 before)
Requests 7,291 (+19.0% vs prior)
Cost per request $0.0205
Errors 285 (3.9%)
What to do about it
1. One model is 56% of your bill (claude-sonnet-4.5, $83.84).
Routing the easy traffic to a cheaper tier is the single highest-leverage
change.
2. gpt-6-astra costs 9.7x your average per request.
$0.1998 vs a blended $0.0205. Check whether its prompts carry large
fixed context. Trimming fixed context pays off on every call.
3. Prompt caching is worth up to $181.43/month on this traffic.注意那个 +88.2% 的环比。这是只看一个当期窗口给不了的信息,通常也是第一个告诉你「有什么变了」的信号。
把真实流量接进来
把你 SDK 的 base URL 指向网关。除了 URL,你的应用什么都不用改,网关会通过 TLS 转发到真正的服务商。
$ export OPENAI_API_KEY=sk-...
$ export ANTHROPIC_API_KEY=sk-ant-...
$ llm-guard serve --port 8080
llm-guard gateway listening on 127.0.0.1:8080
openai -> https://api.openai.com
anthropic -> https://api.anthropic.com然后在你的应用里改一行:
# OpenAI SDK
client = OpenAI(base_url="http://127.0.0.1:8080/v1")
# Anthropic SDK
client = Anthropic(base_url="http://127.0.0.1:8080")
设一个真的能止住花钱的预算
预算按 API key 设置。alert 只告警;block 在达到上限后返回 429。要在事故之前设上限,不是之后 —— 已记录的从发生到被发现的中位时间是约四个半小时,足够花掉很多钱。
$ llm-guard budget set --key prod-agent --daily 30 --action block
$ llm-guard budget list
api key daily monthly action state
prod-agent $30.00 $250.00 block ok
prod-batch $30.00 $700.00 alert ok
prod-web $45.00 $1,200.00 alert ok
staging $25.00 $400.00 alert ok已经有用量数据?可以跳过网关
你不必为了拿到一份诊断而把流量导过任何东西。两家服务商都提供组织级的用量与成本接口,所以一把只读 key 就够了。
从服务商的管理员 API
直接拉用量和成本。不部署、不改代码、只读。
llm-guard import --source anthropic \
--key "$ANTHROPIC_ADMIN_KEY" --days 30
llm-guard import --source openai \
--key "$OPENAI_ADMIN_KEY" --days 30
从文件
CSV 或 JSON。列名做了宽松匹配,所以从服务商后台直接导出的文件通常不用改。
llm-guard import --source file --path usage.csv
llm-guard import --sample # 打印期望的列
导出的数据按天或按小时分桶,所以没有延迟、没有状态码、没有终端用户维度。需要逐请求数据的检查会被跳过而不是猜;量级阈值会相应下调;报告会明确写出哪些检查被省略、以及为什么。
然后把它变成能发给别人的东西
上面的 report 是给你自己的。diagnose 是一份给「对预算负责的那个人」的文档 —— 它先给结论,按每月价值排序每一项发现,展示每个数字背后的证据,并说明这次分析无法告诉你什么。
$ llm-guard diagnose --client "某公司" --out diagnosis.html
Wrote diagnosis.html (18,892 bytes)
Open it in a browser and print to PDF, or send the HTML as-is.
The layout is set for A4 with page breaks; no external assets.每项发现都带金额
每项给一个每月节省区间,而不是一个自信的单点数字,并附上推导它的证据,让你的工程师能核对算术。
每项发现都带置信度
确定 是你数据上的算术。很可能 依赖一个写明的假设。值得一试 是一个假设。报告会说明是哪种。
适用时给出引用
匹配到已记录生产故障模式的发现,会带上模式名、事故数和来源 —— OWASP AISVS C9.1,以及一份收录 63 起已确认事故的目录。
报告的节省是估计值,文档会直接说明它们互相重叠:压缩上下文和提高缓存命中率作用于同一批输入 token,不可能全额同时拿到。一份夸大了节省的诊断会在下一张账单上被拆穿,而之后它里面别的数字也没人信了。
全部命令
十五个子命令。llm-guard <命令> --help 查看任意一个。
| 命令 | 作用 |
|---|---|
doctor | 检查安装、数据库与价格表 |
seed | 载入带真实循环的演示数据 |
serve | 运行网关 |
report | 花费、环比与优先改动项 |
anomalies | 此刻正在失控的花费 |
diagnose | 面向客户的书面诊断 |
dashboard | 自包含 HTML,无 CDN,适合隔离网络 |
import | 从服务商管理员 API 或文件导入用量 |
patterns | 已记录的故障模式及其来源 |
budget | 设置按 key 的日/月上限 |
engage | 冻结基线,之后测量实际发生了什么 |
outcomes | 节省预测历史上兑现了多少 |
export | 把原始行导出为 CSV 或 JSON |
models | 内置价格表 |
cost | 给单次调用计价 |