ProblemHow it worksOpen sourcePricingFAQ

LLMBRAKE

Per-customer cost caps for AI apps.

Put a monthly dollar budget on every customer. LLMBrake switches heavy users to a cheaper model near their cap, blocks them at the cap, and shows what each customer costs you against what they pay.

MIT · coming soon to GitHub·TypeScript·OpenAI-compatible·No prompts stored
Floating cards: illustrative UI, not real data
01 / The problemProject-level limits only

Your provider caps the project. Nobody caps the customer.

OpenAI and Anthropic budgets work at the organization or project level. Your invoice shows one number, not which customer drove it.

A · Surprise bills

One loop, one night

One user scripting your AI feature in a loop can run up real money overnight. You find out when the invoice arrives.

B · Hidden losses

Unprofitable accounts

Your biggest customers may cost more in tokens than they pay you. Without per-customer cost data, you can't prove it or fix the pricing.

C · Rebuilt plumbing

Everyone rebuilds it

Token logging, price tables, streaming usage, counters, upgrade prompts: teams keep writing the same plumbing instead of shipping product.

02 / How it worksVisuals: illustrative, not real data

A small wrapper. Not a proxy.

LLMBrake wraps the client you already use. It calls your provider directly with your own keys, and your prompts never leave your servers.

Wrap once01
import { createGuard } from "llmbrake";

const guard = createGuard({
  budgetUsd: (ctx) => ctx.plan === "pro" ? 20 : 2,
  downgradeAt: 0.8,
  downgradeMap: { "gpt-5.5": "gpt-5.4-mini" },
});
const chat = openai.chat.completions;
const create = guard.wrap(chat.create.bind(chat));
await create({ model: "gpt-5.5", messages },
  { customerId: user.id, plan: user.plan });
// at the cap → throws BudgetExceededError

Wrap once

One line around your OpenAI-compatible create() call. Streaming usage is captured too.

Set budgets02
cus_A71$3.05
cus_C09$10.40
cus_8F2$18.40
cus_2KQ$2.00
$080%: cheaper model100%: blocked
Illustration of how the thresholds work

Set budgets

A monthly USD budget per customer, by plan or by customer. Counters live in memory, Redis or your own store.

Auto-downgrade03
cus_8F2 at 92% of budget
gpt-5.5gpt-5.4-mini
Same request · cheaper model
Illustrative

Auto-downgrade

Near the cap (for example 80%), mapped models switch to cheaper ones, so heavy users keep working while costs slow down.

Block at cap04
BudgetExceededError: customer cus_2KQ
spent $2.00 of $2.00 (2026-10)
code: LLMBRAKE_BUDGET_EXCEEDED
You've hit this month's AI limit. Upgrade for more.Upgrade
Example upgrade prompt

Block at cap

A typed BudgetExceededError lets you show an upgrade prompt instead of eating the cost.

Usage events05
// onUsage(event): metadata only
{
  customerId: "cus_8F2",
  plan: "pro",
  model: "gpt-5.4-mini",
  inputTokens: 1102,
  cachedTokens: 640,
  outputTokens: 182,
  costUsd: 0.00113,
  // no prompt, no completion
}
Illustrative values

Know every call's cost

Each call produces a usage event with tokens and computed cost. Send it to your logs, warehouse, or (soon) the hosted dashboard.

See margin06 · Planned
CustomerAI costRevenueMargin
Customer A$4.10$29.00+$24.90
Customer B$41.70$29.00−$12.70
Mockup with illustrative numbers. Not real customer data.

See margin

A hosted dashboard comparing AI cost to Stripe revenue per customer. In development; this is what founding members fund.

03 / Open sourceMIT licensed SDK · coming soon to GitHub

Yours to keep. Even if we disappear.

The enforcement part runs in your code, under a permissive license. The hosted dashboard is optional.

A

MIT licensed

Read it, fork it, vendor it. Zero runtime dependencies, written in TypeScript.

Coming soon to GitHub
B

Metadata only

Customer ID, plan, model, token counts, cost and timestamps. Never prompts or completions.

C

Fails open

If the budget store is unreachable, calls go through by default, so your product keeps working. Fail closed if you prefer.

04 / PricingNot live yet · no charge

Free SDK. Founding price, reserved.

The SDK is free and open source. The hosted dashboard (shared counters, alerts and the Stripe margin view) is being built now. We'd rather validate demand before building it, so there's no checkout yet: join the waitlist and the founding price is reserved for you.

Founding memberFor early waitlist members
$29 / month
  • Hosted dashboard: cost per customer, feature and model
  • Shared budget counters across servers and serverless
  • Email and Slack alerts near and at caps
  • Stripe revenue vs AI cost per customer
  • Founding price locked if you subscribe after launch
  • Direct line to the builders for feature requests
Reserve the founding price

Nobody is charged: there is no checkout yet. Early waitlist members get $29/month reserved for when the hosted dashboard launches, and you decide then whether to subscribe.

Free waitlist$0
$0

Get the launch email and early access. Early members get the $29/month founding price reserved. Free, and nobody is charged.

One launch email, plus occasional build updates. Unsubscribe any time.

05 / FAQPrivacy first
01Do you see or store my prompts?

No. We store usage metadata only: customer ID, plan, feature tag, model, token counts, computed cost and timestamps. Prompts and completions never leave your servers. The open-source SDK doesn't read message contents at all.

02Is this a proxy in front of OpenAI?

No. The SDK runs inside your app and calls your provider directly with your own keys. The budget check runs before each call, and the cost is recorded after it.

03What happens if your service is down?

By default the SDK fails open: if the budget store is unreachable, calls go through, so your product keeps working. You can choose to fail closed instead.

04Which providers and models are supported?

Any OpenAI-compatible chat completions API, including OpenAI, OpenRouter, Anthropic's OpenAI-compatible endpoint and many gateways. The SDK includes a small price table that you should verify and extend, because provider prices change often.

05How accurate are the costs?

Costs are computed from the token usage the provider returns and your price table. Batch discounts, long-context tiers and regional uplifts aren't modeled yet. Treat the numbers as close estimates and reconcile them with your provider invoice.

06Can I use it without paying?

Yes. The SDK is MIT licensed and works on its own with an in-memory or your own Redis/Postgres store. The planned paid plan is for the hosted dashboard, alerts and the Stripe margin view.

07Who's building this?

A small, independent team building tools for AI app developers. We're validating demand before building the hosted product, which is why there's a free waitlist with a reserved founding price instead of a checkout.

Free waitlist · Nobody is charged

Start braking.

Cap every customer before the next invoice does it for you.