Most teams first notice their LLM bill when it surprises them. A feature that cost almost nothing in testing now produces a monthly invoice that changes every month, and nobody can say which feature, prompt or user caused it. This guide covers LLM cost optimization for teams already running AI features in production. The aim is a bounded, predictable spend without a drop in output quality.
The examples come from AI systems we have built and run at Ranbanka Systems, including our own sales assistant, lead triage and content pipelines. Where a figure is quoted, it comes from those real builds. Everything else is general engineering guidance that you should check against your own provider and traffic.
Why LLM costs drift in production (and why 'optimise later' fails)
LLM spend does not scale with the number of features you ship. It scales with traffic, prompt length and retries. A single chat widget on a busy page can cost more than ten internal tools, and one quiet prompt change can double the tokens on every call.
The same leaks show up again and again:
- Oversized system prompts resent on every call. Long instructions, tool definitions and reference documents are billed again on each request.
- Frontier models doing trivial work. Using your most capable model to label an email as spam is like hiring a senior architect to sort the post.
- Unbounded agent loops. An agent that keeps calling tools until it 'feels done' has no natural ceiling on cost.
- Retries on failure. Automatic retries on timeouts or malformed output can quietly multiply spend during an outage.
- Abuse and bot traffic. Public chat endpoints attract scrapers, testers and people using your assistant as a free general-purpose chatbot.
'Optimise later' fails because these costs are built into the architecture. Once prompts, model choices and agent loops are spread across a codebase, changing them is a refactor, not a config tweak.
So the real goal is not just a lower monthly bill. It is predictable, bounded spend per run and per day. That target leads to four levers, covered below:
- Choose the right model for each task
- Use prompt caching and trim what you send
- Enforce hard daily caps that fail safely
- Measure cost per run
Lever 1: Choose the right model for each task, not one model for everything
Many AI features start as one large prompt sent to one large model. That is fine for a prototype and expensive in production.
Break the feature into steps
Split the workflow into distinct steps, such as classify, extract, draft and review. Then match each step to the cheapest model that meets your quality bar for that step.
- Small or fast models suit routing, triage, classification, intent detection and simple extraction.
- Larger models should be kept for reasoning-heavy work and customer-facing generation, where quality is visible and mistakes are costly.
Parallelise instead of one giant prompt
If a task needs several independent analyses, run them as separate, focused calls in parallel rather than one huge prompt. Shorter calls are easier to price, easier to test and easier to move to a cheaper model one at a time.
We rebuilt our AI Website Audit Engine along these lines. It started as a CrewAI crew and became a LangGraph pipeline. The page and its Google PageSpeed data are fetched once, then five AI analysts work in parallel. An audit now takes about 45 seconds, down from several minutes, at roughly $0.09 per run.
Narrow tasks can be very cheap. Our AI Lead Triage handles every enquiry that reaches the ranbanka.com inbox, from the contact form or the chat. It uses Claude to:
- rate the enquiry hot, warm, cold or spam
- match it to a service and a past project
- suggest qualifying questions
- draft a reply for the team
All of that costs around $0.005 per enquiry, which shows how inexpensive a tightly scoped task can be.
Build an eval set before you downgrade
Before you move any step to a cheaper model, build a small evaluation set. Use real inputs with known-good outputs or clear pass/fail criteria. Run the current model and the candidate model against it. If the cheaper model passes, switch. If not, you have evidence for paying more. Without an eval set, downgrading is guesswork, and quality can slip without anyone noticing.
Lever 2: Prompt caching and trimming what you send
The cheapest token is the one you never send. The next cheapest is one your provider has already processed.
How provider-side prompt caching works
Several LLM providers offer prompt caching. When the start of a prompt (its prefix) matches a recent request, that cached portion is billed at a reduced rate and is often processed faster. Good candidates for the stable prefix are:
- the system prompt
- tool and function definitions
- reference documents, policies or product data
To benefit, put the stable content first and the variable user input last. If you place the user's message or a timestamp near the top, the prefix changes on every call and nothing is cached.
Caveat: caching rules differ by provider. Check minimum cacheable lengths, how long cache entries last and how cached tokens are billed. Your actual savings depend on traffic patterns. A feature with steady, frequent calls benefits far more than one used a few times a day.
Fetch once, share across steps
In multi-step pipelines, do not send the same page or document to every model call by default. Fetch it once, extract what each step needs, and pass along only the relevant parts. The Website Audit Engine's single fetch of the page and PageSpeed data is an example: the data is gathered once for the whole pipeline rather than re-collected for each analyst.
Trim the context
- Retrieve only relevant facts instead of pasting entire knowledge bases into the prompt.
- Cap conversation history. Keep a rolling window or a short summary instead of the full transcript.
- Cap max output tokens for each step. A classifier never needs a 2,000-token answer.
Cache at the application layer too
Provider caching is only one layer. In your own application, you can also cache:
- responses to identical requests
- answers to frequently asked questions
- deterministic intermediate results
If the same question gets the same answer, you should only pay for it once.
Lever 3: Hard daily spend caps that fail safely
Provider dashboards usually let you set account-level limits. Treat those as a backstop, not a control. They are too coarse to protect one feature from another, and hitting them can take down every AI feature at once.
Enforce budgets in your own code
Track spend per feature, per day in your own datastore. Before each model call, check the remaining budget. When the cap is reached, stop or degrade gracefully:
- show a clear fallback message
- queue the work for later
- route the request to a human
Add per-session and per-run limits
Daily caps protect the total. Per-run limits stop one bad request from eating the whole budget. Set limits on:
- tokens per session or per run
- tool calls per run
- agent loop iterations
An agent that hits its iteration limit should stop and report back, not keep going.
Rate-limit public endpoints
Any AI endpoint reachable from the open internet needs rate limiting by IP, session or user. This is your main defence against bots and abuse draining the budget before real users arrive.
What this looks like in practice
Our own AI Sales Assistant on ranbanka.com, built with LangGraph and Claude, runs with a hard $2/day spend cap and still replies in seconds. The Website Audit Engine has a daily cost cap, and AI Lead Triage has a daily budget cap. Each of these features has its own cap rather than relying only on an account-wide limit.
Decide the cap-hit experience in advance
What the user sees when a cap trips is a product decision, not just an engineering one. A chat assistant might offer a contact form. A background job might wait until the next day. Agree on this before launch so the failure mode is calm and on-brand rather than a broken widget.
Lever 4: Measure cost per run, not just the monthly invoice
A monthly invoice tells you what you spent. It does not tell you why. To manage LLM costs, you need data at the level of each call.
Log every call
For every model call, record:
- input tokens, output tokens and cached tokens
- model name
- latency
- feature name and a run ID
Roll up to unit economics
Using those logs, calculate cost per run, cost per user action and, most usefully, cost per business outcome: per lead, per audit, per published article. These are the numbers you can compare against the value each feature creates.
Two examples from our own systems:
- ~$0.09 per audit run for the AI Website Audit Engine
- ~$0.005 per triaged enquiry for AI Lead Triage
Once you know your unit cost, budgeting becomes simple arithmetic instead of guesswork.
Alert on anomalies
A sudden rise in tokens per run is one of the most useful warning signs you can track. It often points to a prompt regression, a retrieval step returning too much context or an agent stuck in a loop. Alert on it before it shows up on the invoice.
Review cost alongside quality
Cost data on its own can push you toward cheaper but worse outputs. Always review it next to your quality evals. Our AI Content Pipeline, which has published 29 blog articles, pairs a 12/12 fact-check eval with human approval on every post, so quality is checked as part of the workflow rather than as an afterthought. Apply the same principle to cost work: whenever you change a model, prompt or context size to save money, re-run your evals and confirm quality held before you ship.
Guardrails that save money as well as risk
Some guardrails reduce cost as well as risk.
- Human approval gates. Both the AI Content Pipeline and our AI Case-Study Writer (6 case studies published) publish only after human approval. Catching a bad output before it reaches customers is much cheaper than fixing it afterwards.
- Confirmation steps. The AI Sales Assistant sends a lead to the team only after the visitor confirms. This helps avoid wasted downstream work on incomplete or accidental submissions. Similarly, AI Lead Triage never emails the visitor. It prepares drafts for the team instead of taking actions by itself.
- Answering from verified data. The Sales Assistant answers visitor questions from verified company data rather than open-ended generation. Grounded answers keep prompts focused and outputs shorter.
- Code-level guards for deterministic rules. The Case-Study Writer blocks confidential client names in code, not with an extra model call. If a rule can be expressed as code, enforce it in code. It adds no model cost per run and behaves the same way every time.
A practical LLM cost optimization checklist and next steps
Use this checklist on an existing AI feature:
- Map each step of the feature to the cheapest model that passes your quality bar
- Build an eval set before changing any model
- Reorder prompts so stable content comes first, for caching
- Trim context and cap max output tokens per step
- Enforce daily per-feature caps and per-run limits in code
- Rate-limit public endpoints
- Log tokens, model, latency and cost per run, tagged by feature
- Set alerts on tokens-per-run anomalies
- Define the user experience when a cap is hit
Start with measurement. You cannot optimise what you are not logging. Once per-call data shows where the money goes, the right fix is often much easier to see.
Where Ranbanka can help
We build AI assistants and pipelines, and we run our own under hard limits: the AI Sales Assistant has a $2/day spend cap, while the Website Audit Engine and Lead Triage each run under a daily cost or budget cap. If you already have an AI feature in production, we can look at its cost profile with you and suggest where the levers above apply. See our AI solutions for the kind of work we do, browse the builds mentioned above in our portfolio, or explore our full range of services.
If your LLM bill is less predictable than you would like, book a free initial consultation. We respond within 24 hours, and we are happy to work under an NDA.