Skip to content
Velum

Managed AI operations

LLM cost optimization for agents running in production

Short answer

LLM cost optimization comes down to sending fewer tokens to cheaper models without hurting task quality: route simple steps to smaller models, cache repeated prompt prefixes, batch work that is not urgent, and trim context to what the task needs. Teams that apply these together commonly cut API spend by 40 to 80 percent. Start by measuring cost per completed task, because you cannot optimise what you only see as one monthly bill.

Key takeaways

  • Cost per completed task is the metric that matters, not the total monthly invoice.
  • Prompt caching bills repeated input tokens at a fraction of the normal rate on major providers.
  • Batch APIs from major providers typically cost about half the standard rate for non-urgent work.
  • Every cost change should be checked against the eval set so savings do not hide quality losses.

Where LLM costs come from and how to monitor them

LLM API costs are driven by input tokens, output tokens and the model tier. In agents, input usually dominates, because each step resends the system prompt, tool definitions, conversation history and retrieved documents. A ten-step agent run can easily send the same 8,000 token prefix ten times. Output tokens are priced higher per token but are usually a smaller share. Retries, loops and oversized retrieval are the hidden multipliers that turn a cheap task into an expensive one.

LLM cost monitoring starts with tagging every request with the agent, the task type and a run ID. Tools such as Langfuse, Helicone or a LiteLLM proxy record tokens and cost per call, so you can roll them up into AI agent cost tracking per completed task. Set alerts on cost per task and daily spend per agent. LLM monitoring in production should show cost next to quality and latency, so a cheaper configuration that fails more often is caught immediately.

LLM cost optimization strategies that work

The biggest lever is model routing. Classification, extraction from clean documents and simple formatting often run well on small, fast models, while planning and difficult reasoning stay on a larger one. Test each step against your eval set and move it down a tier only when scores hold. The second lever is prompt caching: put stable content such as instructions and tool definitions first, so providers can reuse it. Cached input tokens are billed at a small fraction of the normal rate on Anthropic, OpenAI and others.

LLM token cost optimization also means sending less. Retrieve fewer, better chunks instead of whole documents, summarise long histories, and ask for structured output rather than long prose. Use batch APIs for overnight or non-urgent jobs such as bulk document processing, which major providers typically price at about half the standard rate. Cap retries and step counts so a confused agent stops and escalates to a person instead of burning tokens in a loop.

What a normal monthly cost to run an AI agent looks like

For most SMB and mid-market agents, model spend is smaller than people expect. An agent handling a few hundred documents or tickets a day typically costs from tens to a few hundred US dollars a month in API fees. High-volume agents processing tens of thousands of items a day, or agents that do long multi-step research, can reach several thousand dollars a month. Typical cost per task ranges from under one cent for simple classification to a few dollars for long research tasks.

The larger costs are usually around the model: hosting, observability tools, integration upkeep and the people who review exceptions. That is why LLM API cost optimization should be judged against the cost of doing the work by hand. An agent that costs 0.20 dollars per invoice and saves ten minutes of staff time is cheap, even if the monthly bill looks significant. Velum sets a cost per task target during scoping and reports against it monthly.

How it works

  1. 1

    Instrument every call

    We tag each model call with agent, task and run ID so spend can be viewed per completed task, not only per month.

  2. 2

    Find the expensive steps

    We rank steps by cost and look for repeated prefixes, oversized context, loops and retries that inflate spend.

  3. 3

    Route, cache and trim

    We move suitable steps to smaller models, restructure prompts for caching, batch non-urgent work and tighten retrieval.

  4. 4

    Check quality against evals

    Every change runs against the eval set, and we only keep savings where task accuracy and escalation rates hold.

  5. 5

    Launch with limits and approvals

    We set spend alerts and step caps in production, and the agent still waits for a person to approve sends, payments, postings and deletions.

Before and after

TaskBy handWith agents
Cost visibilityOne monthly provider invoiceCost per task, per agent, per day
Model choiceLargest model for every stepSmaller models on 50 to 80 percent of steps
API spend for the same workloadBaselineCommonly 40 to 80 percent lower
Runaway loopsFound on the invoiceStopped by step caps and alerts within minutes

Typical ranges from comparable deployments. Your baseline is measured before anything is built.

Tools it works with

  • Claude
  • OpenAI
  • Langfuse
  • Helicone
  • LiteLLM
  • OpenRouter
  • AWS Bedrock
  • Datadog

Questions people ask

01

How do you reduce LLM API costs?

Route simple steps to smaller models, cache stable prompt prefixes, batch non-urgent jobs and send less context by retrieving fewer, better chunks. Cap retries and steps so agents cannot loop. Test each change against an eval set so you keep quality while cutting spend.

02

How do you track AI agent cost per task?

Tag every model call with the agent name, task type and a run ID, then sum token costs per run. Observability tools like Langfuse or Helicone, or a proxy like LiteLLM, do this automatically. Divide by completed tasks, not attempts, so failed runs count against the agent.

03

Does prompt caching reduce costs?

Yes, often substantially for agents that resend the same instructions and tool definitions on every step. Providers such as Anthropic and OpenAI bill cached input tokens at a fraction of the normal rate and return responses faster. To benefit, keep stable content at the start of the prompt and variable content at the end.

04

When should you use a smaller model?

Use a smaller model for steps with clear, narrow outputs, such as classification, routing, extraction from clean documents and formatting. Keep larger models for planning, ambiguous inputs and multi-step reasoning. Decide per step using your eval scores, not general benchmarks.

05

What is a normal monthly cost to run an AI agent?

For a typical SMB agent handling a few hundred items a day, model API costs usually run from tens to a few hundred US dollars a month. High-volume or research-heavy agents can cost several thousand. Hosting, monitoring and human review often cost more than the model itself, so compare the total against the manual cost of the work.

Start with one workflow.

Thirty minutes. One real process. A practical next step.