Key takeaways
- Cost per completed task is the metric that matters, not the total monthly invoice.
- Prompt caching bills repeated input tokens at a fraction of the normal rate on major providers.
- Batch APIs from major providers typically cost about half the standard rate for non-urgent work.
- Every cost change should be checked against the eval set so savings do not hide quality losses.
Where LLM costs come from and how to monitor them
LLM API costs are driven by input tokens, output tokens and the model tier. In agents, input usually dominates, because each step resends the system prompt, tool definitions, conversation history and retrieved documents. A ten-step agent run can easily send the same 8,000 token prefix ten times. Output tokens are priced higher per token but are usually a smaller share. Retries, loops and oversized retrieval are the hidden multipliers that turn a cheap task into an expensive one.
LLM cost monitoring starts with tagging every request with the agent, the task type and a run ID. Tools such as Langfuse, Helicone or a LiteLLM proxy record tokens and cost per call, so you can roll them up into AI agent cost tracking per completed task. Set alerts on cost per task and daily spend per agent. LLM monitoring in production should show cost next to quality and latency, so a cheaper configuration that fails more often is caught immediately.
LLM cost optimization strategies that work
The biggest lever is model routing. Classification, extraction from clean documents and simple formatting often run well on small, fast models, while planning and difficult reasoning stay on a larger one. Test each step against your eval set and move it down a tier only when scores hold. The second lever is prompt caching: put stable content such as instructions and tool definitions first, so providers can reuse it. Cached input tokens are billed at a small fraction of the normal rate on Anthropic, OpenAI and others.
LLM token cost optimization also means sending less. Retrieve fewer, better chunks instead of whole documents, summarise long histories, and ask for structured output rather than long prose. Use batch APIs for overnight or non-urgent jobs such as bulk document processing, which major providers typically price at about half the standard rate. Cap retries and step counts so a confused agent stops and escalates to a person instead of burning tokens in a loop.
What a normal monthly cost to run an AI agent looks like
For most SMB and mid-market agents, model spend is smaller than people expect. An agent handling a few hundred documents or tickets a day typically costs from tens to a few hundred US dollars a month in API fees. High-volume agents processing tens of thousands of items a day, or agents that do long multi-step research, can reach several thousand dollars a month. Typical cost per task ranges from under one cent for simple classification to a few dollars for long research tasks.
The larger costs are usually around the model: hosting, observability tools, integration upkeep and the people who review exceptions. That is why LLM API cost optimization should be judged against the cost of doing the work by hand. An agent that costs 0.20 dollars per invoice and saves ten minutes of staff time is cheap, even if the monthly bill looks significant. Velum sets a cost per task target during scoping and reports against it monthly.
How it works
- 1
Instrument every call
We tag each model call with agent, task and run ID so spend can be viewed per completed task, not only per month.
- 2
Find the expensive steps
We rank steps by cost and look for repeated prefixes, oversized context, loops and retries that inflate spend.
- 3
Route, cache and trim
We move suitable steps to smaller models, restructure prompts for caching, batch non-urgent work and tighten retrieval.
- 4
Check quality against evals
Every change runs against the eval set, and we only keep savings where task accuracy and escalation rates hold.
- 5
Launch with limits and approvals
We set spend alerts and step caps in production, and the agent still waits for a person to approve sends, payments, postings and deletions.
Before and after
Typical ranges from comparable deployments. Your baseline is measured before anything is built.
Tools it works with
- Claude
- OpenAI
- Langfuse
- Helicone
- LiteLLM
- OpenRouter
- AWS Bedrock
- Datadog