Facebook tracking pixel Skip to main content

Reference · updated 2026-08-27

AI agent pricing

What the models, tools, and platforms behind AI agents actually charge. Every number on this page was read off the vendor's own pricing page and links back to it. Nothing is estimated.

Most cost questions get answered with a single number. A real bill has four layers, and the model is usually the smallest one.

The one number people ask for

$0.0037

per conversation, for inference

Anthropic's own worked example: 10,000 customer-support conversations for about $37 in total. About 3,700 tokens per conversation on Claude Haiku 4.5 at $1/$5 per 1M. Source, checked 2026-08-27.

That figure is real and it is also the reason cost conversations go wrong. Inference is cheap. The invoice that surprises people is built from the three layers above it.

Layer 1

The model

Charged per million tokens, in and out. Every agent pays this whatever sits on top of it. Output costs several times input at every vendor, so an agent that writes long answers costs more than one that reads long documents.

Vendor Model Price Unit Notes Link
Anthropic Claude Haiku 4.5 $1 per 1M input tokens Output is $5 per 1M. The cheapest current Claude tier, and the one Anthropic uses in its own support-agent cost example. source
Anthropic Claude Sonnet 5 $2 per 1M input tokens Output is $10 per 1M. Anthropic states the $2/$10 launch pricing is now standard and the scheduled rise to $3/$15 will not happen. source
Anthropic Claude Opus 5 $5 per 1M input tokens Output is $25 per 1M. source
Anthropic Claude Fable 5 $10 per 1M input tokens Output is $50 per 1M. Top tier. source
OpenAI gpt-5-nano $0.0500 per 1M input tokens Output is $0.40 per 1M. The cheapest row on the page. source
OpenAI gpt-5.6-luna $0.20 per 1M input tokens Output is $1.20 per 1M. Cached input is $0.02. source
OpenAI gpt-5.6-terra $2 per 1M input tokens Output is $12 per 1M. Cached input is $0.20. source
OpenAI gpt-5.6-sol $4 per 1M input tokens Output is $20 per 1M. The page states this promotional rate runs at least through 2026-11-21, so treat it as temporary. source
Google Gemini 2.5 Flash-Lite $0.10 per 1M input tokens Output is $0.40 per 1M. Audio input is priced higher at $0.30. source
Google Gemini 3.5 Flash-Lite $0.30 per 1M input tokens Output is $2.50 per 1M. source
Google Gemini 3.7 Flash $0.75 per 1M input tokens Output is $3.75 per 1M. The page dates a rise to $1.50 / $7.50 on 2027-01-01. source
Google Gemini 2.5 Pro $1.25 per 1M input tokens Output is $10 per 1M including thinking tokens, and both roughly double above a 200k-token prompt. source

Rows are the tiers a build actually chooses between. Each vendor's page carries the full model list.

Layer 1, continued

The multipliers that move the bill

These matter more for agents than for other AI work. An agent re-sends its system prompt and tool definitions on every single turn, so the same tokens get paid for over and over unless caching is on.

Vendor Item Value Unit Notes Link
Anthropic Prompt cache read $0.10 multiple of base input price A cache hit costs 10% of standard input. Writes cost 1.25x (5-minute) or 2x (1-hour), so a 5-minute cache pays for itself after one read. source
Anthropic Batch API $0.50 discount on input and output 50% off both directions, for work that does not need an answer in real time. Not available for interactive sessions. source
Anthropic Web search tool $10 per 1,000 searches Billed on top of tokens. An agent that searches on every turn adds a line item most cost models miss. source
Anthropic US-only inference (data residency) $1.10 multiplier on all token pricing Pinning inference to the US adds 10% across input, output, and cache. source

Decay

What is scheduled to stop being true

Published prices are not permanent, and several current ones carry a date on them. This is the part worth checking before you build a budget on a number you read three months ago.

2026-09-01 (cancelled) Anthropic

Claude Sonnet 5 stays at $2 / $10 per 1M. The rise to $3 / $15 that was scheduled for 2026-09-01 has been cancelled and the launch rate is now standard.

source, checked 2026-08-27
2026-11-21 OpenAI

gpt-5.6-sol is on promotional pricing, stated as available at least through 2026-11-21.

source, checked 2026-08-27
2027-01-01 Google

Gemini 3.7 and 3.6 Flash roughly double: input $0.75 to $1.50, output $3.75 to $7.50, with caching and storage rising in step.

source, checked 2026-08-27
in force Anthropic

Claude 4.7 and later use a newer tokenizer that produces roughly 30% more tokens for the same text. Per-token prices did not change, but cost per task is not comparable across the 4.6 / 4.7 line.

source, checked 2026-08-27

Layer 2

The platform, and its unit

Read the middle column first. What a platform meters decides what happens to your bill when volume grows: a per-seat plan barely moves, a per-task or per-conversation plan moves with every extra customer. Two platforms with the same monthly price can differ by an order of magnitude at scale purely because of this.

Vendor Plan Metered by Price Notes Link
Zapier Professional per task $19.99per month, from Printed as a starting price; the real figure moves with your task tier. Team starts at $69 per month for 25 users. A free tier gives 100 tasks a month. source
Make Core per credit (one module action) $12per month At the default 10,000 credits a month. Pro is $21 and Teams $38 at the same volume. Free tier allows up to 1,000 credits a month with no time limit. source
n8n Starter per full workflow execution, not per step not published Priced in euros even from a US request (EUR 20 a month billed annually, 2,500 executions), so no verified USD figure exists to quote. Metering per whole execution rather than per step is the notable part: a 30-step workflow costs the same as a 3-step one. source
Lindy Plus per seat plus a shared credit pool $29.99per user, per month Includes 3,000 credits per user. Pro is $99.99 and Max $199.99 per user. No free tier. source
Gumloop Pro per credit, plus an 8% orchestration fee $37per month, from The 8% fee stacks on credit consumption, so the sticker is not the effective rate. Only two tiers exist and there is no free plan. source
Salesforce Agentforce, per action per action, via Flex Credits $0.10per action Derived from figures printed on the same page: an action is 20 Flex Credits and credits are $500 per 100,000. Voice actions are 30 credits, so $0.15. A per-conversation option is $2. source
Microsoft Copilot Studio per Copilot Credit $0.0100per credit, pay as you go A prepaid pack is $200 a month for 25,000 credits. Microsoft's own Azure pricing page renders this rate as a dash; the figure is confirmed in the Copilot Credits Guide. source
Google Gemini Enterprise Agent Platform, compute per vCPU-hour $0.0850per vCPU-hour Pure consumption, no seats and no tiers, with the first 50 vCPU-hours a month free. This product was called Vertex AI Agent Builder; the name no longer appears on the pricing page, so older comparisons are stale. source
Relevance AI Self-serve plans Actions and Vendor Credits not published No self-serve price is published at all; the pricing page is a single Enterprise card with a sales contact. Only top-up rates remain documented, at $80 per 1,000 Actions. source
CrewAI Paid tiers per workflow execution not published Nothing purchasable sits between the free tier (50 executions a month) and a custom enterprise quote. There is no published middle. source

Three of these vendors will not tell you the price.

Relevance AI has reduced its public pricing page to a single enterprise card with a sales contact. CrewAI publishes a free tier and a custom quote with nothing purchasable in between. Salesforce prints unit rates but gates every plan behind a form. That is worth recording, because a price you cannot see before a sales call is a price you cannot compare.

Layer 3, voice

Where the headline number lies

If your agent speaks, this is the layer that decides the bill, and it is the one most often quoted wrong. A voice platform's advertised per-minute rate usually covers the platform alone. Bland states exactly that about its competitors on its own pricing page.

Vendor Product Price Unit What it covers Link
Vapi Vapi hosting (Build) $0.0500 per call minute Platform fee only. The page states hosting excludes model provider costs: speech-to-text, the model, and speech synthesis are billed at cost, and transport is charged by the provider. source
Retell AI Voice infrastructure $0.0550 per call minute Retell is unusual in publishing a per-minute rate for every component, which is why its stack is the one that resolves exactly. source
Retell AI Speech synthesis $0.0150 per call minute Rises to $0.040 per minute if you pick ElevenLabs voices. source
Bland AI Start plan $0.14 per talk minute Model, speech-to-text and speech synthesis are included with no token charges. Telephony is billed separately. The only vendor here whose headline number is close to the real one. source
ElevenLabs Agents, additional minute $0.0800 per call minute Speech synthesis, speech-to-text and retrieval are included. The page states the model and any telephony are billed separately on top, both at cost. source
Deepgram Voice Agent API (Standard) $0.0560 per call minute Bundles speech-to-text, model and speech synthesis orchestration, billed on websocket connection time. The page dates a rise to $0.075 on 2026-09-12. Telephony is not included. source

The parts underneath

Whether a platform bundles these or passes them through, somebody pays them. Watch the unit on speech synthesis: it is charged per character, not per minute, which is the reason no honest all-in per-minute quote exists.

Vendor Component Price Unit Notes Link
Deepgram Nova-3 streaming speech-to-text $0.0048 per minute Pay-as-you-go rate. Listening is the cheapest part of a voice agent by a wide margin. source
Deepgram Aura-2 speech synthesis $0.0300 per 1,000 characters Charged per character, not per minute. This is the line that stops any stack from producing a clean per-minute number, because it depends on how much the agent says. source
Twilio US local inbound voice $0.0085 per minute Outbound US is $0.0140. SIP trunking, which is the rate most voice-AI platforms actually hit, is $0.0040 in both directions. source
Twilio US local phone number $1.15 per month Toll-free is $2.15 per month. A fixed cost per number, regardless of call volume. source
OpenAI gpt-realtime-2.1 audio $32 per 1M audio input tokens Audio output is $64 per 1M. OpenAI publishes NO per-minute price for realtime voice; it is token-billed only. Any "$X per minute for OpenAI Realtime" figure is somebody else's conversion. source

The arithmetic

One five-minute call

Priced from the tables above. Two of these resolve to a real number. Two do not, and that is the finding: when speech synthesis bills per character and the model bills per token, the vendor cannot tell you what a call costs until it has happened.

Retell AI $0.55

resolves exactly

Their own default calculator, bring-your-own SIP

Voice infrastructure $0.055 + model $0.04 + speech synthesis $0.015 per minute, telephony $0 on your own SIP. $0.11 per minute.

source
Retell AI $1.35

resolves exactly

Premium voices and a top-tier model

Voice infrastructure $0.055 + ElevenLabs synthesis $0.040 + GPT-5.5 $0.16 + Twilio $0.015. $0.27 per minute, or roughly 2.5x the default.

source
Bland AI $0.74

resolves exactly

Start plan, Twilio local inbound

$0.14 per minute covering model, speech-to-text and synthesis, plus about $0.04 of telephony across five minutes.

source
ElevenLabs $0.40

floor only, cannot resolve

Agents, additional minutes

$0.08 per minute for the platform. The model and telephony are billed at cost on top, so the total depends on which model you attach.

source
Vapi $0.32

floor only, cannot resolve

Build plan, components at cost

A verifiable FLOOR, not a price: hosting $0.25 + speech-to-text $0.024 + telephony $0.043. Speech synthesis is per character and the model is per token, so neither can be resolved without knowing how much the agent says.

source

The same call runs from about $0.55 to $1.35 on one vendor alone, depending only on which voice and which model you attach. Before comparing two platforms, check you are comparing the same stack.

Layer 4

The people

The largest line on most AI budgets, and the one where published numbers are least comparable. A median wage, a total-compensation figure including equity, and a contractor day rate are three different measurements that all get quoted as "what an AI engineer costs". Read the measure column before the number.

Source Role Measure Figure Notes Link
US Bureau of Labor Statistics Data Scientists (SOC 15-2051) median annual wage $120,230per year The 75th percentile is $158,880. Wages only: apply the employer cost loading below for a real comparison. source
US Bureau of Labor Statistics Software Developers (SOC 15-1252) median annual wage $135,980per year The baseline AI roles get compared against. There is NO BLS occupation code for "AI Engineer", so any figure attributed to BLS under that title is invented. source
US Bureau of Labor Statistics Employer cost loading multiplier on wages $1.43x wages, for total employer cost Wages are 69.9% of employer cost and benefits 30.1%. A $120,230 salary costs roughly $172,000 to employ. source
Levels.fyi Machine Learning Engineer median total compensation $279,000per year Not comparable to a BLS wage, because it includes equity. The same site on the same day puts "AI Engineer" at $154,000: an 81% gap between two titles for similar work. source
YunoJuno AI Engineer, freelance average contract rate $70per hour Machine Learning Engineer books at $62. The report is titled 2026 but runs on 2024 to 2025 data. source
Arc.dev AI and ML engineering, freelance average SENIOR contract rate $110per hour, low end of range The range runs to $190. Read the column header: this is senior-only, not a market midpoint, which is most of why it sits about 2x above booking data. source
Clutch AI development firms, US and Canada agency hourly band $50per hour, low end of band The band runs to $99 and the modal global rate is $24 to $49. The same directory shows an identical band for general software development. source
Clutch Average AI project project total $120,594.55per project The most common range is $10,000 to $49,999. The comparable figure for general software projects is $132,480.29, which is HIGHER. At directory level there is no AI premium. source

There is no BLS code for "AI Engineer"

The nearest occupations are Data Scientists, Software Developers, and Computer and Information Research Scientists. Any figure attributed to the Bureau of Labor Statistics under the title "AI Engineer" was invented, because the category does not exist.

Rate guides and booking data differ by about 2x

Survey-based guides put freelance AI work at $110 to $190 an hour, but that column is senior-only. Platforms publishing what contracts actually booked at report $62 to $78. Both are real; they measure different populations.

At agency level there is no AI premium

The same directory shows an identical hourly band for AI firms and general software firms, and a slightly LOWER average project total for AI work. That contradicts most of what is written on the subject.

Context

Whether it pays off

Cost is only half of the question. Named research on returns, with the sample size and the date attached, and with the contested findings marked as contested rather than left out.

BCG September 2025 1,250 senior executives across 68 countries

In AI work, algorithms are about 10% of the effort, technology about 20%, and people, process and organisation about 70%. Among those surveyed, 74% named a shortage of AI talent as a constraint and 63% said scaling costs were hard to control.

source, checked 2026-08-27
IBM Institute for Business Value May 2025 2,000 CEOs across 33 countries, fielded February to April 2025

Only 25% of AI initiatives have delivered the return that was expected of them, and only 16% have scaled across the enterprise.

source, checked 2026-08-27
Gartner June 2025 Poll of 3,412 webinar attendees, January 2025

Over 40% of agentic AI projects will be cancelled by the end of 2027.

Read with care. This is a webinar poll, not a controlled survey. Read it as a signal about expectations, not a measured failure rate.

source, checked 2026-08-27
MIT NANDA July 2025 300+ initiatives, 52 organisations interviewed, 153 leaders surveyed

Widely quoted as finding that 95% of organisations get zero return from generative AI. The same report found externally built systems reached deployment about twice as often as internally built ones.

Read with care. Contested, and included mainly so the caveat travels with the number. The document is labelled preliminary findings, is not peer reviewed, and its own limitations section concedes the build-versus-buy split rests on interviews rather than market data. The methodology repeated in the press, "150 interviews and 350 employee surveys", is a misquote of the figures above.

source, checked 2026-08-27

How to read a quote

When two quotes for the same agent differ by an order of magnitude, they are usually pricing different layers. Ask which of these four a number covers.

  1. Layer 1

    Inference

    Per million tokens, priced above. Predictable, and usually the smallest line.

  2. Layer 2

    Platform and orchestration

    A seat, a task, a credit, or a conversation, depending on the vendor. The unit matters more than the headline price, because it decides what happens to your bill when volume grows.

  3. Layer 3

    Everything the agent touches

    Tool calls, search, telephony if it speaks, and the systems it writes into. Anthropic bills web search at $10 per 1,000 searches on top of tokens; an agent that searches every turn adds a line most models miss.

  4. Layer 4

    The people

    Building it, connecting it to your systems, and keeping it working when a prompt, a model, or an API changes underneath it. Priced above, and the largest line on most budgets. BCG puts algorithms at about 10% of AI effort and people, process and organisation at about 70%.

Method

Every figure was read off the vendor's own published pricing page on the date shown beside it, and each row links back to that page so you can check it. Nothing on this page is estimated, remembered, or averaged from third-party summaries.

Where a vendor does not publish a number, the index says "not published" instead of guessing. That is a finding in itself: a price you cannot see before a sales call is a price you cannot compare.

Prices move. This page carries the date it was last checked at the top, and the section above lists the changes vendors have already dated. If you are reading this well after 2026-08-27, follow the source links before you rely on a number.

Take the data. Every row on this page is downloadable as CSV or JSON, including the source URL and check date for each figure, under CC BY 4.0. Attribution with a link is enough.

Questions

How much does it cost to run an AI agent?

The model itself is usually the smallest line. Anthropic publishes a worked example of 10,000 customer-support conversations on Claude Haiku 4.5 costing about $37 in total, which is roughly a third of a cent per conversation. What makes a real bill large is everything stacked on top: the platform subscription, the tools the agent calls, the telephony if it speaks, and the engineering time to build and maintain it.

Why do two quotes for the same AI agent differ so much?

Because they are usually quoting different layers. One may be quoting inference only, another a platform seat, another a full build with integration and maintenance. Ask any vendor which of the four layers their number covers, and what happens to it at ten times the volume.

Does prompt caching really matter for agents?

More than for most other AI work. An agent re-sends the same system prompt and tool definitions on every turn, so that content is paid for repeatedly. Anthropic prices a cache read at 10% of the standard input rate, with a write at 1.25x for the five-minute cache, which means the cache pays for itself after a single read.

Which published AI prices are about to change?

Three dated changes are live as of this update: OpenAI states gpt-5.6-sol pricing is promotional at least through 2026-11-21, Google has dated a roughly 2x rise on Gemini 3.7 and 3.6 Flash for 2027-01-01, and Anthropic has cancelled the increase that was scheduled for Claude Sonnet 5 on 2026-09-01, making the launch rate standard.

How is this index kept accurate?

Every row records the URL it was read from and the date it was checked, and the page shows that date at the top. Nothing here is estimated or remembered. Where a vendor does not publish a number, the index says so rather than guessing, because "contact sales" is itself useful information when you are comparing options.

Working out what yours would cost?

The layers above tell you what the parts cost. What they cannot tell you is which workflow in your business is worth pointing an agent at first. That is what the free plan is for.