Back to Blog
Cost Management

How n8n AI Agent Costs Add Up Per Execution

The short answer: add every model call

For an n8n AI Agent using your own provider API keys, calculate the AI cost of an execution by adding the billable input and output for every model request inside that run, then adding any separately billed tools or services. A workflow execution, a model request and a tool invocation are different units. One execution can contain several of each.

n8n prices its plans by workflow executions. That platform charge is separate from the model bill in a bring-your-own-key workflow. n8n also advertises AI access without API keys; check the billing arrangement of the node you actually use before assuming a separate provider invoice.

This guide covers the calculation inside an agent run. It does not quote a universal price per agent, because model choice, context, task difficulty and billing route change the answer. Use it to build an estimate, then replace the assumptions with your own recorded usage.

Three meters to keep separate

Keep three columns in your estimate: n8n subscription or hosting, model API usage, and other services. A CRM lookup may use no additional language model, yet still have its own service costs. A second AI model used as a tool adds its own token bill. Your gateway subscription, if you use one, is another platform cost.

Consider a support triage workflow: a trigger starts the run; the agent asks the model which record to fetch; a lookup returns a ticket; the agent asks the model to draft a response. That is one illustrative workflow run with two model requests and one lookup. Counting only the final answer misses the first request.

For a monthly forecast, multiply the measured average AI cost per execution by expected execution volume, then add platform and service charges. Keep failed runs and human rework in view: a low token bill is less useful if the automation rarely completes its job.

A worked example: three turns, one task

An illustrative three-turn support agent with 9,000 uncached input tokens and 1,000 billable output tokens costs $0.028 at the assumed rates below. Add the cost of each turn to estimate the execution. This is not a customer result or a measured benchmark. Assume a text model costs $2 per million uncached input tokens and $10 per million billable output tokens. All three requests use those same assumed rates; no cache discounts, hosted-tool fees, taxes or platform charges are included.

Turn 1 chooses a customer lookup: 2,000 input tokens and 200 output tokens. Cost = (2,000 × $2 + 200 × $10) ÷ 1,000,000 = $0.006.

Turn 2 reads the lookup and requests an order lookup: 3,000 input tokens and 200 output tokens. Cost = (3,000 × $2 + 200 × $10) ÷ 1,000,000 = $0.008.

Turn 3 reads the order and drafts the reply: 4,000 input tokens and 600 output tokens. Cost = (4,000 × $2 + 600 × $10) ÷ 1,000,000 = $0.014.

Total: 9,000 input tokens, 1,000 billable output tokens and $0.028 per execution. At 10,000 executions, that is $280 in model spend under these assumptions. Counting only the final turn would produce $140 and omit half of the modeled bill.

Use your actual model rates in this calculation. If a request has cached input or cache writes, split those into their own priced categories. If output usage already includes reasoning tokens, do not add the reasoning subtotal a second time.

Why tools and memory change the input bill

The Tools Agent uses tool definitions and schemas to let the model choose actions. Instructions, tool descriptions, conversation history and returned records can all contribute to what the model receives. Inspect the actual usage on each request instead of assuming the user’s message is the whole input.

When a lookup returns a large record and that record is included in a later request, you pay for that later input according to the model’s billing rules. If the same history appears on several requests, it can be billed several times; an eligible cache hit may change its rate. Two tools can also be requested together, so tool count alone does not tell you model request count.

Give tools concise descriptions and return the fields the next decision needs. For a ticket response, an order status and delivery date may be more useful than a complete account export. Compare the shortened result against the same tasks before keeping the change: removing a field that the agent needs can cause another lookup or an incorrect answer.

Bound memory to the history the task needs. Test a fresh conversation and a long conversation separately. An average from fresh conversations alone will understate costs if most users return to an established thread.

Retries and repairs: count what actually ran

A retry adds model cost when it causes another billable provider request. Do not assume every error is billed at the successful-call price: a request rejected before generation differs from a completed generation whose response your workflow failed to receive. A timeout by itself does not establish zero provider usage.

In the worked example, repeating the full three-turn path once with identical billable usage would cost $0.056 rather than $0.028. Retrying only the final turn once would cost $0.042. Neither is a prediction of your retry bill: identify where the retry begins and inspect the usage available for each attempt.

n8n’s Auto-fixing Output Parser can call another LLM when the first parser fails. Include that repair request in the estimate. Likewise, trace a specialist AI tool separately from its parent agent; the parent’s final answer is not the specialist’s entire cost.

Before reducing retries, find what they are repairing. An invalid tool argument, an unavailable provider and a malformed final answer need different fixes. For tools that change records or send messages, check whether an action already completed before replaying the run.

Thinking tokens and the Think Tool are different

OpenAI bills reasoning tokens as output tokens, even when those tokens are not visible in the answer. Read the provider’s usage fields and model-specific output limits. A short visible reply does not necessarily imply little billable output.

n8n’s Think Tool invites an agent to reflect through a tool step. It is different from a provider’s internal reasoning or thinking control. It can change the agent’s sequence and context; measure the resulting model requests rather than assigning it a fixed token surcharge.

For the illustrative $10-per-million output rate, an extra 1,000 billable reasoning tokens would add $0.01. If those tokens are already in the reported output total, that $0.01 is already counted. Provider accounting differs, so use the applicable usage definition rather than summing every token field you see.

Test a lower supported reasoning setting only against tasks that still meet your quality requirements. Less thinking can lower usage on some workloads, but another tool loop or a failed answer can erase that saving.

Iteration limits are a control, not a dollar forecast

n8n documents Max Iterations with a default of 10. It limits how many times the model runs while trying to answer. It does not specify a fixed dollar amount: each turn may have a different input size, output length or tool charge.

Choose a bound that allows the normal task to finish, and define what happens when it is reached. Keep node retries, parser repair and any nested AI workflow in the same review. An agent iteration setting should not be treated as a global limit over all of those paths.

For spend enforcement, use a separate budget policy. TokenSense supports budget controls for routed requests, including project and key budgets on Pro and Agency. Allow for requests already in flight and usage that cannot be priced; a configured cap is not a guarantee that the final provider invoice will equal that number.

Measure a small sample before forecasting

1. Duplicate your workflow and choose representative tasks: an easy lookup, a multi-tool task, missing data, a long conversation and a failure that invokes your retry or repair path. Use test records and disable actions that would send real messages or modify live records.

2. Record the workflow execution ID, model and endpoint, every model attempt, input categories, billable output, priced tool usage, latency and final outcome. Keep the pricing date beside the rate used. Mark missing usage or unavailable pricing as unknown.

3. Group model requests by execution, then compare the average with expensive outliers. Also calculate total measured AI spend divided by successfully completed tasks; keep failures in the numerator. That metric reveals a cheap model that repeatedly needs repairs.

4. Change one cost driver at a time: memory length, tool-result size, model, reasoning effort or retry policy. Re-run the same tasks and check correctness before accepting the lower bill. Repeat the estimate when prompts, tools or model versions change.

For routed calls, use TokenSense’s n8n attribution guide to connect request logs to workflows, executions and steps. Calls that bypass the gateway and non-AI service bills need separate accounting.

Unknown cost is not a free execution

A successful response with no calculable cost belongs in an unknown-cost bucket. Do not replace missing usage with zero when totaling an execution or comparing providers. Report the priced subtotal and the number of unpriced attempts separately, then reconcile with provider billing when available.

The n8n node reference explains operation-specific billing metadata, including pricingAvailable and billingReason for image and transcription operations in version 0.1.20. Speech output depends on gateway billing headers. These fields do not make every modality fully priced; check the operation you use.

Investigating a sudden bill change? Use the diagnosis guide after you have the per-execution breakdown. The calculation here tells you what accumulated; the diagnosis helps find what changed.

Start with one workflow

TokenSense is an AI gateway for automation teams that brings routed AI usage and budget controls into one place. Start with a copy of one workflow, connect a Chat-compatible model through the TokenSense Chat Model sub-node, and inspect the attributed request logs. Compare those logs with the execution and provider usage before forecasting the rest of your workflows.

Check model and endpoint compatibility before connecting tools. GPT 6.1 Sol and GPT 6 Astra tool requests require the Responses API; the Chat Model path uses Chat Completions. A successful text-only call is not proof that an agent’s tools will work.

Start free with TokenSense to inspect one workflow’s routed usage. For installation and the current Cloud rollout, follow the setup guide below; package approval and npm publication are different from Cloud availability.

n8n installation and updates