Skip to content
intermediate

Model the Cost of Choosing an Agent or Fixed Workflow

Most teams pick agents or workflows by vibe. The demo looked smart, the diagram looked clean, and nobody wrote down what the choice would cost per run.…

Published 2026-10-03Updated 2026-10-0412 min read
Close-up of ilmenite and yellow sand forming natural abstract patterns on a beach.
Close-up of ilmenite and yellow sand forming natural abstract patterns on a beach. Photo by Thilina Alagiyawanna on Pexels.

Most teams pick agents or workflows by vibe. The demo looked smart, the diagram looked clean, and nobody wrote down what the choice would cost per run. Then production arrives, and the bill shows up in a place nobody was watching: the exception that had no handler, the loop that never hit its exit condition, the wrong action a human had to unwind by hand.

You already know the structural difference. In a fixed workflow, your code decides the path. In an agent, the model decides the path at runtime. That question is settled. The question this article answers is narrower and more useful: given one task, what does that runtime control actually cost?

Flexibility is a property, not a payoff. The payoff depends on how often the task is genuinely ambiguous and how expensive a wrong action is. So we are going to build a small expected-cost model, plug in real numbers, and let the arithmetic argue with your instinct.

Why "Agents Are More Flexible" Is Not a Decision

Flexibility is a property of a design. Cost is a consequence of running it. Those are different things, and conflating them is how teams end up with an agent that improvises beautifully on the 80% of inputs a workflow already handled, and improvises expensively on the 20% that needed a human.

Two cost drivers do most of the work, and most people leave both implicit:

  • The probability that the fixed path fails to cover the input at all.
  • The loss when the system takes an incorrect action.

Everything else — token spend, latency, tool calls — is bookkeeping around those two. Get them wrong and the model is decoration. Get them roughly right and the model tells you which design survives contact with your actual input mix.

Scope check: this is one task, one decision, one expected-cost comparison. Write your assumptions down, because a model you cannot challenge is a model you cannot trust.

The Variables You Need Before Any Comparison

Define notation once, tie every symbol to something you can observe, and the arithmetic later becomes auditable instead of magical.

Task uncertainty. Let p be the probability that an input falls outside what the fixed workflow was designed to handle. This is a property of your input distribution, not of the model. If 20% of inbound requests need judgment the script never anticipated, p = 0.2.

Workflow exception handling. When an out-of-scope input arrives, you route it somewhere. Let c_esc be the cost of that routing — a human review, a fallback path, a queue.

Agent behavior. Let n be the expected number of reasoning and tool iterations per run, and c_step the cost per iteration (tokens, latency, tool calls bundled together). The agent's variable core is n * c_step.

Incorrect-action impact. Let c_err be the loss when the system takes a wrong action, and p_err the probability of a wrong action given the design. Note that c_err is often asymmetric between the two designs: a workflow that confidently misroutes a request may cost more or less than an agent that takes a wrong tool action.

Three assumptions to state plainly before we compute anything:

  1. Costs are per-run averages, not worst cases.
  2. Iterations are independent enough that a mean is a fair summary.
  3. c_err is a single scalar standing in for a distribution of bad outcomes.

Note: That third assumption is the weakest one. We will come back to it, because it is also the one most likely to flip your decision.

Knowledge check

Check your understanding

Answer this question before you continue.

In this model, what does `p` represent?
Misconception Check

Focus: Distinguish uncertainty in the task input distribution from uncertainty introduced by the model.

Deriving Expected Cost for the Fixed Workflow

Build the equation from the branches, not from memory.

A run splits into two cases. With probability 1 - p, the input is in scope and the workflow handles it at baseline cost c_base. With probability p, the input is out of scope and the workflow pays c_base + c_esc — the baseline plus the escape hatch.

So the handling cost is:

E[handling] = (1 - p) * c_base + p * (c_base + c_esc)

Expand and simplify. The c_base terms collapse:

E[handling] = c_base + p * c_esc

Now add the error term. Incorrect actions can happen on either branch, so the expected loss is p_err_wf * c_err:

E[workflow] = c_base + p * c_esc + p_err_wf * c_err

Read it in plain language: baseline cost, plus the cost of exceptions weighted by how often they occur, plus the cost of being wrong weighted by how often you are wrong.

The interpretation matters more than the formula. When p is high, the workflow's cost is dominated by exception handling — it is visibly struggling, which is at least honest. When p is low but p_err_wf is not, the workflow is confidently wrong, and that is the expensive failure mode because nobody notices until downstream.

Knowledge check

Check your understanding

Answer this question before you continue.

Holding the workflow's other inputs fixed, what is the effect of increasing `p` by 0.1 on its expected cost?
Comparison Reasoning

Focus: Explain how the fixed workflow's expected cost responds to a change in the out-of-scope input rate.

Deriving Expected Cost for the Agent

Same discipline, different mechanism. To compare like with like, the agent side has to pay for the same work the workflow side does: a baseline handling cost, plus the cost of runs that fail and need a fallback.

The agent's cost has a variable core: n * c_step. Every iteration spends tokens, waits on latency, and may call a tool. That product is the price of runtime decision-making.

Then the same error term, with its own probability:

E[agent] = c_base + n * c_step + p_err_ag * c_err

Why does the agent get a separate p_err_ag? Because an agent's error probability is usually lower on genuinely ambiguous inputs — it can inspect, retry, and choose a different tool — but it is not zero. Agents fail differently: wrong tool selection, cascading errors where one bad step sends the trajectory somewhere unrecoverable, or a loop that burns budget without converging.

The interpretation: the agent pays a fixed per-run overhead for flexibility. That overhead is only repaid when it converts expensive exceptions or errors into cheap successful runs. If it does not, you are paying a premium for autonomy you are not using.

Knowledge check

Check your understanding

Answer this question before you continue.

An agent uses one additional expected iteration, while `c_step` and all other inputs stay fixed. What changes in the model?
Scenario Interpretation

Focus: Interpret the per-run cost contribution of agent iterations in the expected-cost model.

Worked Example: One Task, Two Numbers

Pick a small, plausible task: classifying and routing inbound support requests. Most are routine. A minority need judgment — a refund dispute, an ambiguous account issue, a request that spans two policies.

Assign values:

VariableValueMeaning
c_base1.0Baseline cost per run
p0.220% of inputs are out of scope
c_esc6.0Cost to route an exception to a human
p_err_wf0.05Workflow wrong-action rate
c_err20.0Loss per incorrect action
n3Expected agent iterations
c_step1.2Cost per iteration
p_err_ag0.02Agent wrong-action rate

Compute the workflow side line by line:

E[workflow] = c_base + p * c_esc + p_err_wf * c_err
            = 1.0 + 0.2 * 6.0 + 0.05 * 20.0
            = 1.0 + 1.2 + 1.0
            = 3.2

The exception term (1.2) and the error term (1.0) contribute almost equally. Neither dominates.

Now the agent side:

E[agent] = c_base + n * c_step + p_err_ag * c_err
         = 1.0 + 3 * 1.2 + 0.02 * 20.0
         = 1.0 + 3.6 + 0.4
         = 5.0

The verdict: the workflow wins on this task, 3.2 to 5.0. The agent's flexibility is real, but at n = 3 iterations it costs more than the exceptions it was hired to prevent.

What the numbers do not capture: maintenance, observability, and the one-time cost of building the exception path in the first place. Those are real, and they usually favor the workflow too — but they are not in the equation, so do not pretend the equation settled them.

Change One Assumption, Flip the Answer

A single number should be able to change the recommendation. If it cannot, you are not modeling, you are rationalizing.

Raise p from 0.2 to 0.5. More inputs fall outside the script.

E[workflow] = 1.0 + 0.5 * 6.0 + 0.05 * 20.0 = 1.0 + 3.0 + 1.0 = 5.0
E[agent]    = 1.0 + 3 * 1.2 + 0.02 * 20.0   = 1.0 + 3.6 + 0.4 = 5.0

The two designs tie. The workflow's exception handling grew to match the agent's iteration overhead.

Raise p to 0.6. Push past the tie.

E[workflow] = 1.0 + 0.6 * 6.0 + 0.05 * 20.0 = 1.0 + 3.6 + 1.0 = 5.6
E[agent]    = 1.0 + 3 * 1.2 + 0.02 * 20.0   = 1.0 + 3.6 + 0.4 = 5.0

The agent now wins. The workflow's exception handling became the dominant cost.

Raise c_err from 20 to 60. Error cost is now the largest term on both sides.

E[workflow] = 1.0 + 0.2 * 6.0 + 0.05 * 60 = 1.0 + 1.2 + 3.0 = 5.2
E[agent]    = 1.0 + 3 * 1.2 + 0.02 * 60   = 1.0 + 3.6 + 1.2 = 5.8

The workflow wins again — not because it is smarter, but because the agent's iteration overhead still outweighs its lower error probability at this error cost. Push c_err higher and the agent's advantage grows.

Raise n from 3 to 8. The agent loop gets longer.

E[agent] = 1.0 + 8 * 1.2 + 0.02 * 20.0 = 1.0 + 9.6 + 0.4 = 11.0

The agent loses badly. An unbounded or poorly capped loop destroys its own advantage.

The general rule the variations reveal: the crossover point is the real output of the model, not the absolute numbers. Solve for the value of p where the two sides are equal, and you have the condition under which your recommendation flips.

Finding the Crossover Point

A cost-versus-p plot shows the workflow cost rising from 2 at p = 0 to 8 at p = 1, while the agent cost stays flat at 5. The lines cross at p = 0.5; the workflow is cheaper to the left and the agent to the right.
The crossover marks the out-of-scope rate above which the agent becomes cheaper under the example’s assumptions.

Set the two equations equal and solve for p. Everything else stays fixed at the example values.

c_base + p * c_esc + p_err_wf * c_err = c_base + n * c_step + p_err_ag * c_err

The c_base terms cancel on both sides:

p * c_esc + p_err_wf * c_err = n * c_step + p_err_ag * c_err

Isolate p:

p * c_esc = n * c_step + (p_err_ag - p_err_wf) * c_err
p = [n * c_step + (p_err_ag - p_err_wf) * c_err] / c_esc

Plug in the example values:

p = [3 * 1.2 + (0.02 - 0.05) * 20.0] / 6.0
  = [3.6 + (-0.6)] / 6.0
  = 3.0 / 6.0
  = 0.5

The crossover is p = 0.5. Below that, the workflow is cheaper. Above it, the agent is cheaper. The tie we saw when we raised p to 0.5 was not a coincidence — it was the boundary.

This is the decision rule the model actually produces. Not "agents are better" or "workflows are better," but: for this task, with these costs, the agent only pays off if more than half your inputs fall outside the script. If your real p is 0.15, the agent is a luxury. If it is 0.7, the workflow is a liability.

Common mistake: Treating the first computed answer as a permanent verdict. It is a snapshot of current assumptions. Change the input mix, change the error cost, change the iteration cap, and the snapshot is stale.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the example values `n = 3`, `c_step = 1.2`, `p_err_ag = 0.02`, `p_err_wf = 0.05`, `c_err = 20`, and `c_esc = 6`, at what value of `p` do the expected costs tie?
Output Prediction

Focus: Calculate the task-uncertainty crossover from the article's example assumptions.

Where the Model Breaks Down

Bound the method honestly before you bet on it.

c_err is the hardest term to estimate. It is also the one most likely to be wrong. Treat it as a range, not a point value, and check whether the decision survives the range. If the answer flips inside the range, you do not have a decision yet — you have a measurement problem.

The model assumes a single task type. Mixed workloads need the comparison run per task class. One number for the whole system hides the task where the agent actually pays off.

Hybrid designs break the clean two-way split. An agent inside a workflow, or a workflow exposed as a tool the agent can call, changes which terms apply. The equations still work, but you have to re-derive them for the actual topology.

Reversibility matters more than the arithmetic when c_err is catastrophic. A low expected cost does not justify an irreversible action. If the wrong action cannot be undone, the expected value is the wrong tool — you want a gate, not a calculation.

Separate what is known from what is inferred. The equations are definitional; they are true by construction. Every input value is an estimate. Measure or bound them before they drive a real decision.

Turning the Calculation Into a Decision Rule

Here is the procedure, compressed to something you can run this week.

  1. Write the four inputs first. p, c_esc, n * c_step, and c_err with both error probabilities. If you cannot estimate one, write a range.
  2. Compute both expected costs. Use the two equations. Show the arithmetic.
  3. Compute the crossover. Solve for the value of p at which the two sides are equal. That value is your decision boundary.
  4. Choose the simpler design when the costs are close. The agent carries maintenance and observability burden the equation does not price. If the numbers are within noise, the workflow wins by default.
  5. Re-run the model whenever the inputs move. Input mix, error cost, iteration cap. Treat the result as a conditional recommendation, not a verdict about agents versus workflows in general.

The point of the model is not to produce a number you can defend in a meeting. It is to make your assumptions visible enough that reality can correct them. Pick one real task, write the four numbers down, and let the arithmetic argue with your instinct. If it loses the argument, you learned something cheap. If it wins, you just avoided paying for flexibility you never needed.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

In the article's example, if `c_err` rises from 20 to 60 while all other values remain fixed, which design has the lower modeled expected cost?
Question 1 of 2Comparison Reasoning

Focus: Use changed error-cost assumptions to compare the two designs without treating a lower error rate as an automatic win.

A task's wrong action could cause catastrophic harm and cannot be undone, even though the model estimates a low expected cost. What does the article recommend?
Question 2 of 2Scenario Interpretation

Focus: Recognize when an expected-cost comparison is insufficient because an incorrect action is catastrophic and irreversible.

References

  1. AI Agents vs Workflows: When to Use Eachredis.io
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.