Model the Cost of Choosing an Agent or Fixed Workflow
Most teams pick agents or workflows by vibe. The demo looked smart, the diagram looked clean, and nobody wrote down what the choice would cost per run.…

Key topics
Most teams pick agents or workflows by vibe. The demo looked smart, the diagram looked clean, and nobody wrote down what the choice would cost per run. Then production arrives, and the bill shows up in a place nobody was watching: the exception that had no handler, the loop that never hit its exit condition, the wrong action a human had to unwind by hand.
You already know the structural difference. In a fixed workflow, your code decides the path. In an agent, the model decides the path at runtime. That question is settled. The question this article answers is narrower and more useful: given one task, what does that runtime control actually cost?
Flexibility is a property, not a payoff. The payoff depends on how often the task is genuinely ambiguous and how expensive a wrong action is. So we are going to build a small expected-cost model, plug in real numbers, and let the arithmetic argue with your instinct.
Why "Agents Are More Flexible" Is Not a Decision
Flexibility is a property of a design. Cost is a consequence of running it. Those are different things, and conflating them is how teams end up with an agent that improvises beautifully on the 80% of inputs a workflow already handled, and improvises expensively on the 20% that needed a human.
Two cost drivers do most of the work, and most people leave both implicit:
- The probability that the fixed path fails to cover the input at all.
- The loss when the system takes an incorrect action.
Everything else — token spend, latency, tool calls — is bookkeeping around those two. Get them wrong and the model is decoration. Get them roughly right and the model tells you which design survives contact with your actual input mix.
Scope check: this is one task, one decision, one expected-cost comparison. Write your assumptions down, because a model you cannot challenge is a model you cannot trust.
The Variables You Need Before Any Comparison
Define notation once, tie every symbol to something you can observe, and the arithmetic later becomes auditable instead of magical.
Task uncertainty. Let p be the probability that an input falls outside what the fixed workflow was designed to handle. This is a property of your input distribution, not of the model. If 20% of inbound requests need judgment the script never anticipated, p = 0.2.
Workflow exception handling. When an out-of-scope input arrives, you route it somewhere. Let c_esc be the cost of that routing — a human review, a fallback path, a queue.
Agent behavior. Let n be the expected number of reasoning and tool iterations per run, and c_step the cost per iteration (tokens, latency, tool calls bundled together). The agent's variable core is n * c_step.
Incorrect-action impact. Let c_err be the loss when the system takes a wrong action, and p_err the probability of a wrong action given the design. Note that c_err is often asymmetric between the two designs: a workflow that confidently misroutes a request may cost more or less than an agent that takes a wrong tool action.
Three assumptions to state plainly before we compute anything:
- Costs are per-run averages, not worst cases.
- Iterations are independent enough that a mean is a fair summary.
c_erris a single scalar standing in for a distribution of bad outcomes.
Note: That third assumption is the weakest one. We will come back to it, because it is also the one most likely to flip your decision.
Knowledge check
Check your understanding
Answer this question before you continue.
Deriving Expected Cost for the Fixed Workflow
Build the equation from the branches, not from memory.
A run splits into two cases. With probability 1 - p, the input is in scope and the workflow handles it at baseline cost c_base. With probability p, the input is out of scope and the workflow pays c_base + c_esc — the baseline plus the escape hatch.
So the handling cost is:
E[handling] = (1 - p) * c_base + p * (c_base + c_esc)
Expand and simplify. The c_base terms collapse:
E[handling] = c_base + p * c_esc
Now add the error term. Incorrect actions can happen on either branch, so the expected loss is p_err_wf * c_err:
E[workflow] = c_base + p * c_esc + p_err_wf * c_err
Read it in plain language: baseline cost, plus the cost of exceptions weighted by how often they occur, plus the cost of being wrong weighted by how often you are wrong.
The interpretation matters more than the formula. When p is high, the workflow's cost is dominated by exception handling — it is visibly struggling, which is at least honest. When p is low but p_err_wf is not, the workflow is confidently wrong, and that is the expensive failure mode because nobody notices until downstream.
Knowledge check
Check your understanding
Answer this question before you continue.
Deriving Expected Cost for the Agent
Same discipline, different mechanism. To compare like with like, the agent side has to pay for the same work the workflow side does: a baseline handling cost, plus the cost of runs that fail and need a fallback.
The agent's cost has a variable core: n * c_step. Every iteration spends tokens, waits on latency, and may call a tool. That product is the price of runtime decision-making.
Then the same error term, with its own probability:
E[agent] = c_base + n * c_step + p_err_ag * c_err
Why does the agent get a separate p_err_ag? Because an agent's error probability is usually lower on genuinely ambiguous inputs — it can inspect, retry, and choose a different tool — but it is not zero. Agents fail differently: wrong tool selection, cascading errors where one bad step sends the trajectory somewhere unrecoverable, or a loop that burns budget without converging.
The interpretation: the agent pays a fixed per-run overhead for flexibility. That overhead is only repaid when it converts expensive exceptions or errors into cheap successful runs. If it does not, you are paying a premium for autonomy you are not using.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: One Task, Two Numbers
Pick a small, plausible task: classifying and routing inbound support requests. Most are routine. A minority need judgment — a refund dispute, an ambiguous account issue, a request that spans two policies.
Assign values:
| Variable | Value | Meaning |
|---|---|---|
c_base | 1.0 | Baseline cost per run |
p | 0.2 | 20% of inputs are out of scope |
c_esc | 6.0 | Cost to route an exception to a human |
p_err_wf | 0.05 | Workflow wrong-action rate |
c_err | 20.0 | Loss per incorrect action |
n | 3 | Expected agent iterations |
c_step | 1.2 | Cost per iteration |
p_err_ag | 0.02 | Agent wrong-action rate |
Compute the workflow side line by line:
E[workflow] = c_base + p * c_esc + p_err_wf * c_err
= 1.0 + 0.2 * 6.0 + 0.05 * 20.0
= 1.0 + 1.2 + 1.0
= 3.2
The exception term (1.2) and the error term (1.0) contribute almost equally. Neither dominates.
Now the agent side:
E[agent] = c_base + n * c_step + p_err_ag * c_err
= 1.0 + 3 * 1.2 + 0.02 * 20.0
= 1.0 + 3.6 + 0.4
= 5.0
The verdict: the workflow wins on this task, 3.2 to 5.0. The agent's flexibility is real, but at n = 3 iterations it costs more than the exceptions it was hired to prevent.
What the numbers do not capture: maintenance, observability, and the one-time cost of building the exception path in the first place. Those are real, and they usually favor the workflow too — but they are not in the equation, so do not pretend the equation settled them.
Change One Assumption, Flip the Answer
A single number should be able to change the recommendation. If it cannot, you are not modeling, you are rationalizing.
Raise p from 0.2 to 0.5. More inputs fall outside the script.
E[workflow] = 1.0 + 0.5 * 6.0 + 0.05 * 20.0 = 1.0 + 3.0 + 1.0 = 5.0
E[agent] = 1.0 + 3 * 1.2 + 0.02 * 20.0 = 1.0 + 3.6 + 0.4 = 5.0
The two designs tie. The workflow's exception handling grew to match the agent's iteration overhead.
Raise p to 0.6. Push past the tie.
E[workflow] = 1.0 + 0.6 * 6.0 + 0.05 * 20.0 = 1.0 + 3.6 + 1.0 = 5.6
E[agent] = 1.0 + 3 * 1.2 + 0.02 * 20.0 = 1.0 + 3.6 + 0.4 = 5.0
The agent now wins. The workflow's exception handling became the dominant cost.
Raise c_err from 20 to 60. Error cost is now the largest term on both sides.
E[workflow] = 1.0 + 0.2 * 6.0 + 0.05 * 60 = 1.0 + 1.2 + 3.0 = 5.2
E[agent] = 1.0 + 3 * 1.2 + 0.02 * 60 = 1.0 + 3.6 + 1.2 = 5.8
The workflow wins again — not because it is smarter, but because the agent's iteration overhead still outweighs its lower error probability at this error cost. Push c_err higher and the agent's advantage grows.
Raise n from 3 to 8. The agent loop gets longer.
E[agent] = 1.0 + 8 * 1.2 + 0.02 * 20.0 = 1.0 + 9.6 + 0.4 = 11.0
The agent loses badly. An unbounded or poorly capped loop destroys its own advantage.
The general rule the variations reveal: the crossover point is the real output of the model, not the absolute numbers. Solve for the value of p where the two sides are equal, and you have the condition under which your recommendation flips.
Finding the Crossover Point
Set the two equations equal and solve for p. Everything else stays fixed at the example values.
c_base + p * c_esc + p_err_wf * c_err = c_base + n * c_step + p_err_ag * c_err
The c_base terms cancel on both sides:
p * c_esc + p_err_wf * c_err = n * c_step + p_err_ag * c_err
Isolate p:
p * c_esc = n * c_step + (p_err_ag - p_err_wf) * c_err
p = [n * c_step + (p_err_ag - p_err_wf) * c_err] / c_esc
Plug in the example values:
p = [3 * 1.2 + (0.02 - 0.05) * 20.0] / 6.0
= [3.6 + (-0.6)] / 6.0
= 3.0 / 6.0
= 0.5
The crossover is p = 0.5. Below that, the workflow is cheaper. Above it, the agent is cheaper. The tie we saw when we raised p to 0.5 was not a coincidence — it was the boundary.
This is the decision rule the model actually produces. Not "agents are better" or "workflows are better," but: for this task, with these costs, the agent only pays off if more than half your inputs fall outside the script. If your real p is 0.15, the agent is a luxury. If it is 0.7, the workflow is a liability.
Common mistake: Treating the first computed answer as a permanent verdict. It is a snapshot of current assumptions. Change the input mix, change the error cost, change the iteration cap, and the snapshot is stale.
Knowledge check
Check your understanding
Answer this question before you continue.
Where the Model Breaks Down
Bound the method honestly before you bet on it.
c_err is the hardest term to estimate. It is also the one most likely to be wrong. Treat it as a range, not a point value, and check whether the decision survives the range. If the answer flips inside the range, you do not have a decision yet — you have a measurement problem.
The model assumes a single task type. Mixed workloads need the comparison run per task class. One number for the whole system hides the task where the agent actually pays off.
Hybrid designs break the clean two-way split. An agent inside a workflow, or a workflow exposed as a tool the agent can call, changes which terms apply. The equations still work, but you have to re-derive them for the actual topology.
Reversibility matters more than the arithmetic when c_err is catastrophic. A low expected cost does not justify an irreversible action. If the wrong action cannot be undone, the expected value is the wrong tool — you want a gate, not a calculation.
Separate what is known from what is inferred. The equations are definitional; they are true by construction. Every input value is an estimate. Measure or bound them before they drive a real decision.
Turning the Calculation Into a Decision Rule
Here is the procedure, compressed to something you can run this week.
- Write the four inputs first.
p,c_esc,n * c_step, andc_errwith both error probabilities. If you cannot estimate one, write a range. - Compute both expected costs. Use the two equations. Show the arithmetic.
- Compute the crossover. Solve for the value of
pat which the two sides are equal. That value is your decision boundary. - Choose the simpler design when the costs are close. The agent carries maintenance and observability burden the equation does not price. If the numbers are within noise, the workflow wins by default.
- Re-run the model whenever the inputs move. Input mix, error cost, iteration cap. Treat the result as a conditional recommendation, not a verdict about agents versus workflows in general.
The point of the model is not to produce a number you can defend in a meeting. It is to make your assumptions visible enough that reality can correct them. Pick one real task, write the four numbers down, and let the arithmetic argue with your instinct. If it loses the argument, you learned something cheap. If it wins, you just avoided paying for flexibility you never needed.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


