Reducing Coding-Agent Costs: Stop Paying a Model to Rename a Variable
How to reduce coding-agent costs by moving deterministic fixes out of the agent loop before optimizing models, prompts or caches.

Most advice on reducing AI costs starts with the model: choose a cheaper one, improve caching or batch more requests together. Those changes can help, but they come after a more basic decision about the work itself. If a deterministic tool can produce the correct result, there is no reason to send that task to a language model.
Compilers, linters and formatters handle this class of work faster, more reliably and at almost no marginal cost. Using a model instead means paying for the same mechanical decision on every run without any guarantee that the agent will make it consistently. Our largest reduction in coding-agent costs came from adding linter rules that could repair their own findings. Continuous integration now fixes those issues before an agent ever sees them, without changing the model at all.
How to decide whether a coding task needs an LLM
Every task given to a coding agent can be classified by asking whether it has one correct answer that can be derived from the source code without understanding the purpose of the feature.
Import ordering, unused variables, formatting, naming conventions, missing awaits and throwing a bare string all fall into that category. The desired result follows from a rule, regardless of what the feature is supposed to do.
Questions about abstractions, failure behaviour or the long-term effect of a data model are different because the answer depends on intent and engineering judgment.
Language models are useful for the second category because they can reason about context and trade-offs. Applying them to the first category buys a slower and non-deterministic version of something established tools already do. The agent may fix the issue, but the platform cannot rely on it happening every time, and mechanical issues still account for a large share of code-review feedback.
What code review comments actually focus on
Research on code review has shown for more than a decade that reviewers spend much of their time on improvements that are not defects.
Bacchelli and Bird classified 570 comments from 200 Microsoft code-review threads for their ICSE 2013 paper. The largest category was code improvements, which included readability, comments, consistency and dead-code removal but explicitly excluded correctness and defects. It contained 165 comments, while defects ranked fourth among nine categories with 78.
Composition of 570 code-review comments from Bacchelli and Bird's 2013 study: 165 code improvements, 78 defects, 12 knowledge transfer, and 315 across six further categories
With 165 code-improvement comments and 78 defect findings, the first category appeared 2.1 times as often. The authors reported that defects represented only one eighth of the comments in their sample and were mostly concerned with superficial, micro-level issues.
The same study asked 873 programmers why they perform reviews, and 44% named finding defects as their main motivation. The outcome did not match the intention: reviewers believed they were looking for bugs, while much of their feedback concerned tidying. One senior developer in the study noted that some reviewers focus on formatting mistakes simply because they are easy to find.
LLM reviewers tend to produce the same kind of feedback because surface-level issues are common, easy to detect and safe to mention. Each comment now incurs model cost, and an implementation agent may then need another turn to address work that a deterministic tool could have handled before the review began.
What Google Tricorder teaches about automated code fixes
Google addressed the same problem years before coding agents existed by integrating static analysis and automated fixes directly into code review.
The Tricorder team describes the system in an ICSE paper by Sadowski, van Gogh, Jaspan, Söderberg and Winter. At the time of that paper, Tricorder produced around 93,000 analysis results per day and was maintained by a team of two or three people. Google's later write-up in Communications of the ACM reports that by January 2018 the system analysed roughly 50,000 code-review changes each day. Reviewers clicked "Please Fix" more than 5,000 times per day, and authors applied automated fixes around 3,000 times per day.
At around 3,000 automated repairs each day, Tricorder removed a large amount of work without requiring a separate request or manual edit. Google's principle was to measure defects corrected rather than warnings shown to developers, captured in the line "Do not just find bugs, fix them." That is also the right standard for deterministic checks in an agent workflow.
To keep automated findings trustworthy, Google admits an analysis into code review only when its effective false-positive rate stays below 10%. A report counts as a false positive whenever the developer chooses not to act on it, so the measure reflects usefulness rather than only technical correctness. An analyser enters probation when the ratio of "Not useful" to "Please fix" clicks exceeds 10%, and it can be disabled above 25%. Across all analysers, the rate was usually around 5%.
The point at which a warning appears also changes whether developers act on it. In Google's survey, developers considered 74% of issues reported at compile time to be genuine problems, compared with 21% when the same type of issue was found in code that had already been checked in. The paper concludes that static analysis depends on careful integration with the developer workflow. For a coding agent, the equivalent of compile time is the period before its turn begins. Once the finding reaches the agent, the platform has already started paying for it.
How to write a safe ESLint autofix rule
The following ESLint rule handles a common review comment: JavaScript code that throws a bare value instead of an Error and therefore loses the stack trace. The fixer allows continuous integration to repair safe cases before they become agent work.
/**
* house/throw-error-instance
*
* Flags `throw <non-Error>` and repairs the cases that can be repaired
* safely: a string literal or template literal becomes `new Error(...)`.
*
* Anything else — an object literal, a number, a computed value — is
* reported WITHOUT a fixer, on purpose. See the note below.
*/
export default {
meta: {
type: "problem",
docs: { description: "throw an Error instance, never a bare value" },
fixable: "code",
schema: [],
messages: {
bare: "Throw an Error instance so the stack trace survives.",
bareNoFix:
"Throw an Error instance. This value carries data, so wrap it " +
"yourself (new Error(msg, { cause })) instead of stringifying it.",
},
},
create(context) {
const src = context.sourceCode;
// Conservative: anything that could evaluate to an Error is left alone.
const mayBeError = (n) =>
n.type === "NewExpression" ||
n.type === "CallExpression" ||
n.type === "Identifier" ||
n.type === "MemberExpression" ||
n.type === "AwaitExpression";
const isStringish = (n) =>
n.type === "TemplateLiteral" ||
(n.type === "Literal" && typeof n.value === "string");
return {
ThrowStatement(node) {
const arg = node.argument;
if (!arg || mayBeError(arg)) return;
if (isStringish(arg)) {
context.report({
node,
messageId: "bare",
fix: (fixer) =>
fixer.replaceText(arg, `new Error(${src.getText(arg)})`),
});
return;
}
context.report({ node, messageId: "bareNoFix" }); // no fixer
},
};
},
};
Running the rule with --fix rewrites a string or template literal while leaving values with unclear semantics untouched:
$ npx eslint sample.js --fix
4:27 error Throw an Error instance. This value carries data, so wrap it
yourself (new Error(msg, { cause })) instead of stringifying it
- if (Number.isNaN(n)) throw `bad port: ${raw}`;
+ if (Number.isNaN(n)) throw new Error(`bad port: ${raw}`);
if (n < 1 || n > 65535) throw { code: "ERANGE", port: n }; // untouched
Leaving the object literal without a fixer avoids changing its meaning. An earlier version wrapped every value in String(...), which satisfied the lint rule but converted throw { code: "ERANGE", port: n } into an Error whose message was only [object Object]. That change silently discarded useful data and would have been applied automatically across every affected file.
Only rewrites that demonstrably preserve the program's meaning should run automatically; every other case should be reported without a fix. This is the machine equivalent of Google's requirement that a warning be easy to understand and have a clear solution. A fixer that is correct 90% of the time is unsafe because the remaining 10% can enter the main branch without anybody reviewing the change.
New rules are easier to refine inside one repository than in a shared package, where disagreements about the correct behaviour affect several teams at once. The implementation will differ by language, but the method stays the same: review a sample of recent comments, identify the most common deterministic issue and write a fixer only for the cases that can be changed safely.
How much an unnecessary agent review round costs
Even when a review round finds only issues that a tool could have fixed, the platform still pays for the full loop. The agent reloads its context, the gate runs again and a person is asked to revisit a pull request that has not changed in a meaningful way. If p is the share of review rounds caused only by those findings, the expected number of rounds grows as 1/(1 − p).
Cost per merged change as a multiplier, against the share of review rounds triggered only by machine-fixable findings: 1.25x at 20%, 1.54x at 35%, and 2.00x at 50%
When p is 0.2, the change costs 1.25 times the useful work. At p = 0.5, the cost doubles even though half of the loop contains no engineering judgment. Because the relationship is non-linear, the waste may look minor in a spreadsheet while becoming substantial across many runs. Teams can estimate their own p by classifying comments from a sample of merged pull requests.
AI cost optimization after automation
Routing work by task class allows cheaper models to handle routine implementation without weakening the parts that shape the whole change. Planning and specification benefit most from a strong model because a weak plan creates problems that downstream agents cannot repair. Routine implementation can use a workhorse model, with stronger models reserved for multi-file or security-sensitive changes. Apparent task size is a poor routing signal because a small-looking fix may turn out to touch several services.
The price per token is not the same as the cost of a completed task. A cheaper model may take more turns or fail more often, so compare cost and completion rate on the same set of tasks. Reasoning effort is a separate setting: planning and diagnosis may justify a higher setting, while a bounded implementation with a clear specification may not. Check the setting the tool actually uses, since its default can differ from the one you assumed.
Retries become less expensive when the unit of work contains fewer dependent steps. If a task contains n steps and each succeeds with probability r, the chance of completing the task is rⁿ. Large issues are therefore paid for repeatedly when runs fail and restart, although splitting work too aggressively introduces its own handoff costs. The task-sizing arithmetic is in our architecture post.
Ordering context carefully can reduce billed input during long sessions by preserving the prompt cache. Cache invalidation has specific triggers, and the price of a cache miss can matter more than the headline token rate. For a fuller treatment, the software factory pillar covers the cache economics in detail, including why rate-card pricing is not enough for planning. Model routing, task sizing and caching should still come after deterministic work has been removed from the agent loop, since none of them helps with a task that should never have reached a model.
The harness matters here because it decides which files and tool results go into each request. Keeping stable instructions ahead of changing context helps preserve a cached prefix; sending a short command excerpt with a path to the full output avoids filling that context with text the agent does not need. In an illustrative factory of 3,000 tasks a month, Matthias Lau models a change from 55% to 90% cache reads that brings a $90,000 monthly bill down to $41,700 at the same token volume. This is a scenario rather than a measured saving, but it shows why cache behaviour is worth checking before switching models.
AI cost management: tracking coding-agent costs by issue and repository
Useful cost data starts with each individual run and connects it to the originating issue and repository. Record the cost, input and output tokens, duration, harness, model and exit reason. That makes it possible to see which repositories and task types are becoming more expensive and whether the costly tail comes from one pathological workflow or a broader change in usage. Separate cached input reads from writes and record the effective reasoning-effort setting; a total input-token count can hide a costly drop in cache reuse.
In practice, a simple view of the last seven and thirty days is often enough when it names the most and least expensive runs. The extremes usually reveal a task that was too large and retried, a loop that should have stopped at a budget limit, or work that should have been handled by a deterministic tool in the first place.
Where not to cut coding-agent costs
Reducing the model used for planning and specification often increases total spend because this relatively small stage determines the quality of all the work that follows. A weak plan creates rework and retries throughout the rest of the pipeline.
Removing sandboxing or network controls saves visible overhead while accepting a much larger security risk. Code written by an agent must be treated as untrusted, and the spec is an attack surface that can influence what the agent executes.
Removing human review from consequential changes shifts cost into production rather than eliminating it. Failures discovered there are more expensive and measured in more serious terms than agent tokens or review time.
How to start reducing coding-agent costs
Start with a sample of merged pull requests and classify each review comment according to whether it required judgment or had one mechanically derivable answer. This gives you p and identifies the most common deterministic issue. Write a rule with a fixer for that issue, require fewer than 10% of findings to be rejected as not useful, and omit the fixer whenever the change is not provably safe. Run the rule in continuous integration before the agent starts, then repeat the review analysis to measure whether p falls.
Sources and further reading
- Bacchelli & Bird, Expectations, Outcomes, and Challenges of Modern Code Review, ICSE 2013 — 570 comments card-sorted, 873 programmers surveyed; the expectation/outcome gap.
- Sadowski, van Gogh, Jaspan, Söderberg & Winter, Tricorder: Building a Program Analysis Ecosystem, ICSE 2015 — the design principles, the effective-false-positive definition, and the not-useful thresholds.
- Lessons from Building Static Analysis Tools at Google, Communications of the ACM — the January 2018 figures, the "do not just find bugs, fix them" argument, and the compile-time vs. checked-in comparison.
- ESLint: Custom Rules — the fixer API used above, including the rules on what a fixer may safely change.
- Matthias Lau, The Hidden Economics of Agentic Software Engineering, Future Coding Day, 24 September 2026 — illustrative cost model and the effects of harness, effort and model choice.
Related articles
- Agentic AI Architecture: The Intent Plane and the Execution Plane — task sizing, and where cost telemetry belongs.
- What is FinOps for AI? — managing AI spend as an operational discipline, metered in tokens.
- What is a software factory? — what a factory costs to run, including cache-read economics.
- Lore: Shared Context Infrastructure for Claude Code — per-run cost attribution in a working platform.
- Wave: Bringing Determinism Back to AI-Assisted Development — deterministic checks that run every time, at pipeline scale.
Your next read

AI-Generated Code in Production: Validation, Accountability, and What Still Needs to Be Built
AI coding agents can fabricate variables and silently remove security checks. The capability is real, validation and governance lag behind.
Michael Czechowski7 mins readSubscribe to Our Bi-Weekly AI Native Newsletter
Lessons from our client work, engineering deep-dives, and the AI Native research worth reading.
Bi-weekly. No spam, unsubscribe anytime.





