

Inside Steve Yegge's Software Factory: Lessons from Running 50 Agents
Steve YeggeIndependent, Creator of Gas Town and Beads
Benedikt Stemmildthackers&wizardsOn This Podcast
Steve Yegge on running a 50-agent software factory: why he burned 40% of it down, and why enterprises should start with production agents, not coding agents.
- Steve burned down 40% of his software factory after its agents fenced themselves into paralysis.
- A newer model acted as an outside consultant and cleaned up a mess the previous model couldn't see.
- "Heresies", wrong beliefs written into docs or code, spread through an agent system and are very hard to remove.
- Agents behave best under mechanical constraints like governors, quotas and pace gates, not appeals to common sense.
- If you don't spend 40% of your effort on code quality, you end up spending 60% of it.
- Faster engineering pushes the bottleneck downstream, onto reviewers, business teams and customers.
- Enterprises should start with narrowly scoped production agents before automating how they build software.
Why Software Factories Burn Down
07:02In this conversation, we examine what happens when a software factory gets sick. Steve's agents, trying to make sure nothing could ever go wrong, added more than a hundred gates until nobody could do anything. A newer model came in like an outside consultant and cut about 40% of the system, keeping only 14 fences. That week showed the other side of the problem. With the fences gone, agents shipped 46 game features during a week when Steve had said no game work. Steve explains why these systems can't be fixed from the inside. He also covers why "heresies", wrong beliefs that get written into docs and code, keep coming back. And he explains why only mechanical constraints like governors, quotas and pace gates keep agents on the rails.
The Factory That Will Never Be Done
20:57In this conversation, we examine why a productive agent fleet can end up busy forever. When Steve asked his factory when it would finish repairing itself, it ran the numbers and answered "never". Every issue it closed opened up to half a new one, because rounds of adversarial review kept surfacing more low-priority work. Rather than delete the backlog, Steve saw the real cause: half a million lines of code in twelve weeks, with too little time spent on technical debt. That's where his rule comes from: spend 40% of your effort on quality, or you'll end up spending 60%. He also explains why models are poor at estimating their own schedules and why planning quality passes remains a manager's job.
Enterprise AI Is Arriving Upside Down
47:50In this conversation, we examine why AI is entering companies in the opposite order to what most leaders expect. Teams trying to automate how software gets built keep running into merge queues, uneven results and too big a blast radius. Meanwhile, agents working ticket queues succeed almost straight away, because each ticket is small, easy to measure and backed by years of historical data. Steve's advice is to stand up one narrowly scoped production agent first and learn your "AI employee" lessons there. Only then should you take on the build side. He also explains why faster engineering overwhelms the business downstream, and why outcomes, not outputs, have to become the measure.
Outcomes Are the New North Star
1:04:55In this conversation, we close on what to measure once agents are doing the work. Steve argues that outputs — tokens spent, pull requests opened, lines written — are what teams reach for because they are easy to count, and that they say almost nothing about whether the business got anything. He points to Mik Kersten's Output to Outcome for the patterns, and to online gambling companies as the surprising example of an industry that has measured engineering outcomes precisely for years. He also offers two consolations. AI is changing jobs rather than taking them, and nobody is as far ahead as the pressure suggests: only the newest and most expensive models are capable of this work, and most companies have them switched off.

