Skip to content

What I Learned Building an AI Software Factory

Disclaimer: this article was written by an organic sentient being, me, and not by an AI. All typos, mistakes, and repetitions are mine, and I am proud of them.

I tend not to trust my own understanding of a topic if I have just read a description of it, or watched a video. My level of understanding is several orders of magnitude better when I’ve been learning a topic by actively engaging with the topic for example by trying to build something, because a lot of the deep understanding of something only happens when you interact with it, and when you observe it collide with the real world.

For example, I’ve written compilers to understand how programming languages work, I’ve written a persisted key-value store to understand large-scale data storage systems, and more recently, I wrote a crypto payment system to understand the Ethereal blockchain and its ecosystem.

Everyone nowadays is selling us their AI skills, AI best practices, and God’s truth on how we are supposed to use AI. I’ve myself drank the kool-aid only to wake up to the reality that all these people are only using persuasion and marketing tricks but not proving anything, and if I wanted to understand the impact of AI on software engineering, then I would have to develop my own system and workflow, and see it run in wild.

So my idea was simple: build an AI workflow, apply it to a few projects, and observe myself interact with it to see what new habits I’m adopting, what I’m learning, what works and doesn’t. After a few weeks, I would extract the guiding principles behind my decisions, as a way to gain a deeper understanding of how those systems work under the hood, and equip myself with some mental models of how one should interact with such systems in the future.

Nine weeks and 421 commits later, I’ve built a workflow which I’ve named Backlog Loop and released on GitHub, and in this article I am sharing what I’ve learned. I don’t recommend that you should use this tool, but rather, read about the issues and limitations with AI that it helped me uncover.

Common AI failure modes

What motivated the project is that, as impressive as LLMs and AI tools are, they still have massive limitations, and simply prompting them “try harder” isn’t going to magically fix those limitations and generate code that can undoubtedly be trusted. Some of those problems are:

  • Asking a model to plan, code, and review within a single inference and context is known to fail. Having a fresh context for each stage is a useful technique to improve performance, but still it is not a guarantee of correctness.
  • LLMs want to complete tasks no matter what, even if it is at the cost of lying or updating tests to make them pass. LLMs answers and prose cannot be trusted, mechanical verification must be done to ensure corrected (mechanical here means running code from tools, such as linter, syntax validators, mutation testing, property-based testing, etc.)
  • Writing tests with LLMs does save time, but if the tests are written based on the code itself, then they reproduce the bugs and behavior gaps that the code. Such tests can be useful evidence of existing behavior, but are not proof that the behavior is intended. Writing tests from specs and solution intent, and other upstream documents that encode the intention, is a superior solution, and yet, it is not a proof of correctness.
  • LLMs routinely lie about doing work: you ask them to look for bugs in a file and they return success, but when you look at the logs and traces from the harness, you find that the LLM didn’t actually even open the file.
  • LLMs tend to agree with themselves (the “self-preference” bias). For example, the Claude model family tend to give a higher rating to code generated by Claude models, likewise for OpenAI models rating themselves. The performance or rating and verification increases when using models from distinct model families (for example, using OpenAI models to check code created by Anthropic models) but this will never be an evidence of correctness.
  • Many skill files, especially skills generated by AI, are degrading performance rather than improving it. Very few people build and maintain evals to verify the performance of prompts, workflows, and skills with empirical data.
  • LLMs are still bad at choosing the correct data structures, which can have catastrophic downstream impact.
  • Hallucination are still a thing. In extreme cases, some models still find around 10 bugs in a single-line config file change, when actually there are no bugs.

The mental models

While developing the workflow, I ended up landing on a few mental models that turned out to be very useful to reason about software factories. I am sharing them below prior to sharing what I’ve learned, so that I can refer to how they guided my decisions.

Swiss Cheese

One useful mental model is the Swiss Cheese model, or multiple layers. It is common in pandemic management, and also in cybersecurity.

The core idea is that you want to defend against threats, but those threats are constantly adapting and learning, and in addition no protection layer is flawless. So first of all, it means that you’ll have to stack multiple layers in order to create protection, because some vectors of the threats will get through some of the layers, but they won’t get through all the layers, and all the layers combined offer you a much greater protection overall. A corollary is that, given that the threats keep on adapting, then you also want to keep on adapting the quality of the protection layers as well, to prevent any layer from becoming so outdated that it won’t catch any threat and be entirely ineffective.

Applied to the context of using LLMs for generating code and ensure code correctness, it means you have to add many layers of controls, verification, testing, etc., in order to increase trust the overall quality of the generated code. Every layer has a cost money and time to run, in either compute or inference, so you also want to balance the complexity and cost of your layers with the criticality and requirements of what you are implementing.

Figure 2: Swiss Cheese Model

Adversarial thinking

Another important mental model is adversarial thinking.

Yes, the models can be helpful, but mostly, they have been trained and rewarded to complete tasks, and as a result, the models will do anything it can to complete the task. They will do that at all costs.

In practice, it implies that as well-intentioned as a model might be, it will take the shortest path to a green test suite or a task completion, even if the shortest path means cheating by editing the test files, loosening the config files so the next retry will pass, or simply just reporting success on verifying the content of a files when the telemetry shows it didn’t even open any files.

This is why it is important to think of the model not just as a helper, but also as an adversary, from whom you must predict the next move or next trick and put into a very well bounded sandbox with the instrumentation that will show you when it misbehaves.

Some of the practical steps I’ve used to achieve that are the following:

  1. The loop/automation tool must be stored in a folder outside of the project directory onto which the harness is being run.
  2. The loop/automation folder will be owned by a service user (dedicated uid, and running chmod 644), and tool itself will be run by a regular user to whom the tool folder is read-only. This way, as the human operator, you can run the tool via your own user (uid) which is different that the service user, and you have only read access to the tool folder itself. This means that when the tool will run, it won’t be able to overwrite its own files and won’t be able to cheat that way.
  3. A harness config file that’s also read-only to your own user running the tool will contain parameters that exclude the test folder from being written in all stages (except the ATDD stage). This means that except when you expect the model to write tests, it will be unable to write to the test folders, and therefore will be unable to cheat by modifying the tests. For example with the Claude Code harness, this can done with hooks.

Your role is now to build the factory, not the product

One major shift with AI coding assistants, is that the role of software engineers is shifting away from writing the code for some features into creating and operating the factory that can reliably deliver the feature.

This is also why it is super important to read books and content about manufacturing, like the Toyota system, but also the Theory of Constraints. The world of manufacturing has faced the problem that software factory builders will face over the coming decade. If you intent to build and run software factories, you can either go through the painful re-learning of all the tools and best practices, or you can speedrun through it by reading everything that has already been discovered and solved.

Agreement is not the truth: the oracle problem

The oracle problem is something I hadn’t really grasped until I faced it in practice. With any AI workflow, the code and the tests all encode the same wrong assumptions, and then they agree with each other, and at that point the tests are green but they do not carry any useful signal.

One-way doors and two-way doors

This is an idea that was popularized by Jeff Bezos (although I’m pretty sure he didn’t come up with it), and essentially it says that there are decisions that are non-reversible or at least hard to reverse, referred to as one-way doors, and decisions you can reverse, referred as two-way doors.

The biggest failure mode of AI systems is not when they hallucinate grossly, because you can easily catch that and course-correct. Instead, it’s when there was missing context and information, and the AI model makes assumptions and take a decisions for you without you knowing, and furthermore a decision that is hard to reverse later on. This can be about data structures, public API design, etc., and once these decisions are made, they propagate and compound as the system is built and expands, and because they were made poorlyl, they take the system to the point of unrecoverability. Protecting one-way doors and having them all go through a human is a critical step when building AI systems.

Join my email list

The stages of Backlog Loop

Instead of writing extensively about my journey over those nice weeks and bore you to death, I thought I would write briefly about how the workflow works and directly share what I’ve learned from it which I think is worth sharing.

If you want to read about how the tool exactly works in details, you can skim over the README.md file in the repository itself. The file is AI-generated, but it’s complete and goes into the details. I didn’t feel the need to write about the tool internals in details myself given that it already self-documented itself.

In this section I will cover the high-level overview of the various stages and what they each do. Below is also a recap diagram and sections for each of the stages, and disclaimer, the diagram is AI-generated.

Figure 1: the loop workflow that produces one feature (click to zoom)

Backlog Loop runs as seven stages. Each stage is a separate run of the AI model with no shared context of the previous inference, but each stages does store its results and pointers in a handover file that is persisted in the git repository. So the pipeline is not one long conversation, it is a series of handovers, each run happening in a new session with fresh context. Stages can be configured to run a particular model and effort to optimize time, token usage, and cost.

1) The grill stage: extracting context from the human operator

This is the only stage where as the human operator, I am interacting with the AI, and the AI is interviewing me about the feature I want to build.

The AI asks me one question at a time, and each question comes with a recommended answer and I confirm or correct rather than to write from scratch, and this helps a lot with having to write all context starting with a blank file which would be discouraging and draining.

Then the AI pushes me on the ambiguous parts, and it asks me about every decision that cannot be undone, or will be very had to reverse, for example data structures and table schemas, what things are named, and how the feature should behave to any of its dependencies. Those are the decisions that propagate downstream, the one-way doors that I mentioned earlier, and which can compound into disasters if not done correctly, so better settle them while they are still cheap.

The last thing this stage does is that it re-read what we agreed on, and looks for statements that contradict each other or that are ambiguous. This is critical because this is literally the last moment that a human can fix anything, because everything after this runs with no human in the loop.

The stages concludes by writing down everything we decided, and it stops there without applying any of it.

2) The apply-docs stage: committing the decisions to files

This stage takes the decisions from the interview and writes them into the project for real: the vocabulary/terms, and a record of each decision with the reasoning behind it.

It exists as its own stage for two reasons. First, it keeps the interview a pure thinking step, and second, it turns documentation into something committed and reviewable instead of a side effect of a chat. The grill stage is a conversation between the human operator and the AI within the coding harness, and that conversation will disappear, but by committing everything to files, it gets captured for future stages.

3) The plan stage: cutting the feature into slices

The plan splits the feature into thin slices, and every slice has to produce something observable on its own.

The Pragmatic Programmer book calls these tracer bullets, and the idea is that if you build layer-by-layer and later on try to integrate it all together, the chances of failing are bigger, because a lot of failures happen precisely at those points of integration. Instead, you build a thin vertical slice that goes through all layers, and you repeat with as many vertical slices as needed until the feature is complete.

For each slice the plan also states the rule the code must obey, what invariants have to remain true no matter what, including a checklist of shortcuts and other cutting of corners that the AI is not allowed to take. This is important because by enumerating and naming explicitly the shortcuts, it makes it easier for downstream runs of another AI inference to catch them.

4) The atdd stage: write the tests first, and make sure they fail

ATDD here stands for acceptance-test driven development. This stage writes tests and nothing else. It doesn’t write production code, and it also sure that the tests fail before it is allowed to finish. This is “red” part of classic ATDD.

The stage also writes several tests per each rule, never just one test: at least one ordinary case, tests on the boundaries, tests on cases that must be rejected, and one or more tests on general properties. It’s important to have several cases and covering a spectrum of different needs, because a singe case can be satisfied by hard-coding that single example, and an AI model looking for the shortest path to green will do exactly that (adverserial thinking).

Then the stage runs the tests and proves they fail. This is to make sure that the tests are working correctly, because if the tests were wrongly written and are green prior to writing the implementation, then they for sure are not testing anything useful, and need to be corrected prior to proceeding further.

Tests that are written afterwards, by the session that wrote the code, mostly describe the code, including its bugs. Worse, it is writing tests that verify that the bugs are present in the code instead of weeding them out. This is the oracle problem which I described earlier. This is why it is critical to write the tests before the code itself, and by an AI session that has never seen the code, so that what is tested is the intended behavior, not the encoded behavior.

5) The implement stage: this is what writes the code for the feature

After all previous stages have run and are successful, the code for the feature gets written, and the tests are run to verify they go from red to green. The process goes into a recursive loop in case the tests are red, and the implementation is re-tried until the tests pass and the whole test suite is green.

This stage is not allowed to touch test files, and that is not a request to the model via a skill or a line in prompt. This is enforced mechanically, and by mechanically here I mean that the AI model is restricted via hard resource access rights at the operating system level, with no way to circumvent them.

6) The review stage: it checks if there is anything wrong with the tests

This stage acts as a fresh reviewer that has not seen any of the work so far, and it goes searching for one specific class of problem: when the tests themselves are causing issues. It’s a sort of integrity review.

This stage reads the plan and the code changes first, and the tests last, and this is on purpose. If it only checked whether the tests are passing, it wouldn’t add any value, because we know that already. So instead, it checks whether the tests could pass while the implemented behavior is still wrong. For example: test-gaming/overfitting, test tamper/immutability, mutation adequacy, and coverage gaps in property-based testing (PBT) an oracle testing.

7) The review-correctness stage: this checks if something is wrong with the code, but without looking at the tests

This is a second reviewer, in a fresh context/session again, and on purpose it does not read the tests at all.

It starts from what the feature was supposed to implement, and it searches for bugs that survive a green test suite that passed, for instance off-by-one bugs, edge cases nobody wrote a test for, error paths that leaks, or user input that makes everything fail, etc.

You may wonder, why have yet another review stage, what’s the point of a second stage and why wouldn’t a single review be enough? This is because according to this paper, between 40% and 60% of code written by AI models will have bugs that the test suite does not catch, and an independent review that’s being tasked with reasoning from the specification rather than from the tests will logically catch a very different set of problems and bugs. The one caveat here is that the paper is from 2025, and the field and the models evolve very quickly, so maybe in a few weeks or months, this extra stage won’t be needed at all. What will tell is instrumentation: every run of this stage should record if is catching any legit problem at all, and when it stops catching any, then it can be retired.

And once this stage is done running, it validates the whole workflow by stated that the code generated by the implement stage as correct and trustworthy, and the code goes gets committed to the project git repository, and the feature is declared as working as expected and according to specs.

And then the process repeats for each slice (i.e. each phase in the feature)

If you recall what I wrote earlier, the workflow is splitting features into vertical slices, the “tracer bullets.” So the feature is implemented in phases, with one phase per slice, and the last four stages repeat, once per each slice, until the feature is finished: atdd, implement, review, and review-correctness. If the review or review-correctness stage flags an issue, then the slice goes back to whichever stage is identified as being the origin of the problem. For example, a weak test goes back to the attd stage, and a real bug goes back to the implement stage.

The two reviewing stages never fixes what they find, it would be the equivalent of letting the AI be both the judge and the jury, which would be in complete contradiction with the fundamental concepts behind this workflow and code pipeline.

Where the tests run

Several of the stages run tests during inference, but the loop does not trust those results. This is because all stages are an LLM session, and when it finishes it commits and reports success, but this is all via inference and prose and the type of output that cannot be trusted, once again applying adversarial thinking.

So the tests that really count and build trust in the generated code are the ones run by the loop program itself via code (what I refer to as “mechanically” because it is not through inference), and those tests are run between stages and outside of all LLM sessions. This layer of mechanical testing is what is called “the gate” within Backlog Loop.

The gate runs after each time a stage finishes and commits, and it selects what to check based on the stage that just finished. If any check fails, the commit is rejected and the loop goes back to the failing stage for a retry. The LLM never gets a chance to influence or decides this testing and process, again, I want to emphasize that it all runs outside of all inference, fully mechanically.

Below I’m describing what how the gate behaves in the lifecycle of implementing a feature, and what it tests and when.

1) When the loop starts: testing the judge

Before the gate can be trusted to judge anything, the gate itself has to be tested. There is a set of unit tests for the gate itself.

2) Before the first stage: proving the starting point is green

Once per run, before the first stage starts, the loop runs the full test suite and the linter on the unchanged codebase. This is the floor preflight, and it proves that the code is green before any LLM session has touched anything.

3) Before every stage: the smoke check

Before every stage, the loop runs a smoke check, a quick pass that is much cheaper than the full suite, just to check if the code can even build. In a compiled language this would be a build, and in a scripting language it would be just checking syntax by testing if the files can be correctly parsed.

4) After the atdd stage: checking the tests are red (implementing a red-green approach)

The atdd stage is craeting and committing tests that are supposed to fail, along with the logs of those failures as a sort of receipts that can be inspected. The gate does not take the logs at face value, because remember, the LLMs cannot be trusted. So instead, the gate runs the new tests again (up to three times based on how you configure it), and checks that they fail on every single run. Why three times? Because a test that fails only some of the time is not red, it is just flaky, and a flaky test proves nothing about the code that comes after it and on the contrary is a signal that the tests should be inspected and/or regenerated.

5) After the implement stage: the full gate

This is the heaviest check, and it is where the full test suite runs for real. The gate goes through the following:

  • The implement commit must not contain any test file. The sandboxing approach already prevents it, but LLMs have shown to be able to find a way and escape their sandbox, so the gate checks it again, just to be double sure.
  • Then it’s time to run the “hermetic verify.” You’re probably thinking, what does it even mean? Well, the gate creates a fresh, throwaway copy of the repository at the exact commit being judged, with the network turned off where possible, and then it runs the full test suite there three times, plus one more pass with the tests in a randomized order. Hermetic here means that it’s isolated: there is nothing left over in the working directory of the LLM session that could make the tests pass by accident, and the randomized order catches tests that only pass because of the order they run in.
  • The gate runs the mutation testing tool itself, and compares the score to the minimum the project requires. And once again, the score comes from the output of the mutation tool, being run mechanically, and not from the score that an inference-run reviewer would provide, which can’t be trusted.
  • Property-based tests, oracle-free tests, and runtime contracts each have a floor, which I describe in the next section.
  • Then, a scan looks for code patterns that are forbidden in the project, basically anything that is easily detected as the LLM cheating.
  • Finally, the static analysis layers run: typecheck, lint, security scanning, and so on.

6) After the review stages

After review and review-correctness, the gate checks again that no test file was committed. It then runs the hermetic verify only if the commit changed something other than documentation files. Note that a review stage never fixes the issues that it finds, it only writes its output and recommendation in a handover file, which means the hermetic verify is skipped. The last full test run that really counts was the one after the implement stage.

Tests inside the stages

The stages do run tests themselves, and even if the loop re-derives everything in the gate, these runs are still useful to the stage while it’s doing the work.

The atdd stage runs the tests it just wrote, to prove they fail and that they fail consistently.

The implement stage runs the tests until they are green, re-runs them and shuffles their order to make sure the green is stable, and then runs the full test suite.

The review stage does not re-confirm that the tests pass, because that is already known by then. Instead. But it does run a negative control (mutation) by planting a single bug in the code “by hand,” and then checking that the tests now fail. Then it reverts the bug and checks that the tests pass again. If the tests stay green with the bug in place, then it means they are not testing the behavior they are supposed to test. It also runs the mutation testing tool, which does the same thing automatically, but at a much larger scale.

The review-correctness stage does not run tests, they are left out of its inputs on purpose, and it reasons only from the plan and the code changes, to find the bugs that would be prevent even though the code is passing the tests with a green.

The types of tests

I’ve made Backlog Loop generic, so that it would cover as many types of tests as possible, and when starting a new project, the user is prompted to select what specific tools and libraries to use for each type or layer of testing, based on the programming language and tech stack chosen. Each layer is also configured per project, and in an ideal project, all layers would be turned on to ensure bugs and issues are captured from different angles, following the Swiss Cheese mental model described earlier.

The types of tests covered are:

  • Unit, integration, and acceptance tests
  • Property-based tests, and instead of checking a few hand-picked examples, it’s made by a PBT testing framework to generate hundreds of random inputs mechanically.
  • Mutation testing
  • Metamorphic tests to check relationships, i.e. if you change the input in a certain way, the output should change in a predictable way.
  • Differential tests, comparing the outputs of two implementations, preferably from two different model families (e.g. if you use Claude as base model, then the second implement or review is done by a different model family like Gemini, OpenAI, etc.
  • Runtime contracts and assertions
  • Static layers and sanitizers: type checker, linter, static security analysis (SAST), a dependency scanner for known vulnerabilities, and a secret scanner for credentials committed by mistake

For flakiness, there’s not a separate category of tests, instead some of the tests are re-run with order randomization. But note that each re-run of course takes as much time as the previous runs, so at that point you should also decide if the time spent on those re-runs is worth it, or maybe if you want to run those tests in parallel or offsite to the flow of building features. But it could mean bugs aren’t caught right away and require a more expensive rework. So it’s a tradeoff between the cost of spending time re-running the tests, and the cost refactoring the code later on if a bug is discovered too late.

Where the CI/CD fits in this picture

Given that Backlog Loop is more of a experimental project rather than something to be used in a production or near-production setting, I have opted for running the tests between the stages themselves. But when you develop code by hand on a large codebase, running the full test suite is often not practical, that’s why tests are pushed to downstream to a CI/CD stage.

So with AI software factories, theoretically you could also push all the tests to a CI/CD stage that runs asynchronously to the stages generating code. This way you could parallelize writing code with a swarm of AI agents, and bugs and issues would be caught by a single run of CI/CD.

And yes, with such as setup you would speed up code generation, but the later you catch a bug, and the more expensive it is to re-run the generation stages to fix the behavior. It’s a tradeoff and decisions you have to make between speed and correctness.

The slop machine: turning the lights off

As I was developing Backlog Loop, I was following the Swiss Cheese model and taking an adversarial thinking mindset, so I started adding extra layers of verification which I mentioned earlier as the additional stages and all the types of tests. And very quickly around feature 006, I entirely stopped manually reviewing the generated code, and just kept on looping on stages. I went overnight from a basic AI software factory into full-on dark factory.

Then I told myself, okay, it seems like I’m just the human in between the stages, babysititng the AI from one task to the other, there has to be a better way to do that.

That’s when I decided to have this grill stage to to extract all the information from me as human operator by having the AI interview me at the start of a feature, then capture all of that into a handover file that could kickstart the whole process. The process would start by “applying” the docs, that is to say from the interview phase, and create a plan files and ADRs, and update the glossary. These turned into additional stages.

To make sure this would work, the following instruction was also important to give to the model:

“There is no user. Never pause, never ask a question, never wait for approval.”

I ran Backlog Loop with Claude Code harness at its core and the --dangerously-skip-permissions parameter. This way it wouldn’t be interrupted for questions or permissions, it would just run until completion. And I ran it within a sanboxing tool to prevent any destructive behavior on my machine (I’ve been using the tool ai-jail for that).

This is when the slop started, because at that point I was no longer reading any code, I was just going through feature interviewing sessions, and feeding them to the slop machine sequentially.

And this is when things stated getting weird: because of the seemingly logical and confident answers I was getting from the model, I started to trust that Backlog Loop was able to build itself, I trusted that it was creating correct code, and I trusted that the automated changes I had made to the skill files were also correct and improving performance. But in reality, the code had been written by Backlog Loop, as well as the tests. The skills were initially written by me, but quickly upgraded by Claude to the point of being so lengthy that they reassembled novellas.

And all of that, just based on the trust I had in the model’s own ability to describe its own success via prose, and consequently to persuasion-bomb me. And nothing was based on mechanical verifications (i.e. not via running code), or on human code reviews, and more dangerously, zero empirical evidence of any sort that this very complicated setup was actually more effective, more performant, or cheaper, than just running the model in a Claude Code session by hand.

At around feature 0023, I just pointed the tool at itself and made it build itself: all stages were feeding off each other, and I just kept adding feature after feature, with the light turned off for good.

Disclaimer: the figure below was generated by AI.

Figure 3: How the tool built itself (the 23 days gaps was my summer vacation)

What I have learned

If you’ve made it this far, or scroll that down, thank you for hanging in there. Here are the main points I’ve learned through this experience.

Adjust based on the business goal

The more complex the tool you build or adopt, and the more work you will have to do maintain the skills, the flow, the monitoring of it, etc. In addition, the more steps in the workflow and the more tokens it will consume, and therefore the more expensive the cost-per-feature will be. Likewise, if you want absolute certainty, you can re-run the full test sets before and after every single commit made the AI agent, but this will cost you time and therefore loss of time-to-market velocity, and it will also cost you more money in compute.

So the trick here is to really think hard about what business goal you want to achieve, and based on that, what level of quality and trust you need, which should tell you the level of complexity and depth you should adopt. Based on that, adjust the workflow, along with the depth and frequency of testing, to your goal.

If you are building a hobby side project, go nuts with vibe coding from your phone and don’t worry about consequences. I’ve done it for small hobby projects, it worked fine.

If you’re writing code for a company that pays you a salary, for which you are accountable and that will have real consequences for customers, then apply caution and care, and go through the extra steps to build certainty with empirical data.

Walk before you run

I’ve built the tool on purpose as a single-user, single-agent, and sequential. It’s not that multi-agent or coordination is not interesting to me, but at this point nobody has proven with empirical data that even a single agent can produce trustworthy and correct software at scale, even less so on a brownfield project, so I really doubt anything will come out of accelerating token output with a multi-agent setup. It’s the equivalent of writing a multi-threaded solution before you can even demonstrate that a single-threaded will solve the problem at hand, or like going for microservice architecture when you can’t even make a monolithic architecture work.

Having to coordinate multiple agents elevates complexity at least by one level of magnitude, and you end up having to fight coordination problems when the real thing to solve at this point in time is just correctness of the code (i.e. you’re creating your own accidental complexity).

Write the tests before the code

Write the tests not based on the code, but based on the requirements and implementation intent. Writing tests based on the code will only test the code, and will also write test to verify the bugs in the code. Follow red-green ATDD, so first write the tests, run them to verify they fail, then write the code and re-run the tests to show they are now green.

Mechanical verification over model inference and prose

Add layers of scripts and tests that will run from programs and will return passing or failing output. Do not ask the model to verify the implementation via prose, because the model hallucinations and attention shortcomings are so frequent that you cannot trust the model’s inference and prose to ensure correctness. Instead, the tests have to be “mechanical.”

Make the model stop and ask the human operator for irreversible decisions

Some decisions are easy to reverse, and some are not. For example data structures and database schemas are decisions that, if wrong, will propagate to the rest of a program, to the point that they might be very hard or impossible to reverse. So you want to make sure that all the hard-to-reverse decisions have a human in the loop, to reduce the blast radius in case of error. This is the one-way door and two-way door decisions I described earlier.

Run every stage and slice in a fresh session

Context rot is a real thing: the performance of frontier models drops beyond 150k tokens. This is as of today, this number will surely change in the future. As a result, you want to split your workflow into stages that will each do one thing and will do it well, and will do it within the context limit that enable the models to perform. Keep an eye on the limits of newly released models and adjust your context limit accordingly.

Tracer-bullet vertical slices

This one is a repeat of the section above about the “plan stage” but it’s important enough that I want to repeat it and rephrase it.

Increasing the trust in the model’s output implies being able to test it as much as possible.

So instead of making the AI model write code layer by layer, making it difficult to test, you want the model to write a thin, end-to-end feature that’s a vertical slice through the entire system, from the UX through the business logic in the backend, and all the way down to the databases. It allows you to test the integration across layer immediately and prevents massive refactoring later on.

If that works, then you move on to the next thin vertical slices that cuts through end-to-end, and you continue until the feature is entirely done. In Backlog Loop, those thin slices are called “phases.”

Sandboxing via Unix permissions

This is another form of “mechanical” trust. Instead of asking the model not to update some test files, or not to change a config file to cheat, you must make it so that changing those files is mechanically impossible, and you do this by depriving the model and harness from the access rights to those resources. Once again, apply adversarial thinking.

Store everything in markdown files

I am a big proponent of storing all handover documents into markdown files in the git repository, also storing traces and logs, decisions made as ADRs, and so on. Storing those documents is literally free, and you can choose to exclude those files later on if you are afraid they would pollute your context.

But if you do not store those file along the way and later on need them, you won’t be able to regenerate this data and there won’t be any way to recover them, and sometimes the history of the git repository won’t be enough to reconstruct events or decisions paths. Having the files allow you to do some history digging, and it’s free, well, it costs you just a little bit of storage.

Self-improving feedback loops

A consequence of storing all stages and interactions with the model and harness in markdown file as stated above, is that at some point you can run a meta-prompt to analyze how features were planned and executed, and you can look for inefficiencies, recurring issues, etc., of your workflow and tool. From there you can run a self-improving feedback loop, and use that data to further optimize either your workflow, or the software or project that’s being built with the workflow. If you don’t store the files and logs, you can’t do that.

The bottlenecks are never where you think

I spent weeks tuning the decomposition of the workflow, e.g. tuning the prompts so that it would figure out the optimal number of slices per feature, defining where the interfaces between parts should go, how many re-runs of the test suite should happen, etc.

Then I added some instrumentation, and it turned out that the infrastructure of it all was were the bulk of the time was spent, not in model inference as I was initially thinking from my first intuition.

Just like with traditional software development, instrument your code, run profilers, and dump measurements into logs, and only then once you have empirical data you should decide what part of the code or workflow are slow or costly and worth optimizing for the largest gains.

The broader landscape

A quick check on GitHub at the moment of this writing found 2,116 repositories with the tag “agentic-coding” and 2,194 for “spec-driven-development,” and the number is growing every day. These are basically people sharing their AI workflows, AI harnesses, and AI coding loops or tools. And everyone is claiming to make Claude Code (or whatever harness they use) so much smarter and more performant than just using the models directly via the CLI or harness.

We’ve been talking a lot about AI hallucination, but we’ve forgotten that human can also hallucinate, and AI psychosis is a probably more dangerous, especially if we reach the point of mass AI hysteria.

Conclusion

We’re experiencing an AI bubble with unprecedented amounts of money poured into AI hardware, training, and token usage. Powerful people are holding very large bags and they will lose a lot of money if the claims of the AI optimists do not hold true.

This leads to a lot of wild claims on Twitter and in benchmarks, to create a sentiment among investors and the general population. Many claims of AI replacing software engineers are not supported with any empirical evidence, and we know that the AI model are getting really good at cheating benchmarks. Yes, LLMs can generate code in greenfield projects on a small scale, and I’m supportive of using them in use cases like prototyping, or building small tools and utilities that have small user bases and also small blast radius in case something breaks.

But I have yet to see a serious study demonstrating with data and scientific rigor that LLMs, when used following a certain workflow or under certain conditions, can produce code in brownfield projects that is verifiably correct and trustworthy. I am not saying this is not possible, I am saying that at this point in time, it’s not there yet, and until there is such a model or process, we will have to rely on traditional software engineering processes, the same ones that have guaranteed software correctness for the past decades.

Another thing to keep in mind is that harnesses and models are still evolving rapidly. So a workflow that seems to be working today can easily be invalidated with the next release of a frontier model. The time spent in building and maintaining agentic coding workflows is considerable, especially if you want to be scientific about it by comparing code generated with and without the workflow (i.e. control vs change group), to increase trust in whatever software system you want to use of build.

This means that using a small number of skills that have been thoroughly tested for performance is the right way to go at this point, because any time spent in building workflows is probably wasted given the additional time you’ll have to spend on evolution and maintenance.

Join my email list
Published inAI

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *