How I gave up writing code without losing the codebase

TL;DR

agents are forbidden from writing docs and comments, scripts run in ci checking that dependency flow is respected, jit docs with artifacts, vibe feature -> test feature -> like feature -> write tests -> refactor but keep tests working

Since absolutely nobody asked me how I use my coding agents, here I am writing a blog post about it.

While developing lisptc, a programming language for self-modifying agents (github), I decided to build it entirely with coding agents, without writing any code by hand.

BUT that doesn’t mean I follow the vibe-coded app path where I prompt for features and let agents do whatever they do these days. It’s I think, I design, they code.

Comic: two builders stand in front of a cracked, propped-up house labelled technical debt, one asking why it takes so long to add a new window

Comic by @vincentdnl

Letting agents write all the code meant I had to build the right harness, process and mindset. I didn’t want to end up two months later with a mess nobody can debug, not even the agents that built it. One bad practice in the code becomes the context for the next agent, which follows it and adds one more. That is the failure mode most vibe-coded codebases fall into, assuming agents are somehow immune to technical debt.

Get the feature right, then refactor

I ask an agent to add a feature and describe how I want it to behave. It builds it, and at this point I don’t care how. We iterate, I try it, I change my mind about parts of it. The only question in this phase is whether I want the feature at all.

Once I’m happy with it, I ask for tests that capture the behavior I just approved. That’s me committing to it. Now I turn to the implementation: I ask the agent to brief me on its changes, visualize them, and list the workarounds it used. I read the parts of the code that matter.

That’s when I decide what to redesign, a better architecture, a different technology, the right abstraction, and the refactor starts. The tests stay in place to catch changes to the behavior I already approved.

is this the feature I want? is this how I want it built? prompt how it behaves try it change my mind tests pin what I liked read it brief, workarounds refactor better design iterate until I like it the tests hold the behavior the first half is about the feature, the second about the code the tests are the hinge between them

Every step of that loop depends on what the agent reads before it starts. The code is most of it, and anything else I leave lying around in the repo competes with it.

What agents read

Comments, docs and AGENTS.md files are all the same decision: what context do I hand an agent, and can it be trusted.

No comments

Code is the only source of truth. I forbid comments because I don’t want an agent relying on an explanation that can drift from the implementation it describes.

A deer held up to a TV interview microphone, captioned NO COMMENT

What I gain is not less reading for the agent, it may well have to inspect more code. What I gain is a single description of the implementation instead of two competing ones.

No permanent docs, just-in-time instead

LLMs love wasting tokens rambling about implementation details, and I never read the docs an agent writes. They fail the same way comments do, with more surface area: the code moves, the file path in the doc doesn’t, and the next agent starts from a wrong assumption and builds on top of it. Long docs also bury the few pointers that were actually useful.

So dev docs become just-in-time, generated directly from the code and scoped to what you need to know. When I want to understand a module I ask an agent to produce an interactive artifact about how it works, with illustrations of the data flow, the dependencies, the pros, cons and tradeoffs. In five minutes I have an HTML page I can explore. I use it to review the design, and I don’t keep it.

AGENTS.md files are routers

AGENTS.md is dev docs written for coding agents, the file one reads when it opens the repo. They nest, so any directory can carry its own and an agent reads the nearest one. Mine hold a map and nothing else: every workspace lists where its things are, so an agent lands on the root file, finds the workspace it needs, and goes straight to the relevant code.

agent starts here AGENTS.md the rules, repo wide dependency flow host ports the session seam no comments typed env only no new md files packages/interpreter AGENTS.md src/lisp.ts the evaluator src/session.ts the seam extensions hook test/helpers.ts builds a bare Interp packages/backend AGENTS.md convex/schema.ts every table and index convex/lib/auth.ts the ownership guards src/agent-repl.ts where extensions compose apps/app AGENTS.md src/routes/ the route tree src/lib/ the state and the glue server/ Nitro, the auth router the root holds the rules, every package carries its own map what is in the package, and which file to open

The root file holds only what is true across the repo, the rules no single package owns. Dependencies point inward. External effects stay behind an interface. An agent can break either one in a single line, which is exactly why they belong there: a rule goes in the root file only if something in CI catches it, with a pointer to where that check lives. Everything package-specific stays in the code, and the package’s AGENTS.md points at it.

What CI enforces

A rule in a markdown file is a suggestion. The root AGENTS.md does forbid comments, but what keeps them out is that the codebase has none, agents copy the conventions of the code they read, and a script fails the build when one slips through anyway.

That script walks the AST instead of grepping, so a // inside a string is left alone. It has a check mode that fails on an added comment and a --fix mode that rewrites the file. I pulled it out into its own repo, no-comments, packaged as a GitHub Action that fails the build and posts the removal as a PR suggestion:

- uses: actions/checkout@v4
- uses: 1hachem/no-comments@v1

The same binary runs on a pre-commit or pre-push hook, if you’d rather the comments never land in the first place.

AGENTS.md files get two checks of their own. The first runs each added diff hunk through Jev, a classifier that takes the rule in plain English and returns a typed judgment. The question is whether the hunk explains how something works. A pointer passes and an implementation detail fails, while the root file’s rules pass because a rule is not an explanation of the code.

Each added hunk is checked on its own: pointers pass, implementation details fail.

The second walks every AGENTS.md in the repo and checks that each backticked path, class name and function name still exists and hasn’t been renamed or removed. When it fails, the agent is instructed to fix the file.

Every backticked path, command, and identifier must still resolve in the repository.

Dependencies get the same treatment. I decide which parts of the codebase can depend on which, and Turborepo’s boundary checks fail on a dependency I’ve forbidden, in CI and before pushing. Custom checks cover the rules they can’t express.

During refactoring I also ask for an illustration of how the modules fit together. This is a simplified example from lisptc, showing the core logic staying separate from external systems:

COMPOSITION apps and front-ends ADAPTERS *-host.ts EXTENSIONS *.ts EVALUATOR lisp.ts environment network SDK subprocess clock filesystem An illustration of module decoupling used during the refactoring of lisptc

What I still inspect myself

Some technical debt passes every check I have. A file gets a little hairier every week and never fails CI, so I built a dashboard for lisptc that answers the questions a check can’t.

It runs its own analysis with fallow, git and the gh GitHub CLI. Where does change concentrate? A hotspot quadrant, and its top corner is my refactor queue. What depends on what? A layered dependency graph and a dependency structure matrix for the imports crossing a layer they shouldn’t, and a separate view for import cycles. What changes together? Temporal coupling arcs pair files that keep landing in the same commits with nothing in the code connecting them. How far does a change reach? Blast radius per pull request, and pull request flow and lifetimes for the ones sitting open too long.

Dashicodes: visualizing technical debt and tracking progress over time.

I compare the results over time to see where it’s going.

fallow also finds duplicated logic and dead code, which fills the other half of that queue. A slightly different request gets a slightly different copy of a function, and now a fix has to reach every copy while the copies drift apart. Replacing an implementation tends to leave the old one behind. Both are signals for a better abstraction or some shared logic.

Scouting agents that run periodically automate most of this, but I still go looking myself.

Finding problems in real conversations

Real usage decides what I improve next. In lisptc I collect conversations with PostHog and have scouting agents look for recurring problems. Sometimes I read the traces myself. They show where people get stuck and give me concrete cases to bring into the next round of features, tests and refactors.

I don’t write code anymore, but I still decide what gets built, what the design should be, and what gets thrown out when the implementation turns out worse than the feature it delivers. Agents write every line of lisptc, and they copy whatever is already in it, good and bad. That is why the harness came first.