Every decision before the first line of code
Most of what my team builds starts the same way. Someone runs a discovery with the client. They write up how the client’s process works today. Sales signs a statement of work. Then an engineer gets handed the pile and is expected to build the thing.
That used to take a team and weeks. The first bot we ran through this new process was scoped at 7 weeks. It was working in 2 days, with one person driving the whole thing: discovery call to contract, contract to spec, spec to bot.
The speed isn’t from the AI writing code faster. It’s from the AI never having to stop and guess.
The pile already answers what the system should do and why. It doesn’t answer how it gets built, out of what, and how anyone knows it’s right. Those questions get answered anyway. They just get answered by whoever is writing that file at that moment.
When that’s a person, you get a decision nobody reviewed. When it’s a coding agent, you get the same thing faster, and with more confidence.
So I built a tool whose whole job is to make those decisions early, out loud, and on the record.
The test I hold every spec to
Could an entry-level engineer who has never heard of this client build it from the document alone?
Not a senior engineer who sat in the discovery call. That person isn’t the one reading it. The reader is new, or it’s a fresh AI session with no memory of anything we discussed. Whatever isn’t in the document doesn’t exist.
That one rule drives everything else:
- A deferred decision isn’t avoided. It’s handed to the builder, with no record and no review.
- Open questions are fine. Open questions without a default are not. The client might take two weeks to answer. The spec has to say what we do in the meantime.
- Precision isn’t a nicety. A reference to a section that moved is a fact that never arrives.
How it works
discovery + signed SOW + client files
-> inputs/ copied verbatim, cited by line
-> BUILDING-BLOCKS what already exists that we should reuse
-> DECISIONS asked one at a time, written down as answered
-> SPEC the file the builder reads
-> cold review a new session reads it with none of the context
A few of those steps look odd until you’ve seen them fail.
The inputs are copied, not linked. Discovery documents get revised after follow-up calls. If the spec cites line 140 of a file that changed, the citation now points at something else and nobody notices. A copy with a recorded hash is a snapshot you can diff later.
The survey comes before the questions. If a shared library already handles the login flow for the site we’re automating, half the questions I’d otherwise ask go away. Existence isn’t coverage, though. The survey has to name the actual functions it will call.
Questions are asked one at a time. I tried batches. A batch of twelve questions at the end gets approved, not decided. One at a time, written down as each is answered, with who decided and when.
The SOW wins on scope. The data wins on facts. The signed SOW is the only document that says what was sold. But it doesn’t overrule a number someone counted in the client’s spreadsheet. When the two disagree, that’s a finding worth surfacing. It usually means the commitment is harder than whoever wrote the SOW knew.
The review loop is the part you wouldn’t think to do
After the spec is written, a completely separate session reads it cold. It gets the spec and the codebase. It doesn’t get our conversation.
This is where most of the value showed up. On the first real project, a reconciliation dashboard for a clinical research client, the cold reviewer found things I’d read past a dozen times:
- The spec said each record’s status had to match its section, but only gave one of the ten status-to-section pairs. A builder would have guessed the other nine.
- Records were matched to a study and a site, but one of the two input files had no site column. The spec never said how the match actually worked.
- “Patient count” meant rows in one place and distinct IDs in another. 593 versus 582.
- A cleanup rule an early draft described as removing “about 9” rows actually removed 213 of 1,017.
None of those are exotic. Every one is the kind of thing that turns into a quiet bug three weeks into a build.
The reviewer also tried to add scope more than once. It suggested a third optional import to make the client’s old tracker easier to migrate. Reasonable on its face, and not something anybody had agreed to. Part of running the loop is knowing which feedback is a gap and which is the reviewer inventing requirements.
It caught itself doing the thing
My favorite line in the repo is in the follow-ups file. An earlier version said the tool would be validated by comparing its output against “a spec built by hand for the same work.” There was no hand-built spec. A session had inferred it existed and written it down as fact.
It’s the exact failure the tool exists to prevent, committed by the tool about itself. I left the correction in on purpose.
The guardrails that aren’t prompts
Rules in a prompt get followed most of the time. So the important ones are also checks that run before every commit:
| Check | Catches |
|---|---|
| Every decision ID cited resolves to the log | A spec citing a decision that was never made |
| Every project’s inputs match their manifest | Source files quietly edited after the fact |
| No unresolvable references in a spec | A pointer the reader can’t follow |
| Specs that mention tracking link their tickets | Work that exists only in the document |
We also handle health data, so the spec has its own rules about what can appear in it. Service codes are fine on their own. Paired with anything that identifies a patient, they aren’t.
Why it’s fast
It’s tempting to read all these rules as overhead. They’re the reason the build took 2 days instead of 7 weeks.
A normal build stalls constantly. The engineer hits a question, pings someone, waits, guesses, and later finds out the guess was wrong. Here, every one of those questions was answered before the build started. The coding agent had nothing to wait on. It just built.
It also changed the front of the process. When you can show a client a working prototype days after the discovery call, the sales conversation is different. The pipeline turns calls into contracts, not just contracts into code.
What I’d tell someone trying this
Ask whether a stranger could build from your spec. Then actually hand it to a stranger, or a fresh session, and watch where it guesses. Every guess is a decision you didn’t make.
And write the questions down as you answer them. The value isn’t the spec. It’s the record of why the spec says what it says.