An Tran Solutions
An Tran Solutions
Back to Blog

Spec-Driven Development Isn't Documentation. It's a Governance System for AI-Written Code

September 30, 20268 min readby An Tran
On this page

Two beliefs are spreading through the software world at the same time.

The first: spec-driven development (SDD) is a return to waterfall. Write a ten-page document before you dare touch code. Slow and outdated.

The second, from the opposite camp: SDD is the cure for vibe coding. Just write a really thorough spec, hand it to the AI, done.

I think both are wrong. And they're wrong in the same place: both treat the spec as a document. What you actually need in the AI era isn't a document. It's a governance system: who decides what, where those decisions are recorded, and which machine checks that the code the AI writes actually follows them.

I used to think "the more detailed the spec, the better." This post is why I don't anymore.

What is vibe coding actually missing?

I've written two posts on vibe coding already, on the cognitive debt it builds and on the stop points it erases. This one takes the next step: looking at it through the lens of governance.

Vibe coding isn't bad because AI writes bad code. It's a way of working in which every technical decision lives in one person's head and in one chat. Close the tab and it's gone. The second person to join the project has no idea why the payment system was built this way. And worst of all: every time you open a new chat session, the project's "policy" gets rewritten from scratch, from memory.

That's especially dangerous wherever the model's defaults don't match your policy. Security is the clearest example. Veracode's 2025 GenAI Code Security Report ran 100+ models through 80 coding tasks and found that models introduced an OWASP Top 10 vulnerability in 45% of cases. In Java, the security failure rate was above 70%; in Python, C# and JavaScript it was 38–45%. Against XSS, the models failed to secure the code in 86% of cases.

The most sobering detail: Veracode found that larger models were not significantly more secure, and security barely improved over time even as the code got better at "working." In other words, this isn't a bug that disappears with the next model. It's a property you have to govern, not wait out.

And you can't govern it by reminding the model in every prompt. The rules have to live somewhere that every session, every person and every agent is required to read.

Vibe coding isn't short on speed or intelligence. It's missing one shared place to record "what we decided," and a mechanism that forces the code to comply.

What SDD is, in plain words

Thoughtworks describes spec-driven development as AI-assisted workflows that "begin with a structured functional specification, then proceed through multiple steps to break it down into smaller pieces, solutions and tasks." In plain English: write down clearly what needs to be done first, then let the AI do it.

Three tools currently represent the space, and each reads SDD differently:

ToolHow it worksWhat stands out
Amazon Kiro3 phases: requirements → design → tasksLean, step by step
GitHub spec-kit4 phases: Specify → Plan → Tasks → ImplementMore orchestration; has a "constitution" of immutable principles
TesslThe spec is what gets maintained; code is a derived artifactThe most radical: humans only edit the spec

GitHub frames the ambition nicely: moving from code as the source of truth to intent as the source of truth. Martin Fowler's site, in a piece by Birgitta Böckeler, splits SDD into three levels, and the level is what really matters:

  1. Spec-first: write a spec to guide the first build, then throw it away.
  2. Spec-anchored: keep the spec and keep using it as the feature evolves.
  3. Spec-as-source: humans only edit the spec; the code is generated.

Böckeler notes: "All SDD approaches and definitions I've found are spec-first, but not all strive to be spec-anchored or spec-as-source." Pay attention to that. A spec you throw away after the first build is just a longer prompt. It isn't governance.

The part SDD fans don't want to hear

I support this direction. But a post that only praises doesn't deserve to call itself contrarian, so here's the evidence against it, and it's serious.

The Thoughtworks Technology Radar (November 2025) places SDD at "Assess": worth exploring, not yet something to adopt broadly. They say plainly that the tools behave very differently depending on task size and type, that "some generate lengthy spec files that are hard to review," and they warn that practitioners may be "relearning a bitter lesson — that handcrafting detailed rules for AI ultimately doesn't scale."

Böckeler's hands-on testing is even more concrete:

  • Agents ignore the spec. In one case, "the agent ignored the notes that these were descriptions of existing classes, it just took them as a new specification and generated them all over again, creating duplicates." A long spec is no guarantee the agent reads it and follows it.
  • Reviewing specs costs more than reviewing code. In her words: "I'd rather review code than all these markdown files."
  • A sledgehammer to crack a nut. She fixed a small bug and the tool generated a whole stack of requirements. One process for every size of job is the wrong process for most jobs.
  • A historical warning. Spec-as-source risks ending up with "the downsides of both MDD and LLMs: Inflexibility and non-determinism." MDD (model-driven development) once promised "draw the diagram and get the software," and failed for similar reasons.

After reading all that, I had to admit: the second belief, "just write a thorough spec," is exactly the one the evidence refutes. A long spec that nobody verifies, that the agent may or may not read, and that humans don't want to review is just a new flavor of vibe coding with extra paperwork.

So is the "SDD is waterfall" camp right? Also no. Waterfall failed because it froze decisions before anyone had learned anything. Here, the loop between spec and code takes minutes, and you fix the spec the moment you see it's wrong. The question isn't whether you have a spec. It's whether that spec is verified and has an owner.

A four-layer governance framework: the spec has to bite

My conclusion: a spec only deserves to exist if something bites when the code violates it. Here are the four layers I use, from lightest to heaviest.

Layer 1: The constitution. One page, only the things that must not be wrong

Not an architecture document. One page of the project's non-negotiable rules, living in the repo where every person and every agent reads it:

CONSTITUTION.md
# Project constitution
 
## Security (non-negotiable)
- Every query that touches user data is parameterized. No string-built SQL.
- All user-supplied output is escaped at the render boundary.
- Secrets never enter the repo or logs.
 
## Architecture
- Server components by default; a client component needs a written reason.
- One price list (quoteCatalog). No second source of prices.
 
## Definition of done
- Every acceptance criterion has a test that fails when it is violated.
- Anything touching money, auth or personal data is human-reviewed line by line.

Notice how the content is chosen: it prioritizes the places where models tend to get it wrong by default (security, per the Veracode data above) and the places where mistakes are expensive (money, authentication, personal data). A constitution longer than two pages is a constitution nobody reads, agents included.

Layer 2: Risk-tiered specs. The paperwork should match the risk

This is the answer to Böckeler's sledgehammer. Don't have one process for everything. Have three sizes:

TierExamplesSpec level
LowCopy fixes, CSS tweaks, one-line bugsNo spec. The commit message is enough
MediumA new page, a new form field, a refactor with test coverageHalf-page spec: goal, constraints, 3–5 acceptance criteria
HighPayments, authentication, user data, public APIsFull spec + line-by-line human review + mandatory tests

The rule: if reviewing the spec costs more than writing the code yourself, that tier doesn't need a spec. Saving paperwork at the low tier is exactly what lets you be strict at the high tier.

Layer 3: Executable acceptance. A spec a machine can read

This is the most important layer, and it's where real SDD differs from performative SDD. Every acceptance criterion in the spec must become a test that goes red when the code is wrong. Without a test, it's a wish, not a spec.

orders.acceptance.test.ts
import { describe, it, expect } from "vitest";
 
describe("SPEC-014 · Cancel order", () => {
  it("AC-1: a customer can cancel only their own order", async () => {
    const res = await cancelOrder({ orderId: "o_1", userId: "someone_else" });
    expect(res.status).toBe(403);
  });
 
  it("AC-2: a shipped order cannot be cancelled", async () => {
    const res = await cancelOrder({ orderId: "o_shipped", userId: "owner" });
    expect(res.status).toBe(409);
  });
});

The two highlighted lines are what I want you to look at: the spec IDs (SPEC-014, AC-1) live right in the test names. That gives you traceability from requirement to evidence and back. When the agent "forgets" AC-2, CI goes red. You don't have to trust whether it read the spec. You only have to trust the test.

This also solves the "agent ignores the spec" problem Böckeler ran into: don't hope the agent complies. Make violations unmergeable.

Layer 4: Control gates and traceability. Who signs, where, and when

Two small things that form the backbone of governance:

  1. The spec travels with the pull request. The same PR changes the spec and the code. Reviewers see how the intent changed before they see how it was implemented. This is also how you reach spec-anchored instead of spec-first: the spec lives alongside the code and doesn't die after the first build.
  2. Every high-tier spec has a named sign-off. Not "the team." One person, by name. When the payment system breaks, the question "who accepted this spec?" should have an answer in three seconds.

Add automated CI gates on top: run every acceptance test, run static security scanning, and block the merge if a high-tier file is touched without a reviewer's sign-off. Machines do the repetitive part; humans keep the judgment.

Four layers, one principle: intent has to live where everyone can read it, and something has to bite automatically when the code drifts from it. Without the second half, a spec is just prose.

Where to start: in a week, not a quarter

Don't install any tool until you've done these three things:

  1. Day 1: write a one-page CONSTITUTION.md for a project that's already running. Only record the things that have burned you, in time or in money.
  2. Days 2–3: pick one high-tier area (payments, login). Write its acceptance criteria as tests, even if the code already exists.
  3. Days 4–5: wire those tests into CI, and make it a rule that high-tier specs ship in the same PR.

Only then evaluate whether Kiro or spec-kit is worth using. They're tools for writing specs; they don't replace the three steps above. And with the Radar placing SDD at "Assess," keeping your process lightweight and tool-independent is itself a form of risk management: you don't get locked into a workflow that's still unsettled.

The cost and the payoff, in money

When every team uses AI, build speed is table stakes. What sets you apart is whether you can prove your code does what you promised. For business clients, that's the difference between "fast" and "trustworthy," and only the second one gets the contract renewed.

If you're a business leader hiring a team to build a website or system "with AI," don't ask "how fast can you go?" Ask these four questions instead:

  1. Where are the non-negotiable rules for my project written down? (You want to see an actual file.)
  2. Do the features that touch my money and my customers' data have acceptance criteria written as tests?
  3. Who is accountable for that part, by name, not "the whole team"?
  4. If I switch vendors, what does the new team read to understand the system?

A team that answers fluently, with real documents, is governing AI-written code, not hoping it's correct. That's how I work. If you'd like to see a real project constitution and acceptance-test suite, just ask.

Vibe coding isn't wrong because it uses AI. It's wrong when it leaves nothing behind but code. Spec-driven development won't save you by making you write more. It saves you when it forces you to write down what you decided, and lets a machine guard it.


Data and quotation sources: Veracode: 2025 GenAI Code Security Report; Thoughtworks Technology Radar: Spec-driven development (Nov 2025, Assess); Birgitta Böckeler, martinfowler.com: Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl (Oct 15, 2025); GitHub Blog: Spec-driven development with AI (Sep 2, 2025); GitHub spec-kit.

Related articles