Spec Driven Development: write the spec before the code, and keep it alive

A request comes in as a three-line ticket: "Automatically remind customers who haven't paid their invoice". Two days later, the code is there, the tests pass, the demo works. Then in code review the first question lands: "What if the customer only paid part of it?". Nobody had asked. The code still does something in that case, by default, without anyone having decided it.
Spec Driven Development (SDD) starts from that observation: the cost of an ambiguity grows at every stage it survives. Found in a spec, it costs a sentence. Found in code review, it costs a rewrite. Found in production, it costs an incident. This article explains what SDD is, why it's back in the spotlight, when it's worth it, how to apply it without drowning in paperwork, and what exists alongside it.
The what: the spec as the source of truth#
The principle fits in one sentence: you first write what the system must do, in a precise and verifiable form, and the code is derived from that description. The spec isn't a form you fill in before coding what you had in mind anyway. It's the reference artifact: when the code and the spec disagree, one of them has a bug, and you need to know which.
In practice, SDD separates three levels we usually mix up in the same ticket or the same conversation:
SPEC PLAN TASKS
What and why ──▶ How ──▶ In what order
(behavior, (architecture, (small steps,
business rules, technical choices, verifiable,
acceptance constraints) each one testable)
criteria)
│ │
└────────────── tests verify the spec ◀────────────┘
- The spec describes observable behavior, from the point of view of the user or the calling system. It says nothing about databases or frameworks.
- The plan turns the spec into technical decisions: which modules are affected, which data model, which constraints (performance, security, compatibility).
- The tasks split the plan into steps small enough to be coded, tested and reviewed one at a time.
That separation isn't new in itself. What is specific to SDD is the order (you don't move on to the plan while the spec still has open questions) and the status of the spec (it stays in the repo, versioned next to the code, and it evolves with it).
Three levels of commitment#
Not everyone means the same thing by "SDD". Birgitta Böckeler, in an analysis published on Martin Fowler's site, distinguishes three levels that help sort things out:
| Level | What happens to the spec | What it implies |
|---|---|---|
| Spec-first | Written before the code, to guide one task | Can be thrown away once the feature ships |
| Spec-anchored | Kept and updated with every change | Becomes living documentation of the behavior |
| Spec-as-source | The only thing a human edits, the code is regenerated | Code is no longer an artifact maintained by hand |
The first level is easy to adopt tomorrow morning. The second takes discipline but pays off the most. The third is still experimental today: it assumes code generation reliable enough that you stop reviewing what comes out, which isn't where I am on real projects.
Why SDD is coming back now#
Writing a spec before coding is what teams already did in the 1990s, with 80-page documents signed off by committees. Agile rightly pushed back against those frozen specs written six months before the first line of code. So why is the topic back?
Because coding agents do what you ask, not what you meant. A human developer given a vague ticket asks questions, talks to the business, remembers last week's meeting. An agent fills the gaps with the most likely assumption and produces code that looks right. Vibe coding (describe vaguely, look at the output, fix by feel) works for a prototype. On a codebase that has to live, it produces exactly the invoice reminder scenario: behavior nobody chose.
The spec gives the agent what a human would have gotten by asking questions. And it gives the reviewer a reference: you no longer review a 400-line diff wondering "is this what we wanted", you check that it implements the criteria in the spec.
Several tools were built around this idea in 2025:
- GitHub Spec Kit, an open source toolkit that drives the cycle through
commands (
/specifyfor the spec,/planfor the technical plan,/tasksfor the breakdown, then implementation), with a "constitution" file holding the project's non-negotiable principles. - Kiro, AWS's IDE, which produces for each feature a
requirements.md(requirements in EARS format), adesign.mdand atasks.md. - Tessl, which explores the spec-as-source level.
None of these tools is required to practice SDD. Three Markdown files in a
specs/ folder are enough, and that's what I show below.
SDD isn't tied to AI. A team with no coding agent at all gets the same benefits: ambiguities resolved early, simpler reviews, tests derived from the criteria. Agents have simply made the cost of a missing spec more visible, and faster to pay.
A complete example: reminders for unpaid invoices#
Let's take the original ticket and walk through the three steps.
Step 1: the spec#
# Spec: automatic reminders for unpaid invoices
## Context
Reminders are sent by hand today, late, and some are forgotten. We want
a first automatic reminder, without ever reminding a customer who has
already paid.
## User story
As the business owner, I want customers who are late on payment to
receive a reminder, so I no longer track due dates by hand.
## Business rules
- R1. An invoice is "overdue" when its due date is more than 7 calendar
days in the past and an amount is still owed.
- R2. A partial payment does not suspend the reminder. The reminder
states the remaining amount owed, not the original amount.
- R3. An invoice is reminded at most once every 14 days.
- R4. No reminder is sent for an invoice flagged "disputed".
- R5. After 3 reminders, the invoice moves to "manual follow-up" and is
no longer reminded automatically.
## Acceptance criteria
- AC1. Given a €1,000 invoice due 8 days ago and unpaid, when the daily
job runs, then a reminder stating €1,000 is sent.
- AC2. Given a €1,000 invoice due 8 days ago with €400 already paid,
when the job runs, then the reminder states €600.
- AC3. Given an invoice reminded 10 days ago, when the job runs, then
no reminder is sent.
- AC4. Given a disputed invoice due 30 days ago, when the job runs,
then no reminder is sent.
- AC5. Given an invoice already reminded 3 times, when the job runs,
then it moves to "manual follow-up" and no reminder is sent.
## Out of scope
- Reminders by SMS or mail.
- Late payment penalties.
- Letting the owner edit the email template.
## Open questions
- [x] Calendar or business days? Calendar (confirmed with the owner).
- [x] Partial payment: remind or not? Yes, on the remaining amount (R2).A few remarks on this document, because this is where most of the value is.
The partial payment question is settled in writing (R2 and AC2). That's exactly the question that would have come up in code review. It surfaced while writing the numbered examples: you can't write AC2 without deciding what happens.
The criteria are examples, not intentions. "The system reminds customers correctly" isn't a criterion: you can't check it. "A reminder stating €600 is sent" is one: a test can prove or disprove it.
Out of scope is explicit. It's the most underrated section. Without it, someone (human or agent) will add late penalties "because it made sense", and the review turns into a negotiation.
Open questions are ticked. As long as a box stays empty, you don't move on to the plan. It's the simplest rule in SDD and the most effective one.
For requirements, the EARS format (Easy Approach to Requirements Syntax) gives a structure that rules out vague sentences: "WHEN an invoice is more than 7 days overdue, THE SYSTEM SHALL send a reminder". The Given / When / Then format, inherited from BDD, does the same job for acceptance criteria. Which one you pick matters less than making each sentence describe a trigger and an observable result.
Step 2: the plan#
The plan is where technical choices appear, and only there.
# Plan: automatic reminders for unpaid invoices
## Approach
- A daily scheduled job (6am) in the `billing` module.
- Eligibility is a pure domain function:
`eligibility(invoice, reminderHistory, today)`, with no database or
system clock access. All rules R1 to R5 live there.
- Sending goes through the existing `NotificationPort`.
## Data model
- New `reminder` table (invoice_id, sent_on, amount_reminded).
- New invoice status `MANUAL_FOLLOW_UP`.
## Constraints
- The job must be idempotent: run twice on the same day, it doesn't
send two reminders (R3 guarantees it if the database insert happens
before sending, in the same transaction).
- No customer data in logs.
## Risks
- Volume: about 2,000 open invoices, no pagination needed.Notice that the plan references the spec's rules (R1 to R5, R3) instead of rewriting them. If a rule changes, it changes in one place.
Step 3: the tasks#
# Tasks
- [ ] T1. `eligibility` function + tests AC1 to AC5 (pure domain)
- [ ] T2. `reminder` table and migration
- [ ] T3. `MANUAL_FOLLOW_UP` status
- [ ] T4. Daily job: loads invoices, calls `eligibility`, records the
reminder then notifies
- [ ] T5. Integration test: two runs on the same day = one reminder sentEach task produces a change that can be reviewed on its own, and each one says how you'll know it's done.
From criterion to test#
The link between spec and code goes through the tests. Each acceptance criterion becomes a test, and the test name points to the criterion:
class ReminderEligibilityTest {
private val today = LocalDate.of(2026, 3, 20)
@Test
fun `AC2 - a partial payment is reminded on the remaining amount`() {
val invoice = invoice(
amount = 1_000.euros,
alreadyPaid = 400.euros,
dueDate = today.minusDays(8),
)
val decision = eligibility(invoice, history = emptyList(), today)
assertThat(decision).isEqualTo(Decision.Remind(amount = 600.euros))
}
@Test
fun `AC4 - a disputed invoice is never reminded`() {
val invoice = invoice(
amount = 1_000.euros,
dueDate = today.minusDays(30),
status = Status.DISPUTED,
)
val decision = eligibility(invoice, history = emptyList(), today)
assertThat(decision).isEqualTo(Decision.Skip)
}
}The day someone asks "why are we reminding a customer who already paid €400?", the answer is traceable: rule R2, criterion AC2, the test with the same name. And if the business changes its mind, you update the spec first, the test next, the code last. It's the same cycle as TDD, described in the article on real TDD with Claude Code, with one extra step upstream: the test no longer comes out of the developer's head, it comes from an example the business signed off on.
When SDD is worth it, and when it's useless#
A spec has a cost: writing time, review time, and maintenance time if you keep it. It pays off when that cost is lower than the cost of the ambiguities it prevents.
| Situation | SDD useful? | Why |
|---|---|---|
| Feature with business rules (billing, permissions, calculations) | Yes | Edge cases are many and expensive to find late |
| Work handed to a coding agent across several files | Yes | The agent fills gaps with guesses |
| Several teams or a client involved | Yes | The spec acts as a contract and avoids "that's not what we agreed" |
| Changing an existing, poorly documented feature | Yes | Writing the spec of the current behavior often reveals bugs |
| One-line bug fix | No | A regression test is enough |
| Throwaway prototype, technical spike | No | The goal is to learn, not to ship precise behavior |
| Purely technical change (version bump, rename) | No | No observable behavior changes |
A simple rule of thumb: if you can't write at least three distinct acceptance criteria for the request, it's probably too small for a spec. If you can't write a single one, it's too vague to be coded, and that's precisely the sign you need to write the spec.
Applying it day to day#
A minimal file layout#
specs/
├── 012-invoice-reminders/
│ ├── spec.md
│ ├── plan.md
│ └── tasks.md
└── 013-accounting-export/
└── spec.md
One folder per feature, numbered, in the repo. Not in Confluence, not in Jira: if the spec isn't versioned with the code, it will drift from the code within weeks.
How a feature unfolds#
- Draft the spec in 20 to 30 minutes. Context, rules, criteria with numbers, out of scope. Let open questions show up instead of settling them alone.
- Get the spec reviewed, not the code. By the business, the PO, or another developer. A one-page spec takes five minutes to review, a 400-line diff doesn't. This is when disagreements get settled, while they only cost a sentence.
- Close the open questions. Each answer becomes a rule or a criterion.
- Write the plan, referencing the spec's rules.
- Split into tasks, each testable and reviewable on its own.
- Implement task by task, starting with the tests that match the criteria.
- In code review, check conformance to the spec on top of code quality. The PR links to the spec, and the new tests carry the criteria names.
When the spec changes along the way#
It will, that's normal. You'll find a case during implementation that nobody saw. The rule: update the spec in the same PR as the code. If the new case changes the expected behavior, it goes through a quick validation again. If you let the code evolve alone, you're back to square one, with a document that lies on top of it.
With a coding agent#
The flow is the same, with two adjustments. First, ask the agent to read the spec and list ambiguities before writing the plan: it's very good at spotting gaps ("what happens if the due date falls on February 29?"). Then have it implement one task at a time, reviewing in between. It's the equivalent of plan mode before any non-trivial task, applied at the scale of a whole feature. The project's stable principles (architecture, conventions, test commands) belong in a context file like CLAUDE.md, not in every spec.
Having an agent write the whole spec, then skimming it to approve, is the fastest way to lose everything SDD brings. The agent will produce a long, well-structured, plausible document, with rules nobody decided. You've moved the ambiguity from the code to the spec without resolving it. The agent can draft, rephrase, look for gaps. The decisions have to come from someone who knows the business.
The traps#
Rebuilding the waterfall without saying so. If each feature's spec takes two weeks and goes through three committees, you've reinvented the 80-page document. An SDD spec covers a feature, not a product, and it's written in hours, not weeks. The cycle stays short and iterative.
Writing pseudo-code in the spec. "Query the invoice table with a LEFT JOIN on payment" has no place in the spec. As soon as the spec talks about implementation, it locks in choices that belong to the plan, and it becomes unreadable for the business.
Unverifiable criteria. "Processing must be fast", "the email must be clear". If no test can fail, it isn't a criterion. "Processing 2,000 invoices takes under 30 seconds" is one.
The dead spec. Written carefully, never updated. Six months later it describes behavior that no longer exists, and someone relies on it. An outdated spec is worse than no spec. If the team doesn't plan to maintain it, it's better to own the spec-first level and archive it once the feature ships.
A spec for everything. Requiring a spec to rename a variable or bump a dependency turns a good practice into bureaucracy, and the team will end up bypassing the rule, including where it was useful.
Forgetting out of scope. Without it, every reviewer projects their own expectations, and the PR becomes the place where you find out what everyone thought was included.
Mixing spec and plan. Putting the "what" and the "how" in the same document prevents changing one without touching the other. The day you replace the daily job with a Kafka event, the spec shouldn't change by a single line.
Other approaches, and how they fit together#
SDD doesn't replace existing practices, it sits upstream of most of them.
| Approach | What it drives | Relationship with SDD |
|---|---|---|
| TDD | Code design, one test at a time | Complementary: the spec's criteria provide the first tests |
| BDD | Behavior, through executable scenarios (Gherkin) | Very close: BDD makes criteria executable, SDD adds plan and tasks |
| DDD | The domain model and language | Complementary: the ubiquitous language makes the spec more precise |
| Contract-first / API-first | The interface contract (OpenAPI, AsyncAPI, Avro schema) | A special case of SDD applied to a technical boundary |
| ADR | An architecture decision and its reasons | Complementary: an ADR explains a choice in the plan, not the behavior |
| Design by contract | Preconditions, postconditions, invariants in code | Same idea, at the level of a function rather than a feature |
| Vibe coding | Nothing written, you iterate on the output | The opposite: useful to explore, risky to ship |
BDD deserves a clarification, because the confusion is common. BDD starts
from Given / When / Then scenarios written with the business, then makes
them executable (with Cucumber, for example). SDD often borrows that format
for its criteria, but doesn't require them to be executable as-is, and it
adds the plan and tasks steps. You can perfectly well do SDD with .feature
files as criteria: it's actually a solid combination.
Contract-first is probably the form of SDD many teams already practice without naming it: write the OpenAPI file before the implementation, have the API consumers validate it, generate the clients. Same principle, applied to an interface instead of business behavior. It's also the spirit of Schema Registry for Kafka: the schema is the contract, the code conforms to it.
A checklist for a good spec#
Before moving on to the plan, I check these points:
- The context explains the problem, not the solution.
- Each business rule is numbered and fits in one or two sentences.
- Each acceptance criterion contains concrete values (amounts, dates, statuses).
- At least one criterion covers an edge case or an error case.
- Out of scope lists what someone might assume is included.
- No open question is left unanswered.
- No implementation term (table, endpoint, class).
- Someone other than the author has reviewed it.
What to take away#
SDD isn't about writing more documentation. It's about moving decisions to the point where they're cheapest: before the code, as precise examples, reviewed by the right people. The spec becomes the shared reference for the business, the developer, the reviewer, and the coding agent if there is one.
To get started, you need neither a tool nor a team process: on the next feature that touches business rules, write one page with numbered rules, five criteria with numbers and an out of scope section, and get it reviewed before opening the IDE. The questions that come up during that review are exactly the ones you would otherwise have found in code review, or in production.


