Spec-Driven AI Tools (Part 4)
Neil Haddley • September 30, 2026
The comparison: what OpenSpec, Spec-Kit, and BMAD each did with the identical calculator and the identical two feature requests — where they genuinely agreed, where they quietly built different products, what it cost, and where the source article held up against running all three myself
This is the last of four posts trying three spec-driven AI development tools against the same fixed task, prompted by Ran the Builder's honest take on three spec-driven AI tools. Part 1 ran OpenSpec, Part 2 ran Spec-Kit, Part 3 ran BMAD — each on a fresh copy of the same flat, four-function, Windows 1.0-style calculator, each given the identical two feature requests, worded the same way every time: a Programmer Mode and a Statistics Mode, both real Windows Calculator features, neither request engineered to hit a trap I already expected.
This post is where that control actually pays off. Three tools, one task, no averaging across different projects — so where they disagree, it is a real disagreement about what to build or how to check it, not noise from testing different things.
The same ambiguities, three honestly different answers
I never specified bit width, negative-number formatting, or the standard-deviation formula. Here is what each tool did with that silence.
| Question | OpenSpec | Spec-Kit | BMAD |
|---|---|---|---|
| Negative numbers in Programmer Mode | Two's-complement, fixed 32-bit word — silently decided, rationale in design.md | Sign-magnitude — silently decided, rationale in research.md | Sign-magnitude — asked me directly, defaulted on "keep it simple" |
| Bitwise ops / word-size selection | Excluded — silent, documented non-goal | Excluded — silent, documented non-goal | Excluded — asked me directly |
| Disable invalid digit keys visually | Yes — DOM-level, "UI communicates what's valid" | No — explicit YAGNI call against the letter of its own requirement | Yes — DOM-level |
| Standard deviation formula | Population (n) — silently decided | Sample (n−1) — silently decided, and its own worked example initially got the math wrong (caught and fixed at planning) | Sample (n−1) — asked me directly |
| Empty-data-set Sum | Errors, same as Average/StdDev | Errors, same as Average/StdDev | Shows 0 — Average/StdDev still error (mathematically the more precise split) |
| Data-point removal | Out of scope, explicit non-goal | In scope — its own User Story 4 | Out of scope, explicit non-goal |
| Statistics + Programmer Mode coexistence | Statistics mode disables Programmer's operators/base toggle | Declared explicitly independent/orthogonal | Mutually exclusive — activating one turns the other off; data list persists across toggles |
| Standard arithmetic operators during Statistics Mode | Disabled | Disabled (a planning-stage decision, not a written requirement) | Left live, deliberately — reviewed, found reachable, kept as matching real Windows Calculator's "compute a value, then Add it" workflow |
Read across a row and the pattern holds: every ambiguity got a real answer, every answer was different at least once, and in three cases (negative-number format, bitwise scope, stddev formula) BMAD is the one tool that actually asked me instead of deciding on its own. I answered "keep it simple" every time I was asked, specifically so I would see each tool's own default rather than my own preference — the asking itself is the finding, not which default came out.
The two most consequential divergences are the operator-disabling ones. Three tools built three different products from "add a Statistics Mode," and only one of them — BMAD, after its review layer checked a finding and rejected it as a deliberate feature rather than a bug — actually lets you compute an expression and capture its result into the data set, which is closer to how the real Windows Calculator Statistics box behaves than either of the other two implementations.
Three verification models, not one
This is where the three tools stopped resembling each other most.
OpenSpec: propose → apply → sync → archive. Once a change is archived, OpenSpec never looks at it again. Nothing in this series' two OpenSpec runs surfaced a bug after the fact, but that is not the same as OpenSpec having no bugs — it is that nothing in its workflow goes looking. I hand-verified both features myself (npm test, then clicking through the browser) precisely because the tool's own workflow stops at archive.
Spec-Kit: /speckit-converge, a step I had to explicitly re-invoke, sometimes several times. For Programmer Mode it took three rounds to reach "Converged" — round one found two test-coverage gaps, round two found a genuine integer-truncation bug (division in Programmer Mode wasn't truncating, so 11 ÷ 10 in binary displayed as 1.1, directly contradicting the integer-only requirement — I reproduced it myself before trusting the claim), round three confirmed clean. For Statistics Mode, one round, zero findings — though the plan step for that same feature had already caught something rarer: a math error in its own earlier /speckit-specify output, a worked example that asserted the wrong standard deviation for a data set (I verified by hand: population stddev of {2,4,4,4,5,5,7,9} is 2, which is what the spec wrongly claimed; sample stddev, what its own requirement actually specified, is ≈2.138090). That is Spec-Kit catching a contradiction in its own prior output before a line of code existed, which is arguably a stronger result than catching the equivalent bug after implementation.
BMAD: a three-layer parallel review (named "Blind Hunter," "Edge Case Hunter," "Verification Gap") that runs automatically inside the same build pass, before anything reaches me. Across both features it found six real, patchable defects — Programmer Mode's activation not truncating a fractional value and a substring-matching bug in digit validation; Statistics Mode's Add silently re-adding a just-computed statistic and an inconsistent display formatter, plus assorted cosmetic fixes — and, just as informative, it correctly rejected several plausible-sounding false findings with a stated reason each, and correctly kept the "operators stay live" behavior as a deliberate feature rather than patching it away. I never saw a broken intermediate version of either feature.
Ranked by how much verification actually happened without my asking for a second pass: BMAD ran automatically every time; Spec-Kit ran only when I re-invoked it, but iterated until clean and caught its own spec's math error along the way; OpenSpec ran zero times after archive, by design.
This is not just our own finding. An independent, much larger comparison — a real multi-phase Python application, a different implementing model (Claude Opus 4.7), nothing to do with our calculator — found the identical asymmetry: OpenSpec's own verify/validate step "mostly found nothing," while Spec-Kit's analyze step "consistently surfaced real issues before implement ran," called "the single best polish helper" that reviewer had used in either framework. Two unrelated tasks, two independent runs, the same gap.
Who can actually read the documentation
A claim worth checking directly against our own artifacts, because it is easy to assert and easy to get backwards: which tool's documentation is written for a person, and which for the agent implementing it.
All three core specs turned out to read the same way — clean prose, no code. OpenSpec's archived spec:
MARKDOWN
1## Requirement: Base mode selection 2The system SHALL provide a control to switch the active calculation base 3among Decimal, Hexadecimal, Octal, and Binary...
Spec-Kit's spec.md, by explicit enforced rule — its own pre-planning checklist requires "No implementation details," "Written for non-technical stakeholders" — reads indistinguishably:
MARKDOWN
1### User Story 1 - Enter and view numbers in another base (Priority: P1) 2A user switches the calculator into Programmer Mode and selects Binary, Octal, 3or Hexadecimal. They enter a number using only the digits valid for that 4base and see it displayed correctly.
Where the three tools actually diverge is what surrounds that clean core, and what happens to it afterward.
Spec-Kit separates audiences into different files, most rigorously of the three — at the cost of volume. Its spec.md stays clean, but four more files sit beside it, and they are unambiguously written for the implementer, not a stakeholder:
MARKDOWN
1# Contract: `CalculatorEngine` public API 2| Field | Type | Description | 3|---|---|---| 4| `mode` | `"standard" \| "programmer"` | **New.** Current calculator mode. |
That split is a real strength — a non-technical reviewer never has to open contracts/ or research.md — but the volume is a real, independently documented cost. A genuine Spec-Kit usage report describes "uncontrolled complexity escalation... 420 lines → 873 lines for similar features," "no guidance on 'appropriate size,'" on real projects — a pattern our two small features didn't trigger, but a fair thing to weigh in.
BMAD keeps everything in one file, marked but not separated. Its SPEC.md puts human intent and engineering detail back to back:
MARKDOWN
1## Intent 2**Problem:** The calculator only operates in decimal; users who need quick 3binary/octal/hex conversions... have no way to switch bases. 4 5## Code Map 6- `calculator-logic.js` -- `CalculatorEngine`: add `base` state (default 10) 7 and `setBase`...
The <frozen-after-approval> tag wrapping the human-owned section is a genuinely useful boundary marker — it tells a reader exactly where to stop if that is all they need — but it is a marker inside one document, not a separate file you could hand a stakeholder without the engineering detail riding along underneath it.
Only OpenSpec produces something that stays true after the work is done. Spec-Kit's and BMAD's specs live in their own per-feature folder and stay there; nothing merges them into a whole-system view. OpenSpec's delta model does exactly that — both archived changes in this series folded cleanly into one evolving openspec/specs/ directory, so a team member arriving cold, months later, can open one file per capability and get an accurate answer with no changelog-archaeology required. That is a different, and arguably more valuable, kind of readability than "clean prose at planning time," and it is the one none of this series' comparisons captured until the archive step actually got tested.
(One external comparison making the rounds claims — using a citation I checked and could not verify — that Spec-Kit's output is "hard to read for non-technical reviewers (BAs/QA)." Our own artifacts don't support that as stated: its spec.md is the most rule-enforced human-safe document of the three. The real, checkable version of that concern is the volume risk above, not the register of the spec itself.)
Operational friction, running all three the same way
None of the three tools' own slash-command skills declared enough pre-authorization to run start-to-finish without me approving something, but the shape of the friction differed:
- OpenSpec pre-authorizes its own CLI calls (Bash(openspec:*) in its command frontmatter), so a non-interactive run gets furthest before hitting a wall — usually only at verification (npm test), which the workflow explicitly expects a human to run anyway.
- Spec-Kit pre-authorizes nothing. Every Bash call, including its own bundled setup scripts, needed approval. I ran those scripts myself and fed the output back, every single step.
- BMAD needed the most: two separate plugins to reach the same starting line the other two reached with one install command, a filesystem-read approval for skill reference files living outside the project (they load from a global plugin cache, not local project files), a one-time subagent-authorization prompt, and the same per-Bash-call wall as Spec-Kit on top of all of it. The heaviest lift of the three to run unattended — none of which reflects on output quality, but a real cost a team would feel from day one.
Driver-agent portability, checked but not run
Everything above used Claude Code throughout, by deliberate choice made at the start of this series. Worth a note on what happens if the driver agent isn't Claude Code — checked directly against each tool's own integration list, not tested end to end the way the rest of this series was:
- Spec-Kit has a genuine, named DeepSeek Harness integration. Its own source is explicit about it: "DSH discovers project skills from .dsh/skills (its native root, highest provider rank)... Project guidance in AGENTS.md at the repo root is loaded automatically by DSH." specify check lists DeepSeek Harness (available), and --integration dsh writes straight into DSH's preferred root, in the format it expects natively.
- OpenSpec reaches it too, through a generic fallback. openspec init --tools agents writes to a shared .agents/skills/ convention — the same one Spec-Kit's own DSH integration notes as an alternate root DSH also reads from — but with no DSH-specific tuning.
- BMAD's usual install path has no DeepSeek Harness support at all. The Claude Code plugin method this whole series used (claude plugin install bmad-method@bmad) is Claude Code/Codex only. BMAD's alternate generic installer can reach the same shared .agents/skills/ convention, but I tried targeting DSH directly and it was flatly rejected — Invalid agents: dsh — out of a list of 70-plus named agents that doesn't include it; universal is the fallback that actually works.
If DeepSeek Harness specifically is your driver agent, that is a real, structural reason to lean Spec-Kit over the other two — not a preference, a support-matrix fact.
Does this support architecture-by-discovery?
A different kind of fit question, and one worth naming honestly before answering it: nothing in this series tested it end to end. Experienced architects often work by proving a "critical move" — a single, real, working slice through the riskiest part of a system, at low level of detail, early — while the report to sponsors stays high-level, and stakeholders get brought along through mockups, prototypes, and conversation rather than a finished document. That is a two-track way of working, and it sits partly outside what any of these three tools are built for: all three want whatever you build to become part of a permanent artifact (OpenSpec's archived living spec, Spec-Kit's per-feature folder, BMAD's frozen SPEC.md), not a disposable spike you prove and discard.
With that said, they do not fit equally badly.
BMAD is the closest philosophical match, and this much is grounded in what we actually saw. Its own delivery loop is explicitly Clarify → Plan → Build and verify → Learn and adjust — "Clarify" as a first-class stage, with an explicit loop back into planning once something is learned. It was also the only one of the three that held a real conversation with us anywhere in this series, stopping to ask about bit width, negative-number representation, and the standard-deviation formula rather than silently producing a finished artifact. Its named personas are built around exactly this dynamic — the PM persona "speaks like a detective interrogating a cold case: short questions, sharper follow-ups"; the UX Designer persona "speaks like a filmmaker pitching the scene before the code exists, painting user stories that make you feel the problem." It also ships bmad-brainstorming, bmad-forge-idea, bmad-product-brief, and bmad-prfaq (an Amazon-style working-backwards press release, a real stakeholder-alignment technique) — skills we never actually ran in this series, so this is inference from its command surface, not tested evidence like the rest of this post. The real cost working against it: BMAD was also the heaviest to set up and the most expensive of the three in our own numbers, which cuts against quickly trying several throwaway things and discarding most of them.
OpenSpec is the best operational fit for rapid, cheap iteration, without being purpose-built for discovery either. It was the cheapest and fastest by a wide margin in our own measurements, and that speed is itself an enabler — propose something small, apply it, see whether it works, propose something different next, all at low cost. It also has an /opsx:explore command, explicitly meant for thinking through an idea before committing to a proposal, which we never invoked in this series. I cannot tell you how well it supports mockup-and-conversation-style discovery specifically, only that it exists and that OpenSpec's low-ceremony posture does not structurally fight against using it that way.
Spec-Kit is the one most likely to get in the way, and this is not speculation — its own ecosystem is the evidence. Its quality checklist requires "all mandatory sections completed" before a spec can move to planning, and the Constitution Check gates planning on every principle passing. That is a process built around knowing what you want, fully, before formalizing it — close to the opposite of proving one critical technical move in isolation while everything else stays deliberately vague. The independent comparison cited earlier in this post noted that the Spec-Kit community built a third-party plugin, TinySpec, described as "the lightweight path that core Spec Kit does not have." The ecosystem itself is confirming that core Spec-Kit does not support this kind of low-commitment exploration — someone had to bolt it on separately.
So: will any of them help? BMAD's conversational posture and stage-based loop is the closest fit, with real caveats on cost and untested skills. Will all three support it? No — Spec-Kit's gates are specifically at odds with it, by its own community's admission. Will any get in the way? Spec-Kit, most clearly, for the same reason.
What it cost
Pulled from the actual session transcripts of every claude invocation across all three tool runs — real token counts, not estimates:
| Cache write | Cache read | Output | Approx. cost | vs. OpenSpec | |
|---|---|---|---|---|---|
| OpenSpec | 1,083,092 | 18,027,680 | 270,815 | ≈$9.02 | 1.0× |
| Spec-Kit | 2,988,806 | 49,760,105 | 664,711 | ≈$24.07 | 2.7× |
| BMAD | 4,414,254 | 40,756,494 | 462,433 | ≈$23.81 | 2.6× |
(Sonnet 5 pricing, standard cache-token multipliers — cache write ≈1.25× input rate, cache read ≈0.1×.) OpenSpec was the clear cheapest in our own runs, and Spec-Kit and BMAD landed close together at roughly 2.6–2.7× OpenSpec's cost.
Two honest caveats before trusting that table too far. First, it measures what it actually cost us to run this investigation, friction included — every non-interactive Bash-approval workaround was an extra round-trip, and BMAD's own multi-agent review architecture spawns more calls by design than a single-agent tool ever would, so part of OpenSpec's advantage here is "asked for fewer workarounds," which is a real advantage but not purely a measure of raw model cost. Second, these numbers are not directly comparable to the source article's own figures — different model generation, different task, different measurement method (per-tool total there, per-session-directory sum here). The article reports OpenSpec at $95 total (the highest implementation cost of its non-BMAD-Full group, in fact), Spec-Kit at $75, BMAD Quick at $85, and BMAD Full — a heavier mode we never exercised, since our bmad-build runs stayed on its lightweight path both times — at $200. Its own stated conclusion: "BMAD Quick, Spec-Kit, and OpenSpec land in the same ballpark on both speed and cost. The differences are noise." Our numbers disagree with the article's exact rank order between OpenSpec and Spec-Kit, but land on a similar shape: three tools within roughly the same order of magnitude, one heavier mode of one tool (BMAD Full, which we did not test) as the real outlier.
A third, independent data point, for calibration: the same OpenSpec-vs-Spec-Kit comparison cited above measured Spec-Kit costing 81% more than OpenSpec and using about 69% more tokens, on a real multi-phase application — a real gap, but nowhere near our own measured 2.6–2.7×. That smaller, independently-measured number is a useful check on our own: it supports reading our own multiplier as inflated by this specific investigation's friction — approval-wall workarounds, a couple of my own mistakes — rather than as a clean measure of what these tools cost to run in general.
Against the source article
Worth stating plainly where independently running all three actually changed my view, rather than confirming what I read first.
"OpenSpec assumes context and adds unstated rationale." This undersells it. Every judgment call I watched OpenSpec make — 32-bit two's-complement, the overflow-wraparound rule, population standard deviation — came with a named alternative it considered and rejected, in a dedicated Decisions section, before a line of code existed. It did not ask me, but it did not hide its reasoning either. "Silent but documented" is a real, different thing from "silent and unstated," and the article's phrasing reads as the latter.
"BMAD: full-lifecycle framework with dedicated elicitation and course-correction workflows." True, and worth a footnote: BMAD has been substantially rewritten (v6, a skills-based architecture, monthly releases with real breaking changes) since whatever version the article likely tested. The classic elicitation/correct-course concepts survive as skills, but the tool that installs today is not the same shape the article's description implies. That is itself a finding — a spec-driven-tools comparison can go stale within months in this space, and BMAD's own changelog is the clearest evidence of that in this whole series.
"OpenSpec: best out-of-the-box experience, parallel work by default." Confirmed, specifically on the low-friction dimension — fewest approval walls, no plugin installs, fastest path from a bare repo to an archived, working change. I would not use "parallel work by default" to describe what I actually saw, though: OpenSpec's tasks.md groups by implementation layer (engine, then UI, then tests), where Spec-Kit's is explicitly organized by user story with [P] parallel markers and a stated MVP-first strategy — if "built for a team to split work" is the claim, Spec-Kit's task structure is the one that actually shows it.
Cost claims generally. The specific multipliers sometimes repeated about this space (that OpenSpec is dramatically cheaper, or that BMAD costs many times more by default) are not supported by either the article's own published figures or by what we measured running all three ourselves. Both datasets show three tools in a broadly similar cost band, with only BMAD's heavier PRD/architecture mode — which neither the article's "BMAD Quick" arm nor our own runs exercised — landing meaningfully higher.
Mainstream standard, lightweight choice, simulated dev team
A framing worth keeping, refined against everything above:
Spec-Kit reads as the most conventionally rigorous. GitHub-backed, the heaviest artifact set of the three (spec.md, plan.md, research.md, data-model.md, contracts/, quickstart.md, tasks.md, plus a pre-planning quality checklist), and the only one with a literal enforced gate — the Constitution Check, a pass/fail table checked against every principle before planning can proceed. Closest to a conventional, audited SDLC of the three, and it earned that with the article's spec math-error catch and the division-truncation bug, both real. It shows in adoption too, not just structure: as of today, Spec-Kit sits at 139,453 GitHub stars against OpenSpec's 70,691 — checked directly, not quoted secondhand — roughly double, a ratio an independent comparison found holding steady since at least May 2026 (96k vs 47k then). "Mainstream" isn't just a vibe here.
OpenSpec reads as the lightweight choice, more precisely: low-ceremony rather than strictly "agile" in the formal sense — nothing about sprints, backlogs, or iteration cadence, just the fastest, cheapest, least-friction path from a plain request to an archived, working, documented change. Its delta-spec model genuinely works: two archived changes accumulated cleanly into a living specs/ directory that read as accurate both times.
BMAD reads as the simulated development team, and not just as branding. Five named personas with distinct stated voices, a literal team = "software-development" field in its own config, and — the concrete evidence, not just the flavor — it is the only tool that stopped and asked me real product decisions the way a PM actually would, and the only one whose review process is explicitly adversarial by design (three differently-named review layers, each checking the others' blind spots) rather than a single pass.
So, which one
Not "which is best" in the abstract — which is what the article itself was right to avoid too. Grounded in what actually happened across six full feature cycles:
- Want the fastest, cheapest path to a working, documented change, and you are comfortable owning your own judgment calls without being asked? OpenSpec. It is fastest and cheapest largely because it skips both things that cost the other two tools time, turns, and money: asking you anything, and checking its own work after the fact. That is a real, honest trade-off, not a flaw — but it means the cost advantage and the verification gap are the same decision, not two separate facts about the tool.
- Want the most rigorous, most conventional process, with an enforced gate and a workflow that keeps checking itself until it can prove correctness? Spec-Kit. It caught the most surprising bug in this whole series — a contradiction in its own earlier output — before any code existed.
- Want to be asked rather than guessed for, and want verification to happen automatically rather than on your own initiative? BMAD. It is the heaviest to set up and the most expensive of the three in our own runs, and that cost bought something real: the only tool that treats "check whether this is actually a bug before fixing it" as seriously as "find bugs."
All three repos are public and unedited: github.com/Haddley/specdriven — main for OpenSpec, spec-kit and bmad branches for the other two — full commit history from the identical baseline through every propose, plan, build, and review, nothing cleaned up after the fact.