Same model. Same repo. Same task.
The comparison that matters is not against a strawman. It is against a modern, well-configured agent with a good prompt, and against a free alternative people actually use.
A · agent alone
auto mode, a good prompt
B · spec kit
the closest free alternative
C · EDIFY
/spec to /verify, with the graph
same model version · same opening prompt · same limits
Published
on purpose.
empty.
Benchmarks
Agent alone, Spec Kit, and EDIFY, on real open-source repositories.
BENCHMARK IN PROGRESS
NO NUMBER UNTIL IT IS REPRODUCIBLE
NO NUMBER UNTIL IT IS REPRODUCIBLE
Methodology and results will be public,
including the runs where EDIFY loses.
benchmarks/ · docs/benchmarks.md
benchmark in progress · not covered yet
Until real, reproducible numbers exist, every cell stays a dash.
A benchmark filled in before the benchmark ran is the fastest way to lose a technical audience.
What has to be true before a number goes up.
Seven conditions. Every one is checkable in the public repository when the runs land.
012 to 3 non-trivial open-source repositories, more than one language
0210 to 20 varied tasks, including a bug with a wrong first hypothesis
03The same model version, opening prompt and limits in every arm
04Success, tests, interventions, retries, tokens, cost, wall-clock time
05Scripts, prompts and raw transcripts published in benchmarks/
06A short video of a visible failure case
07Case studies from design partners come after, not instead
not met yet · each one becomes filled when it is true and checkable
Help run it
Reproduce it,or break it.
Propose a repository, a task, or a flaw in the method. A reproduction report is worth more than a star.