customize.toml
# DO NOT EDIT -- overwritten on every update.
#
# Workflow customization surface for bmad-retrospective. Mirrors the
# agent customization shape under the [workflow] namespace.
[workflow]
# --- Configurable below. Overrides merge per BMad structural rules: ---
# scalars: override wins • arrays (persistent_facts, activation_steps_*): append
# arrays-of-tables with `code`/`id`: replace matching items, append new ones.
# Steps to run before the standard activation (config load, greet).
# Overrides append. Use for pre-flight loads, compliance checks, etc.
activation_steps_prepend = []
# Steps to run after greet but before the workflow begins.
# Overrides append. Use for context-heavy setup that should happen
# once the user has been acknowledged.
activation_steps_append = []
# Persistent facts the workflow keeps in mind for the whole run
# (standards, compliance constraints, stylistic guardrails).
# Distinct from the runtime memory sidecar — these are static context
# loaded on activation. Overrides append.
#
# Each entry is either:
# - a literal sentence, e.g. "All retrospectives must produce SMART action items with named owners."
# - a file reference prefixed with `file:`, e.g. "file:{project-root}/docs/standards.md"
# (glob patterns are supported; the file's contents are loaded and treated as facts).
persistent_facts = []
# Scalar: executed at the end of Phase 5 (Close), after the retrospective
# document is saved and sprint-status is updated. Override wins.
# Leave empty for no custom post-completion behavior.
on_complete = ""
module-manifest.toml
module = "method"
version = "6.13.0-next"
update_source = "github:bmad-code-org/BMAD-METHOD/skills"
knowledge = "`references/help.md` in the `bmad` skill"
references/acceptance-verdict.md
# Decide: Routing and the Acceptance Verdict
Phase 4. Turn the consolidated findings into two outputs: routed action items the human can act on, and an honest verdict on whether the epic met its acceptance criteria. This skill proposes; it does not auto-apply fixes or edit the project spec. The human decides what executes.
## Route each finding
Give every finding two independent dispositions:
- **What to do about this instance** — *fix now*, *defer*, or *accept as-is*. Fix-now findings become action items. Deferred findings carry enough context to be acted on later without re-investigation. Accepted deviations are recorded so later retros stop re-flagging them.
- **What would prevent the next one** — the upstream lesson: spec wording, story sizing, a missing convention or gate, or nothing. This is where a recurring finding becomes a process change rather than a one-off fix.
Findings from sub-agents or the team discussion are unverified reports, not established facts. Before an action item relies on one, re-check it against the primary source — reopen the file, the commit, the spec. A finding whose source does not hold up is dropped, not routed.
## Action items
Compile fix-now findings and process lessons into specific, owned action items. Each names what to change and who owns it. Two kinds are *proposed, not applied* in this version:
- **Remediation** — code fixes are written up as action items (or story-shaped work) for the normal dev loop to execute later. The retrospective does not run the dev loop itself.
- **Spec reconciliation** — where the as-built diverges from the spec, propose the reconciliation as an action item with the evidence attached. The human applies it to the project contract; an uncertain interpretation is never written into the spec automatically.
## Previous-retro follow-through
When a prior retro exists, check whether the action items it committed to were completed. Read `action_items` in `{implementation_artifacts}/sprint-status.yaml` and, for every entry belonging to an earlier epic that is not already `done`, record one line in the retrospective document's Previous-retro follow-through section:
- **How to address the item** — its `id`, exactly as the file spells it. Legacy entries written before ids existed have none; for those, record the item's `epic` (the integer in the file) plus its exact `action` text, character for character. One or the other is what Phase 5 needs to name the item at all.
- **Whether it landed** — with the source that shows it: the commit, the file and line, the test. An item you cannot point at is "no evidence found", not "not done" — the reader must be able to tell a checked item from an unchecked one.
- **The status it argues for** — `done` for a landed item, `in-progress` for one demonstrably underway, or nothing. A proposal, never a write.
That record is exactly what Phase 5's `--set-action-status` offer reads: the selector becomes the JSON, the evidence is what the user is asked to confirm, and the proposed status is written only if they confirm it. A run with no prior retro, or one whose `sprint-status.yaml` is unreadable or carries no `action_items`, records that there was nothing to follow through on — and which of those it was, so a missing file is never mistaken for "no outstanding items."
## The verdict
Judge the final state against the epic's declared acceptance criteria. If the epic declared none, profile the criteria from the diff and stories and mark the verdict as **profiled** rather than declared. Weigh verification results (the Phase 2 behavior check) and unresolved findings. Render one of:
- **Accepted** — criteria demonstrably met in the evidence, no blocking findings open, and **no unfinished stories** for this epic.
- **Accepted-with-open-items** — criteria met, but named findings remain deferred and tracked — still only when every story of this epic is `done`.
- **Rejected** — criteria not met, a blocking finding stands unresolved, **or any of this epic's stories is still not `done`**.
### Unfinished stories
`pending_stories` is authoritative for this epic's incomplete work, whichever mode produced it: sprint-status story keys in file order from `detect-epic`, or `stories.yaml` ids in list order whose artifact status is not `done`. When that list is non-empty:
- The **machine** verdict is **rejected**. Name every unfinished story key in the Acceptance verdict section as the evidence. Do not soften this to accepted-with-open-items: unfinished delivery is not an open finding about a finished epic — the epic itself is incomplete.
- Record the unfinished keys in Epic summary (interactive) or Assumptions (headless) as the Inputs section already requires.
- Headless runs have no human at the console: the document's verdict is **rejected** when `pending_stories` was non-empty. Interactive runs may still let a human override (rule 1 below) after seeing the list.
If the completeness check did not run (no readable `sprint-status.yaml`), do **not** render a rejected or accepted verdict from the absence of data — say the check was unavailable and weigh only the criteria and findings you have.
Three hard rules:
1. A human decision always overrides the machine verdict.
2. An epic that fails its criteria with **no** human decision is recorded as **not accepted** — never as silently accepted.
3. A non-empty `pending_stories` list makes the machine verdict **rejected**, including in headless mode.
The verdict and its evidence carry into the Phase 5 document.
references/aggregate-views.md
# Aggregate Views
Phase 2. An epic is many coding sessions, each validated in isolation; the defects that matter are the ones no single session — and no single diff hunk — could see. Nine sessions each added three hundred lines and none ever saw the 3,000-line class they collectively built. These views are properties of the *whole* change, derived across the full diff range from Phase 1.
Prefer deterministic derivation: a script that measures the codebase is evidence; a model's impression is not. Where you compute a view inline instead of by script, record the narrowed scope. Every observation that becomes a finding carries a source reference — the file, the symbol, the commits. `references/evidence-gathering.md` is authoritative for what every `git_evidence.py` key means, including the commit-level `is_merge` and `stories` (every story id a subject names, so a commit spanning two counts for both) — read it there before deriving anything from the numbers.
## The catalog
- **Architecture delta** — how the dependency structure changed across the epic. Where a language-native dependency tool exists (dependency-cruiser, madge, pydeps, and the like), run it before and after the range and diff the graphs; otherwise derive the module/import graph from the changed files. Look for new cross-cutting dependencies, layering violations, and cycles introduced — structure the code's own conventions would forbid but no single story tripped.
- **Duplication map** — the same problem solved more than one way across stories. Two sessions independently writing near-identical logic, or a helper reimplemented because the second session did not know the first existed.
- **God-class / size growth** — files that grew past a healthy size *over the epic*, invisible per-commit because each session added only a little. The `git_evidence.py` pre-pass (Phase 1) reports `added` / `deleted` / `net` per path in `files` — *change volume*, not a file's absolute size or a per-commit growth rate. Those sums cover the range's **non-merge** commits only, and they are always integers: an unmeasurable revision is left out of them rather than nulling them. Rank on `files`, then open the top of the ranking and read each file's real current size and structure before calling anything a god-class — high net churn makes a file a candidate to inspect, not a verdict on its own. Three qualifiers say how far the ranking can be trusted: `binary_revisions` counts that path's revisions whose churn could not be measured, so its true volume is *at least* what the sums report; `merges_measured` short of `merge_count` means some merges were never measured at all, which caps how complete the ranking can be; and `merge_files` mostly restates churn `files` already counted, so summing the two double counts — but it is not redundant, because a merge's first-parent diff also carries whatever the conflict resolution itself added, code that lives in no non-merge commit and therefore appears in `files` nowhere. So read `merge_files` separately, for the paths whose churn shows up only there, rather than discarding it as double counting. Whether a flagged file is genuinely a god-class or legitimately large stays your judgment.
- **Pattern divergence** — where the epic's code diverges from the conventions the surrounding codebase already established: naming, error handling, test structure, module boundaries. Agents learn conventions by pattern-matching the code, so divergence compounds.
- **Spec-to-implementation reconciliation** — where the as-built diverges from what the epic spec and PRD/architecture described. Requirements silently dropped, added behavior nobody specified, intent reinterpreted between stories. Each divergence is either a defect (fix), an accepted deviation (record so later runs stop re-flagging it), or a spec that should be reconciled to reality (propose in Phase 4).
## Delegation
When sub-agents are available, delegate the derivation: each returns evidence with source refs and checked scope, never a verdict — the parent consolidates and decides. Give each a narrow view and an explicit return format. When sub-agents are unavailable, compute the highest-value views inline (architecture delta and spec reconciliation first) and record which views were narrowed or skipped.
references/evidence-gathering.md
# Evidence Gathering
Phase 1 of the retrospective. Enumerate what the completed epic produced, so every later analysis works from real artifacts instead of memory. Output is an inventory: what exists, what is missing, and the diff range the rest of the retro will read.
## Inventory checklist
Collect what the epic produced and note the source path or range of each:
- **Epic spec** — the epic file under `{planning_artifacts}`, including any declared acceptance criteria. If the spec declares how the epic will be judged, that governs Phase 4; if not, note that the verdict will be profiled from the diff.
- **Story files** — the story specs implemented under this epic (`{implementation_artifacts}`), each carrying its intent and context. These mark the boundaries between coding sessions.
- **Diff range and commits** — the full set of changes the epic introduced. Establish the range from the first and last story commits (or ask the user for it). The range must *include* the first story commit: `A..B` excludes `A`, so use the parent of the first commit as the left endpoint — `<first-commit>^..<last-commit>` — or the whole first story disappears from the diff, the commit attribution, and the verdict evidence. Then run `uv run --no-cache {skill-root}/scripts/git_evidence.py --repo {project-root} --range <range> --stories <story-ids>` to get, as JSON, the per-story commit attribution and the per-file change volume — added / deleted / net across the range — that Phase 2 reads. Record the range explicitly; Phase 2's aggregate views and the `bmad-review` pass both read it. When the range cannot be established, say so and narrow the scope rather than guessing. Read the output keys precisely: each commit carries `is_merge` and `stories` — *every* id its subject names, so a commit spanning two stories counts for both. `files` sums non-merge commits only. `merge_files` is each measured merge's diff against its first parent, so it *restates* the churn that merge brought in plus whatever the conflict resolution added — never add it into `files`, and never read it as merge-introduced work on its own. `merges_measured` counts the merges on the range head's first-parent spine; `merge_count` counts every merge in the range, so a gap between the two means merges went unmeasured. `binary_revisions` is unmeasured churn, not zero churn.
- **Sprint status** — `{implementation_artifacts}/sprint-status.yaml`, for which stories are `done` and the current retro-key state.
- **Previous retrospective** — the prior epic's retro doc, if one exists, so Phase 4 can check whether last epic's action items landed.
- **Session logs** — conversation or session records for the epic's stories, when available. They are the only record of *why* a session took an unexpected turn — what was tried and abandoned. They are also the evidence most likely to be deleted or expire, so capture references now.
## Stories mode
A stories-mode epic is a spec folder. Map it onto the checklist above: `SPEC.md` is the epic spec; `stories.yaml` in list order is the story list, each entry's artifact being the single `stories/<id>-*.md` it names; there is no sprint status; the previous retrospective, when resuming, is `{spec-folder}/RETROSPECTIVE.md`; session logs are unchanged.
The diff range differs. Each story records its own baseline in its artifact frontmatter — `baseline_revision` (deprecated) or `baseline_commit` — so there is no single epic-wide range. The range end is the next story's baseline in list order, which is exact because neither skill adds a commit of its own after the work. For the last story, when nothing records the end, derive it from the history — usually `HEAD`, though not always — and mark it inferred rather than recorded. A baseline that is absent or is not a revision leaves that story with no commit or diff evidence — record that too. Group the stories sharing an identical range and run `git_evidence.py` once per distinct range, passing that group's ids as one comma-separated `--stories` value. No `^` is needed here: unlike the sprint-mode range above, the recorded baseline is already the pre-change commit. Ranges may overlap or diverge; count a shared commit or file change once in the aggregate views while keeping each story's range as its provenance.
## Missing evidence
Evidence availability varies; never hide a gap. Each later analysis declares what it needs and, when that input is absent, records a narrowed scope rather than guessing. A reader of the final retro must always be able to tell **"checked and clean"** from **"never checked."**
- Missing session logs → process-lesson analysis is skipped, and the retro says so.
- No declared acceptance criteria → the verdict is profiled from the diff and stories, flagged as profiled rather than declared.
- Sub-agents unavailable → analyses that would delegate run inline over a narrowed scope, and the narrowing is recorded.
Carry the inventory forward into Phase 2 as the authoritative list of what is available to read.
references/retro-document.md
# Finalize: Retrospective Document and Sprint Status
Phase 5. Finalize the retrospective and update sprint tracking. Two writes: the retrospective document, and the `sprint-status.yaml` update. Stories mode makes only the first.
## The retrospective document
This document is the run's working artifact: it is created as a skeleton once the epic is fixed and filled as each phase completes, so Phase 5 finalizes rather than writes it from scratch. It lives at `{implementation_artifacts}/epic-{{epic_number}}-retro-{date}.md`, as readable markdown; ensure `{implementation_artifacts}` exists. In stories mode it lives at `{spec-folder}/RETROSPECTIVE.md` instead — a fixed name, so a resumed run finds it — and carries the same frontmatter without `epic`, which the folder already names.
Open the document with YAML frontmatter a machine can read without parsing the prose — an epic gate or orchestrator keys off `verdict` to decide whether to hold the next epic:
```
---
epic: {{epic_number}}
date: {date}
verdict: accepted | accepted-with-open-items | rejected
criteria: declared | profiled
headless: true | false
---
```
Keep `verdict` in sync with the Acceptance verdict section below. Do not encode the verdict in the sprint-status retro key — that key's value stays `done` so the existing lifecycle consumers (sprint planning's `optional ↔ done` transition, status TUIs) keep working unchanged.
That holds for a **rejected** epic too: the update below marks the retro key `done` whichever way the verdict went, because `done` there means *the retrospective ran*, not *the epic passed*. The script writes no verdict of any kind into `sprint-status.yaml` — there is no `retro_verdict` key and `--verdict` is only echoed back in the result JSON — so a gate or orchestrator that acts on the verdict **must** read this document's frontmatter. Reading sprint-status alone cannot tell a rejected epic from an accepted one.
Sections:
- **Epic summary** — which epic, the diff range, stories completed, any stories still unfinished (`pending_stories`) that the user accepted retro-ing over, the evidence inventory (what was available, what was missing). Unfinished stories force the machine acceptance verdict to **rejected** (see `references/acceptance-verdict.md`).
- **Findings** — grouped by aggregate view and by lens, each with its source reference and disposition (fix now / defer / accept). This is the record; do not summarize away the provenance.
- **Behavior verification** — what was exercised end to end and what was observed, or an explicit note that runtime behavior was not exercised.
- **Previous-retro follow-through** — if a prior retro exists, whether its action items landed, with evidence, and the selector Phase 5 would need to act on each (`references/acceptance-verdict.md` specifies what to record).
- **Action items** — the routed fix-now items and process lessons, each with an owner. Note which are proposed remediation or spec reconciliations awaiting human application.
- **Acceptance verdict** — accepted / accepted-with-open-items / rejected, whether the criteria were declared or profiled, and the evidence behind the call.
- **Open questions** — what a human answer would materially change, and anything the analyses could not resolve.
- **Assumptions** — in headless runs, every choice made without the user: which epic was selected (invocation or auto-detect), the `detect-epic --epic <N>` (or unflagged) result including any non-empty `pending_stories`, a machine **rejected** verdict forced by unfinished stories or rendered with no human decision, each proposed item. Omit in interactive runs — an interactive run records the same facts where the user confirmed them, in Epic summary.
Do not state time estimates anywhere in the document.
## Sprint-status update
Do not hand-edit `sprint-status.yaml` — its comment blocks and quoting are exactly the write that most often corrupts the file. Use the bundled script, which round-trips through a comment-preserving YAML parser, force-quotes values so punctuation (a leading `#`, a colon) cannot break parsing, and validates the result — restoring the original file untouched if the write does not verify:
```
uv run --no-cache {skill-root}/scripts/sprint_status.py update \
--file "{implementation_artifacts}/sprint-status.yaml" \
--epic {{epic_number}} --set-retro-done \
--add-action '[{"action":"...","owner":"..."}, ...]' \
--ref "{implementation_artifacts}/epic-{{epic_number}}-retro-{date}.md" \
--verdict "<accepted | accepted-with-open-items | rejected>" \
--date "{date}"
```
Keep every value quoted. `--date` is parsed as `MM-DD-YYYY HH:MM` and nothing else — unpadded spellings like `1-2-2026 9:05` are accepted and normalized to the padded form, but a value that does not parse is rejected with `ok: false`, `restored: true` and exit 1, before the file is touched, and the whole update is a no-op. So pass `{date}` only if it is already in that form; otherwise reformat it, or omit the flag entirely and let the script stamp the current time itself. That format carries a space, which is why the flag must be quoted: unquoted, `--date 07-28-2026 14:23` splits into two argv words and dies at argparse (`{"ok": false, "error": "argument error: unrecognized arguments: 14:23"}`, exit 2). `--file` and `--ref` are quoted for the same reason — an `{implementation_artifacts}` path containing a space breaks them exactly the same way.
It sets `development_status["epic-{{epic_number}}-retrospective"]` to `done`, appends one `action_items` entry per proposed item, and bumps `last_updated`. Each appended item carries `status: open`, a stable `id` (`epic-<N>-retro-item-<n>-<slug>` derived from the action text, or the `id` you supply in the JSON), and a `ref` back to this retro document (from `--ref`, or a per-item `ref` in the JSON) — so an orchestrator can dedupe items across re-runs and dispatch each one to its full, sourced finding. `--verdict` is not written into the file; it is echoed back in the result JSON as a signal for consumers. It accepts exactly the frontmatter vocabulary — `accepted`, `accepted-with-open-items`, `rejected` — and any other spelling is rejected (`ok: false`, `restored: true`, exit 1) before the file is touched. Read the JSON it returns:
- `ok: true` → report the retro-key transition, `action_items_added`, `action_items_updated`, and the echoed `verdict`.
- `ok: false` → the file was left untouched (`restored: true`); surface the error, do not hand-edit. `restored: false` means the rollback write also failed and the file may be incomplete — warn the user explicitly.
- `restored` speaks only for a command that may have written. Every `update` failure carries it; `detect-epic` never emits it, because it never writes; and an invocation the parser itself rejects (`argument error: ...`, exit 2) carries neither the key nor a file to speak about. Read a missing `restored` as "nothing was at risk", never as `false`.
- `retro_key_found: false` → the retro key was absent, so nothing was marked done; the document still saved, but tell the user sprint-status needs a manual retro entry.
- `retro_key_found: null` → `--set-retro-done` was not passed, so the key was never looked for. Distinct from `false`, which is a real absence the user needs to be told about.
Moving a *previous* epic's action items off `open` is recorded in the retro document either way. When the Phase 4 follow-through has evidence an item landed, or the user says one did, offer to update the sprint-status entries too and run `--set-action-status` with exactly what the user confirms — that flag is the only supported way to change a status; hand-editing never is. It can be passed in the same invocation as the update above, or run on its own:
```
uv run --no-cache {skill-root}/scripts/sprint_status.py update \
--file "{implementation_artifacts}/sprint-status.yaml" \
--epic {{epic_number}} \
--set-action-status '[{"id":"epic-1-retro-item-1-add-error-handling","status":"done"},{"epic":1,"action":"Exact action text","status":"in-progress"}]'
```
Rules:
- Select an item by its `id`, or — for legacy entries written before ids existed — by `epic` plus the item's exact `action` text. An entry carrying both uses the `id`. Matching is exact: no trimming, no case folding, and `epic` must be a JSON integer. The `--epic` flag does not scope selectors; it only names the retro key and the epic recorded on appended items, so items from any epic are addressable in one call.
- The only statuses are `open`, `in-progress`, and `done`. `bmad-sprint-planning`'s status view counts both `open` and `in-progress` as open action items, so only `done` retires an item from the surfaced list — moving something to `in-progress` records progress, it does not quiet the dashboard.
- Every selector must resolve to exactly one item already in the file. Matching nothing, matching more than one, or colliding with another entry in the same array aborts the whole invocation — `ok: false`, `restored: true`, the file byte-identical and nothing partially applied. "Whole invocation" includes any `--set-retro-done` and `--add-action` passed in the same call: one mistyped selector drops the entire update, so re-run the full command after fixing it rather than assuming the retro key was set.
- Items appended by `--add-action` in the same run are not addressable in that run; they are always written as `open`.
- Only ever apply a status the user confirmed, and in a headless run do not pass this flag at all.
- Success reports `action_items_updated`.
Only ever apply a status the user confirmed: the evidence justifies proposing a transition, and only the user's confirmation justifies writing it. In a headless run do not use this flag at all — record the transitions you would have proposed in the Previous-retro follow-through section and leave the prior items' statuses alone.
## Finish
Report where the document was saved, the verdict, and the action-item count. Then, if `{workflow.on_complete}` is non-empty, follow it as the final terminal instruction before exiting.
references/team-discussion.md
# Team Discussion (opt-in)
An optional discussion layer over Phase 2's findings, off by default. It exists for users who want the retrospective discussed from multiple perspectives, the way a team would. One rule: **the team discusses evidence, never invention.** Agents speak only to findings that carry source references. No agent may describe an event that did not happen or report a pattern the diff does not show.
## When to run it
Only when asked — "discuss it as a team," "run party mode," or similar. A default run never enters this phase.
## How to run it
Invoke **`bmad-party-mode`**, seeded with the consolidated Phase 2 findings and the epic context, so the installed agents react as real subagents with independent thinking rather than a scripted dialogue. Seed it with:
- The findings, each with its source reference, grouped by the aggregate view or lens that produced it.
- The improvements the evidence confirms — real gains, patterns that worked — so positive observations are grounded in fact.
- The epic's acceptance criteria (or the profiled stand-in), so the discussion can weigh the verdict.
- The previous epic's action items and whether they landed, when a prior retro exists, so accountability is grounded in fact.
If `bmad-party-mode` is unavailable, a discussion the user asked for must not silently fail to happen. Run it inline over the same seed — take each perspective yourself, hold every perspective to sourced findings — and record in the retrospective document that the discussion ran inline rather than through the installed agents. Record it as the narrowing it is: one model playing every role loses the independent disagreement that surfaces missed findings. State that in the document rather than omitting it.
Keep the user an active participant and steer toward systemic understanding over blame — the point is which process or convention would have prevented a finding, not who wrote the line. Capture anything the discussion surfaces that the analyses missed; a genuinely new observation becomes a finding only once you can tie it to a source, otherwise it is a question for Phase 4, not a conclusion.
The discussion does not replace Phase 4. Its output feeds the action items and the verdict; it does not render them.
scripts/git_evidence.py
# /// script
# requires-python = ">=3.11"
# ///
"""Measure git commit and file-change evidence over a revision range.
Prints ONLY JSON to stdout. Errors are emitted as JSON to stdout with a
non-zero exit code: 2 for invalid arguments (rejected before git runs),
1 for git or I/O failures. This script only MEASURES — it never judges
acceleration or violations. The model interprets the numbers.
Two git passes. The first lists every commit in the range (merges included)
and sums the per-file churn of the non-merge commits, which is what `files`
reports. The second runs only when the range contains merges and measures
those merges alone, reported separately as `merge_files` — never folded into
`files`, because a merge's diff against its first parent restates the churn
of the commits it merged in, which the first pass already counted.
"""
import argparse
import json
import os
import re
import subprocess
import sys
UNIT_SEP = "\x1f"
# sha, space-separated parents (empty for a root commit), subject.
LOG_FORMAT = f"--format=%H{UNIT_SEP}%P{UNIT_SEP}%s"
def _emit(obj, code=0):
sys.stdout.write(json.dumps(obj))
sys.exit(code)
class JsonArgumentParser(argparse.ArgumentParser):
"""Emit argparse failures on the JSON-only stdout contract, not usage text.
The parser is constructed with ``add_help=False``. The override below covers
``error()``, but ``-h`` never reaches it: the built-in help action calls
``print_help()`` and ``exit(0)`` directly, which would put plain usage text
on stdout with a zero exit and break the JSON-only contract. Removing the
action instead of intercepting it routes ``-h`` through the already-tested
``error()`` path as an ordinary unrecognized argument. The cost is that the
``help=`` strings are unreachable from the CLI; the skill's references carry
the usage a human needs.
"""
def error(self, message):
_emit({"ok": False, "error": f"argument error: {message}"}, 2)
def _parse_numstat_line(line):
# numstat lines: "<added>\t<deleted>\t<path>"; binary files use "-".
parts = line.split("\t")
if len(parts) < 3:
return None
added_raw, deleted_raw, path = parts[0], parts[1], "\t".join(parts[2:])
added = None if added_raw == "-" else int(added_raw)
deleted = None if deleted_raw == "-" else int(deleted_raw)
return added, deleted, path
def _git_log(repo, extra_args, rng):
"""Run one `git log --numstat` pass over `rng` and return its stdout.
`core.quotePath=false` keeps non-ASCII paths as real UTF-8 strings instead
of octal escapes, and `--no-renames` makes a rename an honest delete + add
instead of an unopenable "src/{a => b}" pseudo-path that splits one file's
churn across several keys. Both matter for every pass, so both live here.
`log.diffMerges=separate` is pinned on the command line because it is what
`-m` means: a user or repo config setting it to `off` makes pass 2 emit no
file rows at all, so `merge_files` would come back empty beside a non-zero
`merges_measured` and read as "the merges changed nothing".
"""
cmd = [
"git",
"-c",
"core.quotePath=false",
"-c",
"log.diffMerges=separate",
"-C",
repo,
"log",
"--numstat",
"--no-renames",
*extra_args,
LOG_FORMAT,
rng,
"--", # terminate rev parsing so the range can never match a pathspec
]
try:
# Decode explicitly: git emits UTF-8 path bytes regardless of the
# caller's locale, and a C locale would otherwise decode them as ASCII.
# surrogateescape, not replace: replace maps every invalid byte to the
# same U+FFFD, so two distinct non-UTF-8 paths would collapse into one
# `files` key with their churn silently summed. Lone surrogates survive
# json.dumps (escaped as \udcXX under ensure_ascii) and json.loads.
proc = subprocess.run(
cmd,
capture_output=True,
text=True,
encoding="utf-8",
errors="surrogateescape",
env={k: v for k, v in os.environ.items() if not k.startswith("GIT_")},
)
except Exception as exc: # noqa: BLE001
_emit({"ok": False, "error": str(exc)}, 1)
if proc.returncode != 0:
# stderr can be empty (a signal kill, a quiet failure); the exit code is
# then the only thing left to report, so never emit an empty error.
_emit(
{
"ok": False,
"error": proc.stderr.strip() or f"git exited {proc.returncode}",
},
1,
)
return proc.stdout
def _parse_log(output, stories):
"""Turn one pass's log output into (commits, files_map). Shared by both."""
commits = []
files = {} # path -> {path, _added, _deleted, binary_revisions, commit_count}
seen = set()
counting = True
for raw in output.splitlines():
if UNIT_SEP in raw:
sha, parents, subject = raw.split(UNIT_SEP, 2)
# Under -m, git repeats a merge's header once per parent unless it
# also honours --first-parent (git 2.31+). Count only the first
# block for a sha — git emits parents in order, so that block is
# the first-parent diff either way, and no churn is double counted.
counting = sha not in seen
if not counting:
continue
seen.add(sha)
commits.append(
{
"sha": sha,
"subject": subject,
# Every id the subject names, in --stories order: a commit
# spanning two stories belongs to both. Word-boundary match
# so a story id like "1-2" does not also match "11-2".
"stories": [sid for sid in stories if re.search(rf"\b{re.escape(sid)}\b", subject)],
"is_merge": len(parents.split()) > 1,
}
)
continue
if not counting or not raw.strip():
continue
parsed = _parse_numstat_line(raw)
if parsed is None:
continue
added, deleted, path = parsed
entry = files.get(path)
if entry is None:
# _added/_deleted are running sums over the path's text revisions.
entry = {
"path": path,
"_added": 0,
"_deleted": 0,
"binary_revisions": 0,
"commit_count": 0,
}
files[path] = entry
entry["commit_count"] += 1
if added is None or deleted is None:
# A binary revision is unmeasurable, not zero — count it alongside
# the sums instead of nulling the path's real measured churn.
entry["binary_revisions"] += 1
else:
entry["_added"] += added
entry["_deleted"] += deleted
return commits, files
def _file_list(files):
return [
{
"path": entry["path"],
"added": entry["_added"],
"deleted": entry["_deleted"],
"net": entry["_added"] - entry["_deleted"],
"commit_count": entry["commit_count"],
"binary_revisions": entry["binary_revisions"],
}
for entry in files.values()
]
def main(argv=None):
parser = JsonArgumentParser(
description=(
"Measure commit and per-file change evidence over a git revision range. Measures only; does not judge."
),
add_help=False,
)
parser.add_argument("--repo", default=".", help="Path to the git repo (default: .)")
parser.add_argument("--range", dest="range", help="Revision range REV..REV")
parser.add_argument(
"--stories",
help="Comma-separated story ids to match against commit subjects.",
)
args = parser.parse_args(argv)
stories = []
if args.stories:
# dict.fromkeys dedupes while keeping the caller's order: a repeated id
# would otherwise land twice in a commit's `stories`, double counting
# that commit in any per-story total built from the output.
stories = list(dict.fromkeys(s.strip() for s in args.stories.split(",") if s.strip()))
if not args.range:
_emit(
{
"range": None,
"note": "no range supplied",
"commits": [],
"files": [],
}
)
# Accept only an explicit REV..REV range. Anything else silently measures
# the wrong thing: a leading "-" is consumed by git as an option, a single
# rev logs all history up to it, a bare pathspec logs by path, an empty
# endpoint ("..", "a..", "..b") makes git default that side to HEAD, and a
# three-dot "A...B" is a symmetric difference — a different commit set
# entirely. partition splits at the FIRST "..", so any of those extra-dot
# shapes leaves `right` empty or dot-prefixed.
left, _, right = args.range.partition("..")
if args.range != args.range.strip() or args.range.startswith("-") or not left or not right or right.startswith("."):
_emit(
{
"ok": False,
"error": f"invalid --range {args.range!r}: expected a revision range like REV..REV",
},
2,
)
# Pass 1 — the listing. No extra args, so full topology: every commit in
# the range including merges, which is what per-story attribution reads.
# Merges contribute no numstat rows here, so `files` is non-merge churn.
commits, files = _parse_log(_git_log(args.repo, [], args.range), stories)
merge_count = sum(1 for commit in commits if commit["is_merge"])
# Pass 2 — merge churn, only when there is any. `-m --first-parent
# --min-parents=2` walks the range head's first-parent spine and emits
# exactly one diff-against-first-parent block per merge sitting on it.
# Merges off that spine are counted in merge_count and never measured,
# which is precisely why merges_measured is a separate key: the gap
# between the two is a visible statement that some merges went
# unmeasured. This never folds into `files` — a merge's first-parent diff
# restates the churn of the commits it merged in, which pass 1 already
# counted, so adding it in would double count.
merge_commits, merge_files = [], {}
if merge_count:
merge_commits, merge_files = _parse_log(
_git_log(
args.repo,
["-m", "--first-parent", "--min-parents=2"],
args.range,
),
stories,
)
_emit(
{
"range": args.range,
"commit_count": len(commits),
"merge_count": merge_count,
"merges_measured": len(merge_commits),
"commits": commits,
"files": _file_list(files),
"merge_files": _file_list(merge_files),
"stories_supplied": stories,
}
)
if __name__ == "__main__":
main()
scripts/sprint_status.py
# /// script
# requires-python = ">=3.11"
# dependencies = ["ruamel.yaml>=0.18"]
# ///
"""Detect the current retrospective epic and surgically update sprint-status.yaml.
Prints ONLY JSON to stdout. Errors are emitted as JSON to stdout with a non-zero
exit code. The ``update`` subcommand round-trips the YAML to preserve all comments
and formatting, writes atomically (temp file + ``os.replace``), and restores the
original file bytes on any validation failure.
"""
import argparse
import hashlib
import io
import json
import os
import re
import stat
import sys
import tempfile
from collections import Counter
from collections.abc import Mapping
from datetime import datetime
from ruamel.yaml import YAML
from ruamel.yaml.scalarstring import DoubleQuotedScalarString
STORY_RE = re.compile(r"^(\d+)-\d+[a-z]?-") # trailing [a-z]? matches split-story keys like 2-6a-...
DATE_FORMAT = "%m-%d-%Y %H:%M"
# The authoritative action-item vocabulary, mirrored from bmad-sprint-planning's
# SKILL.md. Anything outside it would render as unknown in the status dashboard.
ACTION_STATUSES = ("open", "in-progress", "done")
# The retro-document frontmatter vocabulary. --verdict is only echoed back, but
# orchestrators branch on the echo, so a free-spelled value ("accepted with open
# items") would silently fall through every branch they write.
VERDICTS = ("accepted", "accepted-with-open-items", "rejected")
def _load_yaml(path):
yaml = YAML(typ="rt")
yaml.preserve_quotes = True
# Pin the emitter to the indentation the sprint-status template ships with.
# Without this, ruamel re-dumps block sequences at its own default offset and
# every write silently de-indents pre-existing, untouched action_items.
yaml.indent(mapping=2, sequence=4, offset=2)
# Pin the dump encoding too: `_dump_bytes` serializes into a BytesIO, so the
# emitter -- not this module -- encodes the bytes that land in the user's
# file. utf-8 is ruamel's current default, but the file is read back as
# utf-8 unconditionally, so state it rather than inherit it.
yaml.encoding = "utf-8"
with open(path, encoding="utf-8") as fh:
data = yaml.load(fh)
return yaml, data
def _emit(obj, code=0):
sys.stdout.write(json.dumps(obj))
sys.exit(code)
def _emit_error(message, code=1, restored=None):
"""Emit a failure on the JSON-only contract.
``restored`` is included only when the caller can speak to the state of the
target file; ``retro-document.md`` teaches callers to read it, so a write-path
failure must never omit it and a read-only subcommand must never invent it.
"""
payload = {"ok": False, "error": message}
if restored is not None:
payload["restored"] = restored
_emit(payload, code)
class JsonArgumentParser(argparse.ArgumentParser):
"""Emit argparse failures on the JSON-only stdout contract, not usage text.
Every parser built from this class is constructed with ``add_help=False``.
The override below covers ``error()``, but ``-h`` never reaches it: the
built-in help action calls ``print_help()`` and ``exit(0)`` directly, which
would put plain usage text on stdout with a zero exit and break the
JSON-only contract for the machine consumer this script exists to serve.
Removing the action instead of intercepting it keeps the fix to one keyword
per parser and routes ``-h`` through the already-tested ``error()`` path as
an ordinary unrecognized argument. The cost is that the ``help=`` strings
are unreachable from the CLI; the skill's references carry the usage a
human needs.
"""
def error(self, message):
_emit({"ok": False, "error": f"argument error: {message}"}, 2)
def _slugify(text, maxlen=40):
text = str(text)
# Unicode-aware: a non-Latin action must keep its own characters in the id
# rather than collapsing to a single placeholder shared by every item.
slug = re.sub(r"[^\w]+", "-", text.lower(), flags=re.UNICODE).strip("-")
slug = slug[:maxlen].strip("-")
if not slug:
# Nothing sluggable (punctuation/emoji only): a short content hash keeps
# the id deterministic and distinct instead of a bare "item".
slug = hashlib.sha256(text.encode("utf-8")).hexdigest()[:8]
return slug
def _selector_label(entry):
"""Human-readable form of a --set-action-status selector, for error text.
An entry carrying both forms is described by its ``id``, because that is the
form resolution actually uses.
"""
if isinstance(entry.get("id"), str):
return f"id={entry['id']!r}"
return f"epic={entry.get('epic')!r} action={entry.get('action')!r}"
def _match_action_items(entry, items):
"""Indices in ``items`` that the selector ``entry`` resolves to.
``id`` wins whenever it is present: the epic/action pair is the fallback for
legacy items written before ids existed, so a caller that copied a whole item
through gets the precise match rather than a text comparison. Matching is
exact equality -- no normalization -- so a file that spells its epic as a
string simply does not match and the caller gets a "no match" error instead of
a silent write to the wrong item. ``bool`` is excluded on the file side too,
since ``True == 1`` in Python.
"""
item_id = entry.get("id")
if isinstance(item_id, str):
return [idx for idx, item in enumerate(items) if isinstance(item, Mapping) and item.get("id") == item_id]
epic_value = entry.get("epic")
action_value = entry.get("action")
return [
idx
for idx, item in enumerate(items)
if isinstance(item, Mapping)
and not isinstance(item.get("epic"), bool)
and item.get("epic") == epic_value
and item.get("action") == action_value
]
def _comment_counts(text):
"""Multiset of the comment lines in ``text``, indentation included.
Keyed by the whole line so that a re-indented comment counts as a loss too:
the guarantee callers are given is comments *and formatting*, and ruamel
re-emits comments at their original column even when the block around them
is re-indented, so an exact key costs nothing in practice.
"""
return Counter(line for line in text.splitlines() if line.lstrip().startswith("#"))
def _load_document(path, restored=None):
"""Load and shape-check the document, reporting every failure as JSON.
Returns ``(yaml, data, dev)``. ``dev`` is the live ``development_status``
mapping when the key exists, otherwise a detached empty mapping -- the key is
never inserted into the document as a side effect of loading.
"""
try:
yaml, data = _load_yaml(path)
except UnicodeDecodeError as exc:
_emit_error(f"{path} is not valid UTF-8: {exc}", 1, restored)
except OSError as exc:
_emit_error(str(exc), 1, restored)
except Exception as exc: # noqa: BLE001 - report any parse error as JSON
_emit_error(str(exc), 1, restored)
if data is not None and not isinstance(data, Mapping):
_emit_error("root document is not a mapping", 1, restored)
dev = data.get("development_status") if data is not None else None
if dev is None:
dev = {}
elif not isinstance(dev, Mapping):
_emit_error("development_status is not a mapping", 1, restored)
return yaml, data, dev
def _retro_status(dev, retro_key, restored=None):
status_value = dev.get(retro_key)
if status_value is not None and not isinstance(status_value, str):
_emit_error(
f"{retro_key} status must be a string or null",
1,
restored,
)
return status_value
def _dump_bytes(yaml, data):
"""Serialize the document to bytes before any file is touched, so a dump
failure cannot leave a partial file anywhere."""
buf = io.BytesIO()
yaml.dump(data, buf)
return buf.getvalue()
def _atomic_write(path, payload, mode=None):
"""Replace ``path``'s contents with ``payload`` atomically.
The bytes land in a temp file alongside the target, are fsynced, take the
target's permission bits (mkstemp creates 0600, which would silently narrow
the file), and only then rename over it -- so a kill or a full disk leaves
the original file intact rather than truncated. ``path`` is resolved through
symlinks first: renaming onto a symlink would detach the link and leave the
real file stale while reporting success. The directory is fsynced too --
best-effort, see below -- so the rename survives a power loss and not just
the bytes.
"""
path = os.path.realpath(path)
directory = os.path.dirname(path) or "."
fd, tmp_path = tempfile.mkstemp(prefix=".sprint-status-", suffix=".tmp", dir=directory)
try:
with os.fdopen(fd, "wb") as fh:
fh.write(payload)
fh.flush()
os.fsync(fh.fileno())
if mode is not None:
os.chmod(tmp_path, mode)
os.replace(tmp_path, path)
except BaseException:
try:
os.unlink(tmp_path)
except OSError:
pass
raise
# The directory sync sits outside the try because once os.replace has
# returned, the new bytes ARE the file: a failure past that point must not
# propagate as a write failure, or the caller would report the original
# "restored" about a write that in fact landed. Skipping it only risks the
# rename not surviving a hard power loss.
try:
dir_fd = os.open(directory, os.O_RDONLY)
try:
os.fsync(dir_fd)
finally:
os.close(dir_fd)
except OSError:
pass
def cmd_detect_epic(args):
# detect-epic never writes, so it reports no "restored" key.
_, _, dev = _load_document(args.file)
done_stories = []
max_epic = None
# Every story key with its epic, in document order, so the pending list can
# be scoped to whichever epic detection lands on without a second pass over
# the mapping. Non-story keys (epic-2, epic-2-retrospective, ...) never enter
# here, because STORY_RE does not match them.
story_keys = []
for key, value in dev.items():
m = STORY_RE.match(str(key))
if not m:
continue
epic_num = int(m.group(1))
story_keys.append((epic_num, key, value))
if value == "done":
done_stories.append(key)
if max_epic is None or epic_num > max_epic:
max_epic = epic_num
# Optional --epic aims the gate at a supplied number (the -H <epic> path)
# instead of auto-picking the highest epic with a done story. Without it,
# behavior is unchanged: detect, then scope pending_stories to that epic.
if args.epic is not None:
if args.epic < 1:
_emit_error(
f"invalid --epic {args.epic} (expected a positive integer)",
1,
)
selected = args.epic
else:
selected = max_epic
if selected is None:
# Uniform shape: pending_stories is always present, even with no epic to
# scope it to, so a caller can read it without branching on epic first.
_emit(
{
"epic": None,
"story_count": 0,
"done_stories": done_stories,
"pending_stories": [],
"retro_key": None,
"retro_status": None,
}
)
# Scoped to the selected epic only -- deliberately unlike done_stories, which
# spans the whole file. A pending story in some *other* epic is not this
# retrospective's business.
selected_keys = [(key, value) for epic_num, key, value in story_keys if epic_num == selected]
pending_stories = [key for key, value in selected_keys if value != "done"]
retro_key = f"epic-{selected}-retrospective"
retro_status = _retro_status(dev, retro_key)
_emit(
{
"epic": selected,
# An epic the file has never heard of returns the same empty
# pending_stories as a finished one; story_count is the key that
# separates "complete" from "nonexistent" (a typo'd --epic), so the
# unfinished-story gate can refuse to read silence as done.
"story_count": len(selected_keys),
"done_stories": done_stories,
"pending_stories": pending_stories,
"retro_key": retro_key,
"retro_status": retro_status,
}
)
def cmd_update(args):
# Every failure below happens before the write is attempted, so the file is
# untouched and "restored": true is the honest report.
untouched = True
# 0. Validate the inputs before anything is mutated or written.
if args.epic < 1:
_emit_error(f"invalid --epic {args.epic} (expected a positive integer)", 1, untouched)
if args.date is not None:
try:
parsed_date = datetime.strptime(args.date, DATE_FORMAT)
except (ValueError, TypeError):
_emit_error(
f'invalid --date {args.date!r} (expected "MM-DD-YYYY HH:MM")',
1,
untouched,
)
# Normalize: strptime also accepts unpadded spellings like
# "1-2-2026 9:05", and writing those through would defeat the point of
# validating the format at all.
last_updated = parsed_date.strftime(DATE_FORMAT)
else:
last_updated = datetime.now().strftime(DATE_FORMAT)
if args.verdict is not None and args.verdict not in VERDICTS:
_emit_error(
f"invalid --verdict {args.verdict!r} (allowed: {', '.join(VERDICTS)})",
1,
untouched,
)
actions = []
if args.add_action:
try:
actions = json.loads(args.add_action)
except json.JSONDecodeError as exc:
_emit_error(f"invalid --add-action JSON: {exc}", 1, untouched)
if not isinstance(actions, list):
_emit_error("--add-action must be a JSON array", 1, untouched)
for item in actions:
if not isinstance(item, dict):
_emit_error("each --add-action item must be an object", 1, untouched)
action_value = item.get("action")
if not isinstance(action_value, str) or not action_value.strip():
# A JSON null/number/object would otherwise be str()'d into a
# literal "None"/"{...}" and written as a real action item.
_emit_error(
"each --add-action item must have a non-empty string action",
1,
untouched,
)
status_updates = []
if args.set_action_status:
try:
status_updates = json.loads(args.set_action_status)
except json.JSONDecodeError as exc:
_emit_error(f"invalid --set-action-status JSON: {exc}", 1, untouched)
if not isinstance(status_updates, list):
_emit_error("--set-action-status must be a JSON array", 1, untouched)
for entry in status_updates:
if not isinstance(entry, dict):
_emit_error("each --set-action-status entry must be an object", 1, untouched)
status_value = entry.get("status")
if not isinstance(status_value, str) or status_value not in ACTION_STATUSES:
_emit_error(
f"invalid --set-action-status status {status_value!r} (allowed: {', '.join(ACTION_STATUSES)})",
1,
untouched,
)
item_id = entry.get("id")
if item_id is not None:
# Present but unusable is an input error, not a silent fallback to
# the epic/action form -- the caller meant to select by id.
if not isinstance(item_id, str) or not item_id.strip():
_emit_error(
"each --set-action-status id must be a non-empty string",
1,
untouched,
)
continue
epic_value = entry.get("epic")
action_value = entry.get("action")
if isinstance(epic_value, bool) or not isinstance(epic_value, int):
_emit_error(
"each --set-action-status entry must have a non-empty string id, "
"or an integer epic and a non-empty string action",
1,
untouched,
)
if not isinstance(action_value, str) or not action_value.strip():
_emit_error(
"each --set-action-status entry must have a non-empty string id, "
"or an integer epic and a non-empty string action",
1,
untouched,
)
# 1. Keep original bytes for restore-on-failure, and the mode to write back
# with -- taken from the open handle so an unlink mid-run cannot leave the
# replacement silently narrowed to mkstemp's 0600.
try:
with open(args.file, "rb") as fh:
original_bytes = fh.read()
original_mode = stat.S_IMODE(os.fstat(fh.fileno()).st_mode)
except OSError as exc:
_emit_error(str(exc), 1, untouched)
try:
original_text = original_bytes.decode("utf-8")
except UnicodeDecodeError as exc:
_emit_error(f"{args.file} is not valid UTF-8: {exc}", 1, untouched)
# Every comment line in the file, not just the leading block: the template
# ships one above action_items, and losing it corrupts the document just the
# same as losing the header.
original_comments = _comment_counts(original_text)
yaml, data, dev = _load_document(args.file, restored=untouched)
if data is None:
_emit_error("empty or invalid YAML document", 1, untouched)
epic = args.epic
retro_key = f"epic-{epic}-retrospective"
# null distinguishes "the flag was not passed" from "the key was absent",
# which is the only case retro-document.md assigns "false" to.
retro_key_found = None
retro_status_before = None
retro_status_after = None
# 2. Optionally set the retrospective status to done (only if key exists).
if args.set_retro_done:
retro_key_found = retro_key in dev
if retro_key_found:
retro_status_before = _retro_status(dev, retro_key, restored=untouched)
dev[retro_key] = "done"
retro_status_after = "done"
# 3. Take the action_items sequence as loaded and shape-check it once; both
# of the steps below operate on this same list.
existing_actions = data.get("action_items")
if existing_actions is not None and not isinstance(existing_actions, list):
# A hand-corrupted file must still fail on the JSON contract, not crash.
_emit_error("action_items in file is not a list", 1, untouched)
items_added = 0
original_action_len = len(existing_actions) if existing_actions is not None else 0
# 4. Optionally transition the status of items already in the file. Selectors
# resolve against action_items *as loaded* and strictly before the
# --add-action append below, which is what makes an item appended in the
# same invocation unaddressable in that run. Every selector is resolved
# before any is applied, so a rejected batch never leaves a partial edit --
# and since nothing has been written yet, the file is still untouched.
status_targets = []
if status_updates:
pool = existing_actions if isinstance(existing_actions, list) else []
claimed = {}
for entry in status_updates:
label = _selector_label(entry)
matches = _match_action_items(entry, pool)
if not matches:
_emit_error(f"no action item matches {label}", 1, untouched)
if len(matches) > 1:
_emit_error(
f"ambiguous --set-action-status selector {label}: {len(matches)} matches",
1,
untouched,
)
idx = matches[0]
if idx in claimed:
# Applying both would overcount action_items_updated, and a
# conflicting pair would surface as a confusing post-write
# validation failure instead of the input error it is.
_emit_error(
f"duplicate --set-action-status targets: {label} and "
f"{claimed[idx]} resolve to the same action item",
1,
untouched,
)
claimed[idx] = label
status_targets.append((idx, entry["status"]))
for idx, new_status in status_targets:
# A plain assignment keeps the item's own scalar style: ruamel's
# CommentedMap re-applies the existing key's style on overwrite, for
# every ScalarString subclass. Pinned by the style tests.
pool[idx]["status"] = new_status
# 5. Optionally append action items.
if actions:
seq = data.get("action_items")
if seq is None:
seq = []
data["action_items"] = seq
for item in actions:
# Stable identity for orchestrator consumers: an id that lets a
# re-run dedupe against prior items, and a ref back to the sourced
# finding in the retro document. Both accept an explicit override.
seq_num = len(seq) + 1
action_text = str(item.get("action", ""))
item_id = item.get("id") or (f"epic-{int(epic)}-retro-item-{seq_num}-{_slugify(action_text)}")
ref = item.get("ref") or (args.ref or "")
entry = {
"id": DoubleQuotedScalarString(str(item_id)),
"epic": int(epic),
"action": DoubleQuotedScalarString(action_text),
"owner": DoubleQuotedScalarString(str(item.get("owner", ""))),
"status": "open",
"ref": DoubleQuotedScalarString(str(ref)),
}
seq.append(entry)
items_added += 1
# 6. Update last_updated.
data["last_updated"] = last_updated
# 7. Serialize, then swap the file atomically.
try:
_atomic_write(args.file, _dump_bytes(yaml, data), original_mode)
except Exception as exc: # noqa: BLE001
# The target is only ever touched by the final rename, so if the write
# raised, the original is still on disk byte-for-byte. Calling _restore
# here would rewrite a file that was never modified -- the one write in
# the program with nothing to gain and a truncated file to lose.
_emit({"ok": False, "error": f"write failed: {exc}", "restored": True}, 1)
# 8. Validate the written file; restore on any failure.
def _fail(msg):
restored = _restore(args.file, original_bytes, original_mode)
_emit({"ok": False, "error": msg, "restored": restored}, 1)
try:
_, reloaded = _load_yaml(args.file)
except Exception as exc: # noqa: BLE001
_fail(f"re-parse failed after write: {exc}")
if reloaded is None:
_fail("re-parse produced empty document after write")
if not isinstance(reloaded, Mapping):
_fail("re-parse produced a non-mapping document after write")
rdev = reloaded.get("development_status") or {}
if args.set_retro_done and retro_key_found:
if not isinstance(rdev, Mapping) or rdev.get(retro_key) != "done":
_fail(f"validation: {retro_key} not set to done after write")
new_action_len = 0
if reloaded.get("action_items") is not None:
new_action_len = len(reloaded.get("action_items"))
if new_action_len != original_action_len + items_added:
_fail(
"validation: action_items length mismatch "
f"(expected {original_action_len + items_added}, got {new_action_len})"
)
if status_targets:
# The recorded indices are still valid: the only other mutation to the
# sequence is an append, and the length check above just confirmed it.
reloaded_actions = reloaded.get("action_items")
if not isinstance(reloaded_actions, list):
_fail("validation: action_items is not a list after write")
for idx, new_status in status_targets:
reloaded_item = reloaded_actions[idx]
if not isinstance(reloaded_item, Mapping) or reloaded_item.get("status") != new_status:
_fail(f"validation: action item at index {idx} is not {new_status!r} after write")
try:
with open(args.file, encoding="utf-8") as fh:
new_text = fh.read()
except (OSError, UnicodeDecodeError) as exc:
_fail(f"re-read failed after write: {exc}")
# Loss-only: a comment may legitimately move or be added (a long quoted value
# can wrap onto a line that begins with '#'), but none may disappear.
lost = original_comments - _comment_counts(new_text)
if lost:
first = next((line for line in original_text.splitlines() if line in lost), None)
_fail(f"validation: comment line lost after write: {first!r}")
_emit(
{
"ok": True,
"retro_key_found": retro_key_found,
"retro_status_before": retro_status_before,
"retro_status_after": retro_status_after,
"action_items_added": items_added,
"action_items_updated": len(status_targets),
"last_updated": last_updated,
"verdict": args.verdict,
}
)
def _restore(path, original_bytes, mode=None):
"""Best-effort restore of the original bytes. Returns True on success so a
caller can surface a restore failure instead of hiding a half-written file.
Atomic for the same reason the primary write is: a truncating rewrite that
dies halfway destroys the very bytes it was trying to put back.
"""
try:
_atomic_write(path, original_bytes, mode)
return True
except Exception as exc: # noqa: BLE001 - best-effort restore
sys.stderr.write(f"restore failed: {exc}\n")
return False
def build_parser():
parser = JsonArgumentParser(
description=(
"Detect the current retrospective epic and surgically update "
"sprint-status.yaml while preserving comments and formatting."
),
add_help=False,
)
sub = parser.add_subparsers(dest="command", required=True)
p_detect = sub.add_parser(
"detect-epic",
help=(
"Find the highest epic with a done story and its retrospective "
"status, or aim the same pending_stories gate at --epic N."
),
add_help=False,
)
p_detect.add_argument("--file", required=True, help="Path to sprint-status.yaml")
p_detect.add_argument(
"--epic",
type=int,
default=None,
help=(
"Optional. Scope the response to this epic number instead of "
"auto-detecting the highest epic with a done story. Orchestrators "
"passing -H <epic> should pass the same number here so pending_stories "
"covers the epic they are about to retro."
),
)
p_detect.set_defaults(func=cmd_detect_epic)
p_update = sub.add_parser(
"update",
help="Surgically update retro status and/or action items.",
add_help=False,
)
p_update.add_argument("--file", required=True, help="Path to sprint-status.yaml")
p_update.add_argument("--epic", required=True, type=int, help="Epic number")
p_update.add_argument(
"--set-retro-done",
action="store_true",
help="Set epic-<N>-retrospective to done if the key exists.",
)
p_update.add_argument(
"--add-action",
help='JSON array of {"action":str,"owner":str,"id"?:str,"ref"?:str} to append.',
)
p_update.add_argument(
"--set-action-status",
help=(
"JSON array of status transitions for action items already in the file. "
'Select each by id -- {"id":str,"status":"open|in-progress|done"} -- or, '
"for legacy items with no id, by epic plus exact action text: "
'{"epic":int,"action":str,"status":...}. An entry carrying both uses the '
"id. Every selector must match exactly one item; any failure aborts the "
"whole invocation and leaves the file untouched."
),
)
p_update.add_argument(
"--ref",
help="Reference (e.g. the retro document path) recorded on each appended action item.",
)
p_update.add_argument(
"--verdict",
help=(
"Acceptance verdict echoed back in the JSON result for orchestrator "
f"consumers. One of: {', '.join(VERDICTS)}."
),
)
p_update.add_argument(
"--date",
help='Value for last_updated (default: now as "MM-DD-YYYY HH:MM").',
)
p_update.set_defaults(func=cmd_update)
return parser
def main(argv=None):
parser = build_parser()
args = parser.parse_args(argv)
args.func(args)
if __name__ == "__main__":
main()
scripts/tests/fixtures/sprint-status-template.yaml
# Sprint Status Template
# This is an EXAMPLE showing the expected format
# The actual file will be generated with all epics/stories from your epic files
# generated: {date}
# project: {project_name}
# project_key: {project_key}
# tracking_system: {tracking_system}
# story_location: {story_location}
# STATUS DEFINITIONS:
# ==================
# Epic Status:
# - backlog: Epic not yet started
# - in-progress: Epic actively being worked on
# - done: All stories in epic completed
#
# Story Status:
# - backlog: Story only exists in epic file
# - ready-for-dev: Story file created, ready for development
# - in-progress: Developer actively working on implementation
# - review: Implementation complete, ready for review
# - done: Story completed
#
# Retrospective Status:
# - optional: Can be completed but not required
# - done: Retrospective has been completed
#
# Action Item Status:
# - open: Committed during a retrospective, not yet addressed
# - in-progress: Actively being worked on
# - done: Completed
#
# WORKFLOW NOTES:
# ===============
# - Epic transitions to 'in-progress' automatically when its first story starts (via build's sprint sync)
# - Stories can be worked in parallel if team capacity allows
# - Developer typically creates the next story after the previous one is 'done' to incorporate learnings
# - Dev moves story to 'review', then runs code-review (fresh context, different LLM recommended)
# - Retrospective appends its action items to action_items; the status view surfaces open ones
# EXAMPLE STRUCTURE (your actual epics/stories will replace these):
# Timestamps use MM-DD-YYYY HH:MM.
generated: 05-06-2025 21:30
last_updated: 05-06-2025 21:30
project: My Awesome Project
project_key: NOKEY
tracking_system: file-system
story_location: "docs/stories"
development_status:
epic-1: backlog
1-1-user-authentication: done
1-2-account-management: ready-for-dev
1-3-plant-data-model: backlog
1-4-add-plant-manual: backlog
epic-1-retrospective: optional
epic-2: backlog
2-1-personality-system: backlog
2-2-chat-interface: backlog
2-3-llm-integration: backlog
epic-2-retrospective: optional
# Action items committed during retrospectives (section created by the retrospective workflow)
action_items:
- epic: 1
action: "Add error-handling review to the code review checklist"
owner: "Charlie"
status: open
scripts/tests/test_git_evidence.py
# /// script
# requires-python = ">=3.11"
# dependencies = ["pytest>=8.0"]
# ///
"""Tests for git_evidence.py — measurement over a real temp git repo.
Run: uv run scripts/tests/test_git_evidence.py
or: uv run --with pytest -m pytest scripts/tests/test_git_evidence.py
"""
import json
import os
import shutil
import subprocess
import sys
import unicodedata
from pathlib import Path
import pytest
SCRIPT = Path(__file__).resolve().parents[1] / "git_evidence.py"
# The fixture commit identity, shared by both git helpers below.
_IDENT = {
"GIT_AUTHOR_NAME": "T",
"GIT_AUTHOR_EMAIL": "t@t",
"GIT_COMMITTER_NAME": "T",
"GIT_COMMITTER_EMAIL": "t@t",
}
def _git_env(repo):
"""The environment every fixture git runs under, layered outward.
Inherit the real environment (PATH above all: git lives in /opt/homebrew,
/usr/local, or a nix store as readily as /usr/bin, and an env holding only
GIT_* vars sends execvp to os.defpath), then strip every ambient GIT_* var
-- GIT_DIR, GIT_WORK_TREE and GIT_CONFIG_COUNT would each silently redirect
or reconfigure the fixture -- and pin identity plus every source git reads
for settings, so nothing on the developer's machine can reach the fixture:
- gitconfig (commit.gpgsign, core.autocrlf, core.hooksPath,
init.defaultBranch). GIT_CONFIG_NOSYSTEM/GIT_CONFIG_GLOBAL cover
git >= 2.32; HOME and XDG_CONFIG_HOME cover older git, and are set to the
repo's parent directory -- always a per-test directory under pytest's
tmp_path -- so nothing is ever planted inside the working tree.
- gitattributes, a separate source GIT_CONFIG_NOSYSTEM does not cover: a
system `* -diff` rule would make numstat call every path binary and take
the churn assertions down with it. GIT_ATTR_NOSYSTEM shuts it out.
- the locale. LC_ALL/LANG are pinned to C, matching _run/_proc's existing
pin, so fixture git's text output cannot vary with the developer's
locale. Inheriting the environment is what makes this pin necessary:
the old four-variable env had no locale in it to inherit.
"""
env = {k: v for k, v in os.environ.items() if not k.startswith("GIT_")}
env.update(_IDENT)
env["GIT_CONFIG_NOSYSTEM"] = "1"
env["GIT_ATTR_NOSYSTEM"] = "1"
env["GIT_CONFIG_GLOBAL"] = os.devnull
env["HOME"] = env["XDG_CONFIG_HOME"] = str(Path(repo).parent)
env["LC_ALL"] = env["LANG"] = "C"
return env
def _json(proc):
"""Parse the JSON-only stdout contract, surfacing a crash instead of hiding
it behind a JSONDecodeError."""
assert proc.stdout, f"empty stdout; stderr was: {proc.stderr}"
assert "Traceback" not in proc.stderr, proc.stderr
return json.loads(proc.stdout)
def _run(*args):
# LC_ALL=C keeps git's error strings in English so assertions on them
# are stable across locales.
proc = subprocess.run(
["uv", "run", str(SCRIPT), *args],
capture_output=True,
text=True,
env={**os.environ, "LC_ALL": "C", "LANG": "C"},
)
return proc.returncode, _json(proc)
def _proc(*args, env=None):
"""Run the script and return the raw process, so a test can assert on the
exit code and stderr together — and so `env` can carry an overlay (a fake
`git` earlier on PATH) that `_run` has no way to pass."""
overlay = {"LC_ALL": "C", "LANG": "C"}
if env:
overlay.update(env)
return subprocess.run(
["uv", "run", str(SCRIPT), *args],
capture_output=True,
text=True,
env={**os.environ, **overlay},
)
def _git(repo, *args):
subprocess.run(
["git", "-C", str(repo), *args],
check=True,
capture_output=True,
env=_git_env(repo),
)
def _git_unchecked(repo, *args):
"""`git` that tolerates a non-zero exit — the conflicting merge in
`_merge_repo` is supposed to fail, and the hand resolution comes after it.
Same environment as `_git`, which runs with check=True."""
return subprocess.run(
["git", "-C", str(repo), *args],
capture_output=True,
text=True,
env=_git_env(repo),
)
def _make_repo(tmp_path):
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q")
(repo / "a.py").write_text("one\ntwo\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "epic-1-1 initial a")
(repo / "a.py").write_text("one\ntwo\nthree\nfour\n")
(repo / "b.py").write_text("x\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "epic-1-2 grow a, add b")
return repo
def test_no_range_returns_empty(tmp_path):
repo = _make_repo(tmp_path)
code, out = _run("--repo", str(repo))
assert code == 0
assert out["range"] is None
assert out["commits"] == [] and out["files"] == []
def test_measures_commits_and_files_with_attribution(tmp_path):
repo = _make_repo(tmp_path)
code, out = _run("--repo", str(repo), "--range", "HEAD~1..HEAD", "--stories", "1-2,1-1")
assert code == 0
assert out["range"] == "HEAD~1..HEAD"
assert out["commit_count"] == 1
# The single commit in range is the second one; attributed to story "1-2".
assert out["commits"][0]["stories"] == ["1-2"]
files = {f["path"]: f for f in out["files"]}
# a.py grew by two lines, b.py added one — measured, not judged.
assert files["a.py"]["added"] == 2 and files["a.py"]["net"] == 2
assert files["b.py"]["added"] == 1
def test_story_attribution_respects_word_boundary(tmp_path):
# Story id "1-2" must NOT match a commit subject mentioning "11-2".
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q")
(repo / "f.py").write_text("a\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
(repo / "f.py").write_text("a\nb\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "epic-11-2 unrelated story")
code, out = _run("--repo", str(repo), "--range", "HEAD~1..HEAD", "--stories", "1-2")
assert code == 0
assert out["commits"][0]["stories"] == []
def test_bad_range_errors_as_json(tmp_path):
repo = _make_repo(tmp_path)
code, out = _run("--repo", str(repo), "--range", "nope..alsonope")
assert code == 1
assert out["ok"] is False and out["error"]
def test_single_rev_range_rejected(tmp_path):
# A single rev is not a range: git would log ALL history up to it and the
# script would report the whole repo as the epic's evidence.
repo = _make_repo(tmp_path)
code, out = _run("--repo", str(repo), "--range", "HEAD")
assert code == 2
assert out["ok"] is False and "invalid --range" in out["error"]
def test_pathspec_range_rejected(tmp_path):
# A path that exists must not be silently consumed as a pathspec.
repo = _make_repo(tmp_path)
code, out = _run("--repo", str(repo), "--range", "a.py")
assert code == 2
assert out["ok"] is False and "invalid --range" in out["error"]
def test_range_shaped_pathspec_forced_to_rev_parse(tmp_path):
# A committed file literally named "a..b" passes the REV..REV shape check;
# without the trailing "--" in the git argv, git silently logs that FILE's
# history with exit 0. The "--" forces rev interpretation, so this must
# error instead of measuring the decoy.
repo = _make_repo(tmp_path)
(repo / "a..b").write_text("decoy\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "add decoy file named like a range")
code, out = _run("--repo", str(repo), "--range", "a..b")
assert code == 1
assert out["ok"] is False and "bad revision" in out["error"]
def test_option_like_range_rejected(tmp_path):
# A range starting with "-" must never reach git, where it would be
# consumed as an option (e.g. --output=... writes an arbitrary file).
repo = _make_repo(tmp_path)
code, out = _run("--repo", str(repo), "--range=--output=evil.txt")
assert code == 2
assert out["ok"] is False and "invalid --range" in out["error"]
assert not (repo / "evil.txt").exists()
def test_degenerate_range_shapes_rejected(tmp_path):
# Shapes that contain ".." but are not REV..REV: git would silently
# default an empty endpoint to HEAD ("..", "a..", "..HEAD"), a leading
# dash must never reach git even when dots are present ("-3..HEAD"),
# unstripped values must not slip past the dash guard, and a three-dot
# range is a symmetric difference — git would measure commits reachable
# from either endpoint but not both, a different evidence set entirely.
repo = _make_repo(tmp_path)
for bad in (
"..",
"a..",
"..HEAD",
"-3..HEAD",
" HEAD~1..HEAD",
"HEAD~1...HEAD",
"a...b",
):
code, out = _run("--repo", str(repo), f"--range={bad}")
assert code == 2, f"accepted {bad!r}"
assert out["ok"] is False and "invalid --range" in out["error"], bad
def test_malformed_args_emit_json_not_usage(tmp_path):
# An unknown flag must still land on the JSON contract, not argparse's
# plain usage text on stderr.
code, out = _run("--bogus-flag")
assert code != 0
assert out["ok"] is False and out["error"]
def test_help_flags_emit_json_not_usage():
# argparse's built-in help action bypasses the error() override entirely --
# it prints usage text on stdout and exits 0, which breaks the JSON-only
# contract for a machine consumer. add_help=False demotes -h to an ordinary
# unrecognized argument, which error() already handles.
for flag in ("-h", "--help"):
proc = _proc(flag)
# Exit 2 specifically: the module docstring reserves 2 for argument
# errors and 1 for git/I-O failures, so collapsing them must fail here.
assert proc.returncode == 2, flag
assert "usage:" not in proc.stdout, flag
out = _json(proc)
assert out["ok"] is False and out["error"], flag
# --- helpers for the fixtures below -----------------------------------------
def _rev(repo, ref):
return _git_unchecked(repo, "rev-parse", ref).stdout.strip()
def _fake_git(tmp_path, body):
"""Write a `git` shim and return the PATH overlay that puts it ahead of the
real binary for the script's own subprocesses (never for the fixtures,
which build their repos through `_git`'s own environment)."""
bindir = tmp_path / "fakebin"
bindir.mkdir()
shim = bindir / "git"
shim.write_text(body)
os.chmod(shim, 0o755)
return {"PATH": f"{bindir}{os.pathsep}{os.environ['PATH']}"}
def _merge_repo(tmp_path):
"""Two story branches merged into the mainline; the second merge conflicts
and is resolved by hand, adding a line neither branch had. Returns
(repo, base_sha) — `base_sha..HEAD` is the epic range."""
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "s.py").write_text("l1\nl2\nl3\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
base = _rev(repo, "HEAD")
_git(repo, "checkout", "-q", "-b", "s1")
(repo / "s.py").write_text("l1\nA\nl3\n")
_git(repo, "commit", "-qam", "epic-1-2 story one")
_git(repo, "checkout", "-q", "main")
_git(repo, "merge", "-q", "--no-ff", "s1", "-m", "merge story 1-2")
_git(repo, "checkout", "-q", "-b", "s2", base)
(repo / "s.py").write_text("l1\nB\nl3\n")
_git(repo, "commit", "-qam", "epic-1-3 story two")
_git(repo, "checkout", "-q", "main")
conflicted = _git_unchecked(repo, "merge", "--no-ff", "s2", "-m", "merge story 1-3")
assert conflicted.returncode != 0, "fixture expected a merge conflict"
(repo / "s.py").write_text("l1\nAB\nl3\nl4\n") # hand resolution
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "merge story 1-3")
return repo, base
def test_rename_yields_two_openable_paths(tmp_path):
# git's default rename detection emits "src/{mod.py => renamed.py}" — an
# unopenable pseudo-path that also splits one file's churn across keys.
# --no-renames makes the rename an honest delete + add.
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "src").mkdir()
(repo / "src" / "mod.py").write_text("l1\nl2\nl3\nl4\nl5\nl6\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
base = _rev(repo, "HEAD")
_git(repo, "mv", "src/mod.py", "src/renamed.py")
_git(repo, "commit", "-qm", "rename mod")
(repo / "src" / "renamed.py").write_text("l1\nl2\nl3\nl4\nl5\nl6\nl7\n")
_git(repo, "commit", "-qam", "add a line after the rename")
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
files = {f["path"]: f for f in out["files"]}
assert not any("=>" in path for path in files), sorted(files)
assert files["src/mod.py"]["added"] == 0
assert files["src/mod.py"]["deleted"] == 6
assert files["src/renamed.py"]["added"] == 7
assert files["src/renamed.py"]["deleted"] == 0
def _accented_repo(tmp_path):
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "src").mkdir()
(repo / "src" / "café.py").write_text("ca\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
base = _rev(repo, "HEAD")
(repo / "src" / "café.py").write_text("ca\ncb\n")
_git(repo, "commit", "-qam", "touch the accented file")
return repo, base
def _assert_accented_path(repo, out):
paths = [f["path"] for f in out["files"]]
assert len(paths) == 1
assert "\\" not in paths[0] and '"' not in paths[0], paths
assert unicodedata.normalize("NFC", paths[0]) == "src/café.py"
# The reported path is a real path: it opens under --repo.
assert (repo / paths[0]).read_text() == "ca\ncb\n"
def test_non_ascii_path_is_a_real_string(tmp_path):
# Without core.quotePath=false git emits "src/caf\303\251.py" — quoted and
# octal-escaped, so the documented "open the ranked files" step cannot.
repo, base = _accented_repo(tmp_path)
_assert_accented_path(repo, _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD")))
def test_non_ascii_path_survives_a_non_utf8_locale(tmp_path):
# Pins the explicit encoding="utf-8" on the subprocess. Modern CPython's
# UTF-8 mode hides its absence even under LC_ALL=C, so the pin only bites
# with UTF-8 mode and C-locale coercion both off — where the interpreter
# default is US-ASCII and git's UTF-8 path bytes fail to decode, taking the
# whole measurement down with them.
repo, base = _accented_repo(tmp_path)
out = _json(
_proc(
"--repo",
str(repo),
"--range",
f"{base}..HEAD",
env={"PYTHONUTF8": "0", "PYTHONCOERCECLOCALE": "0"},
)
)
_assert_accented_path(repo, out)
def test_merge_churn_is_measured_and_counted(tmp_path):
repo, base = _merge_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
assert out["commit_count"] == 4
assert out["merge_count"] == 2
assert out["merges_measured"] == 2
# Both merges' first-parent churn: 1/1 for the clean merge, 2/1 for the
# hand-resolved one (the resolution added a line neither branch had).
merge_files = {f["path"]: f for f in out["merge_files"]}
assert merge_files["s.py"]["added"] == 3
assert merge_files["s.py"]["deleted"] == 2
assert merge_files["s.py"]["net"] == 1
assert merge_files["s.py"]["commit_count"] == 2
def test_merge_commits_listed_but_excluded_from_files(tmp_path):
repo, base = _merge_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
by_subject = {c["subject"]: c for c in out["commits"]}
assert by_subject["merge story 1-2"]["is_merge"] is True
assert by_subject["merge story 1-3"]["is_merge"] is True
assert by_subject["epic-1-2 story one"]["is_merge"] is False
# `files` is the two story commits only — 1/1 each. Folding the merges in
# would double count: their diff restates the churn they merged.
files = {f["path"]: f for f in out["files"]}
assert files["s.py"]["added"] == 2
assert files["s.py"]["deleted"] == 2
assert files["s.py"]["commit_count"] == 2
def test_story_attribution_survives_merges(tmp_path):
repo, base = _merge_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD", "--stories", "1-2,1-3"))
by_subject = {c["subject"]: c for c in out["commits"]}
assert by_subject["epic-1-2 story one"]["stories"] == ["1-2"]
assert by_subject["epic-1-3 story two"]["stories"] == ["1-3"]
def test_off_spine_merge_is_counted_but_not_measured(tmp_path):
# A back-merge of the mainline into a story branch is a merge in the range,
# but its first-parent diff would restate unrelated mainline content as
# epic churn. It stays in merge_count and out of merges_measured, so the
# gap between the two says plainly that a merge went unmeasured.
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "f.txt").write_text("base\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
base = _rev(repo, "HEAD")
(repo / "f.txt").write_text("base\nmain1\n")
_git(repo, "commit", "-qam", "mainline work")
_git(repo, "checkout", "-q", "-b", "s1", base)
(repo / "s1.txt").write_text("s1\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "epic-1-2 story one")
_git(repo, "merge", "-q", "--no-ff", "main", "-m", "back-merge main into s1")
_git(repo, "checkout", "-q", "main")
_git(repo, "merge", "-q", "--no-ff", "s1", "-m", "merge story 1-2")
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
assert out["merge_count"] == 2
assert out["merges_measured"] == 1
merge_files = {f["path"]: f for f in out["merge_files"]}
assert set(merge_files) == {"s1.txt"}
def test_merge_pass_survives_a_hostile_log_diffmerges_config(tmp_path):
# `-m` means "whatever log.diffMerges says", so a user or repo config of
# `off` makes the merge pass emit no file rows: merge_files comes back
# empty beside a non-zero merges_measured and reads as "the merges changed
# nothing". The command-line -c pin beats the config.
repo, base = _merge_repo(tmp_path)
_git(repo, "config", "log.diffMerges", "off")
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
assert out["merges_measured"] == 2
merge_files = {f["path"]: f for f in out["merge_files"]}
assert merge_files["s.py"]["added"] == 3
assert merge_files["s.py"]["deleted"] == 2
def test_distinct_non_utf8_paths_stay_distinct(tmp_path):
# errors="replace" maps every invalid byte to the same U+FFFD, collapsing
# two different files into one `files` key with their churn summed —
# measurement corruption with nothing in the output admitting to it.
overlay = _fake_git(
tmp_path,
"#!/bin/sh\nprintf 'aaaa\\037\\037subj\\n\\n1\\t0\\tsrc/caf\\351.py\\n2\\t0\\tsrc/caf\\377.py\\n'\n",
)
out = _json(_proc("--repo", str(tmp_path), "--range", "a..b", env=overlay))
files = {f["path"]: f for f in out["files"]}
assert len(files) == 2, files
assert sorted(f["added"] for f in files.values()) == [1, 2]
def test_repeated_merge_headers_are_counted_once(tmp_path):
# Pins the dedupe guard in _parse_log. git before 2.31 does not honour
# --first-parent for `-m`'s diff format, so a merge's header repeats once
# per parent with a diff block under each. Without the guard that doubles
# the merge churn and pushes merges_measured above merge_count, inverting
# the invariant the reference documents. Inert on modern git, so it needs
# pre-2.31-shaped output to be exercised at all.
real_git = shutil.which("git")
assert real_git, "git must be on PATH"
repo, base = _merge_repo(tmp_path)
overlay = _fake_git(
tmp_path,
"#!/bin/sh\n"
'for a in "$@"; do\n'
' if [ "$a" = "--min-parents=2" ]; then\n'
" printf 'aaaa\\037p1 p2\\037merge story 1-3\\n\\n2\\t1\\ts.py\\n"
"aaaa\\037p1 p2\\037merge story 1-3\\n\\n5\\t4\\ts.py\\n'\n"
" exit 0\n"
" fi\n"
"done\n"
f'exec "{real_git}" "$@"\n',
)
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD", env=overlay))
# One merge sha, however many blocks git printed for it.
assert out["merges_measured"] == 1
assert out["merges_measured"] <= out["merge_count"]
merge_files = {f["path"]: f for f in out["merge_files"]}
# The first block is the first-parent diff on every git version; the
# repeat must not be added on top of it.
assert merge_files["s.py"]["added"] == 2
assert merge_files["s.py"]["deleted"] == 1
assert merge_files["s.py"]["commit_count"] == 1
def test_valid_range_with_no_commits_keeps_the_full_shape(tmp_path):
# A mis-specified epic range is a valid, empty range. It must still answer
# with every documented key rather than a differently-shaped stub.
repo = _make_repo(tmp_path)
proc = _proc("--repo", str(repo), "--range", "HEAD..HEAD")
out = _json(proc)
assert proc.returncode == 0
assert set(out) == {
"range",
"commit_count",
"merge_count",
"merges_measured",
"commits",
"files",
"merge_files",
"stories_supplied",
}
assert out["commit_count"] == 0
assert out["merge_count"] == 0
assert out["merges_measured"] == 0
assert out["commits"] == [] and out["files"] == [] and out["merge_files"] == []
def test_linear_history_reports_no_merges(tmp_path):
repo = _make_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", "HEAD~1..HEAD"))
assert out["merge_count"] == 0
assert out["merges_measured"] == 0
assert out["merge_files"] == []
def _recording_git(tmp_path, calls):
"""A `git` shim that records each invocation as one \\x1f-delimited record
(one field per argument, so argument boundaries survive) and then execs the
real git."""
real_git = shutil.which("git")
assert real_git, "git must be on PATH"
return _fake_git(
tmp_path,
f'#!/bin/sh\n( printf \'%s\\037\' "$@"; printf \'\\n\' ) >> "{calls}"\nexec "{real_git}" "$@"\n',
)
def _numstat_invocations(calls):
"""The recorded `git log --numstat` invocations, each as an argument list."""
out = []
for line in calls.read_text().splitlines():
argv = [field for field in line.split("\x1f") if field]
if "--numstat" in argv:
out.append(argv)
return out
def test_second_pass_runs_only_when_the_range_has_merges(tmp_path):
# The merge pass is skipped outright on linear history, so the common case
# still costs exactly one `git log` — and the two passes must differ in
# exactly the arguments the design depends on.
calls = tmp_path / "calls.log"
overlay = _recording_git(tmp_path, calls)
merge_args = {"-m", "--first-parent", "--min-parents=2"}
(tmp_path / "linear").mkdir()
(tmp_path / "merged").mkdir()
linear = _make_repo(tmp_path / "linear")
_json(_proc("--repo", str(linear), "--range", "HEAD~1..HEAD", env=overlay))
logs = _numstat_invocations(calls)
assert len(logs) == 1, logs
calls.write_text("")
merged, base = _merge_repo(tmp_path / "merged")
_json(_proc("--repo", str(merged), "--range", f"{base}..HEAD", env=overlay))
logs = _numstat_invocations(calls)
assert len(logs) == 2, logs
# Pass 1 keeps full topology: none of the merge-pass arguments may reach it,
# or the story-branch commits drop out of the listing and attribution dies.
assert merge_args.isdisjoint(logs[0]), logs[0]
assert merge_args.issubset(logs[1]), logs[1]
def test_multi_story_subject_attributes_to_every_match(tmp_path):
# First-match-wins silently dropped the second story from the attribution.
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "f.py").write_text("a\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
(repo / "f.py").write_text("a\nb\n")
_git(repo, "commit", "-qam", "fix seam between 1-2 and 1-3")
out = _json(_proc("--repo", str(repo), "--range", "HEAD~1..HEAD", "--stories", "1-3,1-2"))
# Every match, in --stories order — not whichever id was passed first.
assert out["commits"][0]["stories"] == ["1-3", "1-2"]
# A repeated id must not list the commit twice: any per-story total built
# from `stories` would count it twice.
repeated = _json(_proc("--repo", str(repo), "--range", "HEAD~1..HEAD", "--stories", "1-3,1-2,1-3"))
assert repeated["commits"][0]["stories"] == ["1-3", "1-2"]
assert repeated["stories_supplied"] == ["1-3", "1-2"]
def test_subject_naming_no_story_gets_an_empty_list(tmp_path):
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "f.py").write_text("a\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
(repo / "f.py").write_text("a\nb\n")
_git(repo, "commit", "-qam", "chore: tidy imports")
out = _json(_proc("--repo", str(repo), "--range", "HEAD~1..HEAD", "--stories", "1-2,1-3"))
assert out["commits"][0]["stories"] == []
def test_git_failure_with_empty_stderr_reports_the_exit_code(tmp_path):
# A quiet git failure (signal kill, empty stderr) must not leave the caller
# with `"error": ""` and nothing to report.
overlay = _fake_git(tmp_path, "#!/bin/sh\nexit 3\n")
proc = _proc("--repo", str(tmp_path), "--range", "HEAD~1..HEAD", env=overlay)
out = _json(proc)
assert proc.returncode == 1
assert out["ok"] is False
assert out["error"] == "git exited 3"
def _binary_repo(tmp_path):
"""A binary-only path plus a path that is binary twice and text once."""
repo = tmp_path / "repo"
repo.mkdir()
_git(repo, "init", "-q", "-b", "main")
(repo / "keep.txt").write_text("keep\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "base")
base = _rev(repo, "HEAD")
(repo / "src").mkdir()
(repo / "src" / "x.py").write_bytes(b"\x00\x01bin\n")
(repo / "blob.bin").write_bytes(b"\x00\x01blob\n")
_git(repo, "add", "-A")
_git(repo, "commit", "-qm", "add binary content")
(repo / "src" / "x.py").write_text("one\ntwo\n") # binary -> text: still binary
_git(repo, "commit", "-qam", "x.py becomes text")
(repo / "src" / "x.py").write_text("one\ntwo\nthree\n") # text -> text: measured
_git(repo, "commit", "-qam", "grow x.py")
(repo / "blob.bin").write_bytes(b"\x00\x02blob\n")
_git(repo, "commit", "-qam", "churn the blob")
return repo, base
def test_binary_revisions_no_longer_erase_measured_text_churn(tmp_path):
repo, base = _binary_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
x = {f["path"]: f for f in out["files"]}["src/x.py"]
assert x["added"] == 1 and x["deleted"] == 0 and x["net"] == 1
assert x["binary_revisions"] == 2
assert x["commit_count"] == 3
def test_binary_only_path_reports_zero_sums_and_its_revision_count(tmp_path):
repo, base = _binary_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", f"{base}..HEAD"))
blob = {f["path"]: f for f in out["files"]}["blob.bin"]
assert blob["added"] == 0 and blob["deleted"] == 0 and blob["net"] == 0
assert blob["binary_revisions"] == 2
assert blob["commit_count"] == 2
def test_success_shape_carries_every_documented_key(tmp_path):
repo = _make_repo(tmp_path)
out = _json(_proc("--repo", str(repo), "--range", "HEAD~1..HEAD"))
assert set(out) >= {
"range",
"commit_count",
"merge_count",
"merges_measured",
"commits",
"files",
"merge_files",
"stories_supplied",
}
assert set(out["commits"][0]) == {"sha", "subject", "stories", "is_merge"}
assert set(out["files"][0]) == {
"path",
"added",
"deleted",
"net",
"commit_count",
"binary_revisions",
}
def test_explicit_repo_ignores_ambient_git_dir(tmp_path):
repo = _make_repo(tmp_path)
proc = _proc(
"--repo",
str(repo),
"--range",
"HEAD~1..HEAD",
env={"GIT_DIR": str(tmp_path / "wrong-git-dir")},
)
assert _json(proc)["commit_count"] == 1
if __name__ == "__main__":
sys.exit(pytest.main([__file__, "-q"]))
scripts/tests/test_sprint_status.py
# /// script
# requires-python = ">=3.11"
# dependencies = ["pytest>=8.0", "ruamel.yaml>=0.18"]
# ///
"""Corruption-critical tests for sprint-status.py.
Each test runs the script as a subprocess via ``uv run`` against a temp copy of
an inline fixture, then re-reads the file to assert comments and formatting
survive and punctuation-heavy action values round-trip intact.
Run: uv run scripts/tests/test_sprint_status.py
or: uv run --with pytest --with ruamel.yaml -m pytest scripts/tests/test_sprint_status.py
"""
import importlib.util
import json
import os
import re
import stat
import subprocess
import sys
from pathlib import Path
import pytest
from ruamel.yaml import YAML
SCRIPT = Path(__file__).resolve().parents[1] / "sprint_status.py"
# Vendored copy of bmad-sprint-planning's sprint-status-template.yaml — skills
# must not path into each other's directories (PATH-05). Keep this fixture in
# sync with the source template by hand when it changes.
TEMPLATE = Path(__file__).resolve().parent / "fixtures" / "sprint-status-template.yaml"
FIXTURE = """\
# Sprint Status Tracking
# STATUS DEFINITIONS:
# backlog - not yet started
# ready-for-dev - ready to be implemented
# done - completed
generated: "01-01-2026 09:00"
last_updated: "01-01-2026 09:00"
project: "Demo Project"
project_key: "DEMO"
tracking_system: "file"
story_location: "docs/stories"
development_status:
epic-1: backlog
1-1-user-authentication: done
1-2-account-management: done
epic-1-retrospective: optional
epic-2: backlog
2-1-dashboard: backlog
"""
# Two epics' worth of items, because the flag's headline use is epic N's retro
# closing epic N-1's items: nothing may scope a selector to --epic. The two
# "Scripted item" entries share their action text and differ only by epic, so the
# legacy epic+action selector has to discriminate on the epic to resolve at all.
ACTION_FIXTURE = """\
# Sprint Status Tracking
# STATUS DEFINITIONS:
# open - committed during a retrospective, not yet addressed
generated: "01-01-2026 09:00"
last_updated: "01-01-2026 09:00"
development_status:
1-1-a: done
epic-1-retrospective: optional
2-1-b: done
epic-2-retrospective: optional
# Action items committed during retrospectives
action_items:
- id: "epic-1-retro-item-1-x"
epic: 1
action: "Scripted item"
owner: "Amelia"
status: "open"
ref: "docs/epic-1-retro.md"
- epic: 1
action: "Pre-existing item"
owner: "Charlie"
status: open
- id: "epic-2-retro-item-1-y"
epic: 2
action: "Scripted item"
owner: "Dana"
status: "in-progress"
ref: "docs/epic-2-retro.md"
"""
# Three spellings of the same scalar, to pin that a status write never re-styles
# the line it lands on.
STYLE_FIXTURE = """\
last_updated: "01-01-2026 09:00"
development_status:
1-1-a: done
epic-1-retrospective: optional
action_items:
- id: "double"
epic: 1
action: "Double"
status: "open"
- id: "single"
epic: 1
action: "Single"
status: 'open'
- id: "plain"
epic: 1
action: "Plain"
status: open
"""
def _run(args):
cmd = ["uv", "run", str(SCRIPT), *args]
# LC_ALL=C keeps os.strerror text stable so error-string assertions do not
# depend on the developer's locale.
return subprocess.run(cmd, capture_output=True, text=True, env={**os.environ, "LC_ALL": "C"})
def _module():
"""Import the script as a module, for the few properties that cannot be
triggered through the CLI (a failure after the temp file already exists)."""
spec = importlib.util.spec_from_file_location("sprint_status", SCRIPT)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
def _write_fixture(tmp_path):
target = tmp_path / "sprint-status.yaml"
target.write_text(FIXTURE, encoding="utf-8")
return target
def _write_action_fixture(tmp_path):
target = tmp_path / "sprint-status.yaml"
target.write_text(ACTION_FIXTURE, encoding="utf-8")
return target
def _load(path):
yaml = YAML(typ="rt")
with open(path, encoding="utf-8") as fh:
return yaml.load(fh)
def _json(proc):
"""Parse the JSON-only stdout contract, surfacing a crash instead of hiding
it behind a JSONDecodeError."""
assert proc.stdout, f"empty stdout; stderr was: {proc.stderr}"
assert "Traceback" not in proc.stderr, proc.stderr
return json.loads(proc.stdout)
def test_detect_epic(tmp_path):
target = _write_fixture(tmp_path)
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 0, proc.stderr
out = json.loads(proc.stdout)
assert out["epic"] == 1
assert out["story_count"] == 2
assert out["retro_key"] == "epic-1-retrospective"
assert out["retro_status"] == "optional"
assert set(out["done_stories"]) == {
"1-1-user-authentication",
"1-2-account-management",
}
def test_detect_epic_rejects_typed_retrospective_status_as_json(tmp_path):
fixture = "development_status:\n 1-1-a: done\n epic-1-retrospective: 2026-01-01\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["error"] == "epic-1-retrospective status must be a string or null"
assert "restored" not in out
assert target.read_text(encoding="utf-8") == fixture
# --- pending_stories: the unfinished-epic gate --------------------------------
def test_pending_stories_lists_the_selected_epics_unfinished_keys(tmp_path):
# The gate the skill branches on before Phase 1: an epic whose highest done
# story selected it, but which is not actually finished. Document order, so
# the listing the user confirms matches the file they can open.
fixture = "development_status:\n epic-2: backlog\n 2-1-a: done\n 2-2-b: backlog\n 2-3-c: ready-for-dev\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 2
assert out["pending_stories"] == ["2-2-b", "2-3-c"]
# A non-story key sitting beside them never leaks in: STORY_RE gates entry.
assert "epic-2" not in out["pending_stories"]
def test_pending_stories_is_empty_for_a_complete_epic(tmp_path):
fixture = "development_status:\n 2-1-a: done\n 2-2-b: done\n epic-2-retrospective: optional\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 2
assert out["pending_stories"] == []
def test_pending_stories_ignores_other_epics(tmp_path):
# FIXTURE detects epic 1 while 2-1-dashboard sits at backlog. The key is
# scoped to the *selected* epic -- unlike done_stories, which spans the whole
# file -- so another epic's unfinished work must never block this retro.
target = _write_fixture(tmp_path)
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 1
assert out["pending_stories"] == []
# done_stories keeps its whole-file scope, unchanged.
assert set(out["done_stories"]) == {
"1-1-user-authentication",
"1-2-account-management",
}
def test_pending_stories_present_when_no_epic_is_detected(tmp_path):
# No done story anywhere: the shape stays uniform so a caller can read
# pending_stories without first branching on epic.
fixture = "development_status:\n 1-1-a: backlog\n 2-1-b: ready-for-dev\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] is None
assert out["story_count"] == 0
assert out["pending_stories"] == []
def test_detect_epic_flag_aims_pending_stories_at_a_supplied_epic(tmp_path):
# Auto-detect would pick epic 2 (highest with a done story). --epic 1 aims
# the unfinished-epic gate at the orchestrator's explicit choice instead —
# the -H <epic> path that previously had no pending_stories at all.
fixture = (
"development_status:\n"
" 1-1-a: done\n"
" 1-2-b: backlog\n"
" 1-3-c: review\n"
" 2-1-a: done\n"
" 2-2-b: done\n"
" epic-1-retrospective: optional\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target), "--epic", "1"])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 1
assert out["retro_key"] == "epic-1-retrospective"
assert out["retro_status"] == "optional"
assert out["pending_stories"] == ["1-2-b", "1-3-c"]
# done_stories keeps its whole-file scope.
assert set(out["done_stories"]) == {"1-1-a", "2-1-a", "2-2-b"}
def test_detect_epic_flag_lists_pending_when_no_story_is_done(tmp_path):
# Without --epic, no done story means epic is null. With --epic, an
# unfinished epic that never landed a done story is still addressable —
# every story key of that epic is pending.
fixture = "development_status:\n 3-1-a: backlog\n 3-2-b: ready-for-dev\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target), "--epic", "3"])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 3
assert out["pending_stories"] == ["3-1-a", "3-2-b"]
assert out["retro_key"] == "epic-3-retrospective"
assert out["retro_status"] is None
def test_detect_epic_flag_empty_pending_for_a_complete_supplied_epic(tmp_path):
fixture = "development_status:\n 1-1-a: done\n 1-2-b: done\n 2-1-a: backlog\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target), "--epic", "1"])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 1
assert out["pending_stories"] == []
def test_detect_epic_flag_zero_story_count_marks_a_nonexistent_epic(tmp_path):
# --epic 9 against a file that has no epic-9 stories: pending_stories is
# empty exactly as it is for a finished epic, so story_count is the only
# signal separating "complete" from "typo'd". The gate reads 0 as suspect,
# never as done.
fixture = "development_status:\n 1-1-a: done\n 1-2-b: done\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target), "--epic", "9"])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["epic"] == 9
assert out["story_count"] == 0
assert out["pending_stories"] == []
assert out["retro_status"] is None
# The finished epic it could be confused with reports its real count.
proc = _run(["detect-epic", "--file", str(target), "--epic", "1"])
out = _json(proc)
assert out["story_count"] == 2
assert out["pending_stories"] == []
def test_detect_epic_flag_rejects_non_positive_epic(tmp_path):
target = _write_fixture(tmp_path)
for bad in ("0", "-3"):
proc = _run(["detect-epic", "--file", str(target), "--epic", bad])
assert proc.returncode == 1, proc.stderr
out = _json(proc)
assert out["ok"] is False
assert "epic" in out["error"]
assert "restored" not in out
def test_update_rejects_non_positive_epic(tmp_path):
target = _write_fixture(tmp_path)
for bad in ("0", "-3"):
proc = _run(["update", "--file", str(target), "--epic", bad])
assert proc.returncode == 1, proc.stderr
assert "positive integer" in _json(proc)["error"]
assert target.read_text(encoding="utf-8") == FIXTURE
# --- The JSON-only contract covers the help paths ----------------------------
@pytest.mark.parametrize(
"args",
[["-h"], ["--help"], ["detect-epic", "-h"], ["detect-epic", "--help"]],
ids=["top-short", "top-long", "sub-short", "sub-long"],
)
def test_help_flags_emit_json_not_usage(args):
# argparse's built-in help action bypasses error() -- it prints usage text
# to stdout and exits 0, which is exactly the contract this script sells.
# add_help=False turns -h into an ordinary unrecognized argument instead.
proc = _run(args)
# Exit 2 specifically: the module docstring reserves 2 for argument errors
# and 1 for I/O failures, and retro-document.md teaches callers to tell the
# two apart, so collapsing them must fail here.
assert proc.returncode == 2
assert "usage:" not in proc.stdout
out = _json(proc)
assert out["ok"] is False
assert out["error"]
# An argparse rejection speaks for no file, so it carries no "restored" --
# the same rule test_only_the_write_path_reports_restored pins for exit 1.
assert "restored" not in out
@pytest.mark.parametrize("flag", ["-h", "--help"])
def test_update_help_flag_emits_json_not_usage(tmp_path, flag):
# The update subparser too, driven with its required arguments present so
# nothing but the help flag itself can be what argparse objects to.
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", flag])
assert proc.returncode == 2
assert "usage:" not in proc.stdout
out = _json(proc)
assert out["ok"] is False
assert flag in out["error"]
assert "restored" not in out
# A rejected invocation must not have written anything.
assert target.read_text(encoding="utf-8") == FIXTURE
def test_update_sets_retro_and_appends_action(tmp_path):
target = _write_fixture(tmp_path)
payload = '[{"action":"Fix #42: colons: and # hashes","owner":"Amelia"}]'
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-retro-done",
"--add-action",
payload,
]
)
assert proc.returncode == 0, proc.stderr
out = json.loads(proc.stdout)
assert out["ok"] is True
assert out["retro_key_found"] is True
assert out["retro_status_after"] == "done"
assert out["action_items_added"] == 1
# File must still parse cleanly (punctuation did not corrupt it).
data = _load(target)
assert data is not None
# STATUS DEFINITIONS comment survived.
raw = target.read_text(encoding="utf-8")
assert "STATUS DEFINITIONS" in raw
# Retro status flipped to done.
assert data["development_status"]["epic-1-retrospective"] == "done"
# The action value round-trips with literal '#' and ':' intact.
action = data["action_items"][0]
assert action["action"] == "Fix #42: colons: and # hashes"
assert action["owner"] == "Amelia"
assert action["epic"] == 1
assert action["status"] == "open"
def test_update_rejects_typed_retrospective_status_before_writing(tmp_path):
fixture = "last_updated: 01-01-2026 09:00\ndevelopment_status:\n 1-1-a: done\n epic-1-retrospective: 2026-01-01\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-retro-done",
"--add-action",
'[{"action":"Must not be appended","owner":"Amelia"}]',
]
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert out["error"] == "epic-1-retrospective status must be a string or null"
assert target.read_text(encoding="utf-8") == fixture
def test_detect_epic_matches_split_story_keys(tmp_path):
# A split-story key like 2-6a-... is first-class in BMAD (an oversized story
# split into 2-6a / 2-6b) and must not be invisible to detection — otherwise
# an epic whose only done stories are splits is silently skipped.
fixture = "development_status:\n 1-1-first: done\n 2-6a-split-auth: done\n epic-2-retrospective: optional\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["detect-epic", "--file", str(target)])
assert proc.returncode == 0, proc.stderr
out = json.loads(proc.stdout)
assert out["epic"] == 2
assert "2-6a-split-auth" in out["done_stories"]
assert out["retro_key"] == "epic-2-retrospective"
def test_update_rejects_non_list_action_items(tmp_path):
# A hand-corrupted action_items must fail on the JSON contract, not crash.
fixture = 'development_status:\n 1-1-a: done\n epic-1-retrospective: optional\naction_items: "oops-not-a-list"\n'
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["update", "--file", str(target), "--epic", "1", "--add-action", '[{"action":"x","owner":"y"}]'])
assert proc.returncode == 1
out = json.loads(proc.stdout) # must be JSON, not a traceback
assert out["ok"] is False
assert "action_items" in out["error"]
def test_appended_items_carry_id_and_ref(tmp_path):
target = _write_fixture(tmp_path)
ref = "docs/stories/epic-1-retro-2026-07-21.md"
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-retro-done",
"--add-action",
'[{"action":"Fix the seam","owner":"Amelia"}]',
"--ref",
ref,
"--verdict",
"accepted-with-open-items",
]
)
assert proc.returncode == 0, proc.stderr
out = json.loads(proc.stdout)
assert out["verdict"] == "accepted-with-open-items" # echoed, not written to a key
item = _load(target)["action_items"][0]
assert item["id"].startswith("epic-1-retro-item-1-")
assert item["ref"] == ref
# The retro key value stays "done" — verdict is not encoded into it.
assert _load(target)["development_status"]["epic-1-retrospective"] == "done"
def test_free_spelled_verdict_is_rejected_before_the_file_is_touched(tmp_path):
# The SKILL's prose verdict ("accepted with open items") and the frontmatter
# token (accepted-with-open-items) used to be two spellings of one value; an
# orchestrator branching on the echo would fall through both. Only the
# frontmatter vocabulary passes; anything else fails with the file intact.
target = _write_fixture(tmp_path)
for bad in ("accepted with open items", "ship it", "ACCEPTED"):
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-retro-done", "--verdict", bad])
assert proc.returncode == 1, f"accepted {bad!r}"
out = _json(proc)
assert out["ok"] is False and "--verdict" in out["error"]
assert out["restored"] is True
assert target.read_text(encoding="utf-8") == FIXTURE
def test_explicit_item_id_is_preserved(tmp_path):
target = _write_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--add-action",
'[{"action":"a","owner":"o","id":"custom-id-7"}]',
]
)
assert proc.returncode == 0, proc.stderr
assert _load(target)["action_items"][0]["id"] == "custom-id-7"
@pytest.mark.skipif(
hasattr(os, "geteuid") and os.geteuid() == 0,
reason="root bypasses file permission bits",
)
def test_write_failure_reports_restore_status(tmp_path):
# If the write cannot happen, the caller must be told whether the original
# was restored — a silent failure defeats the script's core guarantee.
#
# The write is atomic (temp file + os.replace), and os.replace needs write
# permission on the *directory*, not on the target — a read-only target is
# now replaceable. Making the containing directory read-only is what blocks
# the write: mkstemp fails, while the restore write to the still-writable
# target succeeds.
holder = tmp_path / "holder"
holder.mkdir()
target = _write_fixture(holder)
os.chmod(holder, 0o555)
try:
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-retro-done"])
finally:
os.chmod(holder, 0o755)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
# The file was never touched: the temp file could not even be created.
assert target.read_text(encoding="utf-8") == FIXTURE
assert [p.name for p in holder.iterdir()] == ["sprint-status.yaml"]
def test_punctuation_does_not_corrupt_file(tmp_path):
# Explicit re-parse guarantee for YAML-breaking punctuation.
target = _write_fixture(tmp_path)
payload = '[{"action":"weird: value # with: hashes","owner":"Bob # Smith"}]'
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--add-action",
payload,
]
)
assert proc.returncode == 0, proc.stderr
# Re-parse must succeed and preserve the literal punctuation.
data = _load(target)
assert data["action_items"][0]["action"] == "weird: value # with: hashes"
assert data["action_items"][0]["owner"] == "Bob # Smith"
# --- Formatting fidelity -----------------------------------------------------
def test_template_round_trip_changes_only_last_updated(tmp_path):
# The repo's own sprint-status template is the shape every generated file
# inherits: 2-space sequence indent and a mid-file comment above
# action_items. An update must touch nothing but last_updated — a re-indent
# of a pre-existing, untouched entry defeats the preservation guarantee that
# motivates "do not hand-edit this file".
source = TEMPLATE.read_text(encoding="utf-8")
target = tmp_path / "sprint-status.yaml"
target.write_text(source, encoding="utf-8")
proc = _run(["update", "--file", str(target), "--epic", "1", "--date", "01-01-2026 09:00"])
assert proc.returncode == 0, proc.stderr
assert _json(proc)["ok"] is True
before = source.splitlines()
after = target.read_text(encoding="utf-8").splitlines()
assert len(before) == len(after)
changed = [(b, a) for b, a in zip(before, after, strict=False) if b != a]
assert len(changed) == 1, changed
assert changed[0][1] == "last_updated: 01-01-2026 09:00"
# The pre-existing action item keeps its 2-space sequence indent.
assert " - epic: 1" in after
def test_mid_file_comment_survives_update(tmp_path):
fixture = (
"# header\n"
"development_status:\n"
" 1-1-a: done\n"
" epic-1-retrospective: optional\n"
"\n"
"# Action items committed during retrospectives\n"
"action_items:\n"
" - epic: 1\n"
' action: "Pre-existing item"\n'
' owner: "Charlie"\n'
" status: open\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-retro-done",
"--add-action",
'[{"action":"New item","owner":"Amelia"}]',
]
)
assert proc.returncode == 0, proc.stderr
raw = target.read_text(encoding="utf-8")
assert "# Action items committed during retrospectives" in raw
assert "# header" in raw
assert ' action: "Pre-existing item"' in raw
def test_legacy_offset_zero_file_is_canonicalized(tmp_path):
# Files the previous version of this script wrote carry action_items at
# column 0. The indent pin re-indents them to the template's shape on the
# next write. That is a deliberate one-time canonicalization, not a silent
# failure: the update still succeeds and no comment is lost.
fixture = (
"# header\n"
"development_status:\n"
" 1-1-a: done\n"
" epic-1-retrospective: optional\n"
"action_items:\n"
'- id: "legacy"\n'
" epic: 1\n"
' action: "written by the old code"\n'
" status: open\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-retro-done"])
assert proc.returncode == 0, proc.stderr
raw = target.read_text(encoding="utf-8")
assert ' - id: "legacy"' in raw
assert ' action: "written by the old code"' in raw
assert "# header" in raw
def test_lost_comment_fails_with_restore(tmp_path):
# A standalone comment inside a flow collection is genuinely dropped by the
# round-trip. The leading-block check never saw it; the full multiset does,
# and the original bytes must come back.
fixture = (
"# header\n"
"development_status:\n"
" 1-1-a: done\n"
" epic-1-retrospective: optional\n"
"tags: [\n"
" # a standalone comment the round-trip drops\n"
' "alpha",\n'
' "beta",\n'
"]\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-retro-done"])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "comment line lost" in out["error"]
assert "a standalone comment the round-trip drops" in out["error"]
assert target.read_text(encoding="utf-8") == fixture
# --- Malformed input stays on the JSON contract ------------------------------
@pytest.mark.parametrize("command", ["detect-epic", "update"])
def test_non_mapping_root_is_json_error(tmp_path, command):
target = tmp_path / "sprint-status.yaml"
target.write_text("- a\n- b\n", encoding="utf-8")
args = ["--file", str(target)] + (["--epic", "1"] if command == "update" else [])
proc = _run([command, *args])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert "root document is not a mapping" in out["error"]
assert target.read_text(encoding="utf-8") == "- a\n- b\n"
@pytest.mark.parametrize("command", ["detect-epic", "update"])
@pytest.mark.parametrize(
"body",
['development_status: "not-a-mapping"\n', "development_status:\n - a\n - b\n"],
ids=["scalar", "list"],
)
def test_non_mapping_development_status_is_json_error(tmp_path, command, body):
target = tmp_path / "sprint-status.yaml"
target.write_text(body, encoding="utf-8")
args = ["--file", str(target)] + (["--epic", "1"] if command == "update" else [])
proc = _run([command, *args])
# update used to report ok:true here while silently doing nothing.
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert "development_status is not a mapping" in out["error"]
@pytest.mark.parametrize("command", ["detect-epic", "update"])
def test_directory_target_is_json_error(tmp_path, command):
target = tmp_path / "a-directory"
target.mkdir()
args = ["--file", str(target)] + (["--epic", "1"] if command == "update" else [])
proc = _run([command, *args])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["error"]
@pytest.mark.skipif(
hasattr(os, "geteuid") and os.geteuid() == 0,
reason="root bypasses file permission bits",
)
@pytest.mark.parametrize("command", ["detect-epic", "update"])
def test_unreadable_target_is_json_error(tmp_path, command):
# The other half of the OSError widening: PermissionError, not just
# IsADirectoryError, has to stay on the JSON contract.
target = _write_fixture(tmp_path)
os.chmod(target, 0o000)
args = ["--file", str(target)] + (["--epic", "1"] if command == "update" else [])
try:
proc = _run([command, *args])
finally:
os.chmod(target, 0o644)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert "denied" in out["error"].lower()
@pytest.mark.parametrize("command", ["detect-epic", "update"])
def test_invalid_utf8_is_json_error(tmp_path, command):
target = tmp_path / "sprint-status.yaml"
target.write_bytes(b"development_status:\n 1-1-a: d\xffone\n")
args = ["--file", str(target)] + (["--epic", "1"] if command == "update" else [])
proc = _run([command, *args])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert "utf-8" in out["error"].lower()
# --- Atomic write ------------------------------------------------------------
def test_atomic_write_failure_leaves_target_byte_identical(tmp_path, monkeypatch):
# The failure the atomic write exists for: something goes wrong after the
# temp file has been written. Nothing may reach the target and no temp file
# may survive. No CLI path reaches here -- a read-only directory fails at
# mkstemp instead -- so this drives the helper directly.
mod = _module()
target = _write_fixture(tmp_path)
def boom(*args, **kwargs):
raise OSError(28, "No space left on device")
monkeypatch.setattr(mod.os, "replace", boom)
with pytest.raises(OSError):
mod._atomic_write(str(target), b"replacement bytes\n", 0o644)
assert target.read_text(encoding="utf-8") == FIXTURE
assert [p.name for p in tmp_path.iterdir()] == ["sprint-status.yaml"]
def test_dir_fsync_failure_after_rename_is_not_a_write_failure(tmp_path, monkeypatch):
# Once os.replace has returned, the new bytes ARE the file. The directory
# fsync that follows is durability polish; if it raised, cmd_update would
# emit "restored": true about a write that in fact landed — the one lie the
# restored contract exists to prevent. Deny opening the directory (the only
# thing _atomic_write opens by path after the rename) and require success.
mod = _module()
target = _write_fixture(tmp_path)
directory = os.path.dirname(os.path.realpath(str(target)))
real_open = os.open
def deny_directory_open(p, *args, **kwargs):
if p == directory:
raise OSError(5, "Input/output error")
return real_open(p, *args, **kwargs)
monkeypatch.setattr(mod.os, "open", deny_directory_open)
mod._atomic_write(str(target), b"replacement bytes\n", 0o644)
assert target.read_bytes() == b"replacement bytes\n"
assert [p.name for p in tmp_path.iterdir()] == ["sprint-status.yaml"]
def test_restore_is_atomic(tmp_path, monkeypatch):
# _restore is the rollback the reference sells as the safety net. A
# truncating rewrite that dies halfway would destroy the very bytes it is
# putting back, which is how a full disk used to corrupt the file.
mod = _module()
target = tmp_path / "sprint-status.yaml"
target.write_text("damaged\n", encoding="utf-8")
def boom(*args, **kwargs):
raise OSError(28, "No space left on device")
monkeypatch.setattr(mod.os, "replace", boom)
assert mod._restore(str(target), FIXTURE.encode("utf-8"), 0o644) is False
# It reported failure honestly and left the file no worse than it found it.
assert target.read_text(encoding="utf-8") == "damaged\n"
assert [p.name for p in tmp_path.iterdir()] == ["sprint-status.yaml"]
def test_symlinked_target_is_written_through(tmp_path):
# os.replace onto a symlink would detach the link and leave the real file
# stale while reporting ok:true.
real = tmp_path / "real-sprint-status.yaml"
real.write_text(FIXTURE, encoding="utf-8")
link = tmp_path / "sprint-status.yaml"
link.symlink_to(real)
proc = _run(["update", "--file", str(link), "--epic", "1", "--set-retro-done"])
assert proc.returncode == 0, proc.stderr
assert link.is_symlink(), "the symlink was replaced by a regular file"
assert _load(real)["development_status"]["epic-1-retrospective"] == "done"
def test_atomic_write_preserves_mode_and_leaves_no_temp_file(tmp_path):
# mkstemp creates 0600; without carrying the target's mode over, every
# update would silently narrow the file.
holder = tmp_path / "holder"
holder.mkdir()
target = _write_fixture(holder)
os.chmod(target, 0o640)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-retro-done"])
assert proc.returncode == 0, proc.stderr
assert stat.S_IMODE(target.stat().st_mode) == 0o640
assert [p.name for p in holder.iterdir()] == ["sprint-status.yaml"]
# --- Result-JSON precision ---------------------------------------------------
def test_retro_key_found_is_null_without_the_flag(tmp_path):
# No development_status key at all: the update must not conjure one, and
# retro_key_found must say "not asked" rather than "absent".
fixture = 'project: "Demo"\nlast_updated: "01-01-2026 09:00"\n'
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(["update", "--file", str(target), "--epic", "1"])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["ok"] is True
assert out["retro_key_found"] is None
assert "development_status" not in target.read_text(encoding="utf-8")
def test_retro_key_found_is_false_when_the_key_is_absent(tmp_path):
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "2", "--set-retro-done"])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["ok"] is True
assert out["retro_key_found"] is False
# Nothing was written into the mapping.
assert "epic-2-retrospective" not in _load(target)["development_status"]
@pytest.mark.parametrize("command", ["detect-epic", "update"])
def test_only_the_write_path_reports_restored(tmp_path, command):
# "restored" speaks to the state of a file the command may have written.
# detect-epic never writes, so inventing the key there would mislead callers
# that branch on it.
target = tmp_path / "sprint-status.yaml"
target.write_text("- a\n- b\n", encoding="utf-8")
args = ["--file", str(target)] + (["--epic", "1"] if command == "update" else [])
out = _json(_run([command, *args]))
assert out["ok"] is False
assert ("restored" in out) is (command == "update")
def test_pre_write_failure_reports_restored(tmp_path):
# retro-document.md teaches callers that ok:false carries restored:true;
# a failure before the write must not read as "the file may be incomplete".
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--add-action", "{not json"])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert target.read_text(encoding="utf-8") == FIXTURE
# --- Action-item validation and identity -------------------------------------
def test_non_latin_action_keeps_its_text_in_the_id(tmp_path):
target = _write_fixture(tmp_path)
payload = json.dumps(
[{"action": "Улучшить обработку ошибок", "owner": "Amelia"}],
ensure_ascii=False,
)
proc = _run(["update", "--file", str(target), "--epic", "1", "--add-action", payload])
assert proc.returncode == 0, proc.stderr
item_id = _load(target)["action_items"][0]["id"]
assert item_id == "epic-1-retro-item-1-улучшить-обработку-ошибок"
def test_unsluggable_action_falls_back_to_a_hash(tmp_path):
target = _write_fixture(tmp_path)
payload = json.dumps([{"action": "!!! 🎉", "owner": "Amelia"}], ensure_ascii=False)
proc = _run(["update", "--file", str(target), "--epic", "1", "--add-action", payload])
assert proc.returncode == 0, proc.stderr
item_id = _load(target)["action_items"][0]["id"]
assert not item_id.endswith("-item")
assert re.fullmatch(r"epic-1-retro-item-1-[0-9a-f]{8}", item_id), item_id
def test_empty_action_is_rejected(tmp_path):
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--add-action", '[{"action":" ","owner":"x"}]'])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "action" in out["error"]
assert target.read_text(encoding="utf-8") == FIXTURE
def test_non_string_action_is_rejected(tmp_path):
# A JSON null would otherwise be str()'d into a literal "None" and written
# as a real action item, which the new emptiness check alone lets through.
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--add-action", '[{"action":null,"owner":"x"}]'])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert target.read_text(encoding="utf-8") == FIXTURE
def test_date_is_normalized_to_the_canonical_format(tmp_path):
# strptime accepts unpadded spellings; writing those through would defeat
# the point of validating the format.
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--date", "1-2-2026 9:05"])
assert proc.returncode == 0, proc.stderr
assert _json(proc)["last_updated"] == "01-02-2026 09:05"
assert _load(target)["last_updated"] == "01-02-2026 09:05"
def test_malformed_date_is_rejected(tmp_path):
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--date", "not-a-date"])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "--date" in out["error"]
assert target.read_text(encoding="utf-8") == FIXTURE
# --- Action-item status transitions ------------------------------------------
def test_set_action_status_applies_both_selector_forms(tmp_path):
# The whole point of the flag: an item written by this script (selected by
# id) and a legacy item that predates ids (selected by epic + exact action
# text) both move off "open" in a single call.
target = _write_action_fixture(tmp_path)
payload = json.dumps(
[
{"id": "epic-1-retro-item-1-x", "status": "done"},
{"epic": 1, "action": "Pre-existing item", "status": "in-progress"},
]
)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-action-status", payload])
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["ok"] is True
assert out["action_items_updated"] == 2
assert out["action_items_added"] == 0
items = _load(target)["action_items"]
assert items[0]["status"] == "done"
assert items[1]["status"] == "in-progress"
# Every other key of both items survived untouched, in place.
assert items[0]["id"] == "epic-1-retro-item-1-x"
assert items[0]["epic"] == 1
assert items[0]["action"] == "Scripted item"
assert items[0]["owner"] == "Amelia"
assert items[0]["ref"] == "docs/epic-1-retro.md"
assert "id" not in items[1]
assert items[1]["epic"] == 1
assert items[1]["action"] == "Pre-existing item"
assert items[1]["owner"] == "Charlie"
# The epic-2 item shares its action text with items[0]; the epic-1 selector
# must not have touched it.
assert items[2]["status"] == "in-progress"
assert items[2]["id"] == "epic-2-retro-item-1-y"
assert len(items) == 3
def test_set_action_status_changes_exactly_one_line(tmp_path):
# The status write is surgical: it must not re-style neighbouring lines, and
# an item whose status was quoted keeps its quoting. last_updated is rewritten
# with the same text it already held, so the whole file differs by one line.
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--date",
"01-01-2026 09:00",
"--set-action-status",
'[{"id":"epic-1-retro-item-1-x","status":"done"}]',
]
)
assert proc.returncode == 0, proc.stderr
before = ACTION_FIXTURE.splitlines()
after = target.read_text(encoding="utf-8").splitlines()
assert len(before) == len(after)
changed = [(b, a) for b, a in zip(before, after, strict=False) if b != a]
assert changed == [(' status: "open"', ' status: "done"')], changed
def test_set_action_status_composes_with_retro_done_and_add_action(tmp_path):
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-retro-done",
"--add-action",
'[{"action":"Brand new item","owner":"Amelia"}]',
"--set-action-status",
'[{"epic":1,"action":"Pre-existing item","status":"done"}]',
]
)
assert proc.returncode == 0, proc.stderr
out = _json(proc)
assert out["ok"] is True
assert out["retro_status_after"] == "done"
assert out["action_items_added"] == 1
assert out["action_items_updated"] == 1
data = _load(target)
assert data["development_status"]["epic-1-retrospective"] == "done"
items = data["action_items"]
assert len(items) == 4
assert items[1]["status"] == "done" # the targeted pre-existing item
assert items[3]["action"] == "Brand new item"
assert items[3]["status"] == "open" # the appended item is always open
def test_set_action_status_cannot_target_an_item_added_in_the_same_run(tmp_path):
# Selectors resolve against action_items as loaded, so the append cannot be
# observed by the same invocation. Silently succeeding here would make the
# flag a back door for writing a non-open status onto a brand-new item.
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--add-action",
'[{"action":"Brand new","owner":"A","id":"brand-new"}]',
"--set-action-status",
'[{"id":"brand-new","status":"done"}]',
]
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "no action item matches" in out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
def test_set_action_status_rejects_unknown_id(tmp_path):
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-action-status",
'[{"id":"not-in-the-file","status":"done"}]',
]
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "not-in-the-file" in out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
def test_set_action_status_rejects_selector_when_action_items_is_absent(tmp_path):
# No action_items key at all must read as "no match", not as a crash.
target = _write_fixture(tmp_path)
proc = _run(
["update", "--file", str(target), "--epic", "1", "--set-action-status", '[{"id":"anything","status":"done"}]']
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "no action item matches" in out["error"]
assert target.read_text(encoding="utf-8") == FIXTURE
def test_set_action_status_rejects_ambiguous_selector(tmp_path):
# Two legacy items with identical epic + action text: guessing between them
# would write the wrong row half the time.
fixture = (
"development_status:\n"
" 1-1-a: done\n"
" epic-1-retrospective: optional\n"
"action_items:\n"
" - epic: 1\n"
' action: "Same text"\n'
' owner: "Charlie"\n'
" status: open\n"
" - epic: 1\n"
' action: "Same text"\n'
' owner: "Dana"\n'
" status: open\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-action-status",
'[{"epic":1,"action":"Same text","status":"done"}]',
]
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "ambiguous" in out["error"]
assert "Same text" in out["error"]
assert "2 matches" in out["error"]
assert target.read_text(encoding="utf-8") == fixture
def test_set_action_status_rejects_two_entries_hitting_the_same_item(tmp_path):
# The id form and the epic/action form can name the same row; applying both
# would overcount action_items_updated and hide a conflicting pair.
target = _write_action_fixture(tmp_path)
payload = json.dumps(
[
{"id": "epic-1-retro-item-1-x", "status": "done"},
{"epic": 1, "action": "Scripted item", "status": "in-progress"},
]
)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-action-status", payload])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "same action item" in out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
def test_entry_with_both_selector_forms_uses_the_id(tmp_path):
# A caller that copied a whole item through supplies both. The id is the
# precise form and wins; the extra keys are ignored, not rejected. The two
# forms are pointed at *different* rows so the precedence is observable:
# the id names items[0], the epic/action pair names items[1].
target = _write_action_fixture(tmp_path)
payload = json.dumps(
[
{
"id": "epic-1-retro-item-1-x",
"epic": 1,
"action": "Pre-existing item",
"owner": "Charlie",
"status": "done",
}
]
)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-action-status", payload])
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 1
items = _load(target)["action_items"]
assert items[0]["status"] == "done" # the id's item
assert items[1]["status"] == "open" # the epic/action item, untouched
def test_set_action_status_rejects_invalid_status(tmp_path):
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-action-status",
'[{"id":"epic-1-retro-item-1-x","status":"closed"}]',
]
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "closed" in out["error"]
# The allowed vocabulary is named so the caller can correct the call.
assert "open, in-progress, done" in out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
def test_set_action_status_rejects_malformed_json(tmp_path):
target = _write_action_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-action-status", "{not json"])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "invalid --set-action-status JSON" in out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
_SELECTOR_SHAPE_ERROR = (
"each --set-action-status entry must have a non-empty string id, or an integer epic and a non-empty string action"
)
@pytest.mark.parametrize(
("payload", "expected_error"),
[
(
'{"id":"epic-1-retro-item-1-x","status":"done"}',
"--set-action-status must be a JSON array",
),
(
'["epic-1-retro-item-1-x"]',
"each --set-action-status entry must be an object",
),
('[{"status":"done"}]', _SELECTOR_SHAPE_ERROR),
(
'[{"id":"","status":"done"}]',
"each --set-action-status id must be a non-empty string",
),
(
'[{"id":42,"status":"done"}]',
"each --set-action-status id must be a non-empty string",
),
(
'[{"epic":"1","action":"Pre-existing item","status":"done"}]',
_SELECTOR_SHAPE_ERROR,
),
(
'[{"epic":true,"action":"Pre-existing item","status":"done"}]',
_SELECTOR_SHAPE_ERROR,
),
('[{"epic":1,"action":" ","status":"done"}]', _SELECTOR_SHAPE_ERROR),
(
# No status key at all: the status validator runs before the selector
# validator, so this is the status branch, not the selector branch.
'[{"epic":1,"action":"Pre-existing item"}]',
"invalid --set-action-status status None",
),
(
'[{"id":"epic-1-retro-item-1-x","status":3}]',
"invalid --set-action-status status 3",
),
],
ids=[
"not-a-list",
"entry-not-an-object",
"no-selector",
"empty-id",
"non-string-id",
"string-epic",
"bool-epic",
"blank-action",
"no-status-key",
"non-string-status",
],
)
def test_set_action_status_rejects_bad_shapes(tmp_path, payload, expected_error):
# Each case pins its own message: collapsing the branches into one generic
# error would leave a caller unable to tell which part of the array is wrong.
target = _write_action_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-action-status", payload])
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert expected_error in out["error"], out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
def test_selectors_are_not_scoped_to_the_epic_flag(tmp_path):
# The flag's headline use: epic 2's retro closing epic 1's items. --epic only
# names the retro key and stamps appended items; scoping selectors to it would
# silently break the documented cross-epic workflow while every same-epic test
# kept passing.
target = _write_action_fixture(tmp_path)
payload = json.dumps(
[
{"id": "epic-1-retro-item-1-x", "status": "done"},
{"epic": 1, "action": "Pre-existing item", "status": "done"},
]
)
proc = _run(["update", "--file", str(target), "--epic", "2", "--set-retro-done", "--set-action-status", payload])
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 2
data = _load(target)
assert data["development_status"]["epic-2-retrospective"] == "done"
items = data["action_items"]
assert items[0]["status"] == "done"
assert items[1]["status"] == "done"
# Epic 2's own item is not swept along.
assert items[2]["status"] == "in-progress"
def test_legacy_selector_discriminates_on_the_epic(tmp_path):
# Two items share the action text "Scripted item" and differ only by epic, so
# an epic-blind text match would be ambiguous -- or worse, silently pick one.
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-action-status",
'[{"epic":2,"action":"Scripted item","status":"done"}]',
]
)
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 1
items = _load(target)["action_items"]
assert items[2]["status"] == "done"
assert items[0]["status"] == "open" # the epic-1 namesake, untouched
def test_non_mapping_action_item_does_not_crash_the_selector(tmp_path):
# A hand-edited scalar in the action_items list must be skipped, not
# AttributeError'd into an empty stdout with a traceback.
fixture = (
"development_status:\n"
" 1-1-a: done\n"
" epic-1-retrospective: optional\n"
"action_items:\n"
' - "a bare string someone hand-edited in"\n'
' - id: "real"\n'
" epic: 1\n"
' action: "Real item"\n'
" status: open\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(
["update", "--file", str(target), "--epic", "1", "--set-action-status", '[{"id":"real","status":"done"}]']
)
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 1
items = _load(target)["action_items"]
assert items[0] == "a bare string someone hand-edited in"
assert items[1]["status"] == "done"
def test_non_mapping_action_item_stays_on_the_json_contract_when_unmatched(tmp_path):
# Same guard, reject path: the scalar must not be dereferenced while looking
# for a selector that is not there.
fixture = "development_status:\n 1-1-a: done\naction_items:\n - 42\n"
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-action-status",
'[{"epic":1,"action":"Nothing","status":"done"}]',
]
)
assert proc.returncode == 1
out = _json(proc) # asserts stdout is JSON and stderr carries no traceback
assert out["ok"] is False
assert out["restored"] is True
assert "no action item matches" in out["error"]
assert target.read_text(encoding="utf-8") == fixture
def test_boolean_epic_in_the_file_does_not_match_epic_one(tmp_path):
# True == 1 in Python, so without the file-side bool guard a hand-edited
# "epic: true" row would be silently rewritten by a selector aimed at epic 1.
fixture = (
"development_status:\n"
" 1-1-a: done\n"
"action_items:\n"
" - epic: true\n"
' action: "Boolean epic"\n'
" status: open\n"
)
target = tmp_path / "sprint-status.yaml"
target.write_text(fixture, encoding="utf-8")
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--set-action-status",
'[{"epic":1,"action":"Boolean epic","status":"done"}]',
]
)
assert proc.returncode == 1
out = _json(proc)
assert out["ok"] is False
assert out["restored"] is True
assert "no action item matches" in out["error"]
assert target.read_text(encoding="utf-8") == fixture
def test_status_write_preserves_every_scalar_style(tmp_path):
# The status write must land on the line without re-styling it, whichever way
# the file spells the scalar.
target = tmp_path / "sprint-status.yaml"
target.write_text(STYLE_FIXTURE, encoding="utf-8")
payload = json.dumps(
[
{"id": "double", "status": "done"},
{"id": "single", "status": "done"},
{"id": "plain", "status": "done"},
]
)
proc = _run(
["update", "--file", str(target), "--epic", "1", "--date", "01-01-2026 09:00", "--set-action-status", payload]
)
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 3
before = STYLE_FIXTURE.splitlines()
after = target.read_text(encoding="utf-8").splitlines()
assert len(before) == len(after)
changed = [(b, a) for b, a in zip(before, after, strict=False) if b != a]
assert changed == [
(' status: "open"', ' status: "done"'),
(" status: 'open'", " status: 'done'"),
(" status: open", " status: done"),
], changed
def test_action_status_vocabulary_is_exactly_the_three():
# bmad-sprint-planning is the authority. Widening this tuple would let the
# script write a value sprint-planning's status view reports as illegal.
assert _module().ACTION_STATUSES == ("open", "in-progress", "done")
def test_empty_status_array_is_a_no_op(tmp_path):
target = _write_action_fixture(tmp_path)
proc = _run(
["update", "--file", str(target), "--epic", "1", "--date", "01-01-2026 09:00", "--set-action-status", "[]"]
)
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 0
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
def test_in_progress_item_transitions_to_done(tmp_path):
# Every other success case starts from "open"; the middle of the lifecycle
# has to work too.
target = _write_action_fixture(tmp_path)
proc = _run(
[
"update",
"--file",
str(target),
"--epic",
"2",
"--set-action-status",
'[{"id":"epic-2-retro-item-1-y","status":"done"}]',
]
)
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 1
assert _load(target)["action_items"][2]["status"] == "done"
def test_action_items_updated_is_always_reported(tmp_path):
# Consumers read the counter unconditionally, so it must be present even when
# the flag was not passed.
target = _write_fixture(tmp_path)
proc = _run(["update", "--file", str(target), "--epic", "1", "--set-retro-done"])
assert proc.returncode == 0, proc.stderr
assert _json(proc)["action_items_updated"] == 0
def test_post_write_status_mismatch_restores(tmp_path, monkeypatch, capsys):
# The last line of defence: the written file is re-parsed and every targeted
# item is checked. No CLI path can fake a mismatch, so the re-parse is
# doctored directly.
mod = _module()
target = _write_action_fixture(tmp_path)
real_load_yaml = mod._load_yaml
calls = {"n": 0}
def flaky(path):
yaml, data = real_load_yaml(path)
calls["n"] += 1
if calls["n"] == 2: # the post-write re-parse
data["action_items"][0]["status"] = "open"
return yaml, data
monkeypatch.setattr(mod, "_load_yaml", flaky)
args = mod.build_parser().parse_args(
[
"update",
"--file",
str(target),
"--epic",
"1",
"--date",
"01-01-2026 09:00",
"--set-action-status",
'[{"id":"epic-1-retro-item-1-x","status":"done"}]',
]
)
with pytest.raises(SystemExit) as excinfo:
mod.cmd_update(args)
assert excinfo.value.code == 1
out = json.loads(capsys.readouterr().out)
assert out["ok"] is False
assert out["restored"] is True
assert "after write" in out["error"]
assert target.read_text(encoding="utf-8") == ACTION_FIXTURE
if __name__ == "__main__":
sys.exit(pytest.main([__file__, "-q"]))
SKILL.md
---
name: bmad-retrospective
description: 'Review a completed epic against the evidence it left behind — spec, stories, diffs, commits, sprint status — and produce a retrospective with sourced findings, action items, and an acceptance decision. Use when the user says "run a retrospective" or "lets retro the epic [epic]". Supports -H/--headless'
---
# Retrospective
Review a completed epic by reading the evidence it left — the epic spec, story files, the full diff, per-story commits, sprint status, and session logs when they exist. An unattended epic run leaves a record; this skill reads that record, surfaces the defects no single story could show, and judges the epic against the criteria it set for itself.
Every finding you report carries a source reference (file, line, commit, or log). A claim you cannot point at — an invented root cause, a pattern the diff does not actually show — is not a finding. Drop it.
## Resolution rules
- Bare paths and `{skill-root}` (e.g. `references/aggregate-views.md`, `scripts/sprint_status.py`) resolve from this skill's installed directory.
- `{project-root}` → the project working directory.
- `{skill-name}` → the skill directory's basename.
## Modes
Interactive by default. With `-H`/`--headless`: skip every confirmation, take the epic from the invocation (falling back to detection only if none was supplied), never open the team discussion, render the verdict on the evidence alone, and record each assumption made without the user (which epic was selected, the machine verdict, each proposed item) into the retrospective document's Assumptions section so the audit trail survives. The Phase 4 acceptance fail-safe still applies in headless runs.
For automation, `-H <epic>` — an explicit epic in headless mode — is the stable orchestrator-facing interface. Pass the same number to `detect-epic --epic <N>` so the unfinished-story gate is script-backed (see Inputs). Epic auto-detection is a human convenience, not an automation contract: unflagged `detect-epic` returns the highest epic with *any* `done` story, and stories-mode projects have no `sprint-status.yaml` to detect from.
## On Activation
Run these in order before the retrospective begins:
1. **Resolve the workflow block.** Run `uv run --no-cache {project-root}/_bmad/scripts/resolve_customization.py --skill {skill-root} --project-root {project-root} --key workflow`. If it fails, resolve `{workflow.*}` yourself by reading `{skill-root}/customize.toml`, then `{project-root}/_bmad/custom/{skill-name}.toml`, then `.user.toml` in that order, merging base → team → user (scalars override, keyed arrays-of-tables merge by `code`/`id`, other arrays append).
2. **Run prepend steps** — execute each entry in `{workflow.activation_steps_prepend}` in order.
3. **Load persistent facts** — treat every `{workflow.persistent_facts}` entry as standing context. `file:` entries are paths/globs under `{project-root}` whose contents load as facts; all others are literal facts.
4. **Resolve config.** Run `uv run {project-root}/_bmad/scripts/resolve_config.py --project-root {project-root} --key core.project_name --key core.output_folder --key modules.bmm.planning_artifacts --key modules.bmm.implementation_artifacts`. `{date}` is the current system datetime. Never state time estimates — AI has changed development speed, so hour/day/week predictions are noise.
5. **Greet and orient** (interactive only). Greet the user, name the epic you are about to retro, and optionally invite their going-in concerns ("anything you want weighted — a story that felt rushed, a risky interaction between two stories?"). Use any answer to focus the Phase 1–2 analysis; it directs attention but never becomes a finding without a source.
6. **Run append steps** — execute each entry in `{workflow.activation_steps_append}` in order.
## Inputs
| Input | Where | Use |
|-------|-------|-----|
| epic | invocation argument, or detected from sprint status | which epic to retro |
| spec folder | invocation argument, or found under the spec roots | the stories-mode epic: `SPEC.md`, ordered `stories.yaml`, `stories/<id>-*.md` |
| sprint status | `{implementation_artifacts}/sprint-status.yaml` | epic detection + final status update |
| architecture / prd | `{planning_artifacts}/*architecture*`, `*prd*` | context for judging as-built vs intended |
| previous retro (optional) | `{implementation_artifacts}/**/epic-{{prev}}-retro-*.md` | check whether last epic's actions landed |
| session logs (optional) | conversation/session records for the epic's stories | process lessons; record the gap when absent |
An epic reaches this skill in one of two shapes, and they are peers. **Sprint mode** reads `sprint-status.yaml`. **Stories mode** reads a spec folder holding `SPEC.md`, an ordered `stories.yaml`, and `stories/<id>-*.md` artifacts — the shape an unattended run leaves behind. Resolve which applies first: a named folder is stories mode whether or not sprint status exists; a named epic number is sprint mode; with neither, use sprint mode when `sprint-status.yaml` exists, and otherwise look for spec folders under `{output_folder}/specs`, `{planning_artifacts}`, and `{implementation_artifacts}`. Ask the user which to retro when there is more than one, and never choose silently; headless, stop and require an explicit folder.
In stories mode, `stories.yaml` in list order is the story list — list order is authoritative, filename sort is not — and each story's `stories/<id>-*.md` frontmatter carries its `status`. `pending_stories` is the ids whose status is not `done`; apply the same completeness gate as below. Then skip to Phase 1: do not read or write sprint status for the rest of the run. The rest of this section is sprint mode.
Determine the epic and its unfinished-story list from `sprint_status.py detect-epic` whenever `{implementation_artifacts}/sprint-status.yaml` is available:
- **Epic supplied** (including the stable `-H <epic>` orchestrator path): run `uv run --no-cache {skill-root}/scripts/sprint_status.py detect-epic --file {implementation_artifacts}/sprint-status.yaml --epic <N>`. The script scopes `pending_stories` to that number even when auto-detect would have picked a different epic, and even when the epic has no `done` story yet. `story_count` is that same scoped count of the epic's story keys: `0` means the file has no such epic at all — a nonexistent epic returns the same empty `pending_stories` as a finished one, so treat `story_count: 0` as a likely mistyped epic number, confirm with the user, and headless, stop and report rather than proceeding.
- **No epic supplied**: run the same command without `--epic` (returns the highest epic with a `done` story). Confirm the detected epic with the user and let them override; in headless mode accept it and record the assumption. If detection returns none, ask the user — or, headless, stop and report.
If the script exits non-zero it emits `{"ok": false, "error": ...}` instead of a detection — the normal path for a stories-mode project with no `sprint-status.yaml`, and for a file that does not parse: surface that error verbatim — or, if the script produced no JSON at all, whatever it wrote to stderr — and ask the user which epic to retro; headless, stop and report. Without a readable sprint-status file there is no `pending_stories` list; record that the completeness check did not run and continue only if the user (or headless Assumptions trail) accepts proceeding without it.
Then check the epic is actually finished before Phase 1. A successful detect carries `pending_stories` — the selected epic's story keys whose status is not `done`, in file order, scoped to that epic alone (an unfinished story in some *other* epic is out of scope for this retrospective). When the list is non-empty, interactively list those stories and ask whether to retro an unfinished epic: if the user declines, stop and report — do not enter Phase 1; if they accept, record the stories they accepted proceeding over in the document's Epic summary. Headless, proceed and record the same list in the Assumptions section — do not invent a confirmation. Either way the list sits in the document, and Phase 4's machine verdict is **rejected** when any story remained unfinished (see `references/acceptance-verdict.md`); a human may override interactively.
## Working state and resumption
The retrospective document is the working artifact, not only the final output. Once the epic is fixed, create it as a skeleton (`references/retro-document.md` names the sections) and write each phase's result into it as you finish — inventory, then findings with sources, then dispositions and verdict. Continuity is re-reading the file.
If a retrospective document for this epic already exists, load it, reconcile its recorded state against the current evidence — the current evidence wins, since commits may have landed and questions may have been answered since — and resume at the first incomplete phase instead of redoing finished ones. In stories mode that document is `{spec-folder}/RETROSPECTIVE.md`, a fixed name so a resumed run finds it; sprint mode keeps its dated `{implementation_artifacts}` filename.
## Flow
Run the phases in order. A default run stops at a written evidence report and verdict; the team discussion in Phase 3 is opt-in.
### Phase 1 — Gather
Enumerate what the epic actually produced and record what is missing. Load `references/evidence-gathering.md` for the inventory checklist, the `git_evidence.py` pre-pass that derives the diff range and per-story commits, and the missing-evidence rule: each later analysis declares what it needs and records a narrowed scope when the evidence is absent, so a reader can always tell "checked and clean" from "never checked."
### Phase 2 — Analyze
Produce findings, each with a source reference, from three angles:
- **Aggregate views** — the defects no single diff hunk shows: architecture delta, duplication map, god-class growth, pattern divergence, spec-to-implementation reconciliation. Load `references/aggregate-views.md` for the catalog and how to derive each (deterministic scripts first).
- **Diff-scope review** — do not reimplement review. Invoke **`bmad-review`** on the epic's diff for the code lenses (adversarial, edge-case, verification-gap), weighting the boundaries between stories, where no single session ever saw both sides. Fold its findings in. If `bmad-review` is unavailable, run those lenses inline over the diff on a narrowed scope and record the narrowing.
- **Behavior check (when the epic changed runtime behavior)** — exercise the changed flows end to end and record what you observed. Passing tests do not substitute for running the system.
Consolidate: merge, dedupe, and provenance-link findings. Drop any finding you cannot tie to a source.
### Phase 3 — Team Discussion (opt-in)
Skip by default; never runs headless. When the user asks to "discuss it as a team," "run party mode," or similar, invoke the skill `bmad-party-mode` seeded with the Phase 2 findings so the installed agents react to real evidence — the god class the diff really grew, the verification gap that is actually there, the wins the evidence confirms. Load `references/team-discussion.md` for how to seed it and keep it grounded. If `bmad-party-mode` is unavailable, run the discussion inline over the Phase 2 findings and record the narrowing. The rule: agents speak only to findings with sources.
### Phase 4 — Decide
- **Action items** — compile fix-now findings and process lessons into specific, owned action items. Fixes and spec reconciliations are *proposed here*, not auto-applied; the human decides what to execute.
- **Acceptance verdict** — judge the final state against the epic's declared acceptance criteria (profile it from the diff and stories if none were declared): **accepted**, **accepted-with-open-items**, or **rejected** — one spelling, everywhere a machine reads it. Unfinished stories in `pending_stories` force the machine verdict to **rejected**. A human decision always overrides. An epic that fails its criteria with no human decision is recorded as *not accepted* — never as silently accepted. Load `references/acceptance-verdict.md` for the rubric, the finding-routing dispositions, and the previous-retro follow-through record — the per-item evidence Phase 5's status offer reads.
### Phase 5 — Finalize
Finalize the retrospective document and update sprint status. Load `references/retro-document.md` for the document's sections and the exact `sprint_status.py update` invocation that marks the retro key `done`, appends the action items, and validates the write. Where the Phase 4 follow-through has evidence a *previous* epic's action item landed, offer `--set-action-status` and pass only the transitions the user confirms — the evidence justifies proposing a transition, and only the user's confirmation justifies writing it; a headless run records the transitions it would have proposed and does not pass the flag at all. In stories mode, finalize `{spec-folder}/RETROSPECTIVE.md` and stop there: no `sprint_status.py` call, no sprint-status file created, and no edits to `SPEC.md`, `stories.yaml`, or any story artifact. Then, if `{workflow.on_complete}` is non-empty, follow it as the final instruction.