agents/openai.yaml
interface:
display_name: "AI Coding Metrics for Engineering Teams"
short_description: "Measures AI coding impact, extension robustness, quality, delivery, and cost"
default_prompt: "Use $dev-ai-coding-metrics to measure AI coding impact, extension robustness, delivery, quality trajectories, cost, and developer experience for pilots, ROI scorecards, and leadership reports."
assets/adoption-survey-template.md
# AI Coding Tools Adoption Survey
**Last Updated**: {{DATE}}
**Owner**: {{NAME}}
**Version**: {{VERSION}}
---
Purpose: measure developer adoption, satisfaction, and friction with AI coding tools. Use as a copy-paste template for your survey platform (Google Forms, Typeform, SurveyMonkey, etc.).
## Administration Instructions
- **Cadence**: run quarterly, at minimum. Run an additional pulse after major tool changes.
- **Response rate target**: >70% of licensed developers. Below 50% makes results unreliable.
- **Anonymity**: responses must be anonymous. Collect only team/role demographics, never names.
- **Distribution**: send via engineering-wide channel. Send 2 reminders (day 3, day 6 of a 10-day window).
- **Time to complete**: target under 8 minutes. Pre-test with 3 developers and adjust.
- **Results sharing**: publish summary to all respondents within 2 weeks of close.
---
## Section 1: Usage (Q1-Q3)
### Q1. How often do you use AI coding tools in your work?
- Response type: **Single choice**
- Options:
- Multiple times per day
- About once per day
- A few times per week
- Rarely (less than once per week)
- Never
- Benchmark: healthy adoption = >60% daily users after 3 months
### Q2. Which AI coding tools do you actively use? (Select all that apply)
- Response type: **Multiple choice**
- Options:
- {{TOOL_1}} (e.g., GitHub Copilot)
- {{TOOL_2}} (e.g., Cursor)
- {{TOOL_3}} (e.g., Claude Code)
- {{TOOL_4}} (e.g., ChatGPT)
- Other: {{FREE_TEXT}}
- None
- Benchmark: track tool distribution shifts quarter over quarter
### Q3. What are your primary use cases? (Select top 3)
- Response type: **Multiple choice (max 3)**
- Options:
- Code completion / autocomplete
- Writing new functions or modules
- Writing tests
- Debugging / fixing errors
- Code review assistance
- Refactoring existing code
- Documentation and comments
- Learning new APIs or languages
- Explaining unfamiliar code
- Other: {{FREE_TEXT}}
- Benchmark: track shifts in use-case mix over time
---
## Section 2: Effectiveness (Q4-Q7)
### Q4. How has AI tooling affected your personal productivity?
- Response type: **Single choice**
- Options:
- Significantly more productive (>20% time savings)
- Somewhat more productive (5-20% time savings)
- No noticeable change
- Somewhat less productive (net time lost on corrections)
- Significantly less productive
- Benchmark: >70% reporting "somewhat" or "significantly" more productive
### Q5. How has AI tooling affected the quality of code you produce?
- Response type: **Single choice**
- Options:
- Noticeably higher quality
- Slightly higher quality
- No change
- Slightly lower quality
- Noticeably lower quality
- Benchmark: <10% reporting lower quality
### Q6. Estimate hours saved per week due to AI coding tools.
- Response type: **Single choice**
- Options:
- 0 hours (no savings)
- 1-2 hours
- 3-4 hours
- 5-6 hours
- 7-8 hours
- More than 8 hours
- Benchmark: median of 3-5 hours at mature adoption
### Q7. For which tasks do AI tools help most vs least?
- Response type: **Two free-text fields**
- "AI tools help most with:" {{FREE_TEXT}}
- "AI tools help least with:" {{FREE_TEXT}}
- Benchmark: qualitative; code for recurring themes
---
## Section 3: Experience (Q8-Q11)
### Q8. Overall, how satisfied are you with the AI coding tools provided?
- Response type: **Likert scale (1-5)**
- 1 = Very dissatisfied
- 2 = Dissatisfied
- 3 = Neutral
- 4 = Satisfied
- 5 = Very satisfied
- Benchmark: target mean >= 3.8
### Q9. How much do you trust the output of AI coding tools?
- Response type: **Likert scale (1-5)**
- 1 = Do not trust at all; review everything line by line
- 2 = Low trust; accept only trivial suggestions
- 3 = Moderate trust; accept after a quick scan
- 4 = High trust; accept most suggestions for familiar code
- 5 = Very high trust; rarely need to modify
- Benchmark: healthy range is 3-4; scores of 5 may indicate under-review risk
### Q10. What frustrates you most about AI coding tools? (Select top 2)
- Response type: **Multiple choice (max 2)**
- Options:
- Irrelevant or wrong suggestions
- Slow response time
- Breaks my flow / context switching
- Security / IP concerns
- Inconsistent quality across languages
- Hard to customize or configure
- Nothing; no frustrations
- Other: {{FREE_TEXT}}
- Benchmark: track top frustration shifts quarter over quarter
### Q11. Does using AI coding tools increase or decrease your cognitive load?
- Response type: **Single choice**
- Options:
- Significantly decreases cognitive load
- Somewhat decreases cognitive load
- No change
- Somewhat increases cognitive load (evaluating suggestions is tiring)
- Significantly increases cognitive load
- Benchmark: <15% reporting increased cognitive load
---
## Section 4: Barriers (Q12-Q14)
### Q12. What prevents you from using AI coding tools more? (Select all that apply)
- Response type: **Multiple choice**
- Options:
- Nothing; I use them as much as I want
- Doesn't work well with my language/framework
- Concerns about code quality/correctness
- Security or compliance restrictions
- Lack of training or onboarding
- Tool is too slow or unreliable
- My workflow doesn't benefit from it
- Team norms discourage use
- Other: {{FREE_TEXT}}
- Benchmark: identify top 3 barriers; track removal progress
### Q13. What would help you get more value from AI coding tools?
- Response type: **Multiple choice (select top 2)**
- Options:
- Better onboarding / training
- Prompt engineering tips and examples
- More context-aware suggestions (repo-level)
- Integration with our specific tools/frameworks
- Faster response times
- Clearer security/compliance guidelines
- Team sharing of effective patterns
- Other: {{FREE_TEXT}}
- Benchmark: feed into training and enablement roadmap
### Q14. Have you received adequate training on AI coding tools?
- Response type: **Single choice**
- Options:
- Yes, comprehensive training
- Yes, basic training (enough to get started)
- Self-taught only
- No training received
- Benchmark: >80% should have at least basic training
---
## Section 5: Open-Ended (Q15)
### Q15. Any other feedback, suggestions, or concerns about AI coding tools?
- Response type: **Free text (optional)**
- {{FREE_TEXT}}
- Analysis: code responses into themes; report top 5 themes with representative quotes
---
## Scoring Methodology
### Aggregate Adoption Score (0-100)
Calculate a weighted composite:
| Component | Weight | Calculation |
|-----------|--------|-------------|
| Usage frequency (Q1) | 30% | Map to 0-100: Never=0, Rarely=25, Few/week=50, Daily=75, Multi-daily=100 |
| Productivity impact (Q4) | 25% | Map to 0-100: Sig. less=0, Somewhat less=25, No change=50, Somewhat more=75, Sig. more=100 |
| Satisfaction (Q8) | 25% | Likert 1-5 mapped to 0-100 |
| Trust (Q9) | 20% | Likert 1-5 mapped to 0-100 |
**Aggregate Score** = (Q1_score x 0.30) + (Q4_score x 0.25) + (Q8_score x 0.25) + (Q9_score x 0.20)
> The items here are unvalidated single measures, and this weighted composite has not been tested for discriminant validity, convergent validity, or reliability. Treat the score as a directional tracking number, not a validated construct — see `software-ux-research/references/survey-design-guide.md` → *Construct Validity* before drawing analytical conclusions from it.
### Interpretation
| Score Range | Label | Action |
|-------------|-------|--------|
| 80-100 | Strong adoption | Maintain; share best practices |
| 60-79 | Healthy adoption | Minor improvements; address top barrier |
| 40-59 | Developing | Targeted training; investigate friction |
| 20-39 | Struggling | Intervention needed; exec sponsorship |
| 0-19 | Not adopted | Reassess tool fit and rollout strategy |
---
## Quarter-over-Quarter Tracking
| Metric | Q{{N-2}} | Q{{N-1}} | Q{{N}} | Trend |
|--------|----------|----------|--------|-------|
| Response Rate (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{TREND}} |
| Aggregate Adoption Score | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{TREND}} |
| Mean Satisfaction (Q8) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{TREND}} |
| Median Hours Saved (Q6) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{TREND}} |
| Top Barrier (Q12) | {{VALUE}} | {{VALUE}} | {{VALUE}} | — |
assets/executive-report-template.md
# AI Coding Metrics Executive Report Templates
**Last Updated**: {{DATE}}
**Owner**: {{NAME}}
**Version**: {{VERSION}}
---
Purpose: two ready-to-fill report templates for communicating AI coding program results to leadership. Template A is a monthly one-pager. Template B is a quarterly deep-dive.
## How to Use
1. Pick the template matching your reporting cadence (monthly or quarterly).
2. Replace all `{{PLACEHOLDER}}` values with actuals.
3. Delete any sections not relevant to your organization.
4. Keep the report under 1 page (Template A) or 4-6 pages (Template B).
5. State the program mode up front: `assistant`, `agent`, or `mixed`.
---
## Template A: Monthly One-Page Summary
```
# AI Coding Program — {{MONTH}} {{YEAR}} Report
Program mode: {{ASSISTANT / AGENT / MIXED}}
## Status: {{GREEN / AMBER / RED}}
### Key Metrics (vs Prior Month)
| Metric | Current | Prior | Delta | Target | Status |
|-----------------------------|------------|-----------|-----------|-----------|-----------------|
| Adoption Rate | {{xx%}} | {{xx%}} | {{+x%}} | {{xx%}} | {{GREEN/AMBER/RED}} |
| Developer Satisfaction | {{x.x/5}} | {{x.x/5}} | {{+x.x}} | {{x.x/5}} | {{GREEN/AMBER/RED}} |
| Avg Hours Saved/Dev/Week | {{x.x}} | {{x.x}} | {{+x.x}} | {{x.x}} | {{GREEN/AMBER/RED}} |
| Deployment Frequency | {{x/wk}} | {{x/wk}} | {{+x}} | {{x/wk}} | {{GREEN/AMBER/RED}} |
| Quality Score (defect rate) | {{x.x}} | {{x.x}} | {{-x.x}} | {{x.x}} | {{GREEN/AMBER/RED}} |
| Monthly ROI | {{xx%}} | {{xx%}} | {{+x%}} | {{xx%}} | {{GREEN/AMBER/RED}} |
{{OPTIONAL: keep the next 3 rows only for agent or mixed programs}}
| Agent Task Completion Rate | {{xx%}} | {{xx%}} | {{+x%}} | {{xx%}} | {{GREEN/AMBER/RED}} |
| Agent PR Merge Rate | {{xx%}} | {{xx%}} | {{+x%}} | {{xx%}} | {{GREEN/AMBER/RED}} |
| Human Takeover Rate | {{xx%}} | {{xx%}} | {{-x%}} | {{xx%}} | {{GREEN/AMBER/RED}} |
### Highlights
- {{HIGHLIGHT_1}}
- {{HIGHLIGHT_2}}
- {{HIGHLIGHT_3}}
### Concerns
- {{CONCERN_1}}
- {{CONCERN_2}}
### Actions Completed This Month
- {{ACTION_COMPLETED_1}}
- {{ACTION_COMPLETED_2}}
### Next Month Focus
- {{ACTION_PLANNED_1}}
- {{ACTION_PLANNED_2}}
### Budget
| Item | Monthly Spend | YTD Spend | Annual Budget | % Used |
|----------------|--------------|-----------|---------------|--------|
| Tool licenses | ${{VALUE}} | ${{VALUE}} | ${{VALUE}} | {{x%}} |
| Training | ${{VALUE}} | ${{VALUE}} | ${{VALUE}} | {{x%}} |
| Other | ${{VALUE}} | ${{VALUE}} | ${{VALUE}} | {{x%}} |
| **Total** | ${{VALUE}} | ${{VALUE}} | ${{VALUE}} | {{x%}} |
Prepared by: {{AUTHOR}} | Distribution: {{AUDIENCE}}
```
---
## Template B: Quarterly Deep-Dive
### 1. Executive Summary
{{QUARTER}} {{YEAR}} — one paragraph summarizing overall program health, biggest wins, top risks, and the single most important recommendation.
> {{EXECUTIVE_SUMMARY_TEXT}}
Overall status: **{{GREEN / AMBER / RED}}**
Program mode: **{{ASSISTANT / AGENT / MIXED}}**
---
### 2. Adoption Progress
| Metric | Q{{N-1}} | Q{{N}} | Delta | Target | Status |
|--------|----------|--------|-------|--------|--------|
| Licensed seats | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} | {{STATUS}} |
| Active daily users | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} | {{STATUS}} |
| Adoption rate (%) | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} | {{STATUS}} |
| Teams fully onboarded | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} | {{STATUS}} |
**Adoption by team**:
| Team | Adoption Rate | Trend | Notes |
|------|--------------|-------|-------|
| {{TEAM_1}} | {{xx%}} | {{UP/FLAT/DOWN}} | {{NOTE}} |
| {{TEAM_2}} | {{xx%}} | {{UP/FLAT/DOWN}} | {{NOTE}} |
| {{TEAM_3}} | {{xx%}} | {{UP/FLAT/DOWN}} | {{NOTE}} |
Chart description: {{DESCRIBE_ADOPTION_TREND_CHART — e.g., bar chart showing weekly active users over 12 weeks}}
---
### 3. Productivity Impact
#### DORA Metrics
| Metric | Q{{N-1}} | Q{{N}} | Delta | Industry Benchmark |
|--------|----------|--------|-------|--------------------|
| Deployment Frequency | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{BENCHMARK}} |
| Lead Time for Changes | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{BENCHMARK}} |
| Change Failure Rate | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{BENCHMARK}} |
| Mean Time to Recovery | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{BENCHMARK}} |
#### SPACE Metrics
| Dimension | Metric | Q{{N-1}} | Q{{N}} | Delta |
|-----------|--------|----------|--------|-------|
| Satisfaction | Developer satisfaction score | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Performance | Cycle time (days) | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Activity | PRs merged per dev per week | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Communication | PR review turnaround (hours) | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Efficiency | Hours saved per dev per week | {{VALUE}} | {{VALUE}} | {{DELTA}} |
---
### 4. Quality Trends
| Metric | Q{{N-1}} | Q{{N}} | Delta | Target |
|--------|----------|--------|-------|--------|
| Defect escape rate | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} |
| Avg cyclomatic complexity (new code) | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} |
| Test coverage (%) | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} |
| Code review rejection rate (%) | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} |
| Security vulnerabilities introduced | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{TARGET}} |
Observations: {{QUALITY_COMMENTARY — e.g., "Defect rate dropped 12% while velocity increased, suggesting AI tools are not trading quality for speed."}}
---
### 5. ROI Analysis
| Component | Q{{N}} Value |
|-----------|-------------|
| Total tool + training cost | ${{VALUE}} |
| Time savings value | ${{VALUE}} |
| Quality savings value | ${{VALUE}} |
| Retention savings value | ${{VALUE}} |
| **Net benefit** | **${{VALUE}}** |
| **ROI (%)** | **{{VALUE}}%** |
| Cumulative ROI (program-to-date) | {{VALUE}}% |
See roi-calculator-template for full methodology.
---
### 5A. Agent Execution (Optional for Agent or Mixed Programs)
| Metric | Q{{N-1}} | Q{{N}} | Delta | Notes |
|--------|----------|--------|-------|-------|
| Task completion rate | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{NOTE}} |
| Acceptance-for-review rate | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{NOTE}} |
| PR merge rate | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{NOTE}} |
| Human takeover rate | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{NOTE}} |
| Reviewer effort per accepted PR | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{NOTE}} |
| Revert rate on merged agent PRs | {{VALUE}} | {{VALUE}} | {{DELTA}} | {{NOTE}} |
---
### 6. Developer Experience
Survey results summary (from adoption-survey-template):
| Metric | Q{{N-1}} | Q{{N}} | Delta |
|--------|----------|--------|-------|
| Survey response rate (%) | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Aggregate adoption score (0-100) | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Mean satisfaction (1-5) | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Mean trust (1-5) | {{VALUE}} | {{VALUE}} | {{DELTA}} |
| Median hours saved / week | {{VALUE}} | {{VALUE}} | {{DELTA}} |
**Top 3 benefits reported**: {{BENEFIT_1}}, {{BENEFIT_2}}, {{BENEFIT_3}}
**Top 3 barriers reported**: {{BARRIER_1}}, {{BARRIER_2}}, {{BARRIER_3}}
---
### 7. Benchmarking
| Metric | Our Org | Industry Median | Industry Top Quartile | Gap |
|--------|---------|-----------------|----------------------|-----|
| Adoption rate | {{VALUE}} | {{BENCHMARK}} | {{BENCHMARK}} | {{GAP}} |
| Hours saved / dev / week | {{VALUE}} | {{BENCHMARK}} | {{BENCHMARK}} | {{GAP}} |
| Developer satisfaction | {{VALUE}} | {{BENCHMARK}} | {{BENCHMARK}} | {{GAP}} |
| ROI (%) | {{VALUE}} | {{BENCHMARK}} | {{BENCHMARK}} | {{GAP}} |
Sources: {{BENCHMARK_SOURCES — e.g., DORA 2025 AI-assisted software development, internal data, benchmark results used only as capability context}}
---
### 8. Risk Register
| # | Risk | Likelihood | Impact | Mitigation | Owner | Status |
|---|------|-----------|--------|------------|-------|--------|
| 1 | {{RISK_1}} | {{H/M/L}} | {{H/M/L}} | {{MITIGATION}} | {{OWNER}} | {{OPEN/MITIGATED/CLOSED}} |
| 2 | {{RISK_2}} | {{H/M/L}} | {{H/M/L}} | {{MITIGATION}} | {{OWNER}} | {{OPEN/MITIGATED/CLOSED}} |
| 3 | {{RISK_3}} | {{H/M/L}} | {{H/M/L}} | {{MITIGATION}} | {{OWNER}} | {{OPEN/MITIGATED/CLOSED}} |
| 4 | {{RISK_4}} | {{H/M/L}} | {{H/M/L}} | {{MITIGATION}} | {{OWNER}} | {{OPEN/MITIGATED/CLOSED}} |
---
### 9. Recommendations and Next Quarter Plan
#### Recommendations
| # | Recommendation | Expected Impact | Effort | Priority |
|---|---------------|-----------------|--------|----------|
| 1 | {{RECOMMENDATION_1}} | {{IMPACT}} | {{EFFORT}} | {{P1/P2/P3}} |
| 2 | {{RECOMMENDATION_2}} | {{IMPACT}} | {{EFFORT}} | {{P1/P2/P3}} |
| 3 | {{RECOMMENDATION_3}} | {{IMPACT}} | {{EFFORT}} | {{P1/P2/P3}} |
#### Next Quarter OKRs
| Objective | Key Result | Target | Owner |
|-----------|-----------|--------|-------|
| {{OBJECTIVE_1}} | {{KR_1}} | {{TARGET}} | {{OWNER}} |
| {{OBJECTIVE_1}} | {{KR_2}} | {{TARGET}} | {{OWNER}} |
| {{OBJECTIVE_2}} | {{KR_3}} | {{TARGET}} | {{OWNER}} |
| {{OBJECTIVE_2}} | {{KR_4}} | {{TARGET}} | {{OWNER}} |
---
Prepared by: {{AUTHOR}}
Reviewed by: {{REVIEWER}}
Distribution: {{AUDIENCE}}
Next report due: {{DATE}}
assets/experiment-design-template.md
# AI Coding Impact Experiment Design Template
**Last Updated**: {{DATE}}
**Owner**: {{NAME}}
**Version**: {{VERSION}}
---
Purpose: plan a controlled experiment to measure the causal impact of AI coding tools on developer productivity, quality, or satisfaction. Fill in all sections before starting the experiment. This template ensures statistical rigor and ethical handling of developer data.
## How to Use
1. Complete the **Experiment Metadata** section to define what you are testing.
2. Define **Population** with treatment and control groups.
3. Choose a **Design** type and document the approach.
4. Build the **Measurement Plan** table with baseline values.
5. Pre-register the **Analysis Plan** before collecting data (prevents p-hacking).
6. Review **Threats to Validity** and document mitigations.
7. Get sign-off on **Ethics and Communication** before launch.
---
## Experiment Metadata
| Field | Value |
|-------|-------|
| Experiment name | {{EXPERIMENT_NAME}} |
| Program mode | {{ASSISTANT / AGENT / MIXED}} |
| Hypothesis | {{SPECIFIC_TESTABLE_HYPOTHESIS — e.g., "Developers using AI coding tools will have 15% shorter cycle times compared to those without, measured over 10 weeks."}} |
| Primary metric | {{PRIMARY_METRIC — e.g., cycle time in hours / agent task completion rate / reviewer effort per accepted PR}} |
| Secondary metrics | {{METRIC_1}}, {{METRIC_2}}, {{METRIC_3}} |
| Duration | {{WEEKS}} weeks (minimum 8 recommended) |
| Start date | {{START_DATE}} |
| End date | {{END_DATE}} |
| Experiment owner | {{OWNER}} |
| Sponsor | {{SPONSOR}} |
| Status | {{PLANNED / RUNNING / COMPLETED / CANCELLED}} |
---
## Population
### Treatment Group (with AI tools)
| Field | Value |
|-------|-------|
| Teams / developers | {{TEAM_NAMES_OR_IDS}} |
| Count (n) | {{COUNT}} |
| Selection criteria | {{HOW_SELECTED — e.g., "Teams working on backend services, matched by size and tech stack"}} |
### Control Group (without AI tools / status quo)
| Field | Value |
|-------|-------|
| Teams / developers | {{TEAM_NAMES_OR_IDS}} |
| Count (n) | {{COUNT}} |
| Selection criteria | {{HOW_SELECTED}} |
### Minimum Sample Size Justification
```
Power analysis parameters:
- Expected effect size (Cohen's d): {{EFFECT_SIZE — e.g., 0.5 (medium)}}
- Significance level (alpha): 0.05
- Statistical power (1 - beta): 0.80
- Test type: {{TWO_SIDED / ONE_SIDED}}
- Minimum n per group: {{CALCULATED_N}}
- Tool used for calculation: {{G*Power / statsmodels / other}}
```
Actual n per group ({{COUNT}}) {{MEETS / DOES NOT MEET}} the minimum sample size requirement.
---
## Design
| Field | Value |
|-------|-------|
| Design type | {{A/B / BEFORE_AFTER / CROSSOVER / MULTIPLE_BASELINE}} |
| Randomization method | {{METHOD — e.g., "Stratified random assignment by team size and tech stack"}} |
| Blinding | {{SINGLE_BLIND (evaluators don't know group) / DOUBLE_BLIND / NONE}} |
| Washout period (crossover only) | {{WEEKS — if applicable}} |
### Design Rationale
{{EXPLAIN_WHY_THIS_DESIGN — e.g., "A/B chosen because crossover is impractical (can't un-learn tool usage). Stratified randomization ensures comparable groups."}}
### Timeline
| Phase | Duration | Activities |
|-------|----------|------------|
| Baseline measurement | {{WEEKS}} weeks | Collect pre-experiment metrics for all groups |
| Intervention | {{WEEKS}} weeks | Treatment group receives AI tools + onboarding |
| Measurement | {{WEEKS}} weeks | Collect metrics from both groups |
| Analysis | {{WEEKS}} weeks | Statistical analysis and report |
| Debrief | {{WEEKS}} week(s) | Share results, decide on rollout |
---
## Measurement Plan
| Metric | Data Source | Collection Frequency | Baseline (Treatment) | Baseline (Control) | Expected Effect Size |
|--------|-----------|---------------------|---------------------|--------------------|---------------------|
| {{PRIMARY_METRIC}} | {{SOURCE}} | {{FREQUENCY}} | {{VALUE}} | {{VALUE}} | {{EFFECT}} |
| {{SECONDARY_METRIC_1}} | {{SOURCE}} | {{FREQUENCY}} | {{VALUE}} | {{VALUE}} | {{EFFECT}} |
| {{SECONDARY_METRIC_2}} | {{SOURCE}} | {{FREQUENCY}} | {{VALUE}} | {{VALUE}} | {{EFFECT}} |
| {{SECONDARY_METRIC_3}} | {{SOURCE}} | {{FREQUENCY}} | {{VALUE}} | {{VALUE}} | {{EFFECT}} |
If `Program mode = AGENT` or `MIXED`, include at least one of:
- task completion rate
- human takeover rate
- PR merge rate
- reviewer effort per accepted task
- revert rate
### Data Collection Checklist
- [ ] Baseline metrics collected for at least {{WEEKS}} weeks before intervention
- [ ] Automated data pipelines tested and validated
- [ ] Manual data collection procedures documented
- [ ] Data storage location: {{LOCATION}}
- [ ] Access restricted to: {{PEOPLE_OR_ROLES}}
---
## Analysis Plan
Pre-register this section before data collection begins.
| Field | Value |
|-------|-------|
| Primary statistical test | {{TEST — e.g., independent samples t-test / Mann-Whitney U / mixed-effects model}} |
| Significance level (alpha) | 0.05 |
| Statistical power (1 - beta) | >= 0.80 |
| Minimum detectable effect | {{VALUE — e.g., "15% reduction in cycle time"}} |
| Multiple comparison correction | {{METHOD — Bonferroni / Holm / Benjamini-Hochberg FDR / none if single primary}} |
| Software / tool for analysis | {{R / Python / SPSS / other}} |
### Decision Criteria
| Outcome | Definition | Action |
|---------|-----------|--------|
| Clear positive | Primary metric significant (p < 0.05) with meaningful effect size | Roll out to all teams |
| Positive trend | Primary metric trends positive but not significant | Extend experiment or expand sample |
| No effect | No significant difference | Investigate barriers; consider tool/process changes |
| Negative | Significant negative impact | Stop rollout; diagnose root cause |
### Interim Analysis (Optional)
- Check at experiment midpoint (week {{MIDPOINT}})
- Stop early only if: {{STOPPING_CRITERIA — e.g., "clear harm detected (p < 0.01 negative effect)"}}
- Adjust alpha for interim look: {{ADJUSTED_ALPHA}}
---
## Threats to Validity
### Internal Validity
| Threat | Risk Level | Mitigation |
|--------|-----------|------------|
| Selection bias | {{H/M/L}} | {{MITIGATION — e.g., "Stratified randomization by team size and stack"}} |
| Hawthorne effect | {{H/M/L}} | {{MITIGATION — e.g., "Both groups told they are being studied"}} |
| Novelty effect | {{H/M/L}} | {{MITIGATION — e.g., "Minimum 8-week duration to let novelty wear off"}} |
| Contamination (control uses tools informally) | {{H/M/L}} | {{MITIGATION — e.g., "License enforcement; periodic compliance check"}} |
| Attrition (developers leave during experiment) | {{H/M/L}} | {{MITIGATION — e.g., "Intent-to-treat analysis"}} |
| History (external events affect one group) | {{H/M/L}} | {{MITIGATION — e.g., "Avoid major release cycles during experiment"}} |
### External Validity
| Threat | Risk Level | Mitigation |
|--------|-----------|------------|
| Generalizability to other teams | {{H/M/L}} | {{MITIGATION — e.g., "Include diverse tech stacks in sample"}} |
| Generalizability to other tools | {{H/M/L}} | {{MITIGATION — e.g., "Document tool-specific features used"}} |
| Time-bound effects | {{H/M/L}} | {{MITIGATION — e.g., "Note tool version; results may not hold as tools evolve"}} |
### Construct Validity
| Question | Assessment |
|----------|-----------|
| Does the primary metric actually measure what we intend? | {{ASSESSMENT}} |
| Could improvements come from something other than AI tools? | {{ASSESSMENT}} |
| Are self-reported metrics (surveys) consistent with observed metrics? | {{ASSESSMENT}} |
| For agent workflows: is benchmark performance distinct from production acceptance? | {{ASSESSMENT}} |
---
## Ethics and Communication
### Developer Consent
| Field | Value |
|-------|-------|
| Consent approach | {{APPROACH — e.g., "Opt-in with written consent form"}} |
| Right to withdraw | {{Yes — developers can leave the experiment at any time without penalty}} |
| Impact on performance reviews | {{None — experiment participation and metrics will not be used in reviews}} |
### Data Anonymization
| Field | Value |
|-------|-------|
| Individual data access | {{WHO — e.g., "Experiment owner only; all reports use team-level aggregates"}} |
| Anonymization method | {{METHOD — e.g., "Developer IDs replaced with random codes before analysis"}} |
| Data retention | {{PERIOD — e.g., "Raw data deleted 90 days after experiment concludes"}} |
### Results Sharing
| Audience | Format | Timing |
|----------|--------|--------|
| Participating developers | Full results + Q&A session | Within 2 weeks of completion |
| Engineering leadership | Executive summary (use executive-report-template) | Within 3 weeks of completion |
| Wider organization | Blog post or all-hands summary | Within 4 weeks of completion |
---
## Sign-Off
| Role | Name | Date | Approved |
|------|------|------|----------|
| Experiment owner | {{NAME}} | {{DATE}} | {{YES/NO}} |
| Engineering sponsor | {{NAME}} | {{DATE}} | {{YES/NO}} |
| Data/privacy lead | {{NAME}} | {{DATE}} | {{YES/NO}} |
| HR (if required) | {{NAME}} | {{DATE}} | {{YES/NO}} |
assets/metric-dashboard-template.md
# AI Coding Metrics Dashboard Template
**Last Updated**: {{DATE}}
**Owner**: {{NAME}}
**Version**: {{VERSION}}
**Program Mode**: {{ASSISTANT / AGENT / MIXED}}
---
Purpose: provide a three-tier dashboard layout so each audience sees the right metrics at the right cadence. Copy the tier that matches your audience, populate placeholders, and wire up data sources.
## How to Use
1. Pick the tier matching your audience (executive, team lead, developer) and the program mode (assistant, agent, mixed).
2. Replace every `{{PLACEHOLDER}}` with your organization's values.
3. Connect each metric to its data source using the mapping table at the bottom.
4. Set alert thresholds in the traffic-light table to match your targets.
---
## Tier 1: Executive Dashboard (C-Suite / VP Engineering)
Refresh cadence: **monthly**. Trend lines: 3-month and 6-month rolling.
### KPI Cards
| # | KPI | Current | Prior Month | 3-Mo Trend | 6-Mo Trend | Target | Status |
|---|-----|---------|-------------|------------|------------|--------|--------|
| 1 | AI Tool ROI (%) | {{ROI_CURRENT}} | {{ROI_PRIOR}} | {{TREND}} | {{TREND}} | {{ROI_TARGET}} | {{GREEN/AMBER/RED}} |
| 2 | Adoption Rate (%) | {{ADOPT_CURRENT}} | {{ADOPT_PRIOR}} | {{TREND}} | {{TREND}} | {{ADOPT_TARGET}} | {{GREEN/AMBER/RED}} |
| 3 | Delivery Change (%) | {{VEL_CURRENT}} | {{VEL_PRIOR}} | {{TREND}} | {{TREND}} | {{VEL_TARGET}} | {{GREEN/AMBER/RED}} |
| 4 | Quality Delta (defect rate change %) | {{QUAL_CURRENT}} | {{QUAL_PRIOR}} | {{TREND}} | {{TREND}} | {{QUAL_TARGET}} | {{GREEN/AMBER/RED}} |
| 5 | Developer Satisfaction (x/5) | {{SAT_CURRENT}} | {{SAT_PRIOR}} | {{TREND}} | {{TREND}} | {{SAT_TARGET}} | {{GREEN/AMBER/RED}} |
| 6 | Monthly Cost per Developer ($) | {{COST_CURRENT}} | {{COST_PRIOR}} | {{TREND}} | {{TREND}} | {{COST_TARGET}} | {{GREEN/AMBER/RED}} |
### Traffic-Light Thresholds
| KPI | Green | Amber | Red |
|-----|-------|-------|-----|
| ROI (%) | > {{GREEN_THRESHOLD}} | {{AMBER_LOW}} - {{AMBER_HIGH}} | < {{RED_THRESHOLD}} |
| Adoption Rate (%) | > {{GREEN_THRESHOLD}} | {{AMBER_LOW}} - {{AMBER_HIGH}} | < {{RED_THRESHOLD}} |
| Delivery Change (%) | > {{GREEN_THRESHOLD}} | {{AMBER_LOW}} - {{AMBER_HIGH}} | < {{RED_THRESHOLD}} |
| Quality Delta (%) | > {{GREEN_THRESHOLD}} | {{AMBER_LOW}} - {{AMBER_HIGH}} | < {{RED_THRESHOLD}} |
| Developer Satisfaction | > {{GREEN_THRESHOLD}} | {{AMBER_LOW}} - {{AMBER_HIGH}} | < {{RED_THRESHOLD}} |
| Monthly Cost per Dev | < {{GREEN_THRESHOLD}} | {{AMBER_LOW}} - {{AMBER_HIGH}} | > {{RED_THRESHOLD}} |
### Layout Sketch
```
+-------------------+-------------------+-------------------+
| ROI (%) | Adoption Rate | Delivery Change |
| [big number] | [big number] | [big number] |
| [sparkline] | [sparkline] | [sparkline] |
+-------------------+-------------------+-------------------+
| Quality Delta | Dev Satisfaction | Monthly Cost/Dev |
| [big number] | [big number] | [big number] |
| [sparkline] | [sparkline] | [sparkline] |
+-------------------+-------------------+-------------------+
| 6-Month Trend Chart (all KPIs overlaid) |
+-----------------------------------------------------------+
```
### Agent Extension (Optional for Agent or Mixed Programs)
| KPI | Current | Prior Month | Target | Status |
|-----|---------|-------------|--------|--------|
| Agent Task Completion Rate (%) | {{VALUE}} | {{VALUE}} | {{TARGET}} | {{STATUS}} |
| Acceptance-for-Review Rate (%) | {{VALUE}} | {{VALUE}} | {{TARGET}} | {{STATUS}} |
| Agent PR Merge Rate (%) | {{VALUE}} | {{VALUE}} | {{TARGET}} | {{STATUS}} |
| Human Takeover Rate (%) | {{VALUE}} | {{VALUE}} | {{TARGET}} | {{STATUS}} |
---
## Tier 2: Team Lead Dashboard
Refresh cadence: **weekly**.
### Metrics Table
| # | Metric | Team: {{TEAM_A}} | Team: {{TEAM_B}} | Team: {{TEAM_C}} | Org Avg | Target | Status |
|---|--------|-------------------|-------------------|-------------------|---------|--------|--------|
| 1 | Deployment Frequency | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 2 | Lead Time for Changes | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 3 | Change Failure Rate (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 4 | Mean Time to Recovery | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 5 | Cycle Time (days) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 6 | PR Throughput (PRs/week) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 7 | PR Review Time (hours) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 8 | Test Coverage (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 9 | AI Tool Adoption (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 10 | Suggestion Accept Rate (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 11 | Agent Task Completion (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
| 12 | Agent PR Merge Rate (%) | {{VALUE}} | {{VALUE}} | {{VALUE}} | {{AVG}} | {{TARGET}} | {{STATUS}} |
### Drill-Down Paths
| Metric | Drill-Down View | Data Source |
|--------|----------------|-------------|
| Deployment Frequency | Deployments by service, by day | {{CI_CD_TOOL}} |
| Lead Time for Changes | Commit-to-deploy timeline | {{CI_CD_TOOL}} + {{SCM}} |
| Change Failure Rate | Failed deploys + rollbacks | {{INCIDENT_TOOL}} |
| MTTR | Incident timeline | {{INCIDENT_TOOL}} |
| Cycle Time | Issue open-to-close breakdown | {{PROJECT_TOOL}} |
| PR Throughput | PR list with status, author, size | {{SCM}} |
| PR Review Time | Review request-to-approval duration | {{SCM}} |
| Test Coverage | Coverage by module/service | {{COVERAGE_TOOL}} |
| AI Tool Adoption | Active users / licensed seats | {{AI_TOOL_ADMIN}} |
| Suggestion Accept Rate | Accepted vs dismissed by language | {{AI_TOOL_TELEMETRY}} |
| Agent Task Completion | Started vs completed runs by workflow | {{AGENT_TELEMETRY}} |
| Agent PR Merge Rate | PRs merged / PRs created by agent | {{SCM}} + {{AGENT_TELEMETRY}} |
### Team Comparison (Anonymized)
Teams are displayed as Team 1, Team 2, etc. unless team leads opt into named display.
| Rank | Team (Anonymized) | Composite Score | Top Strength | Biggest Gap |
|------|-------------------|----------------|--------------|-------------|
| 1 | {{TEAM_ANON}} | {{SCORE}} | {{STRENGTH}} | {{GAP}} |
| 2 | {{TEAM_ANON}} | {{SCORE}} | {{STRENGTH}} | {{GAP}} |
| 3 | {{TEAM_ANON}} | {{SCORE}} | {{STRENGTH}} | {{GAP}} |
---
## Tier 3: Developer Dashboard (Individual / Opt-In Only)
Refresh cadence: **daily** (self-service). Participation is voluntary; no individual data is shared with management.
### Personal Usage Stats
| Metric | This Week | Last Week | 30-Day Avg |
|--------|-----------|-----------|------------|
| Suggestions Shown | {{COUNT}} | {{COUNT}} | {{AVG}} |
| Suggestions Accepted | {{COUNT}} | {{COUNT}} | {{AVG}} |
| Suggestions Rejected | {{COUNT}} | {{COUNT}} | {{AVG}} |
| Accept Rate (%) | {{RATE}} | {{RATE}} | {{AVG}} |
| Estimated Time Saved (hours) | {{HOURS}} | {{HOURS}} | {{AVG}} |
| Top Language Used with AI | {{LANGUAGE}} | {{LANGUAGE}} | — |
| Most Productive Use Case | {{USE_CASE}} | {{USE_CASE}} | — |
### Learning Resources
| Usage Pattern | Suggested Resource |
|--------------|-------------------|
| Low accept rate in {{LANGUAGE}} | {{TRAINING_LINK}} |
| Not using chat/explain features | {{TUTORIAL_LINK}} |
| High reject rate on tests | {{BEST_PRACTICES_LINK}} |
---
## Data Source Mapping
| Metric | Source System | API / Query | Refresh | Owner |
|--------|-------------|-------------|---------|-------|
| ROI | {{FINANCE_TOOL}} + {{AI_TOOL_ADMIN}} | {{API_ENDPOINT_OR_QUERY}} | Monthly | {{OWNER}} |
| Adoption Rate | {{AI_TOOL_ADMIN}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| Velocity (Deployment Freq) | {{CI_CD_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| Lead Time | {{SCM}} + {{CI_CD_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| Change Failure Rate | {{INCIDENT_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| MTTR | {{INCIDENT_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| Cycle Time | {{PROJECT_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| PR Throughput | {{SCM}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| PR Review Time | {{SCM}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| Test Coverage | {{COVERAGE_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Weekly | {{OWNER}} |
| Suggestion Accept Rate | {{AI_TOOL_TELEMETRY}} | {{API_ENDPOINT_OR_QUERY}} | Daily | {{OWNER}} |
| Agent Task Completion | {{AGENT_TELEMETRY}} | {{API_ENDPOINT_OR_QUERY}} | Daily | {{OWNER}} |
| Agent PR Merge Rate | {{SCM}} + {{AGENT_TELEMETRY}} | {{API_ENDPOINT_OR_QUERY}} | Daily | {{OWNER}} |
| Human Takeover Rate | {{AGENT_TELEMETRY}} | {{API_ENDPOINT_OR_QUERY}} | Daily | {{OWNER}} |
| Developer Satisfaction | {{SURVEY_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Quarterly | {{OWNER}} |
| Quality Delta (Defect Rate) | {{BUG_TRACKER}} | {{API_ENDPOINT_OR_QUERY}} | Monthly | {{OWNER}} |
| Cost per Developer | {{FINANCE_TOOL}} | {{API_ENDPOINT_OR_QUERY}} | Monthly | {{OWNER}} |
---
## Implementation Notes
- Start with Tier 1. Add Tier 2 once data pipelines stabilize. Tier 3 is optional.
- Automate data collection before manual entry becomes a bottleneck.
- Review thresholds quarterly; recalibrate as baselines shift.
- Keep developer-level data aggregated at the team level for management views.
assets/roi-calculator-template.md
# AI Coding Tools ROI Calculator
**Last Updated**: {{DATE}}
**Owner**: {{NAME}}
**Version**: {{VERSION}}
---
Purpose: calculate return on investment for AI coding tool adoption. Fill in the input variables, apply the formulas, and use the sensitivity table to stress-test assumptions. Designed to be transferred directly into a spreadsheet.
## How to Use
1. Fill in the **Input Variables** table with your organization's numbers.
2. Walk through the **Calculation Formulas** step by step.
3. Use the **Sensitivity Analysis** to present conservative, moderate, and optimistic scenarios.
4. Determine the **Break-Even** point to set minimum adoption targets.
---
## Input Variables
| Variable | Symbol | Conservative | Moderate | Optimistic | Your Value |
|----------|--------|-------------|----------|------------|------------|
| Number of developers | `N` | {{N}} | {{N}} | {{N}} | {{N}} |
| Average fully-loaded cost ($/year) | `C_dev` | {{150000}} | {{175000}} | {{200000}} | {{VALUE}} |
| Effective hourly cost ($/hour) | `C_hour` | = C_dev / 2080 | = C_dev / 2080 | = C_dev / 2080 | {{VALUE}} |
| AI tool cost per seat ($/month) | `C_seat` | {{19}} | {{19}} | {{19}} | {{VALUE}} |
| Training cost per developer ($, one-time) | `C_train` | {{500}} | {{1000}} | {{1500}} | {{VALUE}} |
| Hours saved per developer per week | `H_saved` | {{2}} | {{4}} | {{6}} | {{VALUE}} |
| Working weeks per year | `W` | 48 | 48 | 48 | {{VALUE}} |
| Quality improvement (% reduction in bug-fix time) | `Q_improve` | {{5%}} | {{10%}} | {{15%}} | {{VALUE}} |
| Annual bug-fix hours per developer | `H_bugfix` | {{200}} | {{200}} | {{200}} | {{VALUE}} |
| Annual developer turnover rate (%) | `R_turnover` | {{15%}} | {{15%}} | {{15%}} | {{VALUE}} |
| Retention improvement (% reduction in turnover) | `R_improve` | {{5%}} | {{10%}} | {{15%}} | {{VALUE}} |
| Replacement cost per developer ($) | `C_replace` | {{75000}} | {{100000}} | {{125000}} | {{VALUE}} |
| Overhead factor (implementation, admin) | `F_overhead` | 15% | 15% | 15% | {{VALUE}} |
---
## Calculation Formulas
Work through these in order. Each formula uses the symbols from the input table.
### Costs
```
Annual tool cost = N x C_seat x 12
Annual training cost = N x C_train
Overhead = (Annual tool cost + Annual training cost) x F_overhead
Total annual cost = Annual tool cost + Annual training cost + Overhead
```
| Cost Component | Formula | Your Calculation |
|---------------|---------|-----------------|
| Annual tool cost | {{N}} x {{C_seat}} x 12 | = ${{VALUE}} |
| Annual training cost | {{N}} x {{C_train}} | = ${{VALUE}} |
| Overhead (15%) | (tool + training) x 0.15 | = ${{VALUE}} |
| **Total annual cost** | sum of above | = **${{TOTAL_COST}}** |
### Benefits
```
Time savings value = N x H_saved x W x C_hour
Quality savings = N x H_bugfix x Q_improve x C_hour
Retention savings = N x R_turnover x R_improve x C_replace
Total annual benefit = Time savings + Quality savings + Retention savings
```
| Benefit Component | Formula | Your Calculation |
|------------------|---------|-----------------|
| Time savings value | {{N}} x {{H_saved}} x {{W}} x {{C_hour}} | = ${{VALUE}} |
| Quality savings | {{N}} x {{H_bugfix}} x {{Q_improve}} x {{C_hour}} | = ${{VALUE}} |
| Retention savings | {{N}} x {{R_turnover}} x {{R_improve}} x {{C_replace}} | = ${{VALUE}} |
| **Total annual benefit** | sum of above | = **${{TOTAL_BENEFIT}}** |
### Returns
```
Net annual benefit = Total annual benefit - Total annual cost
ROI (%) = (Net annual benefit / Total annual cost) x 100
Payback period (months) = Total annual cost / (Net annual benefit / 12)
```
| Return Metric | Formula | Your Calculation |
|--------------|---------|-----------------|
| Net annual benefit | {{TOTAL_BENEFIT}} - {{TOTAL_COST}} | = ${{NET_BENEFIT}} |
| ROI (%) | ({{NET_BENEFIT}} / {{TOTAL_COST}}) x 100 | = {{ROI}}% |
| Payback period | {{TOTAL_COST}} / ({{NET_BENEFIT}} / 12) | = {{MONTHS}} months |
---
## Sensitivity Analysis
ROI (%) at different adoption rates and time-saved scenarios. Assume all other inputs held constant.
| | 2 hrs/wk saved | 4 hrs/wk saved | 6 hrs/wk saved | 8 hrs/wk saved |
|---|----------------|----------------|----------------|----------------|
| **25% adoption** | {{ROI}}% | {{ROI}}% | {{ROI}}% | {{ROI}}% |
| **50% adoption** | {{ROI}}% | {{ROI}}% | {{ROI}}% | {{ROI}}% |
| **75% adoption** | {{ROI}}% | {{ROI}}% | {{ROI}}% | {{ROI}}% |
| **100% adoption** | {{ROI}}% | {{ROI}}% | {{ROI}}% | {{ROI}}% |
How to calculate each cell:
- Effective developers = N x adoption_rate
- Recalculate Total annual benefit using effective developers
- Costs remain fixed (all seats licensed regardless of adoption)
- ROI = (adjusted_benefit - total_cost) / total_cost x 100
---
## Break-Even Analysis
The minimum adoption rate at which the program pays for itself.
```
Break-even adoption rate = Total annual cost / Max annual benefit (at 100% adoption)
```
| Scenario | Total Cost | Max Benefit (100%) | Break-Even Adoption |
|----------|-----------|-------------------|-------------------|
| Conservative (2 hrs/wk) | ${{VALUE}} | ${{VALUE}} | {{VALUE}}% |
| Moderate (4 hrs/wk) | ${{VALUE}} | ${{VALUE}} | {{VALUE}}% |
| Optimistic (6 hrs/wk) | ${{VALUE}} | ${{VALUE}} | {{VALUE}}% |
**Decision rule**: if current adoption rate > break-even rate at the conservative scenario, the investment is justified.
---
## Assumptions and Caveats
- Hours saved are self-reported unless validated by controlled experiment (see experiment-design-template).
- Quality and retention improvements are harder to attribute; use conservative estimates.
- Training cost is year-one only; subsequent years use a lower refresher cost.
- Overhead factor (15%) covers admin time, integration maintenance, and security review.
- This model does not capture second-order effects (e.g., faster time-to-market revenue impact).
---
## Presentation Tips
- Lead with the moderate scenario; use conservative as the floor.
- Show the sensitivity table to demonstrate range of outcomes.
- Highlight the break-even adoption rate: it shifts the question from "should we invest?" to "can we achieve X% adoption?"
- Pair with survey data (adoption-survey-template) for credibility.
data/sample-ai-metrics.json
{
"team_name": "Platform Engineering",
"team_size": 20,
"measurement_period_weeks": 12,
"ai_tooling_monthly_cost": 1200,
"avg_dev_hourly_rate": 95,
"hours_saved_per_dev_per_week": 3.5,
"adoption_pct": 72,
"notes": "12-week pilot using GitHub Copilot (assistant) and Claude Code (agent mode on selected tasks). Baseline established from prior 12-week period.",
"metric_families": {
"adoption": {
"score": 72,
"signals": [
"14 of 20 developers are active weekly (70% WAU)",
"License utilization at 72% — above 60% rollout threshold",
"Feature breadth shallow: 80% use inline completion only; 25% use agent mode",
"Suggestion acceptance rate: 31% (industry median ~27%)",
"Agent invocation rate: 4.2 tasks/dev/week among agent-mode users"
]
},
"delivery": {
"score": 61,
"signals": [
"Cycle time (commit to deploy): reduced 11% vs baseline",
"PR review turnaround unchanged for agent-authored PRs — reviewers spending more time on verification",
"Lead time for changes: no statistically significant change at team level",
"Deployment frequency up 8% — driven by 3 teams, not distributed",
"Onboarding time to first meaningful contribution: -18% for Q1 cohort (n=3, small sample)"
]
},
"quality": {
"score": 54,
"signals": [
"Defect escape rate: no change vs baseline (stable at 1.4 bugs/deploy)",
"Revert rate for AI-assisted PRs: 6.2% vs 4.8% baseline — elevated, watch closely",
"Review burden increased 14% on agent-authored PRs (more rounds requested)",
"Security scan findings on new code: +9% raw count, severity mix unchanged",
"Test coverage delta: +2.1 percentage points — modest positive signal"
]
},
"economics": {
"score": 78,
"signals": [
"Cost per active user: $85.71/month ($1,200 / 14 active users)",
"Annualized tool cost: $14,400",
"Estimated annual hours saved: 3,276 hours (20 devs × 3.5 h/wk × 52 wks × 72% utilization factor)",
"Estimated annual value at $95/hr fully-loaded: $311,220",
"Net annual benefit: ~$296,820 after tool cost",
"Payback period: under 3 weeks"
]
},
"experience": {
"score": 66,
"signals": [
"Tool NPS: +28 (promoters 52%, detractors 24%)",
"Top friction: context window limits, hallucinated API signatures, slow agent mode on large repos",
"Trust burden elevated for agent output — 68% of surveyed devs say they verify agent code more carefully than their own",
"Frustration / give-up rate: 19% of agent sessions abandoned before completion",
"Cognitive load self-report: split — 40% say lower, 35% say higher for complex tasks"
]
},
"agent_execution": {
"score": 49,
"signals": [
"Task completion rate: 61% fully autonomous (agent mode, no human takeover)",
"Human takeover rate: 28% of agent sessions required developer intervention",
"PR merge rate for agent-authored PRs: 74% (vs 88% for human PRs)",
"Post-merge revert/hotfix rate for agent PRs: 8.1% (vs 4.8% baseline) — leading risk signal",
"Reviewer effort per accepted agent task: 1.8x manual PR review time",
"Policy/security exception rate: 2 incidents in period (both low severity)"
]
}
}
}
data/sources.json
{
"metadata": {
"skill": "dev-ai-coding-metrics",
"title": "AI Coding Metrics for Engineering Teams",
"description": "Curated August 2026 source set for measuring coding assistants and coding agents across adoption, delivery, quality, ROI, experience, and agent-execution outcomes.",
"last_updated": "2026-08-21",
"updated": "2026-08-21",
"retrieved": "2026-08-21",
"last_reviewed": "2026-08-21",
"version": "2.4",
"total_sources": 34,
"schema_notes": {
"status": "live | redirect | dead | questionable",
"evidence_tier": "peer_reviewed | official | independent | vendor | commentary"
},
"retired_examples": [
{
"name": "ETH Zurich Codified Context (old ID)",
"former_url": "https://arxiv.org/abs/2603.04692",
"reason": "Wrong paper identifier for the relevant repository-context evidence"
},
{
"name": "Amazon CodeWhisperer resources",
"former_url": "https://aws.amazon.com/codewhisperer/resources/",
"reason": "Replaced by Amazon Q Developer branding and docs"
},
{
"name": "Stripe AI engineering case study",
"former_url": "https://stripe.com/blog/ai-engineering",
"reason": "Unstable or unavailable as a default source"
},
{
"name": "Coinbase AI engineering metrics case study",
"former_url": "https://www.coinbase.com/blog/ai-engineering-metrics",
"reason": "Unstable or unavailable as a default source"
},
{
"name": "Shopify memo via social post",
"former_url": "https://twitter.com/toaboronkay/status/1907841423613473178",
"reason": "Social repost is not a stable primary source"
}
]
},
"categories": {
"foundational_research": [
{
"name": "DORA Research Hub",
"url": "https://dora.dev/research/",
"type": "research-hub",
"description": "Primary home for DORA reports and measurement frameworks.",
"relevance": "Top-level source for engineering performance research and current AI-related report links.",
"status": "live",
"last_verified": "2026-05-17",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["dora", "research", "foundational", "engineering-performance"]
},
{
"name": "DORA 2025: AI-Assisted Software Development",
"url": "https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report",
"type": "report",
"description": "DORA's 2025 dedicated report on AI-assisted software development (Google Cloud). Covers how AI affects code quality, flow, delivery throughput, and delivery stability.",
"relevance": "Primary 2025/2026 framing for conditional AI impact rather than default optimism. Use this URL for the AI-specific report; the dora.dev hub links the broader annual report.",
"status": "live",
"last_verified": "2026-05-17",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["dora", "ai", "delivery", "quality", "2025", "google-cloud"]
},
{
"name": "DORA 2025: AI Capability Model",
"url": "https://dora.dev/ai/capabilities-model/report/",
"type": "framework",
"description": "Capability model describing the organizational conditions that make AI helpful rather than disruptive.",
"relevance": "Useful for separating tool usage from organization readiness and AI-accessibility.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["dora", "capability-model", "readiness", "ai-accessibility"]
},
{
"name": "DORA 2026: ROI of AI-Assisted Software Development",
"url": "https://dora.dev/ai/roi/report/",
"type": "report",
"description": "DORA's 2026 follow-up report (dated 2026.01, widely covered April-May 2026) providing a financial framework for AI investment ROI. Introduces the J-Curve adoption model (learning curve, verification tax, downstream process friction), a scenario-based ROI methodology (conservative/realistic/optimistic), and an 'instability tax' finding: AI adoption raises individual effectiveness and code quality but also change-failure rate in the report's illustrative model (5% to 6%). Also finds AI yields 35-40% gains on simple tasks vs. ~10% on complex legacy code.",
"relevance": "Distinct from the 2025 DORA AI report — cite separately. Directly reinforces this skill's scenario-planning and task-complexity-segmentation rules in roi-framework.md and benchmarking-methodology.md. Treat the illustrative ~39% first-year ROI figure as a vendor scenario model, not a transferable benchmark.",
"status": "live",
"last_verified": "2026-07-11",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["dora", "roi", "j-curve", "instability-tax", "scenario-planning", "2026"]
},
{
"name": "SPACE Framework: A Multidimensional View of Developer Productivity",
"url": "https://queue.acm.org/detail.cfm?id=3454124",
"type": "academic",
"description": "Foundational framework covering satisfaction, performance, activity, communication, and efficiency.",
"relevance": "Prevents over-indexing on raw throughput when evaluating AI tools.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "peer_reviewed",
"add_as_web_search": false,
"tags": ["space", "productivity", "framework", "developer-experience"]
},
{
"name": "METR: Early 2025 AI on Experienced Open-Source Developers (RCT)",
"url": "https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/",
"type": "study",
"description": "Randomized controlled trial (data Feb–Jun 2025, published 2025-07-10) on realistic open-source tasks with experienced developers. Found participants were ~19% SLOWER with early-2025 AI tools, contrary to developer self-predictions.",
"relevance": "Primary counterweight to 'AI always speeds development up' claims; establishes RCT baseline for the early-2025 tooling generation.",
"status": "live",
"last_verified": "2026-05-17",
"evidence_tier": "independent",
"add_as_web_search": true,
"tags": ["metr", "rct", "real-world", "counterpoint", "slowdown", "2025"]
},
{
"name": "METR 2026 Update: Changing Developer Productivity Experiment Design",
"url": "https://metr.org/blog/2026-02-24-uplift-update/",
"type": "study",
"description": "METR update (2026-02-24) stating they believe developers are likely more sped up in early 2026 than early-2025 estimates suggested, but noting their new experiment gives an unreliable signal because 30–50% of participating developers declined to submit tasks they didn't want to do without AI (selection bias).",
"relevance": "Critical for any claim about AI productivity trajectories into 2026: signals likely improvement but with significant methodological caveats. Do not cite as proof of productivity gains without flagging the selection bias.",
"status": "live",
"last_verified": "2026-05-17",
"evidence_tier": "independent",
"add_as_web_search": true,
"tags": ["metr", "rct", "selection-bias", "2026", "productivity-trajectory", "counterpoint"]
},
{
"name": "METR May 2026 Self-Reported AI Productivity Survey",
"url": "https://metr.org/blog/2026-05-11-ai-usage-survey/",
"type": "study",
"description": "Survey of 349 technical workers (87 engineers, 71 researchers, 129 academics/PhD students, 48 founders/managers), conducted Feb–Apr 2026, published 2026-05-11. Median self-reported value-of-work change: 1.4–2x. METR notes significant skepticism: their 2025 RCT found participants overestimated AI time effects by 40 percentage points on average.",
"relevance": "Use when discussing self-reported vs. controlled productivity data. Establishes 2026 trajectory context but must not be cited as controlled evidence of gains. Pair with METR 2025 RCT for balanced framing.",
"status": "live",
"last_verified": "2026-06-09",
"evidence_tier": "independent",
"add_as_web_search": true,
"tags": ["metr", "survey", "self-reported", "2026", "productivity-trajectory", "counterpoint"]
},
{
"name": "Evaluating the Impact of AGENTS.md Files on LLM-Coding Agents",
"url": "https://arxiv.org/abs/2602.11988",
"type": "academic",
"description": "Study of how repository-level agent instruction files affect cost and task performance.",
"relevance": "Use when discussing context engineering as a moderator rather than a guaranteed lift.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "peer_reviewed",
"add_as_web_search": true,
"tags": ["agents-md", "context", "repo-instructions", "agent-performance"]
},
{
"name": "GitHub Copilot RCT (Peng et al. 2023)",
"url": "https://arxiv.org/abs/2302.06590",
"type": "academic",
"description": "Randomized study often cited for Copilot productivity gains on constrained tasks.",
"relevance": "Still useful, but should be interpreted as narrower than current agentic workflows.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "peer_reviewed",
"add_as_web_search": false,
"tags": ["copilot", "rct", "productivity", "historic-baseline"]
},
{
"name": "Productivity Assessment of Neural Code Completion",
"url": "https://arxiv.org/abs/2205.06537",
"type": "academic",
"description": "Study introducing persistence-oriented metrics for code completion usefulness.",
"relevance": "Useful for replacing naive acceptance-rate-only interpretations.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "peer_reviewed",
"add_as_web_search": false,
"tags": ["completion", "persistence", "metrics", "code-completion"]
},
{
"name": "SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks",
"url": "https://arxiv.org/abs/2603.24755v1",
"type": "academic-preprint",
"description": "March 2026 preprint introducing evolving-spec, carried-workspace coding-agent evaluation and trajectory-level structural-erosion and verbosity signals.",
"relevance": "Supports extension-robustness scorecards, checkpoint quality slopes, and late-checkpoint cost/review measurement. Do not transfer paper averages into organizational targets or causal ROI claims.",
"status": "live",
"last_verified": "2026-08-21",
"evidence_tier": "independent",
"add_as_web_search": false,
"tags": ["coding-agents", "iterative-evaluation", "extension-robustness", "structural-erosion", "verbosity", "preprint"]
}
],
"adoption_and_sentiment": [
{
"name": "Stack Overflow Developer Survey 2025",
"url": "https://survey.stackoverflow.co/2025/",
"type": "survey",
"description": "Large annual developer survey with AI adoption and sentiment data.",
"relevance": "Useful for directional benchmarking on awareness, usage, and trust.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "independent",
"add_as_web_search": true,
"tags": ["survey", "adoption", "sentiment", "stackoverflow"]
},
{
"name": "JetBrains Developer Ecosystem Survey 2025",
"url": "https://www.jetbrains.com/lp/devecosystem-2025/",
"type": "survey",
"description": "Developer ecosystem survey with IDE, workflow, and AI assistant adoption data.",
"relevance": "Cross-checks GitHub- or ChatGPT-centric usage assumptions.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "independent",
"add_as_web_search": true,
"tags": ["survey", "jetbrains", "ecosystem", "adoption"]
},
{
"name": "GitHub Octoverse 2025",
"url": "https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/",
"type": "report",
"description": "GitHub's annual report with platform-level development and AI trends.",
"relevance": "Good directional source for broad platform usage shifts, not causal ROI claims.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["github", "octoverse", "platform-trends", "adoption"]
},
{
"name": "Faros AI Engineering Report 2026",
"url": "https://www.faros.ai/research/ai-acceleration-whiplash",
"type": "report",
"description": "Organizational telemetry (not RCT) from 22,000 developers across 4,000 teams. Key figures: average PR size +51.3%; files touched per developer per month +149.9%; PRs merged without review +31.3%; lead time commit to production +480.4%; throughput (epics per developer) approximately +66%; incidents per PR approximately +243%. Reflects throughput-vs-quality divergence at scale.",
"relevance": "Use when illustrating throughput-vs-quality divergence. The PR-merge-without-review (+31.3%) and lead-time (+480.4%) figures are operationally significant — do not cite only throughput. Vendor telemetry, not RCT.",
"status": "live",
"last_verified": "2026-06-09",
"evidence_tier": "vendor",
"add_as_web_search": true,
"tags": ["faros", "telemetry", "throughput", "review-burden", "incidents", "2026"]
},
{
"name": "GitClear: AI and Code Quality Analysis",
"url": "https://www.gitclear.com/coding_on_copilot_data_shows_ais_downward_pressure_on_code_quality",
"type": "research",
"description": "Large-scale code-churn and reuse analysis focused on AI-assisted code quality patterns.",
"relevance": "Useful caution source for rework, churn, and retained-quality discussions.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "independent",
"add_as_web_search": true,
"tags": ["gitclear", "quality", "churn", "retention"]
}
],
"operator_content": [
{
"name": "AI tools are overdelivering: results from our large-scale AI productivity survey | Lenny's Newsletter",
"url": "https://www.lennysnewsletter.com/p/ai-tools-are-overdelivering-results",
"type": "survey",
"description": "Large-scale survey (1,750 respondents) on AI productivity impact across PMs, engineers, designers, and founders. Covers time savings, quality perception, role-specific patterns, and reported downsides.",
"relevance": "Independent cross-role productivity data useful for benchmarking self-reported AI impact alongside controlled studies.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "commentary",
"add_as_web_search": true,
"tags": ["survey", "productivity", "ai-impact", "cross-role", "lenny"]
},
{
"name": "How to measure AI developer productivity in 2025 | Nicole Forsgren (Lenny's Podcast)",
"url": "https://www.lennysnewsletter.com/p/how-to-measure-ai-developer-productivity",
"type": "podcast",
"description": "Nicole Forsgren (creator of DORA and SPACE frameworks) discusses current approaches to measuring AI developer productivity, gaps in existing metrics, and practical guidance for engineering leaders.",
"relevance": "Direct bridge between DORA/SPACE frameworks and AI-era measurement; from the primary author of both frameworks.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "commentary",
"add_as_web_search": true,
"tags": ["dora", "space", "measurement", "ai-productivity", "forsgren", "lenny"]
}
],
"agent_evaluation": [
{
"name": "OpenAI SWE-Lancer",
"url": "https://openai.com/index/swe-lancer/",
"type": "benchmark",
"description": "OpenAI benchmark and evaluation framing for freelance-style software engineering tasks.",
"relevance": "Better fit for agentic software work than narrow completion-only benchmarks.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["swe-lancer", "agents", "benchmark", "task-completion"]
},
{
"name": "RE-Bench",
"url": "https://arxiv.org/abs/2411.15114",
"type": "benchmark",
"description": "Benchmark focused on realistic and economically meaningful software engineering work.",
"relevance": "Useful for benchmark-to-production gap discussions in agent evaluations.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "peer_reviewed",
"add_as_web_search": true,
"tags": ["re-bench", "benchmark", "realistic", "agents"]
},
{
"name": "GitHub Blog: Validating Agentic Behavior",
"url": "https://github.blog/ai-and-ml/generative-ai/validating-agentic-behavior-when-correct-isnt-deterministic/",
"type": "article",
"description": "Discussion of practical validation for nondeterministic agentic behavior in coding workflows.",
"relevance": "Good support for treating benchmark wins as insufficient without workflow-level validation and production evidence.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["benchmarks", "saturation", "github", "evaluation"]
}
],
"measurement_tool_docs": [
{
"name": "GitHub Copilot Admin Metrics",
"url": "https://docs.github.com/en/copilot/concepts/copilot-usage-metrics",
"type": "documentation",
"description": "Current GitHub documentation for Copilot metrics, entitlements, and admin measurement.",
"relevance": "Primary first-party reference for GitHub Copilot adoption and usage definitions.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["github", "copilot", "metrics", "admin", "official"]
},
{
"name": "ChatGPT Enterprise and Edu User Analytics",
"url": "https://help.openai.com/en/articles/10875114-user-analytics-for-chatgpt-enterprise-and-edu-public-beta",
"type": "documentation",
"description": "OpenAI documentation for ChatGPT workspace analytics.",
"relevance": "Useful when measuring workspace adoption separately from API-driven coding workflows.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["openai", "chatgpt", "analytics", "workspace"]
},
{
"name": "Amazon Q Developer",
"url": "https://aws.amazon.com/q/developer/",
"type": "documentation",
"description": "Official Amazon Q Developer landing page replacing historical CodeWhisperer references.",
"relevance": "Use for current naming and product-surface references instead of legacy CodeWhisperer URLs.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["amazon-q", "aws", "developer-tools", "current-product"]
},
{
"name": "OpenTelemetry Documentation",
"url": "https://opentelemetry.io/docs/",
"type": "documentation",
"description": "Vendor-neutral observability framework docs.",
"relevance": "Use when instrumenting custom events, traces, and workflow metadata for internal AI coding systems.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": false,
"tags": ["opentelemetry", "observability", "instrumentation"]
},
{
"name": "PostHog Documentation",
"url": "https://posthog.com/docs",
"type": "documentation",
"description": "Product analytics and experimentation docs.",
"relevance": "Useful for instrumenting internal AI workflow events and experiments.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": false,
"tags": ["posthog", "analytics", "experimentation"]
},
{
"name": "Apache DevLake",
"url": "https://devlake.apache.org/",
"type": "documentation",
"description": "Open-source engineering metrics platform with DORA support.",
"relevance": "Useful for team-level delivery instrumentation and baselining.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": true,
"tags": ["devlake", "dora", "delivery", "open-source"]
}
],
"platform_and_practice": [
{
"name": "DX Platform",
"url": "https://getdx.com/",
"type": "platform",
"description": "Developer experience platform and survey methodology vendor.",
"relevance": "Useful for survey design patterns and DX language, but treat benchmarking claims as vendor evidence.",
"status": "live",
"last_verified": "2026-05-17",
"evidence_tier": "vendor",
"add_as_web_search": true,
"tags": ["dx", "surveys", "developer-experience", "vendor"]
},
{
"name": "DX Core 4: Developer Productivity Measurement Framework",
"url": "https://getdx.com/research/dx-core-4/",
"type": "framework",
"description": "DX's four-dimension developer productivity framework. Covers speed, effectiveness, quality, and impact. A framework for measuring holistic developer productivity beyond narrow throughput metrics.",
"relevance": "Use when the team needs a structured measurement model that goes beyond DORA delivery metrics. Treat specific benchmark figures as vendor evidence pending independent replication.",
"status": "live",
"last_verified": "2026-05-17",
"evidence_tier": "vendor",
"add_as_web_search": true,
"tags": ["dx", "core-4", "framework", "developer-productivity", "measurement"]
},
{
"name": "LinearB Engineering Benchmarks",
"url": "https://linearb.io/platform/engineering-metrics",
"type": "platform",
"description": "Benchmark content for delivery metrics and engineering analytics.",
"relevance": "Useful for directional comparison only; not a substitute for internal controls.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "vendor",
"add_as_web_search": true,
"tags": ["linearb", "benchmarks", "delivery", "vendor"]
},
{
"name": "Swarmia Engineering Metrics",
"url": "https://www.swarmia.com/engineering-metrics/",
"type": "platform",
"description": "Engineering metrics and workflow commentary from a vendor platform.",
"relevance": "Useful for implementation ideas, but treat performance claims as vendor evidence.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "vendor",
"add_as_web_search": true,
"tags": ["swarmia", "engineering-metrics", "vendor"]
},
{
"name": "Grafana Dashboards",
"url": "https://grafana.com/grafana/dashboards/",
"type": "tool",
"description": "Visualization catalog for dashboards and engineering telemetry.",
"relevance": "Useful for packaging delivery and quality metrics once definitions are stable.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": false,
"tags": ["grafana", "dashboards", "visualization"]
},
{
"name": "SonarQube",
"url": "https://www.sonarsource.com/products/sonarqube/",
"type": "tool",
"description": "Static analysis and code quality platform.",
"relevance": "Useful as a quality-side measurement input when tracking AI-generated or AI-assisted code.",
"status": "live",
"last_verified": "2026-05-15",
"evidence_tier": "official",
"add_as_web_search": false,
"tags": ["sonarqube", "quality", "security", "static-analysis"]
}
]
}
}
learnings.consolidated.md
# dev-ai-coding-metrics — Consolidated Learnings
Curated, dated, committed memory for this skill. Pruned from raw `learnings.md` via `agents-skills-feedback-loop/scripts/consolidate.py`. Human-approved.
Cap: 60 entries. When exceeded, promote durable rules to `references/`.
## Filter Override
<!-- Add 2-4 bullets that sharpen what counts as a learning for this skill. Leave empty to use the default filter from agents-skills-feedback-loop/references/learnings-format.md. -->
## Patterns That Work
## Mistakes to Avoid
## Domain Knowledge
## Open Questions
## Consolidated Principles
learnings.md
# dev-ai-coding-metrics — Learnings
## Patterns That Work
## Mistakes to Avoid
## Domain Knowledge
- [2026-07-11] DORA's 2026 ROI report (dora.dev/ai/roi/report/) is distinct from the 2025 AI report: adds a J-Curve dip model, scenario-based ROI, and an 'instability tax' (change-failure rate rising with adoption). Cite separately.
## Open Questions
## Consolidated Principles
references/adoption-metrics.md
# Adoption Metrics for AI Coding Programs
Operational guidance for measuring whether developers are actually using AI coding tools, in which mode, and with what degree of stickiness. This reference covers both assistant-style tools and coding agents.
---
## Table of Contents
- [What Adoption Measurement Should Answer](#what-adoption-measurement-should-answer)
- [The Two Funnels](#the-two-funnels)
- [Assistant Funnel](#assistant-funnel)
- [Agent Funnel](#agent-funnel)
- [Core Adoption Metrics](#core-adoption-metrics)
- [Metric Definitions](#metric-definitions)
- [Assistant vs Agent Interpretation](#assistant-vs-agent-interpretation)
- [Tool-Specific Data Sources](#tool-specific-data-sources)
- [GitHub Copilot](#github-copilot)
- [ChatGPT Enterprise / OpenAI API Usage](#chatgpt-enterprise-openai-api-usage)
- [Claude Code and Similar CLI Agents](#claude-code-and-similar-cli-agents)
- [Amazon Q Developer](#amazon-q-developer)
- [Internal Agents](#internal-agents)
- [Segmentation Rules](#segmentation-rules)
- [Stall Patterns](#stall-patterns)
- [Pattern: Seat Bloat](#pattern-seat-bloat)
- [Pattern: Completion-Only Plateau](#pattern-completion-only-plateau)
- [Pattern: Agent Curiosity Without Trust](#pattern-agent-curiosity-without-trust)
- [Pattern: Mandated Adoption with DX Backlash](#pattern-mandated-adoption-with-dx-backlash)
- [Privacy and Governance](#privacy-and-governance)
- [What to Do Next](#what-to-do-next)
## What Adoption Measurement Should Answer
Adoption is not "how many seats did we buy?" It should answer:
- who is using the tool regularly
- which capabilities they use
- whether usage is voluntary or forced
- whether usage is shallow, sticky, or expanding
- whether agent invocation is producing accepted outcomes
If adoption is weak, do not over-interpret delivery or ROI metrics.
## The Two Funnels
Use different funnels for assistants and agents.
### Assistant Funnel
```
licensed seat
-> activated account
-> weekly active user
-> repeated usage across 4+ weeks
-> feature breadth beyond autocomplete
-> retained/accepted output
```
### Agent Funnel
```
agent access granted
-> task submitted
-> task runs to completion
-> human accepts output for review
-> PR merges
-> low revert / low exception rate
```
Do not collapse those funnels into one "usage" number.
---
## Core Adoption Metrics
| Metric | What It Answers | Good Default Slice |
|--------|-----------------|--------------------|
| License utilization | are purchased seats active | org, team, business unit |
| WAU / MAU or active-user rate | is usage recurring | team, role, seniority band |
| Feature breadth | are teams using more than one surface | completions, chat, edit, review, agent |
| Suggestion acceptance or persistence | is inline output useful enough to keep | language, repo, editor |
| Organic usage ratio | do developers choose the tool without pressure | pilot cohort, optional vs mandated |
| Agent invocation rate | are teams actually trying autonomous flows | workflow, task type |
| Agent completion rate | does invocation turn into finished work | task type, repo, team |
| PR merge rate of AI-created work | does usage create accepted outcomes | agent, reviewer group, task type |
### Metric Definitions
**License utilization**
```
active_users / licensed_seats
```
Use a rolling 28-day window, not point-in-time daily counts.
**Feature breadth**
```
distinct_capabilities_used / capabilities_available
```
Capabilities are tool-specific. For assistants, they usually include completions, chat, edit, code review, docs, test generation. For agents, they include task execution, branch creation, test runs, PR generation, and follow-up remediation.
**Suggestion acceptance**
```
accepted_suggestions / shown_suggestions
```
Treat this as a weak signal unless you can pair it with persistence or retained-code metrics.
**Agent completion rate**
```
tasks_marked_complete / tasks_started
```
Only count "complete" when the task exits the agent workflow successfully, not when a run stops without an error.
**PR merge rate for AI-created work**
```
merged_prs / prs_created_by_tool_or_agent
```
This is often more useful than raw invocation count.
---
## Assistant vs Agent Interpretation
| Pattern | Likely Meaning | Next Check |
|---------|----------------|-----------|
| High seat utilization, low feature breadth | shallow trial behavior | onboarding, docs, workflow fit |
| High acceptance, low persistence | suggestions look good but are later removed | quality and rework metrics |
| High invocation, low completion | agents are interesting but unreliable | handoff, exception, tool stability |
| High completion, low merge rate | agents finish tasks but reviewers reject them | review burden, architecture fit |
| High adoption, flat delivery | tool used mostly on low-value work | task mix, delivery decomposition |
| High adoption, falling satisfaction | usage may be coerced or frustrating | survey, trust burden, give-up rate |
---
## Tool-Specific Data Sources
Use first-party telemetry when available. Tool docs change often, so prefer current admin/usage docs over blog posts and old API examples.
### GitHub Copilot
Default:
- use the current GitHub Copilot admin metrics documentation and dashboards
- prefer the official "About metrics for GitHub Copilot" docs for metric meanings
- treat older REST examples found in blogs and gists as potentially stale
Track:
| Metric | Use |
|--------|-----|
| active users | recurring adoption |
| engaged users by feature | breadth across completions, chat, review, etc. |
| suggestion counts and acceptance | inline relevance |
| editor / language splits | where value is concentrated |
| seat and entitlement data | procurement hygiene |
### ChatGPT Enterprise / OpenAI API Usage
For ChatGPT workspace adoption:
- use admin analytics for active users, engagement trends, and feature usage
- separate chat workspace activity from API-driven agent or tooling activity
For API-driven AI coding usage:
- attribute requests by service, repo, or workflow
- tag assistant-style requests separately from autonomous runs
- pair token usage with accepted outcomes, not just request count
Track:
| Metric | Use |
|--------|-----|
| unique active users | workspace adoption |
| conversations or sessions | recurring engagement |
| requests / tokens by workflow | operational cost and usage split |
| tasks or runs created from API integrations | agent-style automation adoption |
### Claude Code and Similar CLI Agents
Tool surfaces change quickly. Prefer local session logs, audit/billing logs, and SCM events over undocumented assumptions.
Track:
| Metric | Use |
|--------|-----|
| sessions started | top-of-funnel usage |
| tasks completed | outcome-level adoption |
| tools invoked per session | how deeply the agent is used |
| files touched or repos touched | scope of work |
| human feedback rounds | supervision overhead |
| follow-on PRs merged | accepted downstream value |
### Amazon Q Developer
Use current Amazon Q Developer docs and admin reporting, not historical CodeWhisperer naming.
Track:
| Metric | Use |
|--------|-----|
| active users | seat utilization |
| accepted suggestions | inline value |
| chat / assistant usage | breadth |
| security scan or remediation usage | workflow coverage |
### Internal Agents
For in-house agents, define the event model yourself. Minimum recommended events:
- task_created
- task_started
- task_completed
- task_failed
- human_takeover
- pr_opened
- pr_merged
- pr_reverted
- policy_exception
Without this event model, you cannot measure adoption cleanly.
---
## Segmentation Rules
Segment adoption or the averages will mislead you.
Recommended slices:
- team or org unit
- repo or service
- task type: bugfix, test, refactor, feature, incident
- tool surface: completion, chat, edit, review, agent
- seniority band
- language / stack
- optional vs mandated cohort
Avoid cross-team comparisons unless stack, task mix, and delivery environment are roughly comparable.
---
## Stall Patterns
### Pattern: Seat Bloat
Signal:
- high purchased-seat count
- falling active-user ratio
- low repeated usage
Interpretation:
- procurement got ahead of workflow value
- rollout outpaced enablement
Action:
- reclaim seats
- focus on power-user teams
- stop treating purchased seats as success
### Pattern: Completion-Only Plateau
Signal:
- lots of autocomplete usage
- almost no chat, edit, review, or agent usage
Interpretation:
- users see the tool as a typing accelerator, not a workflow tool
Action:
- teach 2-3 high-value patterns by task type
- measure breadth monthly
### Pattern: Agent Curiosity Without Trust
Signal:
- agent invocation rises
- completion or merge rate stays low
Interpretation:
- teams are curious but do not trust the output
Action:
- measure reviewer effort and handoff rate
- narrow agents to smaller task envelopes
### Pattern: Mandated Adoption with DX Backlash
Signal:
- active-user rate looks good
- satisfaction and trust fall
- give-up rate rises
Interpretation:
- usage is being driven by policy rather than value
Action:
- separate voluntary and mandated cohorts
- redesign enablement before expanding the mandate
---
## Privacy and Governance
Adoption dashboards can become surveillance systems unless you set boundaries.
Rules:
1. Report to management at team level by default.
2. Use a minimum team size of 5 before publishing named team comparisons.
3. Keep any individual-level dashboard self-service and invisible to managers.
4. State explicitly that metrics are for tooling and workflow improvement, not performance review.
5. If task-level agent data is sensitive, anonymize before analysis.
Suggested language:
> We track tool and workflow effectiveness at team level. We do not use individual AI-usage metrics for performance management.
---
## What to Do Next
- If the question is "why aren't people using the tool?", pair this file with `developer-experience-metrics.md`.
- If the question is "are agents actually helping?", go to `agent-execution-metrics.md`.
- If the question is "what outcome did adoption create?", go to `productivity-metrics.md` and `quality-metrics.md`.
references/agent-execution-metrics.md
# Agent Execution Metrics
Use this reference when the workflow involves coding agents that can plan, edit files, run tests, open PRs, or complete multi-step tasks. These metrics are different from assistant metrics because the unit of value is usually a **task** or **accepted change**, not a suggestion.
---
## Table of Contents
- [Why Agent Metrics Need Their Own Model](#why-agent-metrics-need-their-own-model)
- [The Agent Funnel](#the-agent-funnel)
- [Core Metrics](#core-metrics)
- [Recommended Event Model](#recommended-event-model)
- [Interpreting the Funnel](#interpreting-the-funnel)
- [Minimum Viable Scorecards](#minimum-viable-scorecards)
- [Scorecard A: Early Agent Pilot](#scorecard-a-early-agent-pilot)
- [Scorecard B: Scaling Agent Workflow](#scorecard-b-scaling-agent-workflow)
- [Scorecard C: Executive Decision](#scorecard-c-executive-decision)
- [Reviewer Burden](#reviewer-burden)
- [Quality and Safety](#quality-and-safety)
- [Benchmark Use](#benchmark-use)
- [Segmentation](#segmentation)
- [Study Designs That Work Well for Agents](#study-designs-that-work-well-for-agents)
- [Executive Summary Template](#executive-summary-template)
- [Anti-Patterns](#anti-patterns)
- [What to Do Next](#what-to-do-next)
## Why Agent Metrics Need Their Own Model
Assistant metrics answer:
- did the tool help a developer in the flow of work?
Agent metrics answer:
- did the workflow complete useful work with acceptable human oversight and acceptable downstream quality?
This means the key questions change:
- what percentage of runs actually finish?
- how often does a human need to intervene?
- how much reviewer effort is required for accepted output?
- how often does accepted output survive after merge?
---
## The Agent Funnel
Measure the full funnel:
```text
task created
-> task started
-> run completes
-> human accepts for review
-> PR opens
-> PR merges
-> no revert / no hotfix / no policy exception
```
If you only measure the first half of the funnel, you will overstate value.
---
## Core Metrics
| Metric | Definition | Why It Matters |
|--------|------------|----------------|
| Task completion rate | completed runs / started runs | top-of-funnel outcome |
| Human takeover rate | runs requiring rescue / started runs | reliability and scope fit |
| Acceptance-for-review rate | outputs accepted for PR or handoff / completed runs | catches superficial "success" |
| PR merge rate | merged PRs / PRs opened by agent | real-world acceptance |
| Revert or hotfix rate | reverted or hotfixed merged PRs / merged PRs | downstream reliability |
| Reviewer effort per accepted task | reviewer minutes, comments, or rounds / accepted task | hidden human cost |
| Policy / security exception rate | exceptions / started runs | safety boundary signal |
| Cost per accepted task | total operating cost / accepted tasks | unit economics |
| Cost per merged PR | total operating + reviewer cost / merged agent PRs | scaling decision |
| Benchmark-to-production gap | benchmark score vs merge / revert outcomes | anti-self-deception metric |
| Extension robustness | retained correctness and acceptable quality across an evolving-spec, carried-workspace checkpoint sequence | reveals deterioration hidden by one-shot completion |
| Structural erosion slope | change in structural-erosion signal across ordered checkpoints | detects complexity concentrating as the workspace evolves |
| Verbosity slope | change in locally defined redundancy/clone signal across ordered checkpoints | detects redundant code accumulating over time |
| Late-checkpoint cost and review burden | runtime cost and reviewer or remediation effort in late/final checkpoints | exposes compounding maintenance cost |
---
## Recommended Event Model
For internal agents, define these events:
- `task_created`
- `task_started`
- `task_completed`
- `task_failed`
- `human_takeover`
- `handoff_accepted`
- `pr_opened`
- `pr_merged`
- `pr_reverted`
- `policy_exception`
- `security_exception`
- `trajectory_checkpoint_completed`
Recommended properties:
- `agent_name`
- `model_name`
- `repo`
- `team`
- `task_type`
- `risk_level`
- `runtime_seconds`
- `input_cost`
- `output_cost`
- `human_rework_minutes`
- `trajectory_id`
- `checkpoint_id`
- `parent_checkpoint_id`
- `spec_version`
- `workspace_identity`
- `workspace_hash`
- `progress_phase`
- `strict_correct`
- `isolated_correct`
- `core_correct`
- `regression_correct`
- `erosion`
- `verbosity`
- `cost`
- `duration`
Without these events, you cannot build reliable agent metrics.
---
## Interpreting the Funnel
| Pattern | Interpretation | Likely Action |
|---------|----------------|---------------|
| high start rate, low completion | workflow is unstable or over-scoped | narrow task envelope |
| high completion, low acceptance-for-review | "complete" does not mean useful | tighten success definition |
| high PR open rate, low merge rate | reviewers do not trust the output | improve guardrails and task selection |
| high merge rate, high revert rate | reviewers are missing real problems | add stronger quality gates |
| good merge rate, rising reviewer effort | agent output is acceptable but expensive | improve prompt/context quality or reduce scope |
| low completion, low takeover | tasks may be abandoned silently | improve failure tagging |
---
## Minimum Viable Scorecards
### Scorecard A: Early Agent Pilot
Use when the goal is to learn whether the workflow is viable.
Track:
- task completion rate
- human takeover rate
- acceptance-for-review rate
- PR merge rate
- reviewer effort per accepted task
- policy / security exceptions
- extension-robustness sequence result for edit-capable agents
- structural erosion and verbosity slopes across checkpoints
### Scorecard B: Scaling Agent Workflow
Use when the pilot is already technically viable.
Track:
- all Scorecard A metrics
- revert / hotfix rate
- cost per accepted task
- cost per merged PR
- completion rate by task type
- benchmark-to-production gap
- late/final checkpoint cost and reviewer or remediation burden
### Scorecard C: Executive Decision
Use when leadership needs a program-level call.
Track:
- accepted tasks per month
- merged PRs per month
- reviewer effort trend
- quality / exception trend
- cost per accepted task
- extension-robustness trend by task class
- late-checkpoint cost and review-burden trend
- recommendation: expand, narrow, or stop
---
## Reviewer Burden
Reviewer burden is often the most important agent metric.
Measure at least one of:
- review time per accepted PR
- review rounds per accepted PR
- requested-changes rate
- substantive comments per accepted PR
Interpretation:
- if output volume rises while reviewer burden rises faster, net value may be negative
- if merge rate is stable and reviewer burden falls, the workflow is improving
Do not claim autonomy gains without reviewer-burden data.
For iterative agent evaluations, segment cost and review or remediation effort by progress phase. Stable aggregate cost can conceal a late-checkpoint spike caused by accumulated design debt.
---
## Quality and Safety
Agent metrics must include safety.
Recommended safety metrics:
- security findings introduced on agent-created code
- policy exceptions or manual overrides
- test failures after agent completion
- incidents or hotfixes attributable to agent-created changes
If a workflow crosses a security or production boundary, this section is mandatory.
---
## Benchmark Use
Use benchmarks as capability evidence only.
Good benchmark uses:
- compare models before a controlled internal test
- identify whether a task envelope is plausibly automatable
- watch capability regressions after a model or prompt change
Bad benchmark uses:
- calling a high benchmark score proof of ROI
- using benchmark wins to skip reviewer-burden measurement
- assuming benchmark gains transfer to messy internal repos
Better framing:
> Benchmarks tell us what the agent may be capable of. Production acceptance tells us whether that capability survives contact with our repos, standards, and reviewers.
For edit, refactor, and migration agents, add at least one evolving-spec sequence with three or more checkpoints. Use a fresh conversation/context at each checkpoint, carry the same agent-created workspace forward, and retain prior regression tests. Report the correctness trajectory beside structural-erosion and verbosity slopes plus late-checkpoint cost and review burden. A one-shot green suite or a planning/quality prompt is not evidence of long-run extension robustness.
SlopCodeBench v1 motivates this measurement shape, but its paper averages are not organizational targets and its trajectory signals do not establish causal ROI. Use `qa-agent-testing` for benchmark protocol, `ai-coding-agents-observability-evals` for lineage and telemetry, and `software-clean-code-standard` for metric interpretation.
---
## Segmentation
Segment agent metrics by:
- task type: bugfix, test, refactor, migration, docs, incident
- repo or service
- risk level
- agent or model version
- reviewer group
Do not compare all task types in one blended acceptance rate.
---
## Study Designs That Work Well for Agents
Best options:
- blind reviewer comparison on paired outputs
- shadow workflow on the same task class
- staggered rollout by task type
- production funnel tracking with mandatory reviewer tags
Useful task-level questions:
- which tasks complete cleanly?
- which tasks complete but fail review?
- which tasks should remain assistant-only rather than agentic?
Cross-reference: `benchmarking-methodology.md`
---
## Executive Summary Template
When summarizing agent performance, use this structure:
1. what task envelope was in scope
2. how many runs started
3. how many produced accepted work
4. what reviewer effort was required
5. what quality and safety signals looked like
6. cost per accepted task or merged PR
7. recommendation on scope expansion or narrowing
Suggested one-line summary:
> The agent completed 61% of scoped tasks, 34% were accepted into review, 22% merged, reviewer effort per merged PR fell after narrowing the task envelope, and post-merge quality remained stable; expand only within the current task class.
---
## Anti-Patterns
| Anti-Pattern | Why It Fails |
|-------------|--------------|
| counting all completed runs as wins | many "successful" runs are not accepted |
| tracking starts without takeovers | hides hidden human rescue |
| using merge rate without revert rate | misses downstream reliability |
| calling benchmark gains business value | benchmark != accepted production work |
| ignoring reviewer cost | overstates ROI materially |
| blending task types | high-volume easy tasks hide hard-task failures |
| using preprint averages as targets | substitutes another benchmark's task/model mix for local evidence |
| reporting only final-checkpoint quality | hides when and how degradation or cost accumulated |
---
## What to Do Next
- For broader team delivery metrics, use `productivity-metrics.md`.
- For cost modeling, use `roi-framework.md`.
- For experiments, use `benchmarking-methodology.md`.
references/benchmarking-methodology.md
# Benchmarking Methodology for AI Coding Impact
Study designs, statistical methods, confound management, and reporting structures for measuring AI coding tool impact with rigor.
---
## Table of Contents
- [A/B Team Comparison](#ab-team-comparison)
- [Team Selection Criteria](#team-selection-criteria)
- [Assignment](#assignment)
- [Duration](#duration)
- [Crossover Design](#crossover-design)
- [Metrics to Collect](#metrics-to-collect)
- [Statistical Analysis](#statistical-analysis)
- [Before/After Study Design](#beforeafter-study-design)
- [Baseline Establishment](#baseline-establishment)
- [Intervention Period](#intervention-period)
- [Interrupted Time Series Analysis](#interrupted-time-series-analysis)
- [Controlling for Seasonal Effects](#controlling-for-seasonal-effects)
- [Multiple Baseline Design](#multiple-baseline-design)
- [Shadow Team Experiments](#shadow-team-experiments)
- [Design](#design)
- [Blind Evaluation](#blind-evaluation)
- [Time and Cost Comparison](#time-and-cost-comparison)
- [Limitations](#limitations)
- [Statistical Rigor](#statistical-rigor)
- [Sample Size Calculations](#sample-size-calculations)
- [Effect Size Estimation](#effect-size-estimation)
- [Confidence Intervals](#confidence-intervals)
- [Multiple Comparison Corrections](#multiple-comparison-corrections)
- [Power Analysis](#power-analysis)
- [Practical vs Statistical Significance](#practical-vs-statistical-significance)
- [Confounding Variable Management](#confounding-variable-management)
- [Developer Skill Level](#developer-skill-level)
- [Task Complexity](#task-complexity)
- [Project Lifecycle Stage](#project-lifecycle-stage)
- [Team Dynamics and Collaboration](#team-dynamics-and-collaboration)
- [External Factors](#external-factors)
- [Tool Version Changes](#tool-version-changes)
- [Reporting Structure](#reporting-structure)
- [Executive Summary](#executive-summary)
- [Detailed Methodology Section](#detailed-methodology-section)
- [Results with Confidence Intervals](#results-with-confidence-intervals)
- [Limitations and Threats to Validity](#limitations-and-threats-to-validity)
- [Recommendations with Evidence Strength](#recommendations-with-evidence-strength)
- [Replication Guidelines](#replication-guidelines)
## A/B Team Comparison
The strongest causal design when you have enough teams. Two comparable groups: one uses AI tools (treatment), one does not (control).
### Team Selection Criteria
Match teams on these dimensions:
| Dimension | How to Match | Tolerance |
|-----------|-------------|-----------|
| Team size | Same +/- 1 person | Strict |
| Avg seniority | Within 1 year of experience | Moderate |
| Tech stack | Same primary language and framework | Strict |
| Project type | Same lifecycle stage (greenfield, growth, maintenance) | Strict |
| Historical velocity | Within 20% of each other (past 3 months) | Moderate |
| Code complexity | Similar cyclomatic complexity per repo | Moderate |
| Team tenure | How long they've worked together | Moderate |
### Assignment
- **Random assignment** is ideal but rarely practical. At minimum, do not let teams self-select.
- **Stratified assignment**: If you have 6+ teams, stratify by project type and randomly assign within strata.
- **Avoid**: Assigning the "best" team to treatment (guarantees confounded results).
### Duration
| Phase | Duration | Purpose |
|-------|----------|---------|
| Pre-measurement baseline | 4 weeks minimum | Establish comparable starting metrics |
| Phase 1 (treatment) | 8 weeks minimum | Treatment team uses AI tools; control does not |
| Washout (optional) | 2 weeks | Both teams pause AI usage |
| Phase 2 (crossover) | 8 weeks | Teams swap: former control gets AI tools |
Total: 20-24 weeks for a crossover design.
**Why 8 weeks minimum**: Weeks 1-3 are adoption ramp (noise). Weeks 4-8 represent stabilized usage. Shorter studies measure novelty, not impact.
### Crossover Design
Each team serves as its own control, which dramatically increases statistical power.
```
Week 1-8 Week 9-10 Week 11-18
Team A: AI tools Washout No AI tools
Team B: No AI tools Washout AI tools
Analysis: Compare each team's AI period to its non-AI period.
Also compare Team A's AI period to Team B's non-AI period.
```
Advantages:
- Controls for team-level confounds (skill, culture, codebase)
- Requires fewer teams for same statistical power
- More palatable than permanently denying tools to a group
Disadvantages:
- Learning effects (Team A may retain AI skills during non-AI period)
- Longer total study duration
- Cannot fully "un-learn" AI-assisted patterns
### Metrics to Collect
Collect all of these for both teams in both periods:
**Primary metrics** (must have):
- Cycle time (commit to deploy)
- Throughput (story points or tasks completed per sprint)
- Defect rate (bugs per 100 commits or per feature)
- Code review turnaround time
**Secondary metrics** (should have):
- Lines of code (context only, not as a productivity measure)
- Test coverage delta
- PR size and frequency
- Developer satisfaction survey (pre and post each phase)
**Tertiary metrics** (nice to have):
- Cognitive load assessment
- Context switch frequency
- Time spent on specific task categories
### Statistical Analysis
- **Primary test**: Mixed-effects model with team as random effect and treatment as fixed effect
- **Simpler alternative**: Paired t-test on team-period averages (if crossover design)
- **Report**: Effect size (Cohen's d), 95% confidence interval, p-value
- **Minimum detectable effect**: Calculate before starting (see Statistical Rigor section)
---
## Before/After Study Design
When you cannot withhold AI tools from any team. Compare the same team's performance before and after AI tool adoption.
### Baseline Establishment
```
BASELINE PERIOD:
Minimum: 8 weeks
Recommended: 12 weeks
Purpose: Establish stable pre-AI metrics
REQUIREMENTS:
- No major process changes during baseline
- No team composition changes
- Normal project work (not a release crunch or lull)
- Collect the same metrics you'll measure post-intervention
```
**Common mistake**: Starting baseline measurement the week you announce AI tools are coming. Anticipation effects distort the baseline.
### Intervention Period
| Phase | Duration | Notes |
|-------|----------|-------|
| Rollout and training | 2-4 weeks | Not counted in post-measurement |
| Ramp-up | 4-6 weeks | Usage stabilizing, some measurement |
| Stabilized usage | 8+ weeks | Primary measurement period |
Do not compare baseline to ramp-up period. Compare baseline to stabilized period only.
### Interrupted Time Series Analysis
The gold standard for before/after designs. Requires frequent metric observations (weekly minimum).
```
APPROACH:
1. Plot weekly metric values for entire study period
2. Mark the intervention point (AI tool rollout)
3. Fit a regression model:
Y = β0 + β1(time) + β2(intervention) + β3(time × intervention) + ε
INTERPRETATION:
β1: Pre-existing trend (was the metric already improving?)
β2: Immediate effect of intervention (level change)
β3: Change in trend after intervention (slope change)
```
If β1 shows the metric was already improving before AI tools, the post-intervention improvement may not be attributable to AI tools.
### Controlling for Seasonal Effects
Software teams have predictable rhythms:
- **Sprint boundaries**: Velocity spikes at sprint end
- **Quarter-end**: Release pressure distorts metrics
- **Holidays**: Reduced capacity
- **Annual planning**: Strategy work displaces coding
- **Hiring cycles**: New team members reduce velocity temporarily
**Mitigation**: Compare same calendar period year-over-year if possible, or use seasonal adjustment in time series models.
### Multiple Baseline Design
Stagger AI tool rollout across teams:
```
Month 1-2 Month 3-4 Month 5-6 Month 7-8
Team A: Baseline AI tools → AI tools AI tools
Team B: Baseline Baseline AI tools → AI tools
Team C: Baseline Baseline Baseline AI tools →
If all teams improve at their specific rollout point (not before),
the effect is more likely causal than coincidental.
```
This is the most rigorous before/after design available without a control group.
---
## Shadow Team Experiments
The most rigorous but most expensive approach: two teams independently complete the same task.
Cross-reference: `dev-context-engineering/references/team-transformation-patterns.md` for the Shadow Team Experiment framework, including setup, evaluation criteria, and documented case studies.
### Design
```
SETUP:
1. Select a meaningful feature or task
- Not trivial (would take traditional team 2-4 weeks)
- Not mission-critical (results used for learning, not shipping)
- Well-defined acceptance criteria
2. Assign two teams:
- AI-Equipped Team: Full access to AI coding tools
- Traditional Team: Standard tooling only
- Teams should NOT know each other's assignment
3. Identical success criteria:
- Feature completeness (checklist)
- Test coverage minimums
- Code quality thresholds (linting, review)
- Documentation requirements
```
### Blind Evaluation
After both teams deliver:
1. **Strip identifying information**: Remove author names, commit messages that mention tools
2. **Independent reviewers**: 3+ senior engineers who were not on either team
3. **Evaluation rubric**:
| Dimension | Weight | Scale |
|-----------|--------|-------|
| Functional correctness | 25% | 1-5 |
| Code quality and maintainability | 25% | 1-5 |
| Test coverage and quality | 20% | 1-5 |
| Architecture decisions | 15% | 1-5 |
| Documentation | 15% | 1-5 |
4. **Record evaluation time**: How long reviewers spend understanding each codebase (proxy for maintainability)
### Time and Cost Comparison
| Metric | AI-Equipped | Traditional | Difference |
|--------|-------------|-------------|------------|
| Calendar time to completion | X days | Y days | Y - X |
| Total person-hours | A hours | B hours | B - A |
| Fully-loaded cost | $C | $D | $D - $C |
| Tool costs | $E | $0 | -$E |
| Net cost difference | | | ($D - $C) - $E |
### Limitations
- **Expensive**: Duplicating work is a hard sell to leadership
- **Artificial**: Developers know it is an experiment (Hawthorne effect)
- **Small sample**: Usually only 1-2 comparisons feasible (low statistical power)
- **Task-dependent**: Results from one task may not generalize
**When it is worth it**: Major investment decisions (enterprise-wide rollout), or when before/after data is unconvincing.
---
## Statistical Rigor
### Sample Size Calculations
Before running any study, determine how many data points you need.
```
SAMPLE SIZE FORMULA (simplified for two-group comparison):
n per group = 2 × ((z_α/2 + z_β) / d)²
Where:
z_α/2 = 1.96 (for 95% confidence)
z_β = 0.84 (for 80% power)
d = expected effect size (Cohen's d)
PRACTICAL MINIMUMS:
Large effect (d = 0.8): ~26 observations per group
Medium effect (d = 0.5): ~64 observations per group
Small effect (d = 0.2): ~394 observations per group
```
For team-level comparisons, "observations" = team-sprints. For individual metrics, "observations" = developer-tasks.
**Reality check**: Most organizations cannot get 394 developer-task observations per condition. This means small effects will not be detectable. Plan for detecting medium-to-large effects only.
### Effect Size Estimation
| Metric | Plausible Effect Size | Cohen's d Category |
|--------|----------------------|-------------------|
| Code generation speed | Large (d = 0.8-1.2) | Large |
| Overall sprint velocity | Small-Medium (d = 0.2-0.5) | Small-Medium |
| Bug rate change | Small (d = 0.1-0.3) | Small |
| Developer satisfaction | Medium (d = 0.4-0.7) | Medium |
| Code review time | Medium (d = 0.3-0.6) | Medium |
Use pilot data to estimate effect sizes for your organization. Vendor claims often correspond to d > 1.5, which is implausibly large.
### Confidence Intervals
Report 95% confidence intervals for every metric. A point estimate without a confidence interval is useless.
```
EXAMPLE:
"AI-assisted developers completed tasks 22% faster
(95% CI: 12% to 32%, p = 0.003)"
NOT:
"AI-assisted developers completed tasks 22% faster"
```
Wide confidence intervals (e.g., -5% to 49%) indicate insufficient data. Do not claim an effect exists when the CI includes zero.
### Multiple Comparison Corrections
When testing many metrics simultaneously, some will appear significant by chance.
| Number of Metrics Tested | Expected False Positives (p < 0.05) |
|--------------------------|--------------------------------------|
| 5 | 0.25 (1 in 4 studies) |
| 10 | 0.50 (1 in 2 studies) |
| 20 | 1.0 (expect at least 1) |
**Bonferroni correction**: Divide significance threshold by number of tests. If testing 10 metrics, use p < 0.005 instead of p < 0.05.
**Better approach**: Pre-register 2-3 primary metrics. Report others as exploratory.
### Power Analysis
```
MINIMUM STANDARD: β = 0.80 (80% power)
Interpretation: If the effect is real, you have an 80% chance of
detecting it. 20% chance of a false negative.
For high-stakes decisions: β = 0.90 (90% power)
PRACTICAL IMPACT:
At 80% power, 1 in 5 real effects go undetected.
At 90% power, 1 in 10 real effects go undetected.
At 50% power (common in underpowered studies), HALF of real
effects go undetected.
```
Run power analysis before the study to determine if you have enough data. If not, either extend the study or accept that you can only detect large effects.
### Practical vs Statistical Significance
A result can be statistically significant but practically meaningless (and vice versa).
```
EXAMPLE OF STATISTICALLY SIGNIFICANT BUT PRACTICALLY MEANINGLESS:
"AI tools reduced cycle time by 12 minutes per task
(95% CI: 3-21 minutes, p = 0.01)"
If tasks take 8 hours on average, 12 minutes is a 2.5% improvement.
Statistically real, but not worth the investment.
EXAMPLE OF PRACTICALLY SIGNIFICANT BUT STATISTICALLY UNCERTAIN:
"AI tools reduced cycle time by 2 hours per task
(95% CI: -0.5 to 4.5 hours, p = 0.11)"
Potentially large effect, but insufficient data to confirm.
Gather more data before deciding.
```
Define practical significance thresholds before the study: "We need at least X% improvement to justify the investment."
---
## Confounding Variable Management
### Developer Skill Level
The most common confound. Higher-skill developers adopt AI tools faster and are more productive regardless.
**Controls**:
- **Stratified analysis**: Report results separately for junior, mid, senior
- **Matched pairs**: Compare each developer to their own baseline (before/after)
- **Covariate adjustment**: Include years of experience in statistical models
- **Random assignment**: If possible, randomly assign developers to conditions within skill strata
### Task Complexity
AI tools excel at routine tasks and struggle with complex ones. If AI-equipped teams get easier tasks, results are confounded.
**Standardized task classification**:
| Complexity Level | Characteristics | Example |
|-----------------|-----------------|---------|
| C1: Routine | Well-defined, repetitive, boilerplate | CRUD endpoint, unit tests |
| C2: Standard | Some judgment required, established patterns | Feature implementation with clear spec |
| C3: Complex | Architecture decisions, multiple systems | API redesign, performance optimization |
| C4: Novel | No precedent, research required | New algorithm, unfamiliar domain |
Report results by complexity level. If AI shows gains only at C1-C2, that is an honest and useful finding.
### Project Lifecycle Stage
| Stage | AI Impact Profile |
|-------|------------------|
| Greenfield | Highest AI leverage; scaffolding, boilerplate, rapid prototyping |
| Active growth | High leverage; feature development, test generation |
| Mature/maintenance | Moderate leverage; bug fixes, refactoring, documentation |
| Legacy rescue | Low leverage; AI lacks context, outdated patterns |
Compare AI impact only within the same lifecycle stage. Greenfield-vs-maintenance comparisons are meaningless.
### Team Dynamics and Collaboration
- **Pair programming teams** may see different AI impact than solo developers
- **Code review culture** affects whether AI-generated bugs are caught
- **Communication overhead** varies with team size and distribution
- **Psychological safety** affects willingness to experiment with new tools
**Control**: Include team dynamics measures in your analysis. At minimum, track team size, co-location, and existing collaboration patterns.
### External Factors
| Factor | How It Confounds | Mitigation |
|--------|-----------------|------------|
| Market pressure / deadline | Teams work harder during crunch | Exclude crunch periods from analysis |
| Org restructuring | Disrupts teams regardless of AI | Note and control for in analysis |
| Major incidents / outages | Divert attention from feature work | Exclude affected sprints |
| New team members | Onboarding reduces velocity | Track team composition stability |
| Tech debt paydown | Planned maintenance skews metrics | Tag and separate from feature work |
### Tool Version Changes
AI tools update frequently. A mid-study tool update can:
- Improve results (making the tool look better than baseline)
- Degrade results (introducing new bugs or changing UX)
- Invalidate comparisons (before and after are measuring different tools)
**Mitigation**: Lock tool versions during study periods if possible. If not, document version changes and note them in the analysis.
---
## Reporting Structure
### Executive Summary
```
TEMPLATE (1 page):
STUDY: [Name/description]
PERIOD: [Start] to [End]
DESIGN: [A/B comparison | Before/after | Shadow team]
TEAMS: [N teams, N developers]
KEY FINDING:
[One sentence: what did we learn?]
PRIMARY METRICS:
| Metric | Treatment | Control | Difference | 95% CI |
|------------------|-----------|---------|------------|---------------|
| [Primary 1] | X | Y | +Z% | [low, high] |
| [Primary 2] | X | Y | +Z% | [low, high] |
RECOMMENDATION: [Continue / Expand / Modify / Pause]
CONFIDENCE: [High / Medium / Low] based on [rationale]
```
### Detailed Methodology Section
Include enough detail for replication:
1. **Study design** with justification for chosen approach
2. **Team/participant selection** criteria and process
3. **Timeline** with phases and durations
4. **Metrics** collected, data sources, collection frequency
5. **Statistical methods** used, software/tools, pre-registration status
6. **Deviations** from original plan and why
### Results with Confidence Intervals
Every metric reported must include:
- Point estimate
- 95% confidence interval
- Sample size
- Effect size (Cohen's d or equivalent)
- p-value (but emphasize CI over p-value)
Present results in tables with consistent formatting. Include visualizations (before/after trend lines, forest plots for multiple metrics).
### Limitations and Threats to Validity
Be explicit. Credibility comes from honesty about limitations.
| Threat | Severity | Mitigation Applied | Residual Risk |
|--------|----------|-------------------|---------------|
| Selection bias | High/Med/Low | [What you did] | [What remains] |
| Hawthorne effect | High/Med/Low | [What you did] | [What remains] |
| Small sample size | High/Med/Low | [What you did] | [What remains] |
| Confounding variables | High/Med/Low | [What you did] | [What remains] |
### Recommendations with Evidence Strength
Rate each recommendation by evidence quality:
| Strength | Meaning | Basis |
|----------|---------|-------|
| **Strong** | High confidence, act on this | Multiple metrics, large effect, narrow CI |
| **Moderate** | Likely true, proceed with monitoring | Consistent direction, moderate CI width |
| **Preliminary** | Suggestive, needs more data | Trend in expected direction, wide CI |
| **Insufficient** | Cannot conclude | Conflicting results, very wide CI, too few observations |
### Replication Guidelines
Enable others to repeat the study:
- Exact tool versions and configurations used
- Survey instruments (full question text)
- Data collection scripts or procedures
- Analysis code or statistical procedures
- Raw data (anonymized) if organizational policy allows
- Known limitations that affect reproducibility
```
REPLICATION CHECKLIST:
[ ] Study design documented with decision rationale
[ ] Team selection criteria and process recorded
[ ] All metrics defined with precise measurement methods
[ ] Data collection timeline and tools specified
[ ] Statistical analysis plan documented before data collection
[ ] Deviations from plan recorded with justification
[ ] Raw data preserved (anonymized) for re-analysis
[ ] Analysis code/procedures documented for reproducibility
```
references/developer-experience-metrics.md
# Developer Experience Metrics for AI Tools
Measurement frameworks for developer satisfaction, cognitive load, tool friction, onboarding, and trust when using AI coding tools.
---
## Table of Contents
- [Satisfaction Surveys](#satisfaction-surveys)
- [Survey Cadence](#survey-cadence)
- [Core Survey Questions (15 items)](#core-survey-questions-15-items)
- [Tool Net Promoter Score (tNPS)](#tool-net-promoter-score-tnps)
- [Benchmarking Sources](#benchmarking-sources)
- [Cognitive Load Measurement](#cognitive-load-measurement)
- [NASA-TLX Adaptation for AI-Assisted Coding](#nasa-tlx-adaptation-for-ai-assisted-coding)
- [Measurement Approach](#measurement-approach)
- [Interpreting Results](#interpreting-results)
- [Tool Friction Indicators](#tool-friction-indicators)
- [Context Switch Frequency](#context-switch-frequency)
- [Give-Up Rate](#give-up-rate)
- [Prompt Retry Rate](#prompt-retry-rate)
- [Suggestion Rejection Patterns](#suggestion-rejection-patterns)
- [Workflow Interruption Frequency](#workflow-interruption-frequency)
- [Developer Skill Atrophy](#developer-skill-atrophy)
- [Onboarding Metrics](#onboarding-metrics)
- [Time to First AI-Assisted PR](#time-to-first-ai-assisted-pr)
- [Time to Self-Sufficiency](#time-to-self-sufficiency)
- [Onboarding Completion Rate](#onboarding-completion-rate)
- [First-Week vs First-Month Usage Patterns](#first-week-vs-first-month-usage-patterns)
- [Mentor Dependency Duration](#mentor-dependency-duration)
- [Trust and Confidence Metrics](#trust-and-confidence-metrics)
- [Trust Calibration](#trust-calibration)
- [Over-Trust Indicators](#over-trust-indicators)
- [Under-Trust Indicators](#under-trust-indicators)
- [Confidence Progression](#confidence-progression)
- [Trust Recovery After AI Failures](#trust-recovery-after-ai-failures)
- [DX Anti-Patterns](#dx-anti-patterns)
- [Mandatory Usage Policies Without Support](#mandatory-usage-policies-without-support)
- [Individual Productivity Tracking and Surveillance](#individual-productivity-tracking-and-surveillance)
- [Comparing Developers by AI Usage](#comparing-developers-by-ai-usage)
- [Ignoring Negative Feedback](#ignoring-negative-feedback)
- [Removing Traditional Tools Before AI Tools Are Ready](#removing-traditional-tools-before-ai-tools-are-ready)
- [Additional Anti-Patterns](#additional-anti-patterns)
## Satisfaction Surveys
### Survey Cadence
| Phase | Cadence | Purpose |
|-------|---------|---------|
| Pilot (first 3 months) | Monthly | Track rapid sentiment changes during rollout |
| Post-pilot (steady state) | Quarterly | Ongoing monitoring without survey fatigue |
| After major tool changes | Ad-hoc (within 2 weeks) | Capture immediate reaction to updates |
**Response rate targets**: 70%+ for actionable data. Below 50% means results are unreliable due to self-selection bias.
Response rate improvement tactics:
- Keep surveys under 5 minutes
- Share results and actions taken from previous surveys
- Have engineering leadership visibly endorse participation
- Send during low-pressure periods (not during sprints or releases)
- Close the loop: "You said X, we did Y"
### Core Survey Questions (15 items)
Rate 1-5 (Strongly Disagree to Strongly Agree):
**Productivity Block**
1. AI tools help me complete coding tasks faster
2. AI tools reduce time I spend on repetitive work
3. I produce higher-quality code when using AI tools
4. AI tools help me learn new patterns and approaches
**Workflow Integration Block**
5. AI tools fit naturally into my existing workflow
6. I can easily switch between AI-assisted and manual coding
7. The AI tools are reliable enough for daily use
8. AI tool suggestions are relevant to my current task
**Trust and Confidence Block**
9. I trust AI-generated code enough to ship it after review
10. I can identify when AI output is incorrect or suboptimal
11. AI tools do not introduce more bugs than they prevent
**Satisfaction Block**
12. Overall, AI tools make my job more enjoyable
13. I would recommend our AI tools setup to a peer at another company
14. I feel supported in learning to use AI tools effectively
15. I prefer working at a company that provides AI coding tools
**Open-ended questions** (include 2-3):
- What is the biggest frustration with our current AI tools?
- Describe a recent situation where AI tools saved you significant time.
- What would you change about how we use AI tools?
### Tool Net Promoter Score (tNPS)
Adapt standard NPS for AI tools:
```
Question: "How likely are you to recommend [AI tool] to a
developer friend?" (0-10 scale)
Scoring:
Promoters (9-10) - Would actively recommend
Passives (7-8) - Satisfied but not enthusiastic
Detractors (0-6) - Would not recommend or would warn against
tNPS = % Promoters - % Detractors
```
| tNPS Range | Interpretation |
|------------|---------------|
| 50+ | Exceptional; tool is a competitive hiring advantage |
| 30-49 | Strong; developers see clear value |
| 10-29 | Moderate; value is real but friction exists |
| 0-9 | Weak; benefits and frustrations roughly balanced |
| Below 0 | Negative; tool is hurting developer experience |
### Benchmarking Sources
- **DX Company** (getdx.com): Developer experience benchmarks across 100+ companies
- **Stack Overflow Developer Survey**: Annual data on tool satisfaction and adoption
- **JetBrains Developer Ecosystem Survey**: IDE and tool usage patterns
- **SPACE Framework** (Microsoft Research): Satisfaction, Performance, Activity, Communication, Efficiency
---
## Cognitive Load Measurement
### NASA-TLX Adaptation for AI-Assisted Coding
The NASA Task Load Index measures perceived workload across 6 dimensions. Adapted for AI coding context:
| Dimension | Original Definition | AI Coding Adaptation | Scale |
|-----------|-------------------|---------------------|-------|
| **Mental Demand** | How mentally demanding was the task? | How much mental effort was needed to formulate prompts, evaluate AI output, and integrate suggestions? | Low (1) to High (7) |
| **Temporal Demand** | How rushed did you feel? | Did AI speed create pressure to work faster than comfortable? Did waiting for AI responses disrupt your flow? | Low (1) to High (7) |
| **Performance** | How successful were you? | How satisfied are you with the quality of the AI-assisted output? | Perfect (1) to Failure (7) |
| **Effort** | How hard did you work? | How much energy did you spend steering AI (prompting, correcting, refining) vs doing the work directly? | Low (1) to High (7) |
| **Frustration** | How frustrated were you? | How much frustration from hallucinations, irrelevant suggestions, tool crashes, or context loss? | Low (1) to High (7) |
| **Trust Burden** | (New dimension) | How much effort did you spend verifying AI output for correctness, security, and style? | Low (1) to High (7) |
### Measurement Approach
**Post-task survey** (best for controlled studies):
- Administer after completing a defined coding task
- Compare AI-assisted vs manual for same task type
- Minimum sample: 20 tasks per condition per developer
**Experience sampling** (best for ongoing monitoring):
- Random prompts 2-3x per day during coding
- Brief (30-second) assessment of current cognitive state
- Captures in-the-moment load rather than recalled load
### Interpreting Results
```
HEALTHY PATTERN:
Mental Demand: 3-4/7 (moderate, engaged but not overwhelmed)
Temporal Demand: 2-3/7 (AI speeds up without creating pressure)
Performance: 2-3/7 (developer is satisfied with output)
Effort: 3-4/7 (some steering needed, but net positive)
Frustration: 1-2/7 (minor friction, manageable)
Trust Burden: 2-3/7 (verification needed but not exhausting)
WARNING PATTERN:
Mental Demand: 5+/7 (prompting is harder than just coding)
Temporal Demand: 5+/7 (AI creates speed expectations)
Effort: 5+/7 (more energy steering than doing)
Frustration: 4+/7 (significant friction)
Trust Burden: 5+/7 (verification negates time savings)
```
---
## Tool Friction Indicators
Operational signals that reveal problems without requiring surveys.
### Context Switch Frequency
Track transitions between AI-assisted and manual coding within a session.
| Metric | How to Measure | Healthy Range |
|--------|---------------|---------------|
| Switches per hour | IDE telemetry or observation | 2-5 |
| Average AI-assisted stretch | Duration before switching to manual | 15+ minutes |
| Switch trigger | Log what caused the switch | Track categories |
Common switch triggers (track distribution):
- AI output not relevant to current task
- Need to think through architecture (AI not helpful)
- AI tool latency or downtime
- Task too complex for AI assistance
- Context window limitations
### Give-Up Rate
Percentage of tasks where a developer starts with AI but finishes manually.
```
Give-up rate = Tasks abandoned mid-AI / Total AI-started tasks
```
| Rate | Interpretation |
|------|---------------|
| < 10% | Excellent; AI is reliably useful |
| 10-25% | Normal; AI handles most tasks but not all |
| 25-40% | Concerning; significant task categories where AI fails |
| > 40% | Critical; tool is not fit for purpose or needs training |
Track give-up rate by task type to identify where AI tools underperform.
### Prompt Retry Rate
Average number of prompts before getting useful output.
| Attempts | Interpretation |
|----------|---------------|
| 1 | Ideal; developer has strong prompting skills |
| 2-3 | Normal; some refinement expected |
| 4-5 | Friction zone; prompting is becoming effortful |
| 6+ | Failure; developer should switch to manual |
### Suggestion Rejection Patterns
Categorize why developers reject AI suggestions:
- **Wrong approach**: AI solved the wrong problem
- **Wrong style**: Correct solution but doesn't match codebase patterns
- **Wrong scope**: Too much or too little code generated
- **Quality issue**: Bugs, security problems, or anti-patterns
- **Stale context**: AI used outdated information
- **Preference**: Developer simply prefers their approach
High rejection rates in specific categories reveal targeted improvement opportunities.
### Workflow Interruption Frequency
Count involuntary interruptions caused by AI tools:
- Tool crashes or freezes
- Unacceptable latency (> 5 seconds for inline suggestions, > 30 seconds for generation)
- Incorrect auto-complete that requires undo
- Distraction from unwanted suggestions
### Developer Skill Atrophy
Distinct from *tool* atrophy (Removing Traditional Tools, below) — this tracks decline in the developer's own ability to debug, design, and write code without AI assistance. AI-coding-heavy teams routinely report this as a personal concern after several months of intensive use (Karpathy, Jan 2026).
| Signal | How to measure | Concerning threshold |
|--------|---------------|---------------------|
| Unassisted-debug time | Time-to-fix on AI-disabled days vs AI-on baseline | > 2x sustained over 4+ weeks |
| Architecture-without-AI confidence | Self-rating on whiteboard/design tasks with no AI | < 6/10 with > 12 months experience |
| Off-keyboard reasoning | Can the developer trace a bug verbally before opening the editor? | Cannot, on previously-fluent areas |
| Library/API recall | Names key APIs in their daily stack from memory | Dependent on autocomplete for routine lookups |
**Mitigation, not measurement-only**: schedule periodic AI-off blocks (one half-day per sprint), keep code review human-led, rotate "from-scratch" exercises. The point is not nostalgia — it is preserving the judgment that lets developers catch AI errors. Pair with `over-trust indicators` below.
---
## Onboarding Metrics
### Time to First AI-Assisted PR
```
Metric: Calendar days from account activation to first merged PR
that used AI tools in its creation.
Target: Within first 5 working days.
```
Track by:
- Role (junior, mid, senior)
- Prior AI tool experience
- Team (some teams enable faster onboarding)
### Time to Self-Sufficiency
When a developer no longer needs help from an AI champion to use tools effectively.
Indicators of self-sufficiency:
- No longer asking "how do I prompt for X?"
- Using tools for multiple task types (not just code completion)
- Helping others with AI tool questions
- Give-up rate below 25%
| Experience Level | Expected Time to Self-Sufficiency |
|-----------------|----------------------------------|
| Senior + prior AI experience | 1-2 weeks |
| Senior + no AI experience | 2-4 weeks |
| Mid-level | 3-6 weeks |
| Junior | 4-8 weeks |
### Onboarding Completion Rate
```
Completion rate = Developers who finished AI training / Total enrolled
```
Target: 90%+. Below 80% indicates training is too long, poorly timed, or not seen as valuable.
### First-Week vs First-Month Usage Patterns
| Metric | Week 1 (healthy) | Month 1 (healthy) |
|--------|------------------|-------------------|
| Daily active usage | 50%+ of workdays | 70%+ of workdays |
| Tasks attempted with AI | 3-5 per day | 8-15 per day |
| Task variety | 1-2 types | 4-6 types |
| Give-up rate | 30-50% (learning) | 15-25% (stabilizing) |
A developer whose Week 1 and Month 1 patterns are identical has stalled. Intervene with coaching.
### Mentor Dependency Duration
Track how long new developers rely on AI champions:
- **Questions asked per week** (should decline over time)
- **Question complexity** (should shift from "how to" to "best approach for")
- **Channel** (direct message vs public channel indicates confidence)
---
## Trust and Confidence Metrics
### Trust Calibration
The goal is calibrated trust: developers review AI output proportionally to its actual error rate.
```
TRUST CALIBRATION MATRIX:
AI Output Correct AI Output Incorrect
┌────────────────────┬────────────────────┐
Developer Accepts │ True Acceptance │ Over-Trust │
│ (desired) │ (dangerous) │
├────────────────────┼────────────────────┤
Developer Rejects │ Under-Trust │ True Rejection │
│ (wasteful) │ (desired) │
└────────────────────┴────────────────────┘
Calibration Score = (True Acceptance + True Rejection) / Total Decisions
Target: > 0.80
```
### Over-Trust Indicators
Signals that developers accept AI output without adequate review:
- **Low review time**: PR review time decreases after AI adoption (should increase or stay flat)
- **Blind acceptance rate**: Accepting suggestions without cursor movement into the generated code
- **Copy-paste without edit**: AI output used verbatim at high rates (> 70%)
- **Bug attribution**: Increase in bugs traced to AI-generated code that passed review
- **Test coverage for AI code**: Lower test coverage on AI-generated sections
### Under-Trust Indicators
Signals that developers waste time rewriting adequate AI output:
- **High rejection rate with low defect rate**: Rejecting suggestions that were actually correct
- **Rewrite rate**: Developer rewrites AI output that is functionally equivalent to what was generated
- **Non-adoption despite training**: Developer avoids AI tools after completing training
- **Manual override pattern**: Consistently turning off AI suggestions in specific file types
### Confidence Progression
Track quarterly or monthly:
```
Survey item: "How confident are you in your ability to effectively
use AI coding tools?" (1-5 scale)
Expected progression:
Month 1: 2.5-3.0 (uncertain, still learning)
Month 3: 3.5-4.0 (developing patterns)
Month 6: 4.0-4.5 (confident in known scenarios)
Month 12: 4.0-4.5 (stable, aware of limitations)
```
A confidence score that exceeds 4.5 may indicate over-confidence rather than mastery.
### Trust Recovery After AI Failures
When AI tools produce a significant failure (major bug, security issue, outage):
| Metric | Measurement |
|--------|-------------|
| Usage dip | % decrease in AI usage in the week following the failure |
| Recovery time | Days until usage returns to pre-failure levels |
| Behavioral shift | Change in review depth post-failure (should increase temporarily) |
| Narrative impact | Does the failure become "folklore" that discourages adoption? |
Track these after every notable AI failure. Typical recovery is 1-3 weeks; if longer, intervene with communication and process improvement.
---
## DX Anti-Patterns
Practices that reliably erode developer experience with AI tools.
### Mandatory Usage Policies Without Support
**What it looks like**: "All developers must use AI tools for code generation" with no training, champions, or adjustment period.
**Why it fails**: Forced adoption without support creates resentment. Developers who struggle feel judged rather than helped.
**Instead**: Set expectations for experimentation, provide training and champions, let adoption grow organically with nudges.
### Individual Productivity Tracking and Surveillance
**What it looks like**: Dashboards showing each developer's AI acceptance rate, lines generated, or time saved.
**Why it fails**: Developers game metrics. Those with lower usage feel surveilled. Trust erodes.
**Instead**: Track metrics at team level only. Individual data is for the individual developer's self-improvement, never for performance reviews.
### Comparing Developers by AI Usage
**What it looks like**: "Developer A accepts 45% of AI suggestions while Developer B only accepts 20%."
**Why it fails**: Acceptance rate depends on task type, code complexity, and coding style. Comparing developers is meaningless and harmful.
**Instead**: Compare teams over time. Compare the same developer's before/after. Never rank developers by AI metrics.
### Ignoring Negative Feedback
**What it looks like**: Dismissing complaints as "resistance to change" or "they just need more training."
**Why it fails**: Negative feedback often identifies real tool limitations. Developers who feel unheard disengage entirely.
**Instead**: Create a structured feedback channel. Categorize feedback (tool issue, training gap, workflow mismatch, personal preference). Act on the first three categories; respect the fourth.
### Removing Traditional Tools Before AI Tools Are Ready
**What it looks like**: Eliminating code snippet libraries, templates, or scaffolding tools because "AI does that now."
**Why it fails**: AI tools have failure modes. Removing fallback options creates single points of failure.
**Instead**: Keep traditional tools available. Let them naturally atrophy as AI tools prove reliable. Remove only after 6+ months of demonstrated redundancy.
### Additional Anti-Patterns
- **One-size-fits-all configuration**: Forcing identical AI tool settings across teams with different needs
- **No opt-out mechanism**: Not allowing developers to disable AI suggestions when doing focused work
- **Measuring only positive outcomes**: Collecting success stories while ignoring or suppressing negative experiences
- **AI tool sprawl**: Deploying 4-5 overlapping AI tools without consolidation guidance
- **Executive demos as proof**: Using curated demos instead of real-world measurement to justify expansion
references/evidence-update.md
# What Changed Since 2025 (as of 2026-08)
_Verified 2026-08-21. Synthesizes METR 2025 RCT, METR Feb 2026 update, METR May 2026 survey, DORA 2025 AI report, DORA 2026 ROI report, DX Core 4, Faros 2026 telemetry, and SlopCodeBench v1._
## SlopCodeBench v1 (extension robustness)
The March 2026 preprint evaluates coding agents across evolving specifications while carrying forward each agent's own workspace and starting each checkpoint without the prior conversation context. Its central measurement implication is that snapshot pass rates can miss degrading extensibility; trajectory-level structural erosion and verbosity reveal a different dimension of performance. Its prompt intervention improved initial quality but did not halt degradation, so planning or quality prompts are not substitutes for longitudinal evaluation. Primary URL: https://arxiv.org/abs/2603.24755v1
Use the paper to justify checkpoint-level extension-robustness, quality-slope, cost, and review-burden measurement. Do not use its model averages as organizational targets or infer that the quality signals cause correctness or ROI outcomes. The reported experiments are Python-only, and the article is a preprint.
## METR RCT (the 2025 baseline)
METR's randomized controlled trial (data Feb–Jun 2025, published 2025-07-10) found experienced open-source developers were **~19% slower** with early-2025 AI tools than without them — the opposite of developer self-predictions. Primary URL: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
## METR 2026 Update (selection bias caveat)
In their 2026-02-24 update, METR stated they believe developers are likely **more sped up** in early 2026 than early-2025 estimates. However, their new experiment is unreliable: **30–50% of participating developers declined to submit tasks they didn't want to do without AI** — a self-selection mechanism that inflates apparent AI benefit. Do not cite as clean evidence of productivity gains. Primary URL: https://metr.org/blog/2026-02-24-uplift-update/
**Combined METR read**: the 2025 generation of tools caused slowdowns in real conditions; the 2026 generation likely improves this, but the measurement problem is now worse, not better.
## DORA 2025 AI-Assisted Report
DORA's dedicated AI-assisted software development report (Google Cloud, 2025) reinforces the conditional-impact model: AI amplifies existing strengths and weaknesses rather than being a universal accelerant. Canonical URL: https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report
## DORA 2026: ROI of AI-Assisted Software Development
Published by Google Cloud's DORA program (report dated 2026.01, widely covered from April-May 2026), this is a distinct follow-up to the 2025 AI-assisted report and should be cited separately, not conflated with it. Canonical URL: https://dora.dev/ai/roi/report/
Three load-bearing contributions for this skill:
1. **The J-Curve model.** Organizations see an initial productivity *dip* after AI rollout before gains materialize, driven by three costs: the learning curve as teams adapt workflows, a "verification tax" from reviewing higher-volume AI-generated output, and downstream process friction (testing, approvals) that has not yet adapted to the new code volume. The report frames this as "the tuition cost of transformation" and recommends budgeting for it explicitly rather than treating a slow first quarter as a rollout failure. This directly supports this skill's existing "before/after" study design defaults (do not compare baseline to the ramp-up period).
2. **Scenario-based ROI, not a single number.** The report's own illustrative model (500-person engineering org) shows a ~39% first-year ROI ($11.6M value against $8.4M investment, ~8-month payback) — but explicitly frames this as one scenario among conservative/realistic/optimistic ranges, reinforcing this skill's existing rule against a single headline ROI figure.
3. **The "instability tax."** AI adoption is associated with higher individual effectiveness and code quality in the report's model, but also with rising delivery instability — the sample model shows change failure rate rising from 5% to 6%, producing a modeled ~$344K negative downtime impact. The report also finds AI yields 35-40% productivity gains on simple, well-scoped tasks but only ~10% on complex legacy code, reinforcing this skill's task-complexity segmentation rule (`benchmarking-methodology.md`, C1-C4 complexity levels).
Use this report to strengthen `roi-framework.md` scenario planning and the anti-gaming checklist's ban on single blended ROI numbers — but treat the specific dollar and percentage figures as an illustrative vendor model, not a transferable benchmark for any given organization.
## DX Core 4
DX Core 4's four dimensions (Speed, Effectiveness, Quality, Business Impact) were formalized and publicly presented at the DX Annual 2026 conference (April 16, 2026), unifying DORA, SPACE, and DevEx into one framework with a Developer Experience Index (DXI) composite under "Effectiveness." Treat this as a later formalization than earlier DX Core 4 commentary; cite 2026, not 2024, as the reference point going forward. Treat specific DX benchmark figures (e.g. "27.4% of production code is AI-generated," "developers save ~3h45m/week") as vendor evidence pending independent replication. URL: https://getdx.com/research/dx-core-4/
## METR May 2026 Self-Reported Survey
Survey of 349 technical workers (87 engineers, 71 researchers, 129 academics/PhD students, 48 founders/managers), conducted Feb–Apr 2026, published 2026-05-11. Median self-reported value-of-work change: **1.4–2x** (retrospective 1.3x for Mar 2025, 2x for Mar 2026, forecast 2.5x for Mar 2027). METR notes significant reasons for skepticism: their 2025 RCT found participants overestimated AI's time effect by 40 percentage points on average. Do not cite these as controlled evidence of productivity gains. Primary URL: https://metr.org/blog/2026-05-11-ai-usage-survey/
**Combined METR read through June 2026**: self-reported gains are rising but consistently overestimated vs. controlled measures; the measurement problem is getting harder, not easier, as willingness to work without AI declines.
## Faros AI 2026 Telemetry
Organizational telemetry (not RCT) from Faros AI, covering 22,000 developers across 4,000 teams. Key figures at highest vs. lowest AI adoption within each org:
- Average PR size: **+51.3%**
- Files touched per developer per month: **+149.9%**
- PRs merged without any review: **+31.3%**
- Lead time commit to production: **+480.4%**
- Throughput (epics per developer): **+66%** (also reported as +66.2%; task completion per developer +33.7%; PR merge rate per developer +16.2%)
- Incidents per PR: **+243%** (also reported as +242.7%)
- Bugs per developer: **+54%**
- Median PR review time: **+441.5%**
- Code churn: **+861%**
The review-bypass, lead-time, and code-churn figures are the most operationally significant. Throughput rises; downstream review capacity, quality gates, and incident load do not keep pace. Source: https://www.faros.ai/research/ai-acceleration-whiplash
## Key implication for measurement
The 2025→2026 period makes the measurement case stronger, not weaker: faster output without paired review-capacity and incident-tracking instrumentation produces a misleading picture. Any scorecard built with this skill should include review burden, PR-merge-without-review rate, and defect-escape metrics alongside delivery speed.
references/productivity-metrics.md
# Productivity and Delivery Metrics for AI-Augmented Teams
Operational guidance for measuring how AI coding assistants and agents affect software delivery. This file keeps the DORA and SPACE framing, but it does **not** assume AI automatically improves throughput.
---
## Table of Contents
- [Evidence Posture](#evidence-posture)
- [The Delivery Stack](#the-delivery-stack)
- [Core Delivery Metrics](#core-delivery-metrics)
- [Recommended Companion Metrics](#recommended-companion-metrics)
- [DORA Metrics Applied Carefully](#dora-metrics-applied-carefully)
- [Deployment Frequency](#deployment-frequency)
- [Lead Time for Changes](#lead-time-for-changes)
- [Change Failure Rate](#change-failure-rate)
- [Mean Time to Recovery](#mean-time-to-recovery)
- [SPACE Applied to AI Workflows](#space-applied-to-ai-workflows)
- [Satisfaction and Well-Being](#satisfaction-and-well-being)
- [Performance](#performance)
- [Activity](#activity)
- [Communication and Collaboration](#communication-and-collaboration)
- [Efficiency and Flow](#efficiency-and-flow)
- [Assistant vs Agent Delivery Metrics](#assistant-vs-agent-delivery-metrics)
- [Assistant-Heavy Teams](#assistant-heavy-teams)
- [Agent-Heavy Teams](#agent-heavy-teams)
- [Benchmark-to-Production Gap](#benchmark-to-production-gap)
- [Study Design Guidance](#study-design-guidance)
- [Before / After](#before-after)
- [Matched A/B or Stratified Assignment](#matched-ab-or-stratified-assignment)
- [Crossover](#crossover)
- [Task-Level Shadow or Blind Review](#task-level-shadow-or-blind-review)
- [Confounds That Frequently Break AI Delivery Analysis](#confounds-that-frequently-break-ai-delivery-analysis)
- [Reporting Rules](#reporting-rules)
- [What to Do Next](#what-to-do-next)
## Evidence Posture
Use this as the starting stance:
- DORA 2025 treats AI as an amplifier of existing system quality and accessibility, not a guaranteed accelerator.
- METR's July 2025 randomized trial found experienced open-source developers were slower on realistic tasks in that setting.
- Benchmark performance and production delivery performance diverge quickly.
Therefore:
- do not start with a positive expected impact
- start with a falsifiable hypothesis and a baseline
- decompose the delivery system instead of using one topline number
---
## The Delivery Stack
AI can affect different parts of delivery differently. Always decompose the path.
```
task assigned
-> active work starts
-> first meaningful artifact
-> PR opened
-> review completed
-> merged
-> deployed
-> production stable
```
Good measurement usually finds that AI helps some segments and leaves others unchanged.
---
## Core Delivery Metrics
| Metric | Default Definition | Why It Matters |
|--------|--------------------|----------------|
| Time to first meaningful artifact | task start -> draft code / plan / PR | catches inner-loop acceleration |
| PR open latency | task start -> PR opened | useful for assistant and agent workflows |
| Review turnaround time | review request -> approval / changes requested | often the hidden bottleneck |
| Lead time for changes | commit -> production | standard DORA metric |
| Cycle time | active work start -> deploy | broad delivery signal |
| Deployment frequency | deploys per service or team per period | detects system-level acceleration |
| Change failure rate | failed deploys / total deploys | speed without stability is not success |
| Mean time to recovery | incident detection -> restored service | tests whether AI helps incident response |
### Recommended Companion Metrics
Pair delivery metrics with:
- defect escape rate
- revert rate
- review burden
- satisfaction or trust burden
Do not publish a delivery-only dashboard.
---
## DORA Metrics Applied Carefully
### Deployment Frequency
Use:
- deploys per team or service per week
- segmented by change type if possible: feature, fix, config, dependency, docs
Interpretation:
- increased deployment frequency can be real improvement
- it can also be caused by smaller PRs, service sprawl, or release process changes
### Lead Time for Changes
Measure:
- commit -> production
- plus a decomposed view:
- task start -> first meaningful artifact
- PR open -> merge
- merge -> production
Interpretation:
- AI often affects the first two segments more than the last one
- if only the first segment improves, infrastructure or approval flow may be the real constraint
### Change Failure Rate
Measure:
- failed deploys / total deploys
- and, when feasible, compare AI-assisted vs non-AI-assisted changes by task type
Interpretation:
- a flat or rising change failure rate can erase delivery gains
- do not treat more deployments as success if failure rate rises
### Mean Time to Recovery
Measure:
- incident detection -> service restored
- optionally split diagnosis time vs fix time
Interpretation:
- AI may help diagnosis, log search, or code search even when it does not improve feature delivery
---
## SPACE Applied to AI Workflows
### Satisfaction and Well-Being
Ask:
- do AI tools reduce toil?
- do they increase pressure or trust burden?
Use `developer-experience-metrics.md`.
### Performance
Measure outcomes, not output volume.
Good:
- time to customer-visible improvement
- bugs fixed
- incidents resolved
- reviewer effort for accepted changes
Avoid:
- lines of code
- raw prompt count
- raw PR count without quality context
### Activity
Useful activity measures:
- PRs merged by task type
- test cases created and retained
- documentation updates completed
- incidents diagnosed with AI assistance
Activity is only useful when paired with performance and quality.
### Communication and Collaboration
AI can shift coordination burden rather than remove it.
Measure:
- review rounds per PR
- comments per accepted AI-generated PR
- follow-up clarification requests
- time spent explaining agent output
### Efficiency and Flow
Measure:
- uninterrupted coding block time
- handoff frequency
- context switching caused by tool failures or retries
This is where many "it feels faster" claims live. Validate them with sampled task data.
---
## Assistant vs Agent Delivery Metrics
### Assistant-Heavy Teams
Good default metrics:
- time to first meaningful artifact
- PR open latency
- review turnaround
- lead time for changes
- defect escape rate
### Agent-Heavy Teams
Good default metrics:
- task completion rate
- human takeover rate
- reviewer effort per accepted task
- PR merge rate
- post-merge revert rate
- cost per accepted change
If the workflow is agent-heavy, see `agent-execution-metrics.md` first and only then roll up to team delivery metrics.
---
## Benchmark-to-Production Gap
Always track the gap between benchmark performance and production acceptance.
| Signal | Why It Matters |
|--------|----------------|
| benchmark score | capability ceiling, often on narrow tasks |
| task completion in your repos | practical usefulness |
| reviewer acceptance | real-world quality threshold |
| revert / hotfix rate | downstream reliability |
Common failure mode:
- benchmark performance rises
- invocation rises
- merge rate stays flat
- review burden rises
That is not a delivery win.
---
## Study Design Guidance
### Before / After
Use when:
- tool is already rolling out
- you cannot withhold access
Requirements:
- 8-12 week baseline
- same metric definitions before and after
- stable team composition where possible
### Matched A/B or Stratified Assignment
Use when:
- you have enough comparable teams
- leadership wants stronger causal evidence
Control for:
- team size
- stack
- project type
- seniority mix
- release cadence
### Crossover
Use when:
- teams object to permanent denial of tooling
- you can tolerate a longer study
### Task-Level Shadow or Blind Review
Use when:
- evaluating coding agents on a narrow workflow
- reviewer acceptance is the key business question
See `benchmarking-methodology.md`.
---
## Confounds That Frequently Break AI Delivery Analysis
| Confound | Failure Mode |
|----------|--------------|
| review policy change | looks like AI sped delivery when policy did |
| team composition change | senior hire or attrition distorts trend |
| release calendar / crunch period | short-term throughput spike misread as tool effect |
| service or repo restructuring | changes deployment frequency mechanically |
| measurement novelty | developers change behavior because they know they are measured |
| tool mandate | adoption rises while satisfaction and quality worsen |
Document confounds in every report.
---
## Reporting Rules
When summarizing delivery impact:
1. report the baseline period
2. state the unit of analysis
3. show at least one quality metric next to each speed metric
4. include sample size and data coverage
5. distinguish measured effects from self-reported effects
6. separate assistant and agent workflows if both are in scope
Suggested one-line summary format:
> Over a 12-week stabilized period, assistant usage increased time-to-first-artifact speed while review time and defect escape were unchanged; agent usage increased PR creation but not merge rate, so net delivery impact remains mixed.
---
## What to Do Next
- For causal rigor, use `benchmarking-methodology.md`.
- For quality guardrails, use `quality-metrics.md`.
- For agent-level operational metrics, use `agent-execution-metrics.md`.
- For leadership decisions, pair this file with `roi-framework.md`.
references/quality-metrics.md
# Code Quality Metrics Under AI Assistance
Operational reference for tracking whether AI coding tools maintain, improve, or degrade code quality. Covers defects, complexity, testing, security, technical debt, and the guardrails required to keep AI-generated code at production standard.
---
## Table of Contents
- [Defect Metrics](#defect-metrics)
- [Core Defect Metrics](#core-defect-metrics)
- [Metric Details](#metric-details)
- [Defect Tracking by AI Involvement](#defect-tracking-by-ai-involvement)
- [Code Complexity Tracking](#code-complexity-tracking)
- [Complexity Metrics](#complexity-metrics)
- [Complexity Drift Monitoring](#complexity-drift-monitoring)
- [Complexity Reduction Strategies for AI Code](#complexity-reduction-strategies-for-ai-code)
- [Test Coverage Impact](#test-coverage-impact)
- [Coverage Metrics](#coverage-metrics)
- [AI-Generated Test Quality Assessment](#ai-generated-test-quality-assessment)
- [Mutation Testing as Quality Gate](#mutation-testing-as-quality-gate)
- [Flaky Test Management](#flaky-test-management)
- [Security Vulnerability Tracking](#security-vulnerability-tracking)
- [Vulnerability Metrics](#vulnerability-metrics)
- [CWE Category Distribution Under AI](#cwe-category-distribution-under-ai)
- [Security Scanning Integration](#security-scanning-integration)
- [OWASP Top 10 Tracking](#owasp-top-10-tracking)
- [Technical Debt Accumulation](#technical-debt-accumulation)
- [Debt Metrics](#debt-metrics)
- [Debt Tracking Dashboard](#debt-tracking-dashboard)
- [Architecture Fitness Functions](#architecture-fitness-functions)
- [Dead Code Management](#dead-code-management)
- [Quality Guardrails for AI-Generated Code](#quality-guardrails-for-ai-generated-code)
- [Mandatory Review Rules](#mandatory-review-rules)
- [Automated Quality Gates (CI/CD Integration)](#automated-quality-gates-cicd-integration)
- [Example quality gate pipeline stages](#example-quality-gate-pipeline-stages)
- [AI-Specific Lint Rules](#ai-specific-lint-rules)
- [Pre-Commit Hooks for Quality Enforcement](#pre-commit-hooks-for-quality-enforcement)
- [.pre-commit-config.yaml additions for AI-assisted development](#pre-commit-configyaml-additions-for-ai-assisted-development)
- [Quality Scorecard Template](#quality-scorecard-template)
## Defect Metrics
### Core Defect Metrics
| # | Metric | Formula | AI-Team Target | Alert Threshold |
|---|--------|---------|---------------|-----------------|
| 1 | Bug Density (per KLOC) | Bugs found / (Lines of code / 1000) | < 2.0 | > 5.0 |
| 2 | Bug Density (per Feature) | Bugs found / Features shipped | < 0.3 | > 1.0 |
| 3 | Defect Escape Rate | Bugs found in production / Total bugs found | < 15% | > 30% |
| 4 | Rework Rate | PRs with follow-up fix within 14 days / Total PRs | < 10% | > 20% |
| 5 | Mean Time to Detect (MTTD) | Average time from defect introduction to detection | < 48 hours | > 2 weeks |
| 6 | Defect Clustering | Gini coefficient of bugs across modules | < 0.6 | > 0.8 |
| 7 | Regression Rate | Bugs introduced by fixes / Total fixes | < 5% | > 15% |
### Metric Details
**Bug Density (per KLOC)** — The foundational quality metric. Compare AI-assisted code vs non-AI code by tagging commits. Important: AI-generated code tends to have higher LOC for equivalent functionality, which can artificially deflate per-KLOC density while total bugs increase. Always track per-feature density alongside per-KLOC.
**Defect Escape Rate** — The most important quality metric for AI adoption. Measures what percentage of defects reach production before being caught. AI code that passes review but fails in production indicates:
- Reviewers rubber-stamping AI output
- Test coverage gaps on AI-generated paths
- AI generating plausible-looking but subtly incorrect code
Track monthly. A rising escape rate after AI adoption demands immediate intervention.
**Rework Rate** — Percentage of merged PRs that require a follow-up PR to fix issues within 14 days (configurable window). AI-assisted PRs with higher rework rates indicate developers are merging AI output without sufficient review.
Calculation:
```
For each merged PR:
1. Find follow-up PRs that modify the same files within 14 days
2. Filter to those with commit messages indicating a fix (heuristic: contains "fix", "bug", "revert", "correct")
3. Rework Rate = count of reworked PRs / total merged PRs
```
**Mean Time to Detect (MTTD)** — Average time between when a defect-introducing commit is merged and when the defect is reported. Shorter MTTD means your quality gates (tests, monitoring, observability) are catching problems early. AI-generated defects often have longer MTTD because they can be semantically correct but logically flawed — tests pass, but behavior is wrong.
**Defect Clustering** — Measures whether bugs concentrate in specific modules or are distributed evenly. A Gini coefficient near 1.0 means all bugs are in a few modules. After AI adoption, watch for new clusters forming in areas where AI was heavily used — this signals the AI model is producing systematically flawed output for certain patterns.
**Regression Rate** — Bugs introduced when fixing other bugs. AI tools can increase regression rate if developers use AI to generate quick fixes without understanding the broader system impact. Track by comparing bug-fix PRs that introduce new bugs vs total bug-fix PRs.
### Defect Tracking by AI Involvement
Tag every bug with the degree of AI involvement in the code that introduced it:
| Tag | Definition | Action if Elevated |
|-----|-----------|-------------------|
| `ai-generated` | Bug is in code primarily written by AI | Review AI prompt/context quality |
| `ai-modified` | Bug is in human code modified by AI | Review AI edit suggestions |
| `ai-reviewed-only` | AI reviewed but didn't write the code | Review AI review accuracy |
| `no-ai` | No AI involvement | Baseline comparison |
---
## Code Complexity Tracking
### Complexity Metrics
| Metric | Tool | What It Measures | AI Concern |
|--------|------|-----------------|------------|
| Cyclomatic Complexity | radon (Python), complexity-report (JS), gocyclo (Go) | Number of independent paths through code | AI generates long functions with many branches |
| Cognitive Complexity | SonarQube, SonarCloud | How difficult code is for humans to understand | AI code can be syntactically simple but semantically confusing |
| Code Duplication Rate | SonarQube, jscpd, PMD CPD | Percentage of code that is duplicated | AI tends to generate similar-but-not-identical blocks |
| Afferent/Efferent Coupling | NDepend, JDepend, deptrac | How interconnected modules are | AI may not respect architectural boundaries |
| Function Length | Linters (ESLint, Pylint) | Lines per function/method | AI generates longer functions than experienced developers |
### Complexity Drift Monitoring
Track these weekly, graphed over time:
**Before/After AI Adoption:**
| Metric | Pre-AI Baseline | 3-Month Post-AI | 6-Month Post-AI | Target |
|--------|----------------|-----------------|-----------------|--------|
| Mean cyclomatic complexity | {baseline} | {measure} | {measure} | ≤ baseline + 10% |
| P90 cyclomatic complexity | {baseline} | {measure} | {measure} | ≤ baseline |
| Code duplication rate | {baseline} | {measure} | {measure} | ≤ baseline |
| Mean function length | {baseline} | {measure} | {measure} | ≤ baseline + 15% |
| Coupling between modules | {baseline} | {measure} | {measure} | ≤ baseline |
**Alert Rules:**
- Mean cyclomatic complexity increases > 10% over rolling 30-day window: investigate
- Code duplication rate increases > 5 percentage points: investigate
- Mean function length increases > 20%: enforce lint rules
### Repeated-Edit Trajectory Signals
For edit-capable coding agents, supplement snapshot complexity metrics with an evolving-spec sequence of at least three checkpoints. Track:
- **Extension robustness** — whether current behavior and all retained prior regression tests remain correct at each checkpoint.
- **Structural erosion slope** — the direction and rate of change in structural erosion across ordered checkpoints.
- **Verbosity slope** — the direction and rate of change in locally defined redundant-pattern and clone-line coverage across ordered checkpoints.
- **Late-checkpoint burden** — operating cost plus reviewer or remediation effort in late/final checkpoints compared with earlier phases.
Use the metric definitions and rule mapping in `software-clean-code-standard/references/code-complexity-metrics.md`. Treat slopes as advisory trend signals: they do not prove correctness, and SlopCodeBench v1 averages are neither universal gates nor organizational targets. Pair them with behavioral tests, regression retention, and human review.
### Complexity Reduction Strategies for AI Code
1. **Constrain AI output** — Include maximum function length and complexity rules in AI context (CLAUDE.md, .cursorrules)
2. **Post-generation refactoring** — Use AI itself to refactor its own output: "Break this function into smaller functions"
3. **Lint enforcement** — CI fails on complexity thresholds, regardless of whether code is AI-generated
4. **Architecture documentation** — Provide AI with module boundaries and coupling rules in context
---
## Test Coverage Impact
### Coverage Metrics
| Metric | Definition | AI-Team Target | Measurement Tool |
|--------|-----------|---------------|-----------------|
| Line Coverage | % of lines executed by tests | > 80% | Istanbul, Coverage.py, JaCoCo |
| Branch Coverage | % of conditional branches exercised | > 70% | Istanbul, Coverage.py, JaCoCo |
| Path Coverage | % of execution paths tested | > 50% | Specialized tools per language |
| Mutation Score | % of code mutations caught by tests | > 60% | Stryker, mutmut, PIT |
| Test-to-Code Ratio | Lines of test code / Lines of production code | 1:1 to 2:1 | LOC count |
| Flaky Test Rate | Tests that pass/fail non-deterministically / Total tests | < 2% | CI history analysis |
### AI-Generated Test Quality Assessment
AI excels at generating tests but the generated tests have systematic weaknesses:
| Strength | Weakness |
|----------|----------|
| High line coverage quickly | Tests often test implementation, not behavior |
| Good at happy-path testing | Weak on edge cases and error paths |
| Fast boilerplate generation | May assert on incidental behavior (brittle tests) |
| Consistent style | May miss domain-specific invariants |
| Good at mimicking existing test patterns | May duplicate test logic instead of using fixtures |
### Mutation Testing as Quality Gate
Standard coverage metrics are insufficient for AI-generated tests. A test suite can have 90% line coverage but fail to catch real bugs because the tests assert on the wrong things.
**Mutation testing** introduces small changes (mutations) to production code and checks whether tests fail. If tests still pass after a mutation, they are not actually verifying behavior.
**Process:**
1. Run mutation testing monthly (or on AI-generated test files)
2. Compare mutation scores: AI-generated tests vs human-written tests
3. Target: AI-generated test mutation score within 10% of human-written tests
4. If gap exceeds 10%: AI tests are superficial and need manual augmentation
**Mutation Score Benchmarks:**
| Score | Interpretation |
|-------|---------------|
| > 80% | Strong test suite — mutations reliably caught |
| 60-80% | Adequate — some gaps in assertion quality |
| 40-60% | Weak — tests present but not catching real bugs |
| < 40% | Superficial — coverage without verification |
### Flaky Test Management
AI-generated tests have a higher flaky rate due to:
- Hardcoded timing assumptions
- Non-deterministic test ordering dependencies
- Implicit environment dependencies
- Snapshot tests that break on style changes
**Tracking:**
- Tag AI-generated tests in test files (comment or metadata)
- Monitor flaky rate separately for AI vs human tests
- Quarantine flaky tests automatically (fail → retry → quarantine if intermittent)
- Target: AI-generated flaky rate ≤ 1.5x human-generated flaky rate
---
## Security Vulnerability Tracking
### Vulnerability Metrics
| Metric | Formula | Target | Alert |
|--------|---------|--------|-------|
| Vuln Introduction Rate | New SAST findings per sprint | Stable or decreasing | > 20% increase over 3 sprints |
| Critical Vuln Rate | Critical/High findings per sprint | 0 critical, < 2 high | Any critical |
| Dependency Vuln Count | Known CVEs in dependencies | < 5 medium, 0 high/critical | Any high/critical |
| Secrets Exposure Incidents | Credentials/tokens committed | 0 | Any occurrence |
| OWASP Top 10 Violations | Violations by category per quarter | Decreasing | Any new category appearing |
| Fix Time (Security) | Days from finding to fix | < 7 days (critical), < 30 days (high) | Exceeding SLA |
### CWE Category Distribution Under AI
AI-generated code tends to introduce specific vulnerability types more frequently:
| CWE Category | Risk Level with AI | Why |
|--------------|-------------------|-----|
| CWE-798: Hardcoded Credentials | High | AI may generate placeholder secrets that reach production |
| CWE-89: SQL Injection | Medium-High | AI may generate string concatenation instead of parameterized queries |
| CWE-79: Cross-Site Scripting | Medium | AI may skip output encoding in generated templates |
| CWE-22: Path Traversal | Medium | AI-generated file operations may not sanitize paths |
| CWE-502: Deserialization | Medium | AI may use unsafe deserialization by default |
| CWE-200: Information Exposure | Medium-High | AI may include verbose error messages with internal details |
| CWE-287: Auth Bypass | Low-Medium | AI may generate auth logic with subtle bypass conditions |
| CWE-330: Weak Randomness | Medium | AI may use Math.random() for security-sensitive operations |
### Security Scanning Integration
**Required scans for AI-generated code:**
| Scan Type | Tool Examples | When | Blocks Merge? |
|-----------|--------------|------|---------------|
| SAST | Semgrep, CodeQL, SonarQube | Every PR | Yes (critical/high) |
| Secret Detection | Gitleaks, TruffleHog, detect-secrets | Pre-commit + PR | Yes (any finding) |
| Dependency Scan | Dependabot, Snyk, Trivy | Daily + PR | Yes (critical) |
| Container Scan | Trivy, Grype | Build pipeline | Yes (critical) |
| License Compliance | FOSSA, Snyk | Weekly | No (advisory) |
### OWASP Top 10 Tracking
Track violations by category per quarter. Identify trends specific to AI-generated code.
| OWASP Category | Tracking Method | AI-Specific Concern |
|---------------|----------------|-------------------|
| A01: Broken Access Control | SAST + manual review | AI may generate RBAC with gaps |
| A02: Cryptographic Failures | SAST | AI may use deprecated algorithms |
| A03: Injection | SAST + DAST | AI may generate unsanitized inputs |
| A04: Insecure Design | Architecture review | AI lacks system-level security context |
| A05: Security Misconfiguration | Config scanning | AI may generate insecure defaults |
| A06: Vulnerable Components | Dependency scan | AI may suggest outdated packages |
| A07: Auth Failures | SAST + pen test | AI-generated auth logic needs extra review |
| A08: Data Integrity Failures | SAST | AI may skip integrity checks |
| A09: Logging Failures | SAST + review | AI may over-log (sensitive data) or under-log |
| A10: SSRF | SAST + DAST | AI-generated HTTP calls may not validate URLs |
---
## Technical Debt Accumulation
### Debt Metrics
| Metric | Definition | Measurement | AI Concern |
|--------|-----------|-------------|------------|
| Technical Debt Ratio | Remediation cost / Development cost | SonarQube | AI generates code faster than debt is addressed |
| Architecture Fitness | Automated architecture rule compliance | ArchUnit, Fitness Functions | AI may violate architectural boundaries |
| API Compatibility Breaks | Breaking changes in internal/external APIs | API diff tools, contract tests | AI may not know API contracts |
| Documentation-to-Code Ratio | Doc lines / Code lines | LOC analysis | AI-generated code often lacks inline documentation |
| Dead Code Accumulation | Unreachable code growth rate | SonarQube, tree-shaking analysis | AI generates code that duplicates existing functions |
| TODO/FIXME Density | Comment markers per KLOC | grep/search | AI may introduce placeholder comments that persist |
### Debt Tracking Dashboard
Track monthly, trend over 12 months:
```
Technical Debt Scorecard — {Month} {Year}
SONARQUBE DEBT RATIO
Current: {n} days
Trend: {↑/↓/→} ({delta} from last month)
Rating: {A/B/C/D/E}
ARCHITECTURE COMPLIANCE
Rules passing: {n} / {total} ({pct}%)
New violations this month: {n}
AI-attributed violations: {n}
CODE FRESHNESS
Files not modified in 12 months: {pct}%
Dead code estimate: {pct}%
Deprecated API usage: {n} call sites
DOCUMENTATION
Public API doc coverage: {pct}%
Inline comment density: {n} per 100 LOC
Stale documentation files: {n}
```
### Architecture Fitness Functions
Automated tests that verify architectural rules are not violated. Critical when AI generates code that may not respect boundaries.
| Rule | Implementation | Frequency |
|------|---------------|-----------|
| Layer dependencies (e.g., UI cannot import DB) | ArchUnit / custom lint | Every PR |
| Module boundary enforcement | Import analysis | Every PR |
| Maximum dependency depth | Dependency graph analysis | Weekly |
| API versioning compliance | Contract tests | Every PR |
| Database migration safety | Migration linter | Every PR |
| Package size limits | Bundle analysis | Every PR |
### Dead Code Management
AI tools accelerate dead code accumulation through two mechanisms:
1. **Replacement without deletion** — AI generates a new implementation; the old one is not removed
2. **Speculative generation** — AI generates helper functions or utility code that is never called
**Detection:**
- Run tree-shaking analysis monthly (for JS/TS: webpack, rollup; for Java: ProGuard; for Python: vulture)
- Track unreachable code percentage over time
- Flag files with zero imports/references
**Prevention:**
- Include "remove dead code" as an explicit step in AI-assisted refactoring workflows
- Configure linters to warn on unused exports, unused variables, unused functions
- Quarterly dead code cleanup sprints
---
## Quality Guardrails for AI-Generated Code
### Mandatory Review Rules
| Rule | Scope | Implementation |
|------|-------|---------------|
| No auto-merge for AI-generated PRs | All AI-assisted PRs | CODEOWNERS or branch protection requiring human approval |
| Security-sensitive paths require 2 reviewers | Auth, crypto, payment, PII handling | CODEOWNERS with 2+ required reviewers for sensitive paths |
| AI-generated tests require human test review | All AI-generated test files | PR label triggers additional review requirement |
| Large AI PRs (> 500 lines) require architecture review | PRs over threshold | Automated PR size check + escalation |
### Automated Quality Gates (CI/CD Integration)
```yaml
# Example quality gate pipeline stages
quality_gates:
- name: lint
blocks_merge: true
rules:
- max_cyclomatic_complexity: 15
- max_function_length: 50
- max_file_length: 500
- no_unused_imports: true
- name: security
blocks_merge: true
rules:
- sast_critical_findings: 0
- sast_high_findings: 0
- secret_detection_findings: 0
- dependency_critical_cves: 0
- name: test_quality
blocks_merge: true
rules:
- line_coverage_minimum: 80
- branch_coverage_minimum: 70
- no_test_skips_without_reason: true
- name: complexity
blocks_merge: false # advisory
rules:
- max_cognitive_complexity: 20
- max_coupling_between_modules: 10
- duplication_rate_max: 5
```
### AI-Specific Lint Rules
Add these rules to your linter configuration when AI tools are in use:
| Rule | Purpose | Severity |
|------|---------|----------|
| No hardcoded strings in auth/crypto paths | Prevent AI-generated placeholder secrets | Error |
| No `console.log` / `print` in production code | AI frequently leaves debug output | Warning |
| No `any` type in TypeScript (strict mode) | AI falls back to `any` when types are complex | Error |
| No string concatenation in SQL | Prevent AI-generated injection vulnerabilities | Error |
| No `TODO` or `FIXME` without issue link | Prevent AI placeholder comments from persisting | Warning |
| No functions exceeding 50 lines | Contain AI-generated monolithic functions | Warning |
| No unused parameters | AI may generate function signatures with extra parameters | Warning |
| Require error handling on async operations | AI may omit error handling for brevity | Error |
### Pre-Commit Hooks for Quality Enforcement
```bash
# .pre-commit-config.yaml additions for AI-assisted development
repos:
- repo: local
hooks:
- id: check-secrets
name: Detect secrets
entry: detect-secrets-hook
language: system
stages: [commit]
- id: check-complexity
name: Complexity check
entry: radon cc --min C --no-assert
language: system
types: [python]
stages: [commit]
- id: check-function-length
name: Function length check
entry: custom-lint --max-function-length 50
language: system
stages: [commit]
```
### Quality Scorecard Template
Use monthly to track overall quality trajectory.
```
Code Quality Scorecard — {Month} {Year}
DEFECTS
Bug density (per KLOC): {n} [{↑/↓/→}] Target: < 2.0
Defect escape rate: {n}% [{↑/↓/→}] Target: < 15%
Rework rate: {n}% [{↑/↓/→}] Target: < 10%
Regression rate: {n}% [{↑/↓/→}] Target: < 5%
COMPLEXITY
Mean cyclomatic complexity: {n} [{↑/↓/→}] Target: ≤ baseline + 10%
Code duplication rate: {n}% [{↑/↓/→}] Target: ≤ baseline
Mean function length: {n} [{↑/↓/→}] Target: ≤ baseline + 15%
Structural erosion slope: {n/checkpoint} [{↑/↓/→}] Advisory; compare with local baseline
Verbosity slope: {n/checkpoint} [{↑/↓/→}] Advisory; compare with local baseline
EXTENSION ROBUSTNESS (EDIT-CAPABLE AGENTS)
Checkpoints passing current + retained regression tests: {n}/{total}
Late-checkpoint operating cost: {amount} [{↑/↓/→}] Compare with early/mid phases
Late-checkpoint review/remediation effort: {minutes} [{↑/↓/→}] Compare with early/mid phases
TESTING
Line coverage: {n}% [{↑/↓/→}] Target: > 80%
Branch coverage: {n}% [{↑/↓/→}] Target: > 70%
Mutation score: {n}% [{↑/↓/→}] Target: > 60%
Flaky test rate: {n}% [{↑/↓/→}] Target: < 2%
SECURITY
New SAST findings/sprint: {n} [{↑/↓/→}] Target: stable or decreasing
Open critical/high vulns: {n} [{↑/↓/→}] Target: 0 critical, < 2 high
Secret exposure incidents: {n} [{↑/↓/→}] Target: 0
TECHNICAL DEBT
SonarQube debt ratio: {rating} [{↑/↓/→}] Target: A
Dead code estimate: {n}% [{↑/↓/→}] Target: < 3%
Architecture violations: {n} [{↑/↓/→}] Target: 0 new
OVERALL QUALITY GRADE: {A/B/C/D/F}
A: All metrics at target
B: 1-2 metrics at alert threshold
C: 3-4 metrics at alert threshold
D: Any metric significantly past alert threshold
F: Quality regression across multiple categories
```
references/roi-framework.md
# ROI Framework for AI Coding Programs
Use this reference when the user needs an investment decision, renewal recommendation, or executive summary. The goal is not to prove AI is valuable. The goal is to estimate whether a specific program creates enough value to justify its cost under realistic assumptions.
---
## Table of Contents
- [Evidence Standard](#evidence-standard)
- [ROI Questions This File Helps Answer](#roi-questions-this-file-helps-answer)
- [Cost Model](#cost-model)
- [1. Direct Tooling Cost](#1-direct-tooling-cost)
- [2. Enablement Cost](#2-enablement-cost)
- [3. Review and Governance Cost](#3-review-and-governance-cost)
- [4. Transition Cost](#4-transition-cost)
- [5. Failure Cost](#5-failure-cost)
- [Benefit Model](#benefit-model)
- [1. Time Saved on Accepted Work](#1-time-saved-on-accepted-work)
- [2. Review Efficiency](#2-review-efficiency)
- [3. Quality Savings](#3-quality-savings)
- [4. Knowledge and Onboarding Gains](#4-knowledge-and-onboarding-gains)
- [5. Strategic Option Value](#5-strategic-option-value)
- [Minimum Viable ROI Formulas](#minimum-viable-roi-formulas)
- [Program ROI](#program-roi)
- [Payback Period](#payback-period)
- [Cost per Accepted Change](#cost-per-accepted-change)
- [Cost per Merged Agent PR](#cost-per-merged-agent-pr)
- [The J-Curve: Do Not Judge ROI From the First Quarter](#the-j-curve-do-not-judge-roi-from-the-first-quarter)
- [Scenario Planning](#scenario-planning)
- [What to Count as Benefit](#what-to-count-as-benefit)
- [Review Cost Is Not Optional](#review-cost-is-not-optional)
- [Benchmark and Research Caveats](#benchmark-and-research-caveats)
- [Executive Reporting Structure](#executive-reporting-structure)
- [Common ROI Failure Modes](#common-roi-failure-modes)
- [Recommended Default Deliverables](#recommended-default-deliverables)
- [For a pilot](#for-a-pilot)
- [For an agent rollout](#for-an-agent-rollout)
- [For renewal](#for-renewal)
- [What to Do Next](#what-to-do-next)
## Evidence Standard
Default:
- do not use a single headline ROI number without a scenario table
- do not use vendor benchmarks as the main evidence base
- do not assume positive net delivery impact
- do not treat benchmark scores as business value
Strong evidence order:
1. internal production data
2. controlled or staggered internal comparisons
3. peer-reviewed or independent external studies
4. first-party tool telemetry docs
5. vendor marketing or consulting estimates
If an assumption comes from level 4 or 5, label it clearly.
---
## ROI Questions This File Helps Answer
- should we buy or renew this tool?
- should we expand from pilot to broad rollout?
- is the agent workflow worth the review overhead?
- is the current program creating enough accepted value for its cost?
---
## Cost Model
Always include visible and hidden costs.
### 1. Direct Tooling Cost
Typical components:
- seat or subscription spend
- premium admin / enterprise tier spend
- API or model inference spend
- sandbox / runner / compute cost for agents
- storage and observability cost for traces, logs, and artifacts
### 2. Enablement Cost
Typical components:
- rollout and onboarding time
- documentation and playbook creation
- champion or enablement owner time
- training sessions and office hours
### 3. Review and Governance Cost
Typical components:
- security and legal review
- admin overhead
- policy maintenance
- reviewer time spent checking AI-produced work
### 4. Transition Cost
Typical components:
- learning curve slowdown
- dual-tool overlap during transition
- workflow churn from changing tools or policies
### 5. Failure Cost
Typical components:
- reverted PRs
- production incidents
- security exceptions
- wasted agent runs
- duplicated work after handoff failure
---
## Benefit Model
Only count benefits that are credible in your environment.
### 1. Time Saved on Accepted Work
Use when you can observe:
- faster delivery on retained work
- lower time to first meaningful artifact
- fewer manual steps for repeated workflows
Do not count time "saved" on output that is later rewritten or rejected.
### 2. Review Efficiency
Use when:
- review cycles decline without quality loss
- reviewer effort per accepted change falls
This is especially important for coding agents. An agent that creates more PRs but consumes more reviewer time may destroy ROI.
### 3. Quality Savings
Use when you can show:
- fewer escaped defects
- fewer hotfixes
- lower revert rate
- fewer security findings on new code
### 4. Knowledge and Onboarding Gains
Use when you can show:
- faster time to first meaningful contribution
- fewer blockers for less familiar codebases
- faster issue triage or codebase orientation
### 5. Strategic Option Value
Use cautiously. This includes:
- faster prototyping
- more experiments per quarter
- faster incident triage
This can matter, but it is easier to overstate than hard savings.
---
## Minimum Viable ROI Formulas
### Program ROI
```text
ROI (%) =
(Total realized benefit - Total program cost)
/ Total program cost
x 100
```
### Payback Period
```text
Payback period (months) =
Initial and rollout cost
/ Average monthly net benefit
```
### Cost per Accepted Change
Useful for mixed or agent-heavy programs.
```text
Cost per accepted change =
total tool + operational cost in period
/ accepted changes in period
```
Accepted change should be one of:
- merged PR
- shipped task
- completed run with accepted human handoff
Pick one and stay consistent.
### Cost per Merged Agent PR
```text
Cost per merged agent PR =
agent operating cost + allocated reviewer cost
/ merged agent-created PRs
```
This is often more decision-useful than top-level ROI during an early agent rollout.
---
## The J-Curve: Do Not Judge ROI From the First Quarter
DORA's 2026 ROI of AI-Assisted Software Development report names a pattern this skill already implies through its baseline and ramp-up rules: AI programs typically show a productivity **dip** before they show a gain. Three costs drive the dip:
1. **Learning curve** — teams are still adapting workflows to the tool.
2. **Verification tax** — reviewers spend more time checking a higher volume of AI-generated output before trusting it.
3. **Downstream friction** — testing, approval, and release processes have not yet adapted to the new code volume, so they become the bottleneck (see `theory-of-constraints-applied.md`).
Treat this dip as "tuition," not failure. Budget for it explicitly in the ROI model rather than pausing or cutting the program after a disappointing first quarter. Do not compare month-1 numbers to the target state; compare them to the pre-rollout baseline and expect the trend line, not the point estimate, to justify the investment. This is why the Study Design Defaults above require an 8+ week stabilized-usage window before judging impact.
## Scenario Planning
Use three scenarios, not one.
| Scenario | Assumptions | Typical Use |
|----------|-------------|-------------|
| Conservative | lower adoption, lower retained value, higher review cost | board or finance review |
| Base | observed adoption and observed quality-adjusted gains | operating plan |
| Upside | higher adoption and validated workflow expansion | planning, not commitment |
For each scenario vary:
- adoption rate
- retained time savings
- reviewer effort
- defect / revert cost
- infrastructure or inference cost
Do not vary only the upside assumptions.
---
## What to Count as Benefit
Count it when:
- the work was accepted, retained, or deployed
- the quality cost is known or bounded
- the evidence comes from observed data or a defensible sample
Do not count it when:
- it comes from a benchmark score only
- it comes from raw suggestion acceptance alone
- it comes from self-report with no operational corroboration
- it comes from a task that produced follow-up rework or policy exceptions
---
## Review Cost Is Not Optional
Most weak AI ROI models ignore reviewer time. Do not.
For assistants:
- measure review rounds
- measure requested-changes rate
- sample reviewer effort on accepted AI-heavy PRs
For agents:
- measure reviewer minutes per PR or per accepted task
- measure takeover and rework after "successful" completion
- allocate reviewer cost into unit economics
If reviewer effort rises faster than accepted value, the program is not scaling cleanly.
---
## Benchmark and Research Caveats
Use external research to bound assumptions, not to replace internal evidence.
Apply these rules:
1. If a study uses simple or lab-style tasks, discount its transferability to your repos.
2. If a benchmark is saturated, treat it as capability evidence, not ROI evidence.
3. If a claim comes from a vendor study, say it is vendor evidence.
4. If internal usage is assistant-heavy but the external study is agent-heavy, do not transfer the number directly.
Suggested language:
> External studies inform the assumption range, but the business case is anchored to our own accepted-work and review-cost data.
---
## Executive Reporting Structure
A good executive summary answers four questions:
1. What did we spend?
2. What accepted value did we get?
3. What risks or hidden costs offset that value?
4. What is the decision recommendation?
Use this sequence:
- program scope and measurement period
- scenario table
- top drivers of value
- top drivers of cost
- quality and safety caveats
- recommendation: expand, hold, narrow, or stop
---
## Common ROI Failure Modes
| Failure Mode | Why It Breaks the Model |
|-------------|--------------------------|
| counting all generated code as value | generated code is not accepted value |
| ignoring rework and revert cost | inflates benefit materially |
| ignoring reviewer labor | especially bad for agents |
| counting benchmark wins as dollars | benchmark != business impact |
| assuming all teams benefit equally | task mix and repo quality matter |
| using a honeymoon period | novelty inflates the early signal |
| mixing assistant and agent costs in one bucket | hides which workflow actually pays |
---
## Recommended Default Deliverables
### For a pilot
- conservative / base / upside ROI table
- cost per active user
- cost per accepted change
- recommendation with caveats
### For an agent rollout
- cost per merged PR
- reviewer effort trend
- takeover rate
- revert / exception rate
- recommendation on task envelope expansion
### For renewal
- trailing 2-3 quarter trend
- adoption by segment
- accepted work per dollar spent
- comparison to next-best alternative or to no-tool baseline
---
## What to Do Next
- For operating metrics, pair this file with `productivity-metrics.md`.
- For agentic unit economics, pair this file with `agent-execution-metrics.md`.
- For experiments, pair this file with `benchmarking-methodology.md`.
- For executive packaging, use `assets/executive-report-template.md` and `assets/roi-calculator-template.md`.
references/theory-of-constraints-applied.md
---
description: Theory of Constraints applied to AI coding metrics. Five Focusing Steps on the developer-flow bottleneck before instrumenting AI lift, throughput accounting for ROI, DBR for code-review queue protection, CRT for stalled rollouts, and evaporating cloud for adoption-vs-quality tensions.
foundation: foundations-theory-of-constraints
last_verified: 2026-05-03
status: stable
---
# Theory of Constraints Applied to AI Coding Metrics
> **Gate before invoking:** Check [`foundations-theory-of-constraints` § When to Apply](../../foundations-theory-of-constraints/SKILL.md#when-to-apply) first. The recipes below assume the foundation is the right tool for the situation; the foundation's skip-conditions route you to a different foundation if not.
_Last verified: 2026-05-03._
See [foundations-theory-of-constraints](../../foundations-theory-of-constraints/SKILL.md) for canonical primitive definitions, playbooks, and worked examples.
## Table of Contents
- [Why TOC Applies to AI Coding Metrics](#why-toc-applies-to-ai-coding-metrics)
- [Primitive Coverage Map](#primitive-coverage-map)
- [Pattern Catalog](#pattern-catalog)
- [P1 — Developer-Flow Bottleneck Identification Before Measuring AI Lift](#p1--developer-flow-bottleneck-identification-before-measuring-ai-lift)
- [P2 — Throughput Accounting for AI-Coding ROI](#p2--throughput-accounting-for-ai-coding-roi)
- [P3 — Drum-Buffer-Rope for Code-Review Queue Protection](#p3--drum-buffer-rope-for-code-review-queue-protection)
- [P4 — Evaporating Cloud for Adoption-vs-Quality Tensions](#p4--evaporating-cloud-for-adoption-vs-quality-tensions)
- [P5 — Current Reality Tree for Stalled AI-Tool Rollouts](#p5--current-reality-tree-for-stalled-ai-tool-rollouts)
- [P6 — Subordinating Non-Bottleneck Metrics to Avoid Local Optimisation](#p6--subordinating-non-bottleneck-metrics-to-avoid-local-optimisation)
- [Anti-Pattern Catalog](#anti-pattern-catalog)
- [A1 — Lines-of-Code Dashboards](#a1--lines-of-code-dashboards)
- [A2 — Isolated Tool-Adoption Metrics Ignoring Delivery](#a2--isolated-tool-adoption-metrics-ignoring-delivery)
- [A3 — Celebrating PR Throughput When Review Is the Bottleneck](#a3--celebrating-pr-throughput-when-review-is-the-bottleneck)
- [A4 — AI-Quality Gates Measured Without Baseline](#a4--ai-quality-gates-measured-without-baseline)
- [A5 — Cost-of-Delay Ignored in AI-Coding Investment Cases](#a5--cost-of-delay-ignored-in-ai-coding-investment-cases)
- [Recipes](#recipes)
- [R1 — Diagnose Your AI-Coding Bottleneck in 5 Questions Before Instrumenting](#r1--diagnose-your-ai-coding-bottleneck-in-5-questions-before-instrumenting)
- [R2 — ROI Scorecard Structured by Throughput / OE / Inventory](#r2--roi-scorecard-structured-by-throughput--oe--inventory)
- [R3 — Pilot-to-Rollout Metric Plan Using CRT and FRT](#r3--pilot-to-rollout-metric-plan-using-crt-and-frt)
- [Composition](#composition)
- [Sources](#sources)
---
## Why TOC Applies to AI Coding Metrics
An AI-coding programme is a multi-step delivery system. Writing code is one step; review, merge, QA, deploy, and production validation are the rest. AI assistants and agents can accelerate individual steps — but system throughput (features shipped, bugs resolved, value delivered) is set by the slowest link, not by the fastest tool.
This creates a measurement trap: teams instrument where AI is visible (code generation, acceptance rate, lines accepted) rather than where the constraint lives. If code review is the bottleneck, an AI tool that doubles authoring speed doubles the queue in front of review without increasing delivery. The metric looks great; the system gets worse.
TOC gives AI metrics three things cost-accounting or velocity-only measurement cannot:
1. **A reason to measure the system before measuring the tool** — 5FS identifies the constraint first; AI lift is only meaningful if it moves the constraint.
2. **A financial frame** — Throughput Accounting (T, I, OE) evaluates AI investment by delivery impact, not by licence cost or lines-of-code output.
3. **A conflict-resolution method** — Evaporating Cloud dissolves the recurring tension between adoption speed and quality gates without compromising either.
The primitives below are domain-specific applications of the canonical TOC tools. Full playbooks are in [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/).
---
## Primitive Coverage Map
| Primitive | # | Applied in Patterns / Recipes |
|-----------|---|-------------------------------|
| Five Focusing Steps | 1 | P1, R1 |
| Drum-Buffer-Rope | 2 | P3 |
| Throughput Accounting | 3 | P2, A1, A2, A5, R2 |
| Evaporating Cloud | 4 | P4 |
| Current Reality Tree | 5 | P5, R3 |
| Future Reality Tree | 6 | R3 |
| Policy Constraints | 10 | P5, A3 |
Full primitive playbooks: [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/)
---
## Pattern Catalog
### P1 — Developer-Flow Bottleneck Identification Before Measuring AI Lift
**Primitive**: #1 Five Focusing Steps → [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/01-five-focusing-steps.md`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/01-five-focusing-steps.md)
**When to use.** Before instrumenting any AI coding tool — assistant or agent — run 5FS on the team's delivery pipeline to identify where work actually stalls. Measuring AI lift before knowing the constraint produces metrics that are accurate about the tool and irrelevant to the system.
**The problem it solves.** Adoption metrics and code-generation stats tell you about one step. If that step is not the constraint, improving it does not increase throughput. A team that ships AI-assisted code 40% faster but whose review queue grows 40% longer has not improved delivery; it has moved the queue.
**5FS applied to a delivery pipeline:**
1. **Identify** — Map the full path from task-start to production. Measure cycle time and queue time at each stage: specification → code → review → QA → deploy. The constraint is the stage where WIP accumulates — not the slowest average stage, but the stage whose queue grows while upstream stages drain.
Concrete signals:
- PRs waiting for review for > 1 day while authoring is < 4 hours → review is the constraint.
- Code complete in hours but QA backlog spans days → QA is the constraint.
- Story cycle times long but individual stage times short → a handoff or approval policy is the constraint.
2. **Exploit** — Before deploying AI tooling, squeeze maximum throughput from the constraint with existing resources: pair-review on high-complexity PRs, async review SLA agreements, review checklist standardisation, draft-PR workflow to surface early feedback.
3. **Subordinate** — Confirm that any AI tooling investment is directed at the constraint stage. If review is the constraint, the highest-value AI application is review assistance (automated code review, AI-generated PR summaries, AI-flagged diff risks) — not code generation at the authoring stage.
4. **Elevate** — If exploitation is insufficient, invest in constraint capacity: hire reviewers, rotate review duty, adopt AI review tooling with real constraint impact. Evaluate elevation by whether the constraint stage's queue clears — not by whether the AI tool's utilisation metrics look good.
5. **Repeat** — After the constraint is broken, re-run the measurement pass. The constraint shifts. Authoring may become the new bottleneck once review is unblocked; that is the right time to invest in code-generation AI.
**Failure mode to avoid.** Measuring AI acceptance rate and code-generation speed as primary outcomes before identifying the delivery constraint. These are local metrics; they cannot tell you whether system throughput improved.
---
### P2 — Throughput Accounting for AI-Coding ROI
**Primitive**: #3 Throughput Accounting → [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/03-throughput-accounting.md`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/03-throughput-accounting.md)
**When to use.** When building or reviewing an AI coding investment case, pilot ROI scorecard, or renewal decision. Throughput Accounting reframes the ROI question from "what does the tool cost?" to "does the tool move the constraint and increase system throughput?"
**The problem it solves.** Cost-accounting-based AI ROI models calculate: (seat cost) vs. (time saved × hourly rate). This conflates operating expense with throughput. A tool that saves 10 developer-hours per week but does not move the delivery constraint has T/IU ≈ 0: the system ships no more software.
**TOC financial triad applied to AI coding:**
| TOC term | AI-coding meaning |
|----------|------------------|
| Throughput (T) | Rate at which working software reaches production and generates business value: features shipped, bugs resolved, deployment frequency at the constraint |
| Investment (I) | In-flight work: PRs open, tasks in-progress, agent runs not yet merged |
| Operating Expense (OE) | All costs to keep the system running: seat licences, infrastructure, review time, rework, model API costs |
**T/CU (Throughput per Constraint Unit)** becomes: delivery throughput increase per unit of constraint-stage time consumed by AI tooling. To compute:
1. Identify the constraint stage (from P1).
2. Measure baseline throughput at the constraint: PRs merged/week, features shipped/sprint.
3. After AI rollout, measure the same throughput signal at the constraint.
4. Divide throughput delta by incremental OE (total cost of the AI programme).
**Decision rules:**
- An AI tool that increases authoring speed but does not increase constraint-stage throughput has T/IU ≈ 0. Licence cost is pure OE with no T return.
- An AI tool that reduces review burden at the constraint (automated code review, AI PR summaries, risk triage) directly increases constraint throughput. T/IU is positive and measurable.
- An AI tool that reduces rework (better quality at generation time, catching defects before review) reduces OE and may increase T if rework was consuming constraint time.
**Cost-of-delay framing for investment cases.** When building the ROI case for leadership, anchor on cost of delay: what is the throughput cost per sprint of not addressing the constraint? If the constraint prevents two features per sprint from shipping, and each feature is worth £X in revenue or cost-avoidance, the investment case is: does this AI tool justify its total OE by removing or reducing that constraint? Seat-cost comparisons against engineer-hour savings miss this entirely.
**Failure mode to avoid.** Calculating ROI as (hours saved × hourly rate) when hours saved are in non-constraint authoring time. This produces a compelling number that measures local efficiency at a step that does not gate delivery.
---
### P3 — Drum-Buffer-Rope for Code-Review Queue Protection
**Primitive**: #2 Drum-Buffer-Rope → [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/02-drum-buffer-rope.md`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/02-drum-buffer-rope.md)
**When to use.** When AI coding tools have increased authoring speed and the code-review queue has grown correspondingly — or is projected to grow. DBR protects the review constraint from being flooded by upstream AI-accelerated output.
**The problem it solves.** Without WIP control, increasing authoring speed through AI directly increases review queue depth. More PRs arrive than reviewers can process; each PR waits longer; review quality degrades under pressure; defect escape rises. The AI tool appears to help delivery; the system outcome is worse review, higher rework, and no throughput gain.
**DBR translation for code review:**
| DBR term | Code-review meaning |
|----------|-------------------|
| Drum | Reviewer capacity: PRs reviewable per day at target quality without reviewer burnout |
| Buffer | WIP cap on the review queue: the maximum number of open PRs allowed before authoring intake is paused |
| Rope | The mechanism that limits new PR submission to the drum rate: a WIP limit enforced in the project board, a PR-per-author daily cap, or an AI triage gate that queues low-priority PRs |
**Mechanic.**
1. **Set the drum.** Measure reviewer capacity: how many PRs can the review team process per day at the team's target quality standard (not under pressure)? This is the drum rate. Example: a team of 3 reviewers, each reviewing 2 PRs/day at quality = 6 PRs/day.
2. **Set the buffer.** Determine the maximum acceptable review queue depth before throughput degrades. Rule of thumb: 1.5× the drum rate. At 6 PRs/day drum rate, buffer = 9 open PRs before the intake rope triggers.
3. **Set the rope.** Configure the intake limit. Options:
- A WIP limit label on the project board: no more than N PRs can move to "Ready for Review" simultaneously.
- An AI triage gate: AI-assisted pre-review flags PRs as HIGH/MEDIUM/LOW review priority; LOW-priority PRs enter a deferred queue rather than the main review column when the buffer is reached.
- A team norm: authors who finish a PR while the buffer is full pick up a review before opening the next PR.
4. **Subordinate authoring to review.** When the buffer is full, authors stop opening new PRs and shift to reviewing. This is the hardest cultural change: it feels like slowing down, but it is the correct TOC response to a flooded constraint. Measure it: the team's WIP at review should not grow week-over-week.
**Metric to track.** Review queue depth (P85 age of open PRs awaiting review) is the primary health signal. If it is growing after AI tooling adoption, the rope is not set or not respected. AI-accelerated authoring without DBR at review is a local optimisation that harms the system.
---
### P4 — Evaporating Cloud for Adoption-vs-Quality Tensions
**Primitive**: #4 Evaporating Cloud → [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/04-evaporating-cloud.md`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/04-evaporating-cloud.md)
**When to use.** When an AI coding rollout is deadlocked between two camps: those pushing for broad adoption now versus those requiring quality gates, safety reviews, or security sign-off before expansion. Both positions are well-reasoned; the team is in a compromise loop that satisfies neither.
**The problem it solves.** Compromise between adoption speed and quality gates typically produces a slow rollout with gates that are not rigorous enough to satisfy the quality camp, and not fast enough to satisfy the adoption camp. The underlying assumption that makes the two positions appear irresolvable is never surfaced.
**Cloud structure:**
```
Shared Goal (A): Capture AI coding productivity gains without introducing quality or security regressions
Requirement B (adoption camp): Maximise tool adoption to realise throughput benefits
→ Prerequisite D: Roll out AI coding tools to all teams with minimal friction
Requirement C (quality camp): Protect code quality, security posture, and developer trust
→ Prerequisite D′: Require quality gate review, security assessment, and baseline measurement before any team onboards
Conflict: D and D′ appear mutually exclusive — broad rollout and gated rollout cannot both be true.
```
**Arrow challenges:**
| Arrow | Assumption | Challenge |
|-------|-----------|-----------|
| B → D | Productivity benefit requires broad adoption | Controlled rollout to high-readiness teams first captures most throughput benefit; marginal benefit from forcing adoption on resistant teams is small |
| C → D′ | Quality protection requires a gate before any adoption | Quality measurement can run in parallel with adoption if a rollback plan is in place; gates before every team are not the only way to protect quality |
| A → B | Throughput benefit requires rapid adoption | Throughput benefit is a function of adoption at the constraint, not adoption breadth; targeted adoption at the constraint stage yields T gain faster than broad shallow rollout |
| A → C | Protecting quality requires slowing adoption | Automated quality baselines (defect escape rate, rework rate, review burden) can be established quickly and monitored continuously — gating does not have to be a manual checkpoint |
**Injection.** Progressive adoption with automated quality monitoring: roll out to constraint-stage teams first (P1) with automated baseline metrics running from day one. Quality gates are continuous dashboards, not one-time approval checkpoints. Security review runs in parallel on the first cohort, not as a prerequisite to all cohorts. If a quality metric degrades beyond a threshold, the rollout pauses — automatically, not by committee.
This satisfies B (adoption proceeds at the highest-value constraint stages) and C (quality is protected by continuous automated monitoring and automatic pause, not by manual gates that throttle adoption).
**Signal to apply.** The rollout has been "in discussion" for more than two sprint cycles; the two camps keep restating the same positions; solutions are half-measures that both camps accept reluctantly.
---
### P5 — Current Reality Tree for Stalled AI-Tool Rollouts
**Primitive**: #5 Current Reality Tree → [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/05-current-reality-tree.md`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/05-current-reality-tree.md)
**When to use.** An AI coding tool rollout has stalled: adoption is plateauing, usage is not sustaining, or the programme has been running for months without measurable delivery impact. Multiple explanations have been offered (training, change management, context quality, management support) but none has resolved the problem.
**The problem it solves.** Stalled rollouts have multiple visible symptoms — low acceptance rate, team resistance, reverting to old workflows — that are treated as independent problems with independent solutions. The CRT traces all symptoms to a single Core Problem, allowing one injection to unblock the entire rollout rather than patching symptoms one by one.
**Mechanic.**
Step 1: Collect 5–8 UDEs from the rollout. Write each as a concrete, negative, observable outcome:
```
UDE 1: "AI suggestion acceptance rate has been below 20% for 8 weeks despite training"
UDE 2: "Developers report suggestions are irrelevant to the codebase context"
UDE 3: "Three teams disabled the plugin after the first week and have not re-enabled it"
UDE 4: "Review burden increased after rollout — reviewers report more AI-generated noise in diffs"
UDE 5: "Engineering leads are not tracking AI usage in their team metrics"
UDE 6: "Pilot teams cannot articulate a business outcome the tool has changed"
```
Step 2: Build "If…Then" chains. Trace from UDEs toward a shared root:
```
IF the tool receives minimal codebase context (IDE setup incomplete, no project .clinerules or CLAUDE.md)
THEN suggestions are semantically irrelevant to the actual codebase → UDE 2
THEN developers do not accept suggestions → UDE 1
THEN developers disable the plugin → UDE 3
IF authoring speed increases but review WIP is uncapped
THEN diff noise increases → UDE 4
THEN reviewers report negative experience → reinforces UDE 3
IF no T-level metric is tracked (only acceptance rate)
THEN leads cannot connect usage to outcomes → UDE 5, UDE 6
THEN the programme cannot make a delivery-impact case → continued stall
```
Step 3: Identify the Core Problem. In most stalled rollouts, the core is one of two causes:
- **Context gap**: the tool was deployed without configuring the context layer (codebase conventions, project memory, tool instructions) that makes suggestions relevant. Fix: invest in context engineering before re-expanding adoption.
- **Metric mismatch**: the programme is measured on tool-adoption metrics (acceptance rate, seat activation) rather than delivery-system metrics (constraint throughput, defect escape, rework rate). Fix: redefine success criteria as T-level outcomes before the next rollout phase.
Step 4: Design the injection and validate with a Future Reality Tree before implementing, confirming that the injection resolves all UDE chains without introducing new undesirable effects.
**Policy constraint variant.** If the Core Problem traces to a policy — "AI tools require security review before use on any codebase containing PII" — apply policy-constraint analysis (primitive #10) before treating it as a fixed obstacle. Challenge the assumption: does the policy need to apply to all AI tools equally, or can a risk-tiered policy allow lower-risk tools to proceed while the security review covers high-risk agent capabilities?
---
### P6 — Subordinating Non-Bottleneck Metrics to Avoid Local Optimisation
**Primitive**: #1 Five Focusing Steps (step 3 — subordinate) → [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/01-five-focusing-steps.md`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/01-five-focusing-steps.md)
**When to use.** When the AI coding programme has an established metric set and teams are optimising individual metrics without improving delivery. Subordination ensures that non-bottleneck metrics are treated as supporting signals, not targets, once the constraint is identified.
**The problem it solves.** Once the constraint is identified, every other step in the pipeline should be subordinated to it: managed to keep the constraint fed and protected, not optimised independently. Metric systems that incentivise non-constraint steps to maximise their own output create local optimisation at the cost of system throughput — the classic Goldratt "local optima" failure.
**Application to an AI coding metric set:**
After running P1, the constraint is known. Example: review is the constraint.
| Metric | Constraint relationship | Correct treatment |
|--------|------------------------|-------------------|
| AI suggestion acceptance rate | Non-constraint (authoring) | Report as context; do not set improvement targets that drive authoring volume above review capacity |
| PRs merged per sprint | Constraint-stage output | Primary throughput signal; set targets and track trends here |
| PR review cycle time | Constraint efficiency | Drive improvement directly; this is where AI investment should focus |
| Lines of AI-generated code | Non-constraint | Eliminate from the dashboard or treat as a diagnostic, never a goal |
| Rework rate post-merge | Quality signal (affects constraint) | Track — rework that lands in the constraint queue inflates constraint OE |
| Deployment frequency | Downstream of constraint | Tracks whether constraint improvements are reaching production; secondary signal |
**Subordination rule.** If a metric is for a non-constraint step, it has one valid purpose: confirming the step is feeding the constraint adequately. It should never be a primary KPI or improvement target until the constraint shifts to that step.
---
## Anti-Pattern Catalog
### A1 — Lines-of-Code Dashboards
**Primitives implicated**: #3 Throughput Accounting, #1 Five Focusing Steps
**Description.** Engineering leadership publishes a dashboard tracking lines of code generated by AI tools, lines accepted, and AI-authored code as a percentage of total commits. These metrics become the primary evidence of AI programme success.
**Why it fails.** Lines of code is a classic cost-accounting proxy. It measures a byproduct of authoring, not throughput. Under Throughput Accounting, T is the rate at which working software reaches production and generates value — not the rate at which lines are written. An AI tool that generates 5,000 lines per sprint that never merge (still in review, failing tests, reverted) has produced T = 0 despite impressive LoC metrics.
Worse: optimising for LoC incentivises high-volume low-quality generation — the exact pattern that floods the review constraint and degrades system throughput.
**Fix.** Replace LoC dashboards with delivery-system metrics: PRs merged/week at the constraint stage, review cycle time, rework rate, deployment frequency. AI tool contribution shows up in these metrics when and only when it is actually moving the system. If the tool is not showing up in delivery metrics, it is not at the constraint.
---
### A2 — Isolated Tool-Adoption Metrics Ignoring Delivery
**Primitives implicated**: #3 Throughput Accounting
**Description.** The AI coding programme is measured entirely on adoption signals: seat activation, daily active users, prompts per user, acceptance rate. Delivery metrics — cycle time, merge rate, defect escape, deployment frequency — are tracked separately and not connected to the AI programme's reporting.
**Why it fails.** Adoption metrics measure I (inventory in the system — developers using the tool) and OE (are we getting usage out of the seats we are paying for?). They do not measure T. A programme that achieves 80% seat activation and 35% acceptance rate but has no measurable effect on delivery cycle time or defect escape has demonstrated usage, not value.
Disconnecting adoption from delivery metrics makes it structurally impossible to detect the failure mode where the tool is being used but the constraint is elsewhere and delivery is unchanged.
**Fix.** Pair every adoption metric with at least one delivery metric scoped to the same team and time window. Adoption without delivery movement is a signal to investigate the constraint, not a success story.
---
### A3 — Celebrating PR Throughput When Review Is the Bottleneck
**Primitives implicated**: #10 Policy Constraints, #1 Five Focusing Steps
**Description.** After AI tooling adoption, the programme celebrates an increase in PR-open rate (number of PRs opened per sprint) as evidence of productivity improvement. Review queue depth, review cycle time, and merge rate are not tracked.
**Why it fails.** PR-open rate is a non-constraint output metric when review is the bottleneck. Increasing the number of PRs opened without increasing review capacity floods the constraint. Each additional PR that opens adds to review queue depth, increases the age of all open PRs, and forces reviewers into time-pressure review that degrades quality.
A visible success metric (more PRs, more activity) disguises a system deterioration (review backlog, lower quality, higher rework). This is a policy constraint variant: the implicit policy of "more PRs = more productivity" throttles throughput by ignoring the constraint.
**Fix.** Track PR-open rate only in conjunction with merge rate, review cycle time, and review queue depth. If open rate rises and merge rate does not, the system is accumulating WIP at review. Apply P3 (DBR) to protect the review constraint before reporting open rate as a success metric.
---
### A4 — AI-Quality Gates Measured Without Baseline
**Primitives implicated**: #1 Five Focusing Steps, #3 Throughput Accounting
**Description.** A team adopts AI coding tools and introduces a quality gate — a defect escape threshold or rework rate ceiling — to confirm AI-generated code meets quality standards. However, the gate is set based on intuition or vendor benchmarks, with no pre-rollout baseline from the team's own delivery data.
**Why it fails.** Without a baseline, the quality gate cannot distinguish between:
- Quality that was already at or above the gate threshold before AI (the gate adds no signal).
- Quality that was below the gate and the AI is making it worse (the gate fires but cannot attribute causation without a baseline).
- Quality that was below the gate and the AI is improving it (passes the gate, but the programme cannot prove the improvement is attributable to the tool).
Gates without baselines are compliance theatre: they create a sense of measurement rigour without producing actionable signal.
**Fix.** Establish an 8-week pre-intervention baseline for all quality metrics in the scorecard before any AI tool is switched on in a team. The minimum baseline period is enforced because week-to-week variance in defect escape and rework routinely exceeds the signal size of AI tooling effects (see the SKILL.md measurement rules). Quality gate thresholds are set from actual baseline distributions, not benchmarks.
---
### A5 — Cost-of-Delay Ignored in AI-Coding Investment Cases
**Primitives implicated**: #3 Throughput Accounting
**Description.** The AI coding investment case compares tool licence cost against developer time saved (a cost-accounting frame). The cost of delay — the throughput lost per sprint by not addressing the delivery constraint — is never computed or presented.
**Why it fails.** A cost-accounting investment case produces a threshold: "the tool pays for itself if it saves N hours per developer per month." This frame is indifferent to whether the hours saved are at the constraint or not, and it ignores the throughput cost of the status quo.
If the constraint is causing two features per sprint to slip, and each feature represents £Y in revenue or cost-avoidance, the opportunity cost of inaction is 2 × £Y per sprint. An AI investment that removes the constraint may have a T/IU that is 5–10× the cost-savings frame — or it may have T/IU ≈ 0 if it addresses a non-constraint step. The cost-accounting frame cannot tell you which.
**Fix.** Build the investment case in two frames simultaneously: (a) the standard cost-savings frame for stakeholders who think in OE reduction, and (b) the throughput frame: what constraint does this investment address, by how much, and what is the throughput value of breaking that constraint? Present both. The throughput frame is the one that drives actual business impact; the cost-savings frame is the one that passes finance approval.
---
## Recipes
### R1 — Diagnose Your AI-Coding Bottleneck in 5 Questions Before Instrumenting
**Goal.** Identify the delivery-system constraint before designing an AI coding metric plan, so the metric set is anchored to the right stage of the pipeline.
**Inputs.** Team's current delivery process (task-start to production), at least 4 weeks of PR and deployment history.
**Step 1: Map the delivery pipeline.**
List every stage from task-start to production: specification, authoring, self-review, PR open, peer review, QA/CI, merge, deploy, production validation. For each stage, measure average cycle time and average queue time (time waiting to enter the stage after the prior stage completes).
→ verify: you have cycle time and queue time for at least 4 stages; data spans at least 4 weeks.
**Step 2: Answer the 5 diagnostic questions.**
Ask these in order. Stop when the first answer points to a bottleneck:
1. **Where does WIP accumulate?** Count open items by stage at the end of each week. The stage with the most consistently-high WIP count is the constraint candidate.
2. **Where is queue time longest?** If a stage has a short average cycle time but a long queue time (the work waits to enter the stage longer than it takes to complete it), the constraint is there — or immediately downstream.
3. **Does adding AI authoring speed increase the WIP at another stage?** If so, the downstream stage is the constraint. Adding more authoring speed will worsen it.
4. **Where do handoff delays occur?** Manual approval gates, batch release windows, or single-reviewer dependencies are policy constraint signals.
5. **What does the team complain about?** "We write code fast but can't get it reviewed" → review is the constraint. "Reviews are fast but QA takes forever" → QA is the constraint. "Deploys are infrequent despite code being ready" → deploy process is the constraint.
→ verify: at least one question points unambiguously to a stage. If not, extend the measurement window to 8 weeks.
**Step 3: Name the constraint.**
State it explicitly: "The constraint in this delivery pipeline is [stage]. Evidence: [queue time / WIP count / team feedback]." Document this as the anchor for the metric plan.
→ verify: the constraint is a named stage, not a vague description like "everything is slow."
**Step 4: Design AI investments for the constraint.**
Once the constraint is named, identify which AI capabilities address it:
- Review constraint → AI code review, PR summary generation, AI-assisted diff triage.
- QA constraint → AI test generation, AI-assisted defect detection, automated regression coverage.
- Authoring constraint (rare; only if authoring is truly the bottleneck) → code generation, AI refactoring, AI documentation.
- Deploy constraint → AI-assisted pipeline repair, AI incident triage, auto-rollback signal.
→ verify: every AI investment in the plan has a line connecting it to the named constraint stage.
**Step 5: Define the primary metric.**
The primary metric is the throughput signal at the constraint: PRs merged per week (if review is the constraint), deployments per week (if deploy is the constraint), QA-passed tickets per sprint (if QA is the constraint). Secondary metrics (acceptance rate, LoC, suggestion frequency) are diagnostic context, not success criteria.
→ verify: the primary metric is a constraint-stage throughput signal, not a tool-usage signal.
**Output.** A one-page constraint diagnosis document: constraint stage, evidence, AI investment map, and primary metric. This document replaces a vendor benchmark as the baseline for the AI coding programme.
---
### R2 — ROI Scorecard Structured by Throughput / OE / Inventory
**Goal.** Build an AI coding ROI scorecard that measures programme value in throughput terms rather than cost-savings terms, suitable for both engineering and finance audiences.
**Inputs.** Constraint diagnosis from R1, 8-week delivery baseline, AI programme costs (licence, infrastructure, context engineering, training, ongoing maintenance).
**Step 1: Establish the three TOC accounts.**
Map every measurable signal in the programme to T, I, or OE:
| Signal | Account | Measurement |
|--------|---------|-------------|
| Features shipped per sprint at constraint | T | Constraint-stage merge rate; deployment frequency |
| Defects reaching production | T (negative) | Defect escape rate; severity-weighted bug count |
| Reduction in rework cycles | T + OE | Rework PRs as % of total; rework hours |
| Open PRs / in-flight tasks at sprint end | I | WIP count at constraint stage at sprint end |
| AI tool licence cost | OE | Monthly seat cost |
| Model / API / infrastructure cost (agents) | OE | Token cost, compute cost per merged PR |
| Developer review time on AI-generated diffs | OE | Review hours at constraint stage |
| Training and onboarding time | OE | Hours per developer onboarded |
→ verify: every metric in the scorecard is assigned to exactly one account; no metric appears in both T and OE.
**Step 2: Compute the baseline and target T/IU.**
Baseline T: features shipped per sprint at the constraint stage, averaged over 8 pre-intervention weeks.
Total programme investment (I in the ROI sense, not the TOC inventory sense): sum of all OE increments attributable to the AI programme per sprint.
Baseline T/IU: T ÷ incremental OE. This is the throughput return on AI programme spend.
Target T/IU: set the minimum acceptable T/IU before programme renewal. Example: "We will renew if T/IU ≥ 1.5× baseline after 12 weeks."
→ verify: T/IU is computed from the constraint-stage throughput signal, not from a blended "productivity score."
**Step 3: Structure the scorecard.**
| Dimension | Signal | Baseline | 4-week | 8-week | 12-week | Target |
|-----------|--------|----------|--------|--------|---------|--------|
| **T — Delivery** | Constraint-stage merge rate | — | | | | +15% |
| **T — Quality** | Defect escape rate | — | | | | ≤ baseline |
| **T — Flow** | Review cycle time (P85) | — | | | | −20% |
| **I — WIP** | Open PRs at sprint end | — | | | | ≤ 5 |
| **OE — Tool** | Licence + infra cost/sprint | — | | | | fixed |
| **OE — Review** | Review hours at constraint | — | | | | −10% |
| **T/IU** | T delta ÷ incremental OE | 1.0× | | | | ≥ 1.5× |
**Step 4: Add the cost-of-delay row.**
For finance stakeholders, add a cost-of-delay calculation: estimated value of features not shipped per sprint at the pre-AI baseline, compared to the post-AI target. Express in revenue or cost-avoidance terms appropriate to the business. This row answers the question "what is the cost of not acting?" without replacing the T/IU row.
→ verify: the scorecard has at least one T signal, one I signal, one OE signal, and a T/IU composite. No LoC metric appears on the scorecard.
**Step 5: Track and review cadence.**
Review the scorecard at 4, 8, and 12 weeks post-rollout. Before week 8, treat all signals as directional — sample size is insufficient for causal claims (see SKILL.md study design defaults). A T/IU that is flat or declining by week 8 is the signal to run P5 (CRT) on the rollout before the 12-week renewal decision.
---
### R3 — Pilot-to-Rollout Metric Plan Using CRT and FRT
**Goal.** Design a metric plan for an AI coding pilot that produces credible evidence for rollout decisions, using CRT to surface the real problems the pilot must test and FRT to validate that the proposed rollout will resolve them.
**Inputs.** Proposed pilot scope (team, AI tool, duration), delivery system constraint from R1, existing concerns from stakeholders (quality, security, adoption, review burden).
**Step 1: Build a pre-pilot CRT from existing UDEs.**
Before the pilot begins, collect the undesirable effects stakeholders are attributing to the current (no-AI) state:
```
Example UDEs:
UDE 1: "Feature cycle time has grown from 3 to 5 days over the past two quarters"
UDE 2: "Senior developers spend > 30% of time in code review"
UDE 3: "Onboarding new developers to the codebase takes 8+ weeks"
UDE 4: "Defect escape rate has increased since the team doubled in size"
UDE 5: "Two experienced developers are leaving; knowledge transfer is blocking delivery"
```
Trace "If…Then" chains to the Core Problem. In this example: *Codebase context is not accessible to developers at authoring and review time — relying instead on synchronous knowledge transfer via senior developers — making review the constraint and senior-developer time the critical resource.*
The Core Problem tells you what the pilot must test: not "does AI generate code?" but "does AI coding tooling reduce senior-developer review burden and accelerate codebase onboarding without increasing defect escape?"
→ verify: the Core Problem is a single statement that, if resolved, would weaken at least three of the UDEs.
**Step 2: Design pilot metrics from the CRT.**
Each UDE that the AI tool claims to address becomes a pilot metric. Each metric needs a baseline and a success threshold:
| UDE addressed | Metric | Baseline period | Success threshold |
|---------------|--------|----------------|-------------------|
| UDE 1 (cycle time) | Feature cycle time (P85) | 8 weeks pre-pilot | −15% vs. baseline |
| UDE 2 (review burden) | Senior-developer review hours/sprint | 8 weeks pre-pilot | −20% vs. baseline |
| UDE 3 (onboarding) | Time to first solo PR for new devs | Last 3 cohorts | −25% vs. cohort average |
| UDE 4 (defect escape) | Defects per sprint in production | 8 weeks pre-pilot | ≤ baseline |
| UDE 5 (knowledge transfer) | Context-quality score (codebase Q&A eval) | Pre-pilot benchmark | Measurable improvement |
→ verify: every metric maps to a named UDE; no metric is a tool-usage signal without a delivery or quality link.
**Step 3: Build a pre-rollout FRT.**
Before committing to full rollout, build a Future Reality Tree: assume the pilot injection (AI coding tooling with context engineering) is applied at scale.
```
Injection: AI coding tools with full codebase context deployed to all teams
→ IF context-quality reduces irrelevant suggestions
AND AI-assisted review reduces senior-developer review time
THEN review cycle time decreases → resolves UDE 1
THEN senior developers have capacity for deeper design review → resolves UDE 2
THEN new developers can get faster answers from the tool → resolves UDE 3
Check for Negative Branch Reservations:
→ Could AI suggestions create more noise at review if context is incomplete?
Mitigation: context engineering is a prerequisite to rollout; measured by suggestion acceptance rate during pilot as a proxy for relevance.
→ Could reducing review burden inadvertently reduce the mentor relationship that senior developers provide?
Mitigation: track developer experience metric on mentorship satisfaction; if it degrades, review-assistance scope is adjusted.
```
→ verify: the FRT resolves all Core Problem-linked UDEs; every Negative Branch Reservation has a named mitigation with a metric.
**Step 4: Set the pilot-to-rollout decision criteria.**
State explicitly: the rollout proceeds if and only if, by the end of the pilot:
- T is at or above threshold for the two highest-priority UDE metrics.
- Defect escape rate is ≤ baseline (quality gate with baseline, not intuition — see A4).
- No Negative Branch Reservation has materialised without a mitigation in place.
If these criteria are not met, run P5 (CRT) on the pilot results before redesigning the rollout scope.
**Output.** A pilot design document containing: the pre-pilot CRT, the metric set with baselines and thresholds, the FRT with Negative Branch Reservations and mitigations, and the explicit pilot-to-rollout decision criteria. This replaces vendor pilot templates as the measurement foundation for AI coding programmes.
---
## Composition
| Workflow | Entry primitive | Secondary | Close with |
|----------|----------------|-----------|-----------|
| New AI programme design | P1 (5FS) → constraint identification | P2 TA for investment framing | R1 + R2 for metric plan |
| Review queue overloaded after AI adoption | P3 DBR → set rope | P1 to confirm review is the constraint | A3 as diagnostic check |
| Rollout adoption plateau | P5 CRT → Core Problem | P4 EC if conflict is sustaining the Core Problem | R3 FRT before redesigning rollout |
| Leadership ROI case | P2 TA → T/IU calculation | A5 as frame for cost-of-delay row | R2 for full scorecard |
| Adoption-vs-quality deadlock | P4 EC → surface the sustaining assumption | P1 to redirect adoption to the constraint | P6 to subordinate non-constraint metrics |
**Never start with metrics before identifying the constraint.** Applying Throughput Accounting (P2) or building a scorecard (R2) before running 5FS (P1) produces a metric set that optimises for the wrong stage. 5FS is always step zero.
---
## Sources
These sources underpin the TOC primitives applied here. Full citation list is in
[`../../foundations-theory-of-constraints/references/primitives-overview.md`](../../foundations-theory-of-constraints/references/primitives-overview.md).
- Goldratt, E.M. & Cox, J. (1984). *The Goal*. North River Press. — Origin of 5FS and throughput accounting; the core argument that system throughput is set by the constraint, not average performance.
- Goldratt, E.M. (1990). *The Haystack Syndrome*. North River Press. — Throughput Accounting formalisation: T, I, OE and the inversion of cost-accounting logic applied in P2 and R2.
- Goldratt, E.M. (1994). *It's Not Luck*. North River Press. — Evaporating Cloud and the "challenge every assumption" discipline applied in P4.
- Cox, J.F. & Spencer, M.S. (1998). *The Constraints Management Handbook*. CRC Press. — Policy constraint detection and DBR mechanics applied in P3 and P5.
- Dettmer, H.W. (2007). *The Logical Thinking Process*. ASQ Quality Press. — CRT, EC, and FRT construction methodology applied in P5 and R3.
- Kim, G., Behr, K. & Spafford, G. (2013). *The Phoenix Project*. IT Revolution Press. — TOC applied to IT operations and software delivery; the "three ways" and WIP control in development pipelines.
- Forsgren, N., Humble, J. & Kim, G. (2018). *Accelerate*. IT Revolution Press. — Empirical evidence on delivery performance metrics; grounds the constraint-stage throughput signals in P2 and R2 in research-validated outcomes.
- McKinsey & Company. (2023). *The economic potential of generative AI*. McKinsey Global Institute. — Baseline productivity research framing; grounds the throughput-vs-cost-savings distinction in A5.
- Primitive playbooks in [`../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/`](../../foundations-theory-of-constraints/assets/templates/theory-of-constraints/) — canonical per-primitive definitions, failure modes, and worked examples for all 11 TOC primitives.
scripts/extract_github_events.py
"""
extract_github_events.py — GitHub telemetry extractor for AI coding metrics.
Pulls pull-request or commit data from the GitHub REST API using only the
Python standard library (urllib + json). Outputs CSV to stdout or a file.
Usage:
python extract_github_events.py pulls --repo owner/name --since 2025-01-01
python extract_github_events.py commits --repo owner/name --since 2025-01-01 --output out.csv
Requirements:
GITHUB_TOKEN environment variable must be set (classic PAT or fine-grained
token with repo:read scope). Works with public repos without a token but
rate limits are lower (60 req/h vs 5000 req/h).
Rate-limit policy:
Reads X-RateLimit-Remaining on every response. When remaining < 20, sleeps
until X-RateLimit-Reset (UNIX timestamp returned by the API).
Pagination:
Follows RFC 5988 Link headers (rel="next") automatically.
"""
from __future__ import annotations
import argparse
import csv
import io
import json
import math
import os
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime, timezone
from typing import Any, Generator, Iterator
# ---------------------------------------------------------------------------
# HTTP helpers
# ---------------------------------------------------------------------------
GITHUB_API = "https://api.github.com"
_TOKEN = os.environ.get("GITHUB_TOKEN", "")
def _headers() -> dict[str, str]:
h = {
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": "2022-11-28",
}
if _TOKEN:
h["Authorization"] = f"Bearer {_TOKEN}"
return h
def _check_rate_limit(response_headers: Any) -> None:
"""Sleep until rate-limit resets if remaining budget is low."""
remaining = response_headers.get("X-RateLimit-Remaining")
reset_at = response_headers.get("X-RateLimit-Reset")
if remaining is not None and int(remaining) < 20:
if reset_at is not None:
reset_ts = int(reset_at)
now_ts = int(time.time())
wait = max(reset_ts - now_ts + 2, 1)
print(
f"[rate-limit] Remaining={remaining}; sleeping {wait}s until reset.",
file=sys.stderr,
)
time.sleep(wait)
def _get(url: str) -> tuple[Any, Any]:
"""HTTP GET; returns (parsed_json, http.client.HTTPResponse headers)."""
req = urllib.request.Request(url, headers=_headers())
try:
with urllib.request.urlopen(req, timeout=30) as resp:
_check_rate_limit(resp.headers)
return json.loads(resp.read()), resp.headers
except urllib.error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
raise RuntimeError(f"HTTP {exc.code} from {url}: {body}") from exc
def _paginate(url: str) -> Generator[Any, None, None]:
"""Yield all pages from a paginated GitHub endpoint."""
while url:
data, headers = _get(url)
yield data
link_header = headers.get("Link", "")
url = _parse_next_link(link_header)
def _parse_next_link(link_header: str) -> str | None:
"""Parse RFC 5988 Link header, return URL for rel=next or None."""
if not link_header:
return None
for part in link_header.split(","):
parts = [p.strip() for p in part.split(";")]
if len(parts) == 2 and parts[1] == 'rel="next"':
return parts[0].strip("<>")
return None
# ---------------------------------------------------------------------------
# Shared utilities
# ---------------------------------------------------------------------------
def _iso(ts: str | None) -> datetime | None:
if not ts:
return None
return datetime.fromisoformat(ts.replace("Z", "+00:00"))
def _hours_between(a: datetime | None, b: datetime | None) -> str:
if a is None or b is None:
return ""
delta = (b - a).total_seconds() / 3600
return f"{delta:.2f}"
def _open_output(path: str | None) -> io.TextIOWrapper:
if path:
return open(path, "w", newline="", encoding="utf-8")
return sys.stdout
# ---------------------------------------------------------------------------
# Subcommand: pulls
# ---------------------------------------------------------------------------
PULLS_FIELDS = [
"pr_number",
"author",
"opened_at",
"merged_at",
"additions",
"deletions",
"changed_files",
"review_count",
"time_to_first_review_h",
"time_to_merge_h",
]
def _fetch_pr_details(repo: str, pr_number: int) -> dict[str, Any]:
url = f"{GITHUB_API}/repos/{repo}/pulls/{pr_number}"
data, _ = _get(url)
return data
def _fetch_pr_reviews(repo: str, pr_number: int) -> list[dict[str, Any]]:
url = f"{GITHUB_API}/repos/{repo}/pulls/{pr_number}/reviews?per_page=100"
reviews: list[dict] = []
for page in _paginate(url):
reviews.extend(page)
return reviews
def cmd_pulls(args: argparse.Namespace) -> None:
repo = args.repo
since = args.since # ISO date string, e.g. "2025-01-01"
output_path = args.output
url = (
f"{GITHUB_API}/repos/{repo}/pulls"
f"?state=closed&sort=created&direction=desc&per_page=100"
)
out = _open_output(output_path)
try:
writer = csv.DictWriter(out, fieldnames=PULLS_FIELDS, lineterminator="\n")
writer.writeheader()
for page in _paginate(url):
stop = False
for pr in page:
opened_at = _iso(pr.get("created_at"))
if opened_at and opened_at.isoformat() < since:
stop = True
break
# Merged PRs only
if not pr.get("merged_at"):
continue
pr_number = pr["number"]
author = (pr.get("user") or {}).get("login", "")
merged_at = _iso(pr.get("merged_at"))
# Detail endpoint for additions/deletions/changed_files
detail = _fetch_pr_details(repo, pr_number)
additions = detail.get("additions", "")
deletions = detail.get("deletions", "")
changed_files = detail.get("changed_files", "")
# Reviews
reviews = _fetch_pr_reviews(repo, pr_number)
review_count = len(reviews)
first_review_at = None
if reviews:
first_review_at = min(
(_iso(r.get("submitted_at")) for r in reviews if r.get("submitted_at")),
default=None,
)
writer.writerow(
{
"pr_number": pr_number,
"author": author,
"opened_at": pr.get("created_at", ""),
"merged_at": pr.get("merged_at", ""),
"additions": additions,
"deletions": deletions,
"changed_files": changed_files,
"review_count": review_count,
"time_to_first_review_h": _hours_between(opened_at, first_review_at),
"time_to_merge_h": _hours_between(opened_at, merged_at),
}
)
if stop:
break
finally:
if output_path:
out.close()
# ---------------------------------------------------------------------------
# Subcommand: commits
# ---------------------------------------------------------------------------
COMMITS_FIELDS = [
"sha",
"author",
"date",
"additions",
"deletions",
]
def _fetch_commit_detail(repo: str, sha: str) -> dict[str, Any]:
url = f"{GITHUB_API}/repos/{repo}/commits/{sha}"
data, _ = _get(url)
return data
def cmd_commits(args: argparse.Namespace) -> None:
repo = args.repo
since = args.since
output_path = args.output
url = (
f"{GITHUB_API}/repos/{repo}/commits"
f"?since={since}T00:00:00Z&per_page=100"
)
out = _open_output(output_path)
try:
writer = csv.DictWriter(out, fieldnames=COMMITS_FIELDS, lineterminator="\n")
writer.writeheader()
for page in _paginate(url):
for commit in page:
sha = commit["sha"]
author_login = (
(commit.get("author") or {}).get("login")
or (commit.get("commit", {}).get("author") or {}).get("name", "")
)
date = (commit.get("commit", {}).get("author") or {}).get("date", "")
detail = _fetch_commit_detail(repo, sha)
stats = detail.get("stats", {})
additions = stats.get("additions", "")
deletions = stats.get("deletions", "")
writer.writerow(
{
"sha": sha,
"author": author_login,
"date": date,
"additions": additions,
"deletions": deletions,
}
)
finally:
if output_path:
out.close()
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
prog="extract_github_events",
description=(
"Extract GitHub PR or commit telemetry to CSV. "
"Set GITHUB_TOKEN env var for authenticated requests (5000 req/h vs 60)."
),
)
sub = parser.add_subparsers(dest="subcommand", required=True)
# pulls subcommand
pulls_p = sub.add_parser(
"pulls",
help="Export merged PR metrics: number, author, timings, review count.",
)
pulls_p.add_argument("--repo", required=True, metavar="OWNER/NAME", help="e.g. torvalds/linux")
pulls_p.add_argument(
"--since",
required=True,
metavar="YYYY-MM-DD",
help="Include PRs opened on or after this date.",
)
pulls_p.add_argument("--output", metavar="FILE", help="Write CSV here (default: stdout).")
# commits subcommand
commits_p = sub.add_parser(
"commits",
help="Export commit metrics: sha, author, date, additions, deletions.",
)
commits_p.add_argument("--repo", required=True, metavar="OWNER/NAME")
commits_p.add_argument("--since", required=True, metavar="YYYY-MM-DD")
commits_p.add_argument("--output", metavar="FILE", help="Write CSV here (default: stdout).")
return parser
def main() -> None:
parser = build_parser()
args = parser.parse_args()
if args.subcommand == "pulls":
cmd_pulls(args)
elif args.subcommand == "commits":
cmd_commits(args)
else:
parser.print_help()
sys.exit(1)
if __name__ == "__main__":
main()
scripts/README.md
# roi_calculator.py
Stdlib-only Python CLI for AI coding metrics analysis. No external dependencies — runs with any Python 3.9+ installation.
## Purpose
Gives engineering leaders and productivity teams fast, reproducible answers to three core questions:
1. **ROI** — How much time and money is the program saving? What is the payback period and annualized ROI?
2. **Score** — How healthy is adoption across all 6 metric families? What is the composite grade?
3. **Report** — A full Markdown dashboard combining scorecard, ROI, and per-family signal details.
## Quick Start
Run from the `dev-ai-coding-metrics/` directory:
```bash
# ROI: time saved, cost saved, payback period, annualized ROI %
python scripts/roi_calculator.py roi --input data/sample-ai-metrics.json
# Scorecard: 6-family scores, Strong/Developing/Weak ratings, health grade
python scripts/roi_calculator.py score --input data/sample-ai-metrics.json
# Scorecard with all key signals printed per family
python scripts/roi_calculator.py score --input data/sample-ai-metrics.json --signals
# Full Markdown report (prints to stdout)
python scripts/roi_calculator.py report --input data/sample-ai-metrics.json
# Full Markdown report written to file
python scripts/roi_calculator.py report --input data/sample-ai-metrics.json --output /tmp/ai-metrics-report.md
```
## JSON Input Format
All subcommands read from a structured JSON file. Required fields:
```json
{
"team_name": "Platform Engineering",
"team_size": 20,
"measurement_period_weeks": 12,
"ai_tooling_monthly_cost": 1200,
"avg_dev_hourly_rate": 95,
"hours_saved_per_dev_per_week": 3.5,
"adoption_pct": 72,
"notes": "12-week pilot. Baseline established from prior 12-week period.",
"metric_families": {
"adoption": { "score": 72, "signals": ["..."] },
"delivery": { "score": 61, "signals": ["..."] },
"quality": { "score": 54, "signals": ["..."] },
"economics": { "score": 78, "signals": ["..."] },
"experience": { "score": 66, "signals": ["..."] },
"agent_execution": { "score": 49, "signals": ["..."] }
}
}
```
| Field | Used by | Notes |
|---|---|---|
| `team_size` | roi | Number of developers in the program |
| `ai_tooling_monthly_cost` | roi | Total monthly spend on AI coding tools ($) |
| `avg_dev_hourly_rate` | roi | Fully-loaded hourly rate per developer ($) |
| `hours_saved_per_dev_per_week` | roi | Self-reported time saved per developer per week |
| `measurement_period_weeks` | score, report | Duration of the measurement window |
| `team_name` | all | Display label in output headers |
| `notes` | report | Free-text context shown in report header |
| `metric_families` | score, report | Object with per-family `score` (0-100) and `signals` array |
See `data/sample-ai-metrics.json` for a complete example with realistic values for a 20-person team.
## Subcommand Reference
```
python scripts/roi_calculator.py roi --help
python scripts/roi_calculator.py score --help
python scripts/roi_calculator.py report --help
```
## Scoring: Rating Bands
| Rating | Score Range | Interpretation |
|---|---|---|
| Strong | 80–100 | Healthy signal; sustain and expand |
| Developing | 60–79 | Progress visible; gaps remain |
| Weak | 0–59 | Requires focused intervention |
## Health Grade Scale
| Grade | Composite Score | Interpretation |
|---|---|---|
| A | 90–100 | Excellent across all families |
| B | 80–89 | Strong program with minor gaps |
| C | 70–79 | Mixed results; address weak families |
| D | 60–69 | Below threshold; reassess design |
| F | 0–59 | Program underperforming; consider reset |
## ROI Calculation Notes
The ROI subcommand uses:
- **Weekly hours saved** = `team_size × hours_saved_per_dev_per_week`
- **Annual value** = weekly hours saved × 52 × `avg_dev_hourly_rate`
- **Annual net savings** = annual value − (`ai_tooling_monthly_cost` × 12)
- **Payback period** = monthly tool cost ÷ weekly value of time saved
- **Annualized ROI %** = (annual net savings ÷ annual tool cost) × 100
Hours-saved inputs are self-reported estimates. Apply a conservative discount (50% is a reasonable starting point) before presenting to leadership. Triangulate with delivery and quality metrics from the scorecard.
scripts/roi_calculator.py
#!/usr/bin/env python3
"""
AI coding metrics ROI calculator — stdlib-only CLI tool.
Subcommands:
roi — Calculate ROI: time saved, cost saved, payback period, annualized ROI %
score — Score adoption across 6 metric families with a health grade (A/B/C/D/F)
report — Full metrics dashboard report in Markdown
Usage:
python scripts/roi_calculator.py roi --input data/sample-ai-metrics.json
python scripts/roi_calculator.py score --input data/sample-ai-metrics.json
python scripts/roi_calculator.py report --input data/sample-ai-metrics.json
python scripts/roi_calculator.py report --input data/sample-ai-metrics.json --output report.md
"""
from __future__ import annotations
import argparse
import datetime
import json
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Optional
# ---------------------------------------------------------------------------
# Constants
# ---------------------------------------------------------------------------
# 6 metric families and their display labels
FAMILY_KEYS = [
"adoption",
"delivery",
"quality",
"economics",
"experience",
"agent_execution",
]
FAMILY_LABELS = {
"adoption": "Adoption",
"delivery": "Delivery",
"quality": "Quality",
"economics": "Economics",
"experience": "Experience",
"agent_execution": "Agent Execution",
}
# Score thresholds for rating bands
RATING_BANDS = [
(80, "Strong"),
(60, "Developing"),
(0, "Weak"),
]
# Grade thresholds (applied to composite score 0-100)
GRADE_BANDS = [
(90, "A"),
(80, "B"),
(70, "C"),
(60, "D"),
(0, "F"),
]
WEEKS_PER_YEAR = 52
MONTHS_PER_YEAR = 12
# ---------------------------------------------------------------------------
# Data models
# ---------------------------------------------------------------------------
@dataclass
class FamilyScore:
key: str
label: str
score: int
rating: str
signals: list[str]
@dataclass
class ScoreResult:
families: list[FamilyScore]
composite: float
grade: str
summary: str
@dataclass
class RoiResult:
team_size: int
hours_saved_per_dev_per_week: float
weekly_hours_saved: float
monthly_hours_saved: float
annual_hours_saved: float
avg_dev_hourly_rate: float
monthly_value_saved: float
annual_value_saved: float
ai_tooling_monthly_cost: float
ai_tooling_annual_cost: float
monthly_net_savings: float
annual_net_savings: float
payback_weeks: float
annualized_roi_pct: float
measurement_period_weeks: int
# ---------------------------------------------------------------------------
# Core calculation functions
# ---------------------------------------------------------------------------
def classify_rating(score: int) -> str:
for threshold, rating in RATING_BANDS:
if score >= threshold:
return rating
return "Weak"
def classify_grade(composite: float) -> str:
for threshold, grade in GRADE_BANDS:
if composite >= threshold:
return grade
return "F"
def calc_family_scores(data: dict) -> list[FamilyScore]:
families_raw = data.get("metric_families", {})
results: list[FamilyScore] = []
for key in FAMILY_KEYS:
family_data = families_raw.get(key, {})
score = int(family_data.get("score", 0))
signals = family_data.get("signals", [])
rating = classify_rating(score)
results.append(FamilyScore(
key=key,
label=FAMILY_LABELS[key],
score=score,
rating=rating,
signals=signals,
))
return results
def calc_score(data: dict) -> ScoreResult:
families = calc_family_scores(data)
composite = sum(f.score for f in families) / len(families) if families else 0.0
grade = classify_grade(composite)
strong_count = sum(1 for f in families if f.rating == "Strong")
weak_count = sum(1 for f in families if f.rating == "Weak")
if grade in ("A", "B"):
summary = f"Program is performing well across most families ({strong_count} Strong). Sustain and expand."
elif grade == "C":
summary = f"Mixed results. Focus investment on Weak families ({weak_count} flagged). Address quality and agent risks."
elif grade == "D":
summary = f"Below threshold. Significant gaps in {weak_count} families. Reassess program design before expanding."
else:
summary = "Program is underperforming across the board. Consider a structured reset with a focused pilot scope."
return ScoreResult(
families=families,
composite=round(composite, 1),
grade=grade,
summary=summary,
)
def calc_roi(data: dict) -> RoiResult:
team_size = int(data.get("team_size", 0))
hours_saved = float(data.get("hours_saved_per_dev_per_week", 0.0))
hourly_rate = float(data.get("avg_dev_hourly_rate", 0.0))
monthly_cost = float(data.get("ai_tooling_monthly_cost", 0.0))
period_weeks = int(data.get("measurement_period_weeks", 12))
weekly_hours_saved = team_size * hours_saved
monthly_hours_saved = weekly_hours_saved * (MONTHS_PER_YEAR / WEEKS_PER_YEAR) * WEEKS_PER_YEAR / MONTHS_PER_YEAR
# Simpler: monthly = weekly * (52/12)
monthly_hours_saved = weekly_hours_saved * WEEKS_PER_YEAR / MONTHS_PER_YEAR
annual_hours_saved = weekly_hours_saved * WEEKS_PER_YEAR
monthly_value = monthly_hours_saved * hourly_rate
annual_value = annual_hours_saved * hourly_rate
annual_cost = monthly_cost * MONTHS_PER_YEAR
monthly_net = monthly_value - monthly_cost
annual_net = annual_value - annual_cost
# Payback period in weeks: time until cumulative savings cover first month of cost
# (i.e. how many weeks of savings equal one month of tool cost)
if weekly_hours_saved > 0 and hourly_rate > 0:
weekly_value = weekly_hours_saved * hourly_rate
payback_weeks = monthly_cost / weekly_value if weekly_value > 0 else float("inf")
else:
payback_weeks = float("inf")
# Annualized ROI % = (annual net savings / annual cost) * 100
annualized_roi_pct = (annual_net / annual_cost * 100) if annual_cost > 0 else 0.0
return RoiResult(
team_size=team_size,
hours_saved_per_dev_per_week=hours_saved,
weekly_hours_saved=weekly_hours_saved,
monthly_hours_saved=round(monthly_hours_saved, 1),
annual_hours_saved=round(annual_hours_saved, 1),
avg_dev_hourly_rate=hourly_rate,
monthly_value_saved=round(monthly_value, 2),
annual_value_saved=round(annual_value, 2),
ai_tooling_monthly_cost=monthly_cost,
ai_tooling_annual_cost=annual_cost,
monthly_net_savings=round(monthly_net, 2),
annual_net_savings=round(annual_net, 2),
payback_weeks=round(payback_weeks, 1),
annualized_roi_pct=round(annualized_roi_pct, 1),
measurement_period_weeks=period_weeks,
)
# ---------------------------------------------------------------------------
# Formatting helpers
# ---------------------------------------------------------------------------
def fmt_money(value: float, prefix: str = "$") -> str:
if value >= 1_000_000:
return f"{prefix}{value / 1_000_000:.2f}M"
if value >= 1_000:
return f"{prefix}{value / 1_000:.1f}K"
return f"{prefix}{value:.0f}"
def fmt_hours(value: float) -> str:
return f"{value:,.1f} hrs"
def fmt_pct(value: float) -> str:
return f"{value:.1f}%"
def fmt_weeks(value: float) -> str:
if value == float("inf"):
return "N/A"
return f"{value:.1f} weeks"
def print_separator(width: int = 62, char: str = "-") -> None:
print(char * width)
def print_kv(label: str, value: str, width: int = 36) -> None:
print(f" {label:<{width}} {value}")
def grade_bar(score: int, width: int = 20) -> str:
filled = round(score / 100 * width)
return "[" + "#" * filled + "." * (width - filled) + "]"
# ---------------------------------------------------------------------------
# Subcommand: roi
# ---------------------------------------------------------------------------
def cmd_roi(args: argparse.Namespace) -> None:
data = _load_json(args.input)
r = calc_roi(data)
team_name = data.get("team_name", "Engineering Team")
print()
print(f"=== ROI ANALYSIS: {team_name} ===")
print_separator()
print_kv("Team size", f"{r.team_size} developers")
print_kv("Measurement period", f"{r.measurement_period_weeks} weeks")
print_kv("Hours saved / dev / week", fmt_hours(r.hours_saved_per_dev_per_week))
print_kv("Avg dev rate (fully loaded)", fmt_money(r.avg_dev_hourly_rate) + "/hr")
print_kv("AI tooling cost (monthly)", fmt_money(r.ai_tooling_monthly_cost))
print_separator()
print(" TIME SAVED")
print_kv(" Weekly hours saved (team)", fmt_hours(r.weekly_hours_saved))
print_kv(" Monthly hours saved (team)", fmt_hours(r.monthly_hours_saved))
print_kv(" Annual hours saved (team)", fmt_hours(r.annual_hours_saved))
print_separator()
print(" FINANCIAL IMPACT")
print_kv(" Monthly value of time saved", fmt_money(r.monthly_value_saved))
print_kv(" Annual value of time saved", fmt_money(r.annual_value_saved))
print_kv(" Annual tool cost", fmt_money(r.ai_tooling_annual_cost))
print_kv(" Annual net savings", fmt_money(r.annual_net_savings))
print_separator()
print(" ROI SUMMARY")
print_kv(" Payback period", fmt_weeks(r.payback_weeks))
print_kv(" Annualized ROI", fmt_pct(r.annualized_roi_pct))
print()
print(" Note: Time-saved inputs are self-reported estimates. Validate")
print(" against delivery metrics before presenting to leadership.")
print()
# ---------------------------------------------------------------------------
# Subcommand: score
# ---------------------------------------------------------------------------
def cmd_score(args: argparse.Namespace) -> None:
data = _load_json(args.input)
result = calc_score(data)
team_name = data.get("team_name", "Engineering Team")
period = data.get("measurement_period_weeks", "?")
print()
print(f"=== ADOPTION SCORECARD: {team_name} ===")
print(f" Measurement period: {period} weeks")
print_separator()
col_w = [18, 7, 14, 22]
headers = ["Family", "Score", "Rating", "Health bar"]
header_row = " " + " ".join(f"{h:<{w}}" for h, w in zip(headers, col_w))
print(header_row)
print_separator()
for f in result.families:
bar = grade_bar(f.score)
row = " " + " ".join([
f"{f.label:<{col_w[0]}}",
f"{f.score:<{col_w[1]}}",
f"{f.rating:<{col_w[2]}}",
f"{bar:<{col_w[3]}}",
])
print(row)
print_separator()
print(f" Composite score: {result.composite:.1f} / 100")
print(f" Health grade: {result.grade}")
print()
print(f" Assessment: {result.summary}")
print()
if args.signals:
print(" --- Key Signals by Family ---")
for f in result.families:
if f.signals:
print(f"\n [{f.label}]")
for sig in f.signals:
print(f" • {sig}")
print()
# ---------------------------------------------------------------------------
# Subcommand: report
# ---------------------------------------------------------------------------
def cmd_report(args: argparse.Namespace) -> None:
data = _load_json(args.input)
team_name = data.get("team_name", "Engineering Team")
team_size = data.get("team_size", 0)
period = data.get("measurement_period_weeks", 0)
notes = data.get("notes", "")
adoption_pct = data.get("adoption_pct", None)
roi = calc_roi(data)
score_result = calc_score(data)
lines: list[str] = []
a = lines.append
a("# AI Coding Metrics Dashboard")
a("")
a(f"**Team:** {team_name} ")
a(f"**Team size:** {team_size} developers ")
a(f"**Measurement period:** {period} weeks ")
a(f"**Report date:** {datetime.date.today().isoformat()} ")
if notes:
a(f"**Notes:** {notes} ")
a("")
# --- Scorecard ---
a("---")
a("")
a("## Scorecard")
a("")
a(f"**Composite score:** {score_result.composite:.1f} / 100 ")
a(f"**Health grade:** {score_result.grade} ")
a("")
a("| Family | Score | Rating | Signals (sample) |")
a("|--------|-------|--------|-----------------|")
for f in score_result.families:
first_signal = f.signals[0] if f.signals else "—"
# Truncate long signals for table readability
if len(first_signal) > 70:
first_signal = first_signal[:67] + "..."
a(f"| {f.label} | {f.score} | {f.rating} | {first_signal} |")
a("")
a(f"> {score_result.summary}")
a("")
# --- ROI Analysis ---
a("---")
a("")
a("## ROI Analysis")
a("")
a("| Metric | Value |")
a("|--------|-------|")
a(f"| Hours saved / dev / week | {roi.hours_saved_per_dev_per_week} hrs |")
a(f"| Weekly hours saved (team) | {fmt_hours(roi.weekly_hours_saved)} |")
a(f"| Annual hours saved (team) | {fmt_hours(roi.annual_hours_saved)} |")
a(f"| Annual value of time saved | {fmt_money(roi.annual_value_saved)} |")
a(f"| Annual tool cost | {fmt_money(roi.ai_tooling_annual_cost)} |")
a(f"| Annual net savings | {fmt_money(roi.annual_net_savings)} |")
a(f"| Payback period | {fmt_weeks(roi.payback_weeks)} |")
a(f"| Annualized ROI | {fmt_pct(roi.annualized_roi_pct)} |")
a("")
a("> **Assumption note:** Hours-saved inputs are self-reported estimates.")
a("> Triangulate with delivery metrics before citing to leadership.")
a("> A conservative 50% discount on self-reported hours is a reasonable starting adjustment.")
a("")
# --- Per-family signals ---
a("---")
a("")
a("## Metric Family Details")
a("")
for f in score_result.families:
a(f"### {f.label} — {f.score}/100 ({f.rating})")
a("")
if f.signals:
for sig in f.signals:
a(f"- {sig}")
else:
a("- No signals recorded.")
a("")
# --- Measurement rules reminder ---
a("---")
a("")
a("## Measurement Rules Applied")
a("")
a("1. Baseline established before rollout (12-week reference period).")
a("2. Assistant and agent metrics tracked separately.")
a("3. Every speed metric paired with at least one quality metric.")
a("4. Results aggregated at team level — no individual surveillance.")
a("5. Self-reported inputs labeled as estimates, not evidence.")
a("")
report_text = "\n".join(lines)
output_path = getattr(args, "output", None)
if output_path:
Path(output_path).write_text(report_text, encoding="utf-8")
print(f"Report written to: {output_path}")
else:
print(report_text)
# ---------------------------------------------------------------------------
# JSON loader
# ---------------------------------------------------------------------------
def _load_json(path: str) -> dict:
p = Path(path)
if not p.exists():
print(f"Error: File not found: {path}", file=sys.stderr)
sys.exit(1)
try:
with p.open(encoding="utf-8") as f:
return json.load(f)
except json.JSONDecodeError as exc:
print(f"Error: Invalid JSON in {path}: {exc}", file=sys.stderr)
sys.exit(1)
# ---------------------------------------------------------------------------
# CLI wiring
# ---------------------------------------------------------------------------
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
prog="roi_calculator",
description="AI coding metrics ROI calculator — stdlib only. No pip install required.",
)
subparsers = parser.add_subparsers(dest="command", metavar="SUBCOMMAND")
subparsers.required = True
# --- roi ---
p_roi = subparsers.add_parser(
"roi",
help="Calculate ROI: time saved, cost saved, payback period, annualized ROI pct.",
description=(
"Reads team size, hours saved per dev per week, tooling cost, and hourly rate "
"from the input JSON and outputs weekly/monthly/annual savings, net benefit, "
"payback period in weeks, and annualized ROI percentage."
),
)
p_roi.add_argument(
"--input",
metavar="JSON_FILE",
required=True,
help="Path to metrics JSON file (e.g. data/sample-ai-metrics.json).",
)
p_roi.set_defaults(func=cmd_roi)
# --- score ---
p_score = subparsers.add_parser(
"score",
help="Score adoption across 6 metric families with a health grade (A/B/C/D/F).",
description=(
"Reads per-family scores from the input JSON and outputs a scorecard table "
"with Strong/Developing/Weak ratings, a composite 0-100 score, and a health grade."
),
)
p_score.add_argument(
"--input",
metavar="JSON_FILE",
required=True,
help="Path to metrics JSON file (e.g. data/sample-ai-metrics.json).",
)
p_score.add_argument(
"--signals",
action="store_true",
default=False,
help="Print all key signals for each family after the scorecard table.",
)
p_score.set_defaults(func=cmd_score)
# --- report ---
p_report = subparsers.add_parser(
"report",
help="Generate a full Markdown metrics dashboard report from a JSON input file.",
description=(
"Combines scorecard, ROI analysis, and per-family signal details into a single "
"Markdown report. Prints to stdout by default; use --output to write to a file."
),
)
p_report.add_argument(
"--input",
metavar="JSON_FILE",
required=True,
help="Path to metrics JSON file (e.g. data/sample-ai-metrics.json).",
)
p_report.add_argument(
"--output",
metavar="OUTPUT_FILE",
help="Write Markdown report to this file instead of stdout.",
)
p_report.set_defaults(func=cmd_report)
return parser
def main() -> None:
parser = build_parser()
args = parser.parse_args()
args.func(args)
if __name__ == "__main__":
main()
SKILL.md
---
name: dev-ai-coding-metrics
description: "Measures AI coding impact and extension robustness. Use when tracking delivery, quality trajectories, cost, experience, pilots, scorecards, or leadership reporting."
compatibility: Portable core. Works on Claude Code and Codex.
version: "1.2"
last_validated: 2026-08-21
---
# AI Coding Metrics
Measures coding assistants and coding agents without collapsing results into vanity metrics or one blended score.
The critical distinction is **mode**: assistants help inline or in chat; agents execute multi-step work and need task-level measurement. Do not measure them as if they were the same thing.
## When to Use This Skill
| Trigger | Example |
|---------|---------|
| Designing a pilot or rollout scorecard | "We're rolling out Copilot to 200 engineers — what do we measure?" |
| Diagnosing usage-up / outcomes-flat | "Seat utilization is 80% but PR throughput is unchanged" |
| Comparing assistant vs. agent workflows | "Should we instrument these separately?" |
| Building an ROI model or leadership report | "Finance wants a renewal decision by Q3" |
| Designing an experiment better than vendor benchmarks | "We can't trust the vendor's numbers — how do we run our own study?" |
## Defaults
| Rule | Rationale |
|------|-----------|
| Start from the decision, not the telemetry available | Prevents instrument-what-is-easy bias |
| Separate assistant and agent funnels | Mixing hides which workflow drives results |
| Pair every speed metric with quality + experience | Speed alone is misleading |
| Aggregate at team level | Individual dashboards become surveillance |
| Treat benchmarks as capability signals, not business KPIs | Benchmark gaps do not equal production gaps |
## Workflow
1. Define the decision.
2. Pick the program mode: assistant, agent, or mixed.
3. Build the minimum viable scorecard.
4. Choose the study design.
5. Produce one deliverable.
## ASCII Flow
```text
AI coding metrics request
-> decision to support: buy, renew, improve, prove, or diagnose
-> split mode: assistant, agent, or mixed
-> select scorecard families: adoption, delivery, quality, economics, experience
-> choose study design and baseline window
-> collect team-level and task-level evidence
-> report confidence, sample size, and confounds
-> deliver ROI model, dashboard, experiment plan, or executive report
```
## Quick Reference
## Decision to Deliverable Map
| Decision | Default Output |
|----------|----------------|
| buy, renew, or cut a tool | ROI model plus executive report |
| improve adoption | adoption metrics plus survey |
| prove delivery impact | productivity metrics plus experiment plan |
| check quality drift | quality metrics plus dashboard |
| understand trust or friction | developer-experience metrics plus survey |
| evaluate coding agents | agent-execution metrics plus experiment plan |
## Program Modes
| Mode | Unit of Analysis | Primary Emphasis |
|------|------------------|------------------|
| assistant | developer-day, team-week, repo-month | adoption, delivery, quality, experience |
| agent | task, PR, workflow run | task success, merge, revert, review burden, cost per accepted change |
| mixed | team-week plus task-level samples | separate the two funnels before combining results |
## Metric Families
Use the smallest scorecard that can answer the decision:
| Family | What It Tells You |
|--------|-------------------|
| adoption | whether usage is real and sustained |
| delivery | whether software flow is faster where AI actually touches the path |
| quality | whether speed gains are offset by defects, rework, review burden, or declining extension robustness |
| economics | whether the value justifies tool and operating cost |
| experience | whether developers trust the tool and want to keep using it |
| agent execution | whether autonomous workflows succeed in production, not just in demos |
## Study Design Defaults
Minimum baseline: **8 weeks** of pre-intervention data. Two-week baselines produce noisy causal inference — week-to-week variance in PR throughput, review lag, and defect escape routinely exceeds the signal size of AI tooling effects.
| Situation | Design |
|-----------|--------|
| new pilot, no control group | before/after with ≥8 weeks baseline |
| enough comparable teams | matched A/B or stratified assignment |
| teams resist permanent denial of tools | crossover design |
| agent workflow change on one task family | task-level shadow comparison or reviewer-blind evaluation |
| leadership wants a fast answer | balanced scorecard with explicit caveats, not a causal claim |
## Measurement Checklist
Use before publishing any AI coding report:
- [ ] Baseline established (≥8 weeks before intervention)
- [ ] Assistant and agent funnels tracked separately
- [ ] Every speed metric paired with at least one quality metric
- [ ] Sample size, confidence level, and study design stated
- [ ] Confounds documented (team changes, release pressure, policy changes)
- [ ] Vendor evidence labeled as vendor evidence
- [ ] Usage measured after stabilization (not week-1 novelty period)
- [ ] Review burden and rework cost included in ROI model
- [ ] Edit-capable agents measured across evolving-spec checkpoints, including late-checkpoint cost and quality slopes
- [ ] Aggregated at team level (no manager-visible individual dashboards)
## Current Evidence Posture (as of 2026-08-21)
| Claim | Evidence | Caveat |
|-------|----------|--------|
| AI amplifies existing strengths and weaknesses | DORA 2025 AI report; conditional-impact model confirmed | Not a universal accelerant |
| Experienced developers ~19% slower with early-2025 tools (RCT) | METR July 2025 RCT, realistic open-source tasks | Specific to early-2025 tooling generation |
| METR believes developers more sped-up in 2026 than 2025 | METR Feb 2026 update | 30-50% of participants declined no-AI tasks (selection bias); unreliable signal |
| Self-reported: median 1.4-2x value of work from AI (2026) | METR May 2026 survey, n=349 | Self-report; METR found 40pp gap between perceived and actual gains in 2025 study |
| Throughput +66%, PR review time +441%, incidents per PR +243% | Faros AI 2026 telemetry, 22k devs / 4k teams | Organizational telemetry, not RCT; PRs merged without review up +31% |
| DORA 2025: 90% of developers use AI daily | DORA 2025 AI report | Adoption does not equal delivery impact |
| Modeled first-year AI ROI ~39% (500-person org); adoption raises change-failure rate (5%->6%), an "instability tax" | DORA 2026 ROI of AI-Assisted Software Development report (Apr 2026) | Vendor-modeled scenario, not a cross-org RCT; treat the 39% figure as an illustrative scenario, not a universal benchmark |
| AI yields 35-40% gains on simple tasks but ~10% on complex legacy code | DORA 2026 ROI report | Reinforces task-complexity segmentation already required by this skill's study design defaults |
| DX Core 4 unifies DORA + SPACE + DevEx into 4 dimensions (Speed, Effectiveness, Quality, Business Impact) | DX Core 4, formalized publicly Apr 2026 | Vendor framework; specific benchmarks need independent replication |
| One-shot pass rates can miss degradation across repeated agent edits | SlopCodeBench v1, Mar 2026 preprint | Python experiments only; trajectory signals are not correctness proofs or universal targets |
## Anti-Gaming Checklist
Reject a scorecard or report if any of the following apply:
- [ ] Single blended AI productivity score mixing usage, speed, sentiment, and quality
- [ ] Seat activation or prompt volume cited as delivery impact
- [ ] Cross-team comparison without controlling for stack, task mix, staffing, or release pressure
- [ ] Measurement period is <8 weeks or includes week-1 novelty window
- [ ] Vendor benchmark cited as production ROI evidence
- [ ] Review burden excluded from ROI model
- [ ] Individual-level AI usage visible to managers
- [ ] Directional before/after movement stated as causal without controlled design
- [ ] SlopCodeBench averages or trajectory signals used as organizational targets or causal ROI evidence
## Navigation
**References**
- [references/adoption-metrics.md](references/adoption-metrics.md) — assistant and agent adoption funnels, metric definitions, stall patterns, privacy rules
- [references/productivity-metrics.md](references/productivity-metrics.md) — DORA and SPACE applied to AI workflows, delivery stack decomposition, confound management
- [references/quality-metrics.md](references/quality-metrics.md) — defect, complexity, test, security, and technical debt metrics with targets and alert thresholds
- [references/roi-framework.md](references/roi-framework.md) — full cost model (including review burden), benefit model, scenario planning, executive report structure
- [references/developer-experience-metrics.md](references/developer-experience-metrics.md) — satisfaction surveys, cognitive load, friction indicators, trust calibration, DX anti-patterns
- [references/agent-execution-metrics.md](references/agent-execution-metrics.md) — agent funnel, core metrics, reviewer burden, scorecards for pilot / scaling / executive decisions
- [references/benchmarking-methodology.md](references/benchmarking-methodology.md) — A/B, before/after, crossover, shadow designs; statistical rigor; confound management
- [references/theory-of-constraints-applied.md](references/theory-of-constraints-applied.md) — bottleneck identification before instrumenting, throughput accounting for ROI, DBR for review-queue protection, CRT for stalled rollouts, evaporating cloud for adoption-vs-quality tensions
- [references/evidence-update.md](references/evidence-update.md) — load when citing current research: METR RCT (2025 baseline), METR 2026 update (selection-bias caveat), DORA 2025 AI report, DX Core 4, Faros 2026 telemetry
**Assets and data**
- [assets/metric-dashboard-template.md](assets/metric-dashboard-template.md)
- [assets/adoption-survey-template.md](assets/adoption-survey-template.md)
- [assets/roi-calculator-template.md](assets/roi-calculator-template.md)
- [assets/executive-report-template.md](assets/executive-report-template.md)
- [assets/experiment-design-template.md](assets/experiment-design-template.md)
- [data/sources.json](data/sources.json)
- [data/sample-ai-metrics.json](data/sample-ai-metrics.json)
**Scripts**
- [scripts/roi_calculator.py](scripts/roi_calculator.py)
- [scripts/README.md](scripts/README.md)
## Cross-References
- [../dev-context-engineering/SKILL.md](../dev-context-engineering/SKILL.md)
- [../ai-agents/SKILL.md](../ai-agents/SKILL.md)
- [../qa-observability/SKILL.md](../qa-observability/SKILL.md)
- [../product-management/SKILL.md](../product-management/SKILL.md)
## Fact-Checking
- Known bugs, regressions, framework/compiler/runtime footguns, and version-specific crash or workaround guidance must be verified against current primary web sources before being treated as current fact.
- Verify current research claims, benchmark status, and vendor telemetry specifics before final advice.
- Prefer peer-reviewed, official, and first-party telemetry docs over social or vendor marketing claims.
- If live verification is unavailable, mark current-evidence claims as unverified.
## Learnings Loop
Before applying this skill on a non-trivial task, read `learnings.consolidated.md` in this directory (and `learnings.md` if present).
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to `learnings.md` via `agents-skills-feedback-loop/scripts/append_learning.py`. Do not modify `SKILL.md` itself.