references/audience-segmentation.md
---
title: Audience Segmentation — ICP Scoring, Persona Matrices, Segmentation Frameworks
domain: growth-strategy
level: 3
skill: growth-strategy
---
# Audience Segmentation Reference
> **Scope**: Audience segmentation frameworks, ICP (ideal customer/reader profile) scoring templates, and persona matrices for growth strategy decisions. Use when determining which audience to prioritize, how to tailor content, or when a channel is underperforming and audience mismatch may be the cause.
> **Version range**: Framework-agnostic — does not depend on platform-specific features.
> **Generated**: 2026-04-09 — validate demographic and behavioral data against your analytics, not assumptions.
---
## Overview
Audience segmentation fails when it stops at demographics ("software engineers, 25-45, male"). That describes a population, not a buyer or reader. Effective segmentation identifies the specific problem, moment of need, and decision trigger that bring a segment to your content or product. A segment without a problem is a demographic slice, not an audience. ICP scoring transforms gut-feel audience definitions into prioritized, measurable segments with clear content implications.
---
## ICP (Ideal Reader/Customer Profile) Scoring Matrix
Identify 3-5 candidate audience segments and score them. Do not collapse all segments into one persona — that is the most common segmentation mistake.
| Dimension | Weight | Segment A | Segment B | Segment C | Scoring Guide |
|-----------|--------|-----------|-----------|-----------|---------------|
| Problem acuity | 4× | ___ | ___ | ___ | How urgently do they need to solve this? |
| Willingness to pay | 3× | ___ | ___ | ___ | For paid products/newsletters — ability + willingness |
| Reachability | 3× | ___ | ___ | ___ | Can you find them and get their attention cost-effectively? |
| Content fit | 3× | ___ | ___ | ___ | Does your format, voice, and depth match their consumption habits? |
| Referral potential | 2× | ___ | ___ | ___ | Will they recommend to others? (B2B scores higher — procurement chains) |
| Retention likelihood | 2× | ___ | ___ | ___ | Will they stay engaged long-term or is this one-time interest? |
| Size | 1× | ___ | ___ | ___ | Is the segment large enough to sustain growth? |
| **Weighted Total** | **18×** | **___** | **___** | **___** | |
**Problem acuity scoring guide**:
- 9-10: Actively searching for a solution, experiencing pain daily
- 7-8: Aware of the problem, looking for better approaches intermittently
- 5-6: Knows a problem exists, not urgently seeking resolution
- 3-4: Would benefit from your content but isn't actively looking
- 1-2: No recognized need — content would need to create awareness of the problem first
**Worked example** (technical blog for platform engineers):
| Dimension | Weight | IC SWEs | Engineering Mgrs | CTOs |
|-----------|--------|---------|------------------|------|
| Problem acuity | 4× | 8 | 6 | 7 |
| Willingness to pay | 3× | 4 | 7 | 9 |
| Reachability | 3× | 8 | 5 | 3 |
| Content fit | 3× | 9 | 6 | 4 |
| Referral potential | 2× | 6 | 8 | 7 |
| Retention | 2× | 7 | 7 | 5 |
| Size | 1× | 9 | 6 | 3 |
| **Total** | | **161** | **124** | **100** |
Decision: Primary audience = individual contributor SWEs (highest ICP score). Secondary = engineering managers (lower reachability but higher monetization potential). CTOs are aspirational-adjacent — pitch to, don't write for.
---
## Persona Matrix Template
One persona per prioritized segment. Each persona needs a problem statement and a "jobs to be done" framing.
```
## Persona: [Descriptive Name, Not "Marketing Mary"]
### Context
- Role: [specific job title, not "manager"]
- Experience level: [years in role, seniority]
- Company context: [size, industry, stage]
### Problem Statement
- Primary frustration: [the specific recurring problem they need solved]
- Current workaround: [what they do today that is inefficient]
- Cost of the problem: [time wasted, money lost, stress, risk]
### Jobs to Be Done
- Functional job: [the task they need to accomplish]
- Emotional job: [how they want to feel after accomplishing it]
- Social job: [how they want to be perceived by peers/managers]
### Content Consumption Habits
- When do they read: [commute, lunch, dedicated learning time, reactive]
- Format preference: [long-form deep dives, quick tips, video, audio]
- Discovery pattern: [search engines, Twitter, Slack communities, colleagues]
- Depth expectation: [fundamentals or advanced practitioner level]
### Content Triggers
- What problem makes them search: [the exact scenario that leads them to look for content]
- What headline would stop their scroll: [format: "[specific problem] + [specific solution/outcome]"]
- What makes them subscribe: [concrete example of content or offer that earns the email]
### Anti-fit signals
- [This person will NOT engage if: ...]
- [Content that drives them away: ...]
```
---
## Segmentation Decision Framework
Use this to choose between segmentation strategies when starting from scratch.
| Segmentation Approach | Best For | Avoid When | Example |
|-----------------------|----------|------------|---------|
| **Problem-based** | Clear pain point, well-defined problem space | Problem is too broad or abstract | "Go engineers fighting production incidents at 3 AM" |
| **Role-based** | B2B content, professional tool adoption | Roles vary too much within the category | "Engineering managers doing first performance reviews" |
| **Experience-level** | Tutorial content, skill progression | Audience spans multiple levels within one article | "Developers learning Kubernetes for the first time" |
| **Context-based** | Situational content, migration/decision moments | Contexts are too diverse to address meaningfully | "Teams migrating from AWS to GCP" |
| **Belief-based** | Opinion content, community building | Audience holds heterogeneous beliefs | "Developers who believe fast iteration beats big-design-upfront" |
---
## Audience Validation Checklist
Before investing in a segment, validate these signals:
```
[ ] Can you find them in a specific online community? (subreddit, Slack, Discord, forum)
[ ] Are they actively asking questions in those communities? (problem acuity signal)
[ ] Does any existing content for this segment have visible engagement? (demand signal)
[ ] Can you name 5 real people who match this persona?
[ ] Have you interviewed at least 2 people in this segment about their problems?
[ ] Is there a measurable metric you can track that proves this segment is responding?
[ ] Does your content format match how this segment consumes content? (text vs video vs audio)
```
Passing < 4 of these = segment is unvalidated hypothesis, not audience.
Passing 5-7 = validated enough to test with 30 days of content.
---
## Segment Prioritization by Content Stage
Different segments become important at different audience sizes.
| Audience Size | Focus Segment | Reasoning |
|---------------|--------------|-----------|
| 0-500 subscribers | Narrow, specific, high-acuity | Only highly relevant content earns email in small audiences |
| 500-5,000 | Add adjacent segments | Enough core to support exploring edge cases and adjacent roles |
| 5,000-20,000 | Formalize segment content tracks | Segment-specific series, dedicated tags, email sequences |
| 20,000+ | Spin off niches or add paid tiers | High-value segments justify premium content; low-value segments justify free |
**Counter-intuitive truth**: The smallest, most specific audience is the easiest to grow from zero. "Go engineers building internal developer platforms" is a smaller search volume than "Go engineers" but also has near-zero competition and high problem acuity. Niche first, expand second.
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Demographic persona without a problem
**What it looks like**: "Our audience is software engineers, 28-40, living in urban areas, interested in career growth."
**Why wrong**: This describes who they are, not why they need your content. Two engineers with identical demographics can have wildly different content needs based on their specific problem and job context.
**Do instead**: Anchor every persona to a named frustration and a current workaround. "Senior engineers frustrated by slow CI who are manually running local test suites" is a persona. "Senior engineers, 30-45" is not.
**Fix**: Every persona must include a "Primary frustration" and a "Current workaround." If you cannot fill those in, the persona is not finished.
### Building for the audience you have instead of the audience you want
**What it looks like**: Analytics shows 80% of current readers are junior developers, so all content targets juniors — even though the creator's goal is to build authority with senior engineers.
**Why wrong**: Optimizing for current audience composition entrenches current positioning. If the ICP is senior engineers, the content must serve senior engineers even when they are <20% of current readers.
**Do instead**: Score ICP fit against the target you are building toward, not the readers you already have. Write for senior engineers even when they are 15% of current traffic, and accept a 12-month lag before composition shifts.
**Fix**: Segment ICP scoring is forward-looking. Compare your highest-scored ICP against your current audience composition. If there is mismatch, the strategy task is to shift composition over 12 months, not to serve the current composition.
### Persona proliferation
**What it looks like**: 7 detailed personas, all with equal priority, representing every possible reader type.
**Why wrong**: If everything is a priority, nothing is a priority. Content written to serve 7 personas simultaneously serves none of them distinctively.
**Do instead**: Maintain exactly one primary persona and at most one secondary. Every piece of content maps to one of those two. Remaining candidates go on a watch list for future quarters, not into the active rotation.
**Fix**: Maximum 2 active personas at any time. One primary (80% of content) and one secondary (20% of content). Additional personas go on the watch list for future quarters.
---
## Detection Commands Reference
```bash
# Find community size for audience validation (manual — check these sources)
# Reddit: https://www.reddit.com/search/?q={topic} (look for r/ subreddits, member count)
# Slack: search "join {niche} slack community" for public communities
# Twitter: search "{niche} {problem}" and filter for engagement
# Estimate search demand for problem-based segment
# Manual: Google Keyword Planner or ahrefs/semrush for "{problem} {tool}" search volume
# Signal: >100 monthly searches = validated demand. <10 = emerging or too niche.
```
---
## See Also
- `channel-evaluation.md` — channel scoring matrices and CAC/LTV models
- `skills/competitive-intel/references/market-positioning.md` — how to position for specific segments against competitors
references/channel-evaluation.md
---
title: Channel Evaluation — CAC/LTV Models, Content Funnel Stages, Channel Matrices
domain: growth-strategy
level: 3
skill: growth-strategy
---
# Channel Evaluation Reference
> **Scope**: Channel evaluation matrices, CAC/LTV models for content and product channels, and content funnel stage mapping. Covers the decision of which 2-3 channels a solo creator or small team should invest in, and how to measure whether those channels are working.
> **Version range**: Platform-agnostic — specific platform metrics (impressions, reach algorithms) change; the evaluation framework does not.
> **Generated**: 2026-04-09 — validate platform-specific CAC benchmarks against current data.
---
## Overview
Most growth strategy failures are channel diffusion failures: too many channels, each receiving insufficient investment to reach escape velocity. Organic channels (SEO, community, referral) need 6-12 months of consistent investment before producing measurable returns. Solo creators who spread across 5 channels simultaneously give each channel 4 months — exactly the minimum needed before any signal emerges. Channel evaluation helps you pick fewer channels and sustain them long enough to learn.
---
## Channel Evaluation Matrix
Score each channel before committing investment. Score 1-10 per dimension. Customize weights for your constraint.
| Channel | Reach Potential | Audience Fit | Creator Fit | Time to Signal | CAC | Compounding | Total Score |
|---------|----------------|--------------|-------------|----------------|-----|-------------|-------------|
| **Weight** | **2×** | **3×** | **2×** | **2×** | **3×** | **3×** | **/15 dims** |
| SEO / organic search | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| Email newsletter | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| Twitter/X | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| LinkedIn | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| YouTube | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| Podcast | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| Reddit / communities | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
| Paid (Meta/Google) | ___ | ___ | ___ | ___ | ___ | ___ | **calc** |
**Scoring guides**:
- **Reach Potential** (1-10): Size of addressable audience on this channel relative to your niche
- **Audience Fit** (1-10): Is your specific audience active and engaged on this channel?
- **Creator Fit** (1-10): Do you enjoy producing this format? Can you sustain it for 12 months?
- **Time to Signal** (1-10): 10 = signal in weeks (paid ads), 1 = signal in 12+ months (SEO from zero)
- **CAC** (1-10): 10 = near-zero CAC (referral), 1 = very high CAC (paid acquisition in competitive niche)
- **Compounding** (1-10): Does it compound over time? 10 = SEO builds permanently; 1 = paid stops instantly when you stop paying
**Worked example** (technical blog for Go developers):
| Channel | Reach | Fit | Creator | Signal | CAC | Compound | Score |
|---------|-------|-----|---------|--------|-----|----------|-------|
| SEO | 7 | 9 | 7 | 2 | 9 | 10 | **144** |
| Email | 5 | 8 | 9 | 5 | 8 | 7 | **126** |
| Twitter/X | 8 | 6 | 5 | 7 | 7 | 4 | **102** |
| YouTube | 8 | 5 | 3 | 4 | 7 | 8 | **97** |
Decision: Invest in SEO + email (top 2). Twitter as lightweight amplifier (< 2 hrs/week). YouTube deferred until email list hits 1,000.
---
## CAC / LTV Model for Content Channels
For content-based businesses, "customer" = subscriber/reader. Adapt definitions for your business model.
### CAC Calculation
```
## CAC by Channel
### SEO
- Content creation cost per post: ___ hours × $___/hour = $___
- Monthly organic signups attributed to SEO: ___
- Monthly CAC (SEO): $___/month ÷ ___ signups = $___/subscriber
### Email (direct)
- Monthly email campaign cost: $___
- New subscribers from email campaigns: ___
- Monthly CAC (email): $___/month ÷ ___ signups = $___/subscriber
### Social (Twitter/LinkedIn/etc.)
- Hours/month × $___/hour = $___
- New subscribers attributed to social: ___
- Monthly CAC (social): $___/month ÷ ___ signups = $___/subscriber
### Paid Acquisition
- Ad spend per month: $___
- Paid subscribers acquired: ___
- Monthly CAC (paid): $___/month ÷ ___ signups = $___/subscriber
```
### LTV Calculation
```
## LTV by Subscriber Type
### Free subscriber
- Revenue per free subscriber: $0 direct
- Indirect value: [referrals × conversion rate × paid LTV] + [sponsored post CPM equivalent]
- Estimated LTV: $___
### Paid subscriber (if applicable)
- Monthly or annual revenue: $___/subscriber
- Average retention: ___ months
- LTV: $___/month × ___ months = $___
### LTV:CAC Ratio by Channel
Channel | LTV | CAC | Ratio | Verdict
SEO | $___ | $___ | ___:1 | ___
Email | $___ | $___ | ___:1 | ___
Social | $___ | $___ | ___:1 | ___
```
**Benchmark ratios**:
- LTV:CAC < 1:1 = destroying value, stop this channel
- LTV:CAC 1-3:1 = borderline; only continue if compounding effect is strong
- LTV:CAC 3-10:1 = healthy; sustain and optimize
- LTV:CAC > 10:1 = exceptional; increase investment
---
## Content Funnel Stage Mapping
Map your content to funnel stages. A common failure is creating only awareness content with no conversion path.
| Stage | Goal | Content Types | Success Metric | Conversion Action |
|-------|------|--------------|----------------|-------------------|
| **Awareness** | Reach new people | SEO posts, viral social, podcast appearances | Unique visitors, impressions | — |
| **Interest** | Deepen engagement | Long-form content, tutorials, case studies | Time on page, scroll depth, return visits | Subscribe to email |
| **Consideration** | Build trust | Email series, deep dives, comparison guides | Email open rate, click-through | Follow CTA, book call |
| **Conversion** | Monetize | Product pages, sales emails, webinars | Revenue, conversion rate | Purchase |
| **Retention** | Keep and grow | Community, updates, exclusive content | Churn rate, referral rate | Upgrade, referral |
**Funnel audit** — classify your last 10 pieces of content:
```
Content Item | Funnel Stage | Has Conversion CTA? | Next Step for Reader?
1. | | |
2. | | |
...
```
Common finding: 80% of content is Awareness. No Interest or Consideration content means the funnel leaks. Readers who want to go deeper have nowhere to go.
---
## Channel Sustainability Check
Before committing to a channel, verify you can sustain it for 6 months minimum.
| Channel | Weekly Hours Required | Monthly Cost | Can Sustain 6 months? |
|---------|-----------------------|-------------|----------------------|
| SEO blog (1 post/week) | 6-10 hours | $0-50 (tools) | ___ |
| Email newsletter (1x/week) | 2-4 hours | $20-100 (provider) | ___ |
| Twitter (daily) | 1-2 hours | $0 | ___ |
| YouTube (1 video/week) | 8-15 hours | $0-200 (equipment) | ___ |
| Podcast (1 episode/week) | 4-8 hours | $20-50 (hosting) | ___ |
**Capacity reality check**:
```
Available hours/week for growth work: ___
Total hours required for selected channels: ___
Buffer (should be 20%+): ___
If selected channel hours > 80% of available hours: remove a channel.
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Platform-native metrics as success metrics
**What it looks like**: "We got 50 retweets!" — but zero new email subscribers, zero traffic, zero revenue.
**Why wrong**: Platform metrics (likes, impressions, follower count) measure engagement with the platform, not growth of YOUR asset. A viral tweet that doesn't move subscribers is entertainment, not growth.
**Do instead**: Report only metrics that connect to your owned asset (email subscribers, direct traffic) or revenue. Track platform metrics separately as leading indicators, never as primary success measures.
**Fix**: Track only metrics that connect to your owned channel (email subscribers, direct traffic) or revenue. Report platform metrics only as leading indicators of owned metrics.
### Optimizing for CAC without measuring LTV
**What it looks like**: "Reddit drives the cheapest subscriber acquisition at $0.50 each."
**Why wrong**: If Reddit subscribers have 50% lower LTV (because they signed up for one specific post and have low topic affinity), the channel is less efficient than it appears.
**Do instead**: Measure 90-day engagement rate per acquisition channel and multiply it by CAC to get channel-adjusted CAC. A channel with $2 CAC and 80% engagement beats a channel with $0.50 CAC and 10% engagement.
**Fix**: Track subscriber cohort behavior by acquisition channel. Measure 90-day engagement rate per channel. CAC × engagement rate gives channel-adjusted CAC.
### Treating all funnel stages as the same
**What it looks like**: Every piece of content ends with "check out my other posts." No direct subscribe CTAs, no product mentions, no differentiated asks by content type.
**Why wrong**: Awareness content with hard conversion CTAs drives away top-of-funnel readers. Consideration content without conversion paths wastes intent. Each stage needs its own CTA.
**Do instead**: Assign one primary CTA per funnel stage before writing the content. Awareness gets a subscribe prompt. Interest gets a resource download. Consideration gets a trial or demo link. Conversion gets a purchase path.
**Fix**: Assign one primary CTA per funnel stage. Awareness = subscribe. Interest = download resource. Consideration = book a call or trial. Conversion = purchase.
---
## See Also
- `audience-segmentation.md` — ICP scoring, persona matrices, audience segmentation for channel targeting
- `skills/growth-strategy/SKILL.md` — full growth strategy workflow with 90-day planning
references/competitive-mapping.md
---
title: Competitive Landscape Mapping — Templates and Feature Comparison Matrices
domain: competitive-intel
level: 3
skill: competitive-intel
---
# Competitive Landscape Mapping Reference
> **Scope**: Competitive landscape mapping templates, feature comparison matrices, and competitive intelligence collection frameworks. Use at the start of any competitive analysis to structure the landscape before diving into individual competitor profiles.
> **Version range**: Framework-agnostic — platform and product specifics change; the mapping structure does not.
> **Generated**: 2026-04-09 — competitive landscapes change fast; date-stamp all data you collect.
---
## Overview
Most competitive analyses drown in detail about individual competitors and never produce a map of the overall landscape. A landscape map reveals structure that individual profiles miss: where the market is clustered, where it is empty, which segments are contested vs. undefended, and what the dominant competitive dimensions are. Build the map before profiling individuals.
---
## Competitive Landscape Map Template
**Step 1: Identify the competitive dimensions**
Before placing anyone on the map, choose the two dimensions that best explain why customers choose between options. Bad dimensions: "quality" and "price" (every market). Good dimensions are specific to the category.
| Market Type | Dimension Examples |
|-------------|-------------------|
| SaaS tools | Feature depth vs. ease of use; Self-serve vs. sales-led; SMB vs. Enterprise |
| Content / media | Technical depth vs. accessibility; Broad vs. niche; Free vs. paid |
| Services | Speed vs. quality; Specialization vs. breadth; Async vs. synchronous |
| Consumer products | Premium vs. value; Established vs. emerging; B2C direct vs. retail |
**Step 2: Map the landscape**
```
## Competitive Landscape: [Market/Category Name]
## Date: [YYYY-MM-DD] — refresh every 6 months
### Competitive Dimensions
- X-axis: [Dimension 1] — Low (left) to High (right)
- Y-axis: [Dimension 2] — Low (bottom) to High (top)
### Players Mapped
| Player | X-axis position (1-10) | Y-axis position (1-10) | Tier | Notes |
|--------|------------------------|------------------------|------|-------|
| Us | ___ | ___ | — | Current position |
| [Competitor A] | ___ | ___ | Direct | |
| [Competitor B] | ___ | ___ | Direct | |
| [Competitor C] | ___ | ___ | Adjacent | |
| [Competitor D] | ___ | ___ | Aspirational | |
| [Emerging E] | ___ | ___ | Emerging | |
### White Space Analysis
- Underserved quadrant: [describe position that has demand but no strong player]
- Overcrowded quadrant: [describe where competition is densest]
- Our target position: [where we aim to be vs. where we are]
```
**Worked example** (indie technical blog in Go/cloud-native space):
X-axis: Accessibility (1 = expert-only, 10 = beginner-friendly)
Y-axis: Technical depth (1 = surface overview, 10 = production-grade detail)
| Player | Accessibility | Depth | Tier |
|--------|--------------|-------|------|
| Our blog | 4 | 8 | — |
| GoByExample.com | 9 | 3 | Adjacent |
| Official Go Blog | 3 | 7 | Aspirational |
| Dave Cheney's blog | 3 | 9 | Aspirational |
| Ardan Labs | 4 | 9 | Direct |
| Random medium posts | 7 | 3 | Adjacent |
White space: High accessibility + high depth (position 7-8, 7-8). No current player makes production-grade content accessible to mid-level engineers. That is the target position.
---
## Competitor Tier Classification
| Tier | Definition | Research Depth | Review Cadence |
|------|-----------|----------------|----------------|
| **Direct** | Same audience, same problem, same format/channel | Full profile (see template) | Monthly |
| **Adjacent** | Same audience, different approach | 1-paragraph summary | Quarterly |
| **Aspirational** | Larger player in a broader version of your space | Strategy extraction only | Quarterly |
| **Emerging** | New entrant showing directional signals | Watch list only | Quarterly |
| **Legacy** | Established player in decline | Monitor for opportunities | Semi-annually |
**Direct competitor cap**: Limit full analysis to 3 direct competitors max. Above 3, you are doing competitive research tourism instead of strategy.
---
## Feature Comparison Matrix
Use when technical/product capability comparison is the primary decision. Adapt feature list to your market.
| Feature / Capability | Us | Competitor A | Competitor B | Notes |
|---------------------|-----|--------------|--------------|-------|
| **Core Features** | | | | |
| [Feature 1] | ✓ Full | ✓ Full | ✗ Missing | |
| [Feature 2] | ~ Partial | ✓ Full | ✓ Full | |
| [Feature 3] | ✗ Missing | ✓ Full | ✓ Full | Roadmap Q3 |
| **Differentiating Features** | | | | |
| [Our strength] | ✓ Full | ✗ Missing | ✗ Missing | Core differentiator |
| [Their strength] | ✗ Missing | ✓ Full | ~ Partial | |
| **Price / Packaging** | | | | |
| Entry price | $0 | $29/mo | $49/mo | |
| Enterprise tier | Contact | $500/mo | $800/mo | |
| Free tier | Yes | No | Yes (limited) | |
| **Support / Community** | | | | |
| Documentation quality | High | Medium | Low | |
| Community size | Small | Large | Medium | |
| Support SLA | None | 48h | 24h | |
**Legend**: ✓ Full = complete implementation. ~ Partial = limited or early. ✗ Missing = not offered. ? = unknown.
**Decision use**: A feature matrix reveals where you are competitive, where you are weak, and what you can win on. Never use it to claim superiority on features you are "partial" on.
---
## Competitor Activity Tracker
Track what competitors actually do, not just what they claim. Update this monthly.
```
## [Competitor Name] — Activity Log
### Product Activity
- Date | Change | Signal
- [YYYY-MM-DD] | Launched [feature] | [what it means for their strategy]
- [YYYY-MM-DD] | Deprecated [feature] | [what it means]
- [YYYY-MM-DD] | Changed pricing to [model] | [impact on our positioning]
### Content Activity
- Date | Topic/Format | Engagement | Signal
- [YYYY-MM-DD] | [post title] | [high/med/low] | [what topics they're doubling down on]
### Company Activity
- Date | Event | Signal
- [YYYY-MM-DD] | Hired [role] | [capability they're building]
- [YYYY-MM-DD] | Raised [amount] | [runway and ambition signal]
- [YYYY-MM-DD] | Partnership with [company] | [market expansion or channel signal]
### Sentiment Signals
- Most praised (G2/Reddit/HN): [what users love]
- Most criticized: [what users complain about — your opportunity]
- Recent sentiment shift: [positive/negative trajectory]
```
---
## Market Structure Analysis
Classify your market structure before setting positioning strategy.
| Structure | Characteristics | Implications for Strategy |
|-----------|----------------|--------------------------|
| **Monopoly** | One dominant player (>60% share) | Differentiate or find underserved niche; don't compete head-on |
| **Duopoly** | Two dominant players sharing the market | Third-player positioning: both incumbents have weaknesses; find the gap |
| **Oligopoly** | 3-5 significant players | Specialization wins; compete in a narrow segment where you can be #1 |
| **Fragmented** | Many small players, no dominant leader | Consolidation opportunity OR signal of unmet need without viable solution |
| **Emerging** | Market being defined in real-time | Speed matters; first-mover into clearly defined niche wins |
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Analyzing features, ignoring momentum
**What it looks like**: Competitor B has fewer features than us today. No further analysis.
**Why wrong**: Momentum matters more than current state. A competitor shipping 10 features/quarter with high user satisfaction will surpass a feature-rich stagnant product within 12-18 months.
**Detection**: Check competitor's changelog/release notes. If releases are accelerating and quality is high, treat them as stronger than their current position suggests.
**Do instead**: Add a velocity column to your competitive tracker. Rate each competitor's development and distribution momentum separately from current capability. A competitor with low capability but high velocity is a higher-priority threat than a capable but stagnant one.
**Fix**: Add "velocity" to your competitive tracker. Rate each competitor's development and distribution momentum separately from their current capability.
### Mapping on irrelevant dimensions
**What it looks like**: Positioning map uses "enterprise vs. SMB" and "cloud vs. on-premise" when all relevant players are SMB-focused SaaS. Everyone clusters in one quadrant.
**Why wrong**: If dimensions don't separate the competitors, the map reveals nothing.
**Do instead**: Before drawing the map, verify that your chosen dimensions separate at least 50% of plotted players. If everyone clusters in one quadrant, discard the dimensions and try axes that actually discriminate.
**Fix**: Require that the chosen dimensions separate at least 50% of plotted players. If everyone clusters, try different dimensions.
### Static competitive landscape
**What it looks like**: Competitive analysis document dated 18 months ago, still being cited in strategy discussions.
**Why wrong**: Companies pivot, funding changes strategies, new entrants emerge, and products get discontinued. A stale map actively misleads.
**Detection**: Check the date on any competitive document. Over 12 months = must refresh core data before using. Over 18 months = discard and redo.
**Do instead**: Date-stamp every competitive document and set a calendar reminder to refresh it. Over 12 months old: refresh core data before citing. Over 18 months: redo the analysis from scratch.
---
## Detection Commands Reference
```bash
# Check competitor content velocity (RSS feed)
curl -s "https://{competitor-domain}/feed.xml" | grep -c "<item>"
# Higher count = more content output
# Check competitor's web presence (Wayback Machine for changes)
# Manual: https://web.archive.org/web/*/{competitor-domain} — compare snapshots
# Check competitor's GitHub activity (for OSS or dev tool companies)
curl -s "https://api.github.com/orgs/{competitor-org}/repos" | \
python3 -c "import json,sys; repos=json.load(sys.stdin); \
[print(r['name'], r['pushed_at']) for r in sorted(repos, key=lambda x: x['pushed_at'], reverse=True)[:5]]"
# Check Hacker News mentions
curl -s "https://hn.algolia.com/api/v1/search?query={competitor}&tags=story&hitsPerPage=5" | \
python3 -c "import json,sys; d=json.load(sys.stdin); [print(h['title'], h['created_at']) for h in d['hits']]"
```
---
## See Also
- `market-positioning.md` — positioning map templates, differentiation scoring, win/loss analysis
- `skills/strategic-decision/references/strategic-frameworks.md` — Porter's Five Forces for market structure
references/content-funnel.md
---
title: Content Funnel — Content Types Per Stage, Conversion Metrics, Content-to-Revenue Attribution
domain: growth-strategy
level: 3
skill: growth-strategy
---
# Content Funnel Reference
> **Scope**: Content type classification per funnel stage, conversion metrics and benchmarks, and content-to-revenue attribution models for solo creators and small content teams. Extends the funnel stage mapping in channel-evaluation.md with specific content formats, measurable triggers, and attribution methodology. Does NOT cover paid acquisition funnels — this reference covers organic content only.
> **Version range**: Platform-agnostic — attribution models and content types apply across blogging, newsletters, YouTube, and podcasting.
> **Generated**: 2026-04-09 — validate conversion benchmarks against your own cohort data after 90 days of operation.
---
## Overview
A content funnel without conversion mechanics is a media operation, not a growth strategy. Publishing content that builds traffic but has no clear path from reader to subscriber, or subscriber to customer, produces audience that does not compound into revenue. The most common gap: creators invest heavily in Awareness content (SEO, viral posts) and have almost no Consideration content (deep dives, comparison guides) — so readers who are ready to buy or subscribe have nowhere to go. This reference maps each funnel stage to specific content types, success metrics, and conversion actions.
---
## Content Funnel Stage Reference
### Stage 1: Awareness
**Goal**: Reach people who have the problem but don't know you exist.
| Content Type | Format | Length | Primary Distribution | Trigger for Stage 2 |
|-------------|--------|--------|---------------------|---------------------|
| SEO problem post | Long-form article | 1,500-3,000 words | Search engines | Internal link to deeper content |
| Social hook thread | Twitter/LinkedIn thread | 5-12 posts | Algorithmic feed | Profile bio CTA or link in bio |
| Guest post | Article on host publication | 800-1,500 words | Host's audience | Byline with link to your subscribe page |
| Short-form video | YouTube Shorts, Reels | 30-90 seconds | Algorithmic feed | Profile bio or pinned comment |
| Community answer | Reddit, Slack, Discord reply | 200-500 words | Community traffic | Link to longer resource |
| Podcast appearance | Guest interview | 30-60 minutes | Host's subscribers | Show notes link to subscribe page |
**Success Metrics for Awareness**:
| Metric | What It Measures | Target Benchmark |
|--------|-----------------|-----------------|
| Unique visitors from organic | New audience reached | Growing month-over-month |
| Impressions (social) | Reach on platform | Platform-specific; track trend not absolute |
| Click-through rate (social → site) | Content relevance | Twitter: 1-3%, LinkedIn: 0.5-1.5% |
| Scroll depth (articles) | Content quality signal | >50% average = content is relevant |
| New visitor % | How much is discovery vs. return | >60% for awareness-stage content |
**Conversion action**: Subscribe to email list (primary). Follow on social (secondary — lower value, harder to monetize).
---
### Stage 2: Interest
**Goal**: Convert first-time visitors into engaged readers who return.
| Content Type | Format | Length | Primary Distribution | Trigger for Stage 3 |
|-------------|--------|--------|---------------------|---------------------|
| Tutorial / how-to | Article or video | 2,000-5,000 words | SEO + email | "Related series" CTA |
| Case study | Detailed narrative | 1,500-3,000 words | Email + SEO | Outcome + "here's the framework" CTA |
| Resource roundup | Curated list with commentary | 1,000-2,000 words | Email + SEO | Opinionated commentary builds trust |
| Newsletter issue | Email | 500-1,500 words | Email subscribers | Consistent publication builds habit |
| Deep-dive explainer | Article or video | 3,000-6,000 words | SEO (long-tail) | Subscribe CTA at midpoint and end |
| Template / starter kit | Download | N/A | Gated or freemium | Free value → email capture |
**Success Metrics for Interest**:
| Metric | What It Measures | Target Benchmark |
|--------|-----------------|-----------------|
| Return visitor % | Content value creating habit | >20% of monthly visitors |
| Email open rate | Subscriber relationship quality | 35-50% (solo creators), 25-35% (larger lists) |
| Time on page (deep dives) | Content depth engagement | >4 minutes for 3,000+ word posts |
| Pages per session | Content exploration | >1.8 pages per session for returning visitors |
| Email click-through rate | Content relevance to subscribers | 3-8% per issue |
**Conversion action**: Click through to Consideration content (product comparison, pricing page, booking link). Upgrade from free to paid tier.
---
### Stage 3: Consideration
**Goal**: Move engaged readers toward a purchase, trial, or high-intent action.
| Content Type | Format | Length | Primary Distribution | Trigger for Conversion |
|-------------|--------|--------|---------------------|------------------------|
| Comparison guide | "X vs Y" article | 2,000-4,000 words | SEO (high-intent queries) | "Try our approach" CTA |
| Process breakdown | Behind-the-scenes methodology | 1,500-2,500 words | Email to subscribers | "Work with me / buy" CTA |
| FAQ / objection handler | Article or email | 800-1,500 words | Email (mid-series) | Direct answer + conversion CTA |
| Webinar / office hours | Live or recorded | 30-60 minutes | Email to subscribers | Enrollment or pitch at end |
| Email nurture sequence | Series of 4-8 emails | 300-600 words each | Automated after subscribe | Final email = conversion offer |
| Testimonial / social proof | Stories from real users | 500-1,000 words | Email + landing page | Lower objection → buy link |
**Success Metrics for Consideration**:
| Metric | What It Measures | Target Benchmark |
|--------|-----------------|-----------------|
| Webinar attendance rate | Email list engagement quality | 20-40% of registrants |
| Email sequence completion rate | Nurture series effectiveness | >50% complete all emails |
| "X vs Y" post conversion rate | High-intent traffic quality | 2-8% → trial/purchase action |
| Sales page visit rate (from email) | Content building buying intent | Track which emails drive page visits |
| Trial start rate | Consideration → Conversion handoff | 1-5% of subscribers who reach this stage |
**Conversion action**: Purchase, trial enrollment, or booking a call.
---
### Stage 4: Conversion
**Goal**: Close the purchase or commitment.
| Content Type | Format | Length | Primary Distribution | Conversion Action |
|-------------|--------|--------|---------------------|------------------|
| Sales / landing page | Web page | 800-2,000 words | All upstream CTAs | Purchase |
| Conversion email | Direct offer | 400-800 words | Segmented subscribers | Purchase link click |
| Free trial onboarding | In-product + email | Short (in-app) | New trial users | First meaningful action |
| Limited-time offer | Email + landing page | 300-600 words | Full list | Purchase before deadline |
**Success Metrics for Conversion**:
| Metric | What It Measures | Target Benchmark |
|--------|-----------------|-----------------|
| Sales page conversion rate | Offer + audience fit | 1-5% (cold), 5-15% (warm email list) |
| Conversion email conversion rate | Offer relevance to subscribers | 0.5-3% of list per launch |
| Trial-to-paid conversion rate | Product quality + onboarding | 15-40% (varies widely by product type) |
| Revenue per subscriber (RPS) | Overall funnel efficiency | Track and improve monthly |
---
### Stage 5: Retention
**Goal**: Extend customer lifetime and generate referrals.
| Content Type | Format | Purpose | Distribution |
|-------------|--------|---------|--------------|
| Onboarding series | Email sequence | Activate new customers | Automated post-purchase |
| Advanced use case content | Articles or video | Expand product usage | Email to customers |
| Changelog / release notes | Email or in-app | Demonstrate ongoing value | All customers |
| Community access | Slack, Discord, forum | Peer connection + retention | Customers only |
| Annual review / retrospective | Email or article | Reinforce value delivered | All customers |
**Success Metrics for Retention**:
| Metric | What It Measures | Target Benchmark |
|--------|-----------------|-----------------|
| Monthly churn rate | Product-market fit + onboarding quality | <3% monthly (SaaS); <5% (newsletter paid) |
| Net Revenue Retention (NRR) | Expansion vs. churn | >100% = growing without new customers |
| Referral rate | Word-of-mouth growth | 10-20% of customers refer at least one person |
| Product usage depth | Onboarding effectiveness | Increase from month 1 to month 3 |
---
## Content-to-Revenue Attribution Models
### Model 1: Last-Touch Attribution (Default, Simple)
Every conversion is credited to the last piece of content the customer interacted with before purchasing.
```
## Last-Touch Attribution Log
For each paying customer, record:
- Purchase date: ___
- Last content touchpoint before purchase: [post title / email subject / URL]
- Content type: [Awareness / Interest / Consideration]
- Channel: [organic search / email / social]
Aggregate monthly:
Content Item | Touches Before Purchase | Conversions Attributed | Revenue
___ | ___ | ___ | $___
```
**Limitation**: Overcredits Consideration content (closest to purchase) and undercredits Awareness content that brought the customer in. Use as a starting point, not a final answer.
### Model 2: First-Touch Attribution
Every conversion credited to the first content interaction.
```
## First-Touch Attribution Tracking
Requires: UTM parameters on all content links, or tracking pixel from day one.
For each subscriber:
- First content touchpoint: [URL, source, medium]
- Acquisition date: ___
- Eventual conversion: [Yes / No]
- Revenue generated: $___
Monthly rollup:
Source Content | Subscribers Acquired | Eventual Conversions | Conversion Rate | Revenue Attributed
___ | ___ | ___ | ___% | $___
```
**Use case**: Tells you which Awareness content produces the subscribers most likely to eventually convert. Drives SEO investment decisions.
### Model 3: Cohort Attribution (Most Actionable for Solo Creators)
Group subscribers by the month they joined. Track their revenue contribution over 12 months.
```
## Cohort Revenue Attribution
| Cohort | Month 0 Subscribers | Month 3 Revenue | Month 6 Revenue | Month 12 Revenue | LTV |
|--------|---------------------|-----------------|-----------------|------------------|-----|
| Jan 2025 | 120 | $240 | $480 | $840 | $7.00/subscriber |
| Feb 2025 | 95 | $190 | $380 | $665 | $7.00/subscriber |
| Mar 2025 | 150 | $450 | $900 | $1,575 | $10.50/subscriber |
March cohort LTV is 50% higher. Investigate: what content drove March sign-ups?
```
**Insight extraction**: If one cohort's LTV is consistently 50%+ higher than adjacent months, the content that drove acquisition that month is producing higher-quality subscribers. Scale it.
### Model 4: Content ROI Model
Measure return on individual content investments.
```
## Content ROI Calculation: [Post/Video/Series Title]
### Investment
- Creation time: ___ hours × $___/hour = $___
- Distribution cost (social ads, tools): $___
- Total investment: $___
### Return (over 12 months)
- Subscribers acquired (attributed, first-touch): ___
- Subscribers' average LTV: $___
- Revenue attributed: ___ × $___ = $___
### ROI
Content ROI = (Revenue attributed - Investment) / Investment × 100
= ($___ - $___) / $___ × 100
= ___%
Payback period: ___ months
```
**Benchmark**: A content ROI of 200-500% over 12 months is healthy for organic content. Under 100% means the content is not earning its creation cost. Over 1,000% indicates content that should be promoted actively.
---
## Funnel Audit Template
Run this quarterly to find funnel leaks.
```
## Funnel Audit: [Quarter]
### Content Inventory by Stage
Classify every piece of content published this quarter:
| Stage | Content Count | % of Total |
|-------|-------------|------------|
| Awareness | ___ | ___% |
| Interest | ___ | ___% |
| Consideration | ___ | ___% |
| Conversion | ___ | ___% |
| Retention | ___ | ___% |
| Total | ___ | 100% |
### Common Imbalance Findings
- Awareness > 70%: Funnel leaks at Interest. No depth content for engaged readers.
- Interest = 0%: Newsletter or returning reader content missing.
- Consideration = 0%: No content for readers who are ready to buy.
- Conversion content exists but no CTAs in Awareness content: Funnel is siloed.
### Conversion Rate Audit by Stage
Awareness → Interest (subscribe rate): ___% (benchmark: 1-3% for blog)
Interest → Consideration (click through to offer): ___% (benchmark: 5-15% of subscribers)
Consideration → Conversion (trial/purchase): ___% (benchmark: 1-5%)
Overall funnel rate: ___% (benchmark: 0.1-0.5% of total visitors)
### Action Items from Audit
- Missing stage to fill: ___
- Lowest-converting transition to improve: ___
- Highest-performing content to amplify: ___
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Awareness content without subscribe CTA
**What it looks like**: SEO articles driving 5,000 unique visitors/month with no email subscribe prompt, no content upgrade, no follow CTA.
**Why wrong**: Traffic without capture is rented attention. If the newsletter/product is not capturing subscribers from Awareness content, growing traffic does nothing for the business.
**Do instead**: Place a subscribe CTA at the midpoint and end of every Awareness piece. Add a content upgrade (a downloadable checklist tied to the post topic) to convert at 3-5x the rate of a generic subscribe prompt.
**Fix**: Every Awareness piece needs one primary CTA (subscribe to email) in at least 2 locations: midpoint and end of content. Content upgrades (downloadable checklist related to the post) convert at 3-5× higher than generic subscribe prompts.
### Consideration content gated behind subscribe when audience is unaware
**What it looks like**: Comparison guide ("our product vs. Competitor X") is email-gated. But audience doesn't know the product exists yet.
**Why wrong**: High-intent comparison content is searched by people not yet in your funnel. Gating it prevents discovery by the exact audience most likely to convert.
**Do instead**: Publish high-intent comparison and alternative content freely so search can drive discovery. Reserve gating for lower-intent assets (templates, checklists) where the audience is already aware and needs a nudge.
**Fix**: Consideration content targeting high-intent search queries (X vs. Y, alternative to Z) should be freely accessible and optimized for search. Gate lower-intent content (templates, checklists) instead.
### Single attribution model driving all decisions
**What it looks like**: Team uses last-touch attribution exclusively. SEO posts driving first-touch subscribers are consistently "credited" with zero revenue, so SEO investment is cut.
**Why wrong**: Last-touch attribution systematically undervalues Awareness content. A reader who finds you via a blog post, subscribes, reads 6 newsletters, and then buys after a conversion email credits the conversion email — but the blog post started the relationship.
**Do instead**: Run first-touch and last-touch attribution models in parallel. Use first-touch to evaluate acquisition content. Use last-touch to evaluate conversion content. Use cohort attribution for the long-run view of what actually built the relationship.
**Fix**: Run both first-touch and last-touch attribution models. Use first-touch to evaluate acquisition content investment and last-touch to evaluate conversion content. Cohort attribution is the most honest long-run view.
### Revenue attributed to content that did not exist at time of purchase
**What it looks like**: Post published this month credited with converting subscribers who joined 14 months ago.
**Why wrong**: The post could not have influenced subscribers who joined before it existed.
**Do instead**: Anchor first-touch attribution to the subscriber's original acquisition touchpoint. In cohort attribution, group by the subscriber's join date, not their purchase date, so the analysis reflects what content was actually available to influence them.
**Fix**: First-touch attribution must use the subscriber's original acquisition touchpoint. Cohort attribution must use the subscriber's cohort join date, not their purchase date, as the grouping variable.
---
## See Also
- `channel-evaluation.md` — CAC/LTV models and basic funnel stage mapping by channel
- `audience-segmentation.md` — ICP scoring to determine which audience segments to build the funnel for
- `skills/growth-strategy/SKILL.md` — full growth strategy workflow including 90-day planning templates
references/csuite.md
# C-Suite Decision Support
Umbrella skill for all executive decision-making: CEO-level strategy, CTO-level technology choices, CMO-level growth planning, competitive intelligence, and project evaluation. Each domain loads its own reference files on demand -- this skill detects the mode, loads the right references, and executes the appropriate framework.
**Scope**: Business decisions with meaningful consequences. Use decision-helper for technical architecture micro-choices, domain agents for code, voice-writer for content, and systematic-debugging for debugging.
---
## Mode Detection
Classify the user's request into exactly one mode before proceeding. If the request spans multiple modes, choose the primary one and note the secondary.
| Mode | Signal Phrases | Role Lens |
|------|---------------|-----------|
| **STRATEGY** | Market entry, partnerships, resource allocation, opportunity, "should I/we", strategic pivots, investment | CEO |
| **TECHNOLOGY** | Build vs buy, vendor, SaaS, tech stack, architecture, adopt, technology choice | CTO |
| **GROWTH** | Content strategy, audience, SEO, marketing, brand, community, positioning, channel | CMO |
| **COMPETITIVE** | Competitor, competition, market landscape, differentiation, positioning against, market share | Cross-role |
| **EVALUATION** | Feasibility, effort estimate, ROI, priority, go/no-go, viability, "is it worth it" | Cross-role |
---
## Reference Loading Table
Load references based on the detected mode. Load only the references required by the mode.
| Signal | Mode | Reference |
|--------|------|-----------|
| Market entry, partnerships, resource allocation, opportunity | STRATEGY | `references/strategic-frameworks.md`, `references/decision-matrices.md` |
| Build vs buy, vendor, SaaS, tech stack, architecture | TECHNOLOGY | `references/tco-framework.md`, `references/vendor-evaluation.md` |
| Content, audience, SEO, marketing, brand, community | GROWTH | `references/audience-segmentation.md`, `references/channel-evaluation.md` |
| Competitor, market landscape, positioning, differentiation | COMPETITIVE | `references/competitive-mapping.md`, `references/market-positioning.md` |
| Feasibility, effort, ROI, priority, go/no-go | EVALUATION | `references/feasibility-scoring.md`, `references/roi-frameworks.md` |
---
## Instructions
### Mode: STRATEGY (CEO)
**Framework**: FRAME -> ANALYZE -> DECIDE
**Phase 1: FRAME** -- Convert the user's question into a structured decision with clear stakes and timeline.
- Name the actual decision (users present symptoms; the real decision is broader)
- Identify irreversibility -- reversible decisions deserve less analysis
- Set a time horizon -- 3-month and 3-year decisions need different frameworks
- Classify the decision type: Expansion, Partnership, Allocation, Pivot, or Timing
- Get the user to state: options (2-4), default path risk, deadline, and what makes it hard
**Gate**: Decision framed as one sentence. Options listed (2-4). Type classified.
**Phase 2: ANALYZE** -- Evaluate each option through multiple lenses with evidence.
For each option, assess: Upside (best realistic + expected outcome), Downside (worst realistic + recovery path + irreversible losses), Requirements (resources, assumptions, dependencies), Opportunity Cost (what you cannot do).
Separate facts from assumptions. Quantify where possible. Load reference files for scoring matrices and strategic frameworks.
**Gate**: All options analyzed. Facts and assumptions labeled. Opportunity costs explicit.
**Phase 3: DECIDE** -- Synthesize into a clear recommendation.
- Apply the reversibility test: one-way doors need high confidence; two-way doors can act faster with a checkpoint
- Produce: Recommendation (one sentence), Confidence (High/Medium/Low), Why this option (2-3 reasons), What must be true (invalidating assumptions), First move (48-hour action), Revisit trigger
- State explicitly what would change the recommendation
**Gate**: Recommendation stated. First action identified. Revisit trigger set.
---
### Mode: TECHNOLOGY (CTO)
**Framework**: SCOPE -> EVALUATE -> RECOMMEND
**Phase 1: SCOPE** -- Define the capability needed, stripped of solution bias.
- Start with the need, not the product ("we need reliable async delivery" not "we need Kafka")
- Quantify hard requirements (latency, throughput, compliance)
- Identify the real driver (build vs buy is sometimes "convince management" or "hire someone")
- List actual options: build from scratch, build on OSS, buy SaaS, buy + customize, do nothing
**Gate**: Capability defined without solution bias. Options enumerated. Hard requirements quantified.
**Phase 2: EVALUATE** -- Score options on dimensions that matter for technology decisions.
- Total cost of ownership at Year 3, not sticker price (the "free" OSS needing a full-time engineer is expensive)
- Score on: Fit (5), TCO (4), Operational burden (4), Team capability (3), Lock-in risk (3), Time to value (3), Flexibility (2)
- Apply the build-vs-buy heuristic: core competency, requirements stability, team capacity, timeline, scale, compliance
Load `references/tco-framework.md` for TCO templates and `references/vendor-evaluation.md` for vendor scorecards.
**Gate**: TCO estimated. Dimensions scored. Build-vs-buy heuristic applied.
**Phase 3: RECOMMEND** -- Deliver a clear recommendation with reasoning.
- Present the weighted scoring matrix
- State: Decision, Confidence, Why this option, Watch-for risks, Migration path, First step
- Define exit criteria: when to reconsider for each option type
**Gate**: Recommendation stated. Exit criteria defined. First step identified.
---
### Mode: GROWTH (CMO)
**Framework**: ASSESS -> STRATEGIZE -> PLAN
**Phase 1: ASSESS** -- Understand current state before recommending.
- Audit: publications, content volume, existing audience, active channels, performance data
- Identify the binding constraint: Discovery, Content, Conversion, Retention, or Capacity
- Creator capacity is the binding constraint -- recommend what one person can sustain
**Gate**: Current state audited. Binding constraint identified.
**Phase 2: STRATEGIZE** -- Design an approach matching capacity and constraint.
- Solve the constraint, not everything -- address one binding constraint well
- Prefer compound strategies (SEO, evergreen, community) over one-shot campaigns
- Recommend maximum 3 active channels with format, cadence, success metric, and effort estimate
Load `references/audience-segmentation.md` for ICP scoring and `references/channel-evaluation.md` for channel matrices.
**Gate**: Strategy selected. Maximum 3 channels. Effort estimated against capacity.
**Phase 3: PLAN** -- Convert strategy into a 90-day executable plan.
- Define one primary metric and 2-3 secondary metrics
- Break into 30-day phases: Foundation (days 1-30), Execution (31-60), Evaluate (61-90)
- Set explicit abandon criteria, pivot triggers, and double-down conditions
**Gate**: 90-day plan with checkpoints. Primary metric defined. Abandon criteria explicit.
---
### Mode: COMPETITIVE
**Framework**: MAP -> ANALYZE -> POSITION
**Phase 1: MAP** -- Build a structured picture of the competitive landscape.
- Define the competitive arena: what you compete on, who you serve, where you compete
- Tier competitors: Direct (full analysis), Adjacent (positioning only), Aspirational (strategy extraction), Emerging (watch list)
- Map the landscape before zooming in -- analyzing one competitor in isolation misses gaps
**Gate**: Arena defined. Competitors identified and tiered. At least 2 direct competitors mapped.
**Phase 2: ANALYZE** -- Extract actionable intelligence from behavior, not surface impressions.
- Focus on what they DO, not what they SAY (pricing, launches, cadence reveal strategy)
- Analyze for gaps, not imitation -- find what competitors miss or do poorly
- For each direct competitor: product/content analysis, audience analysis, strategy signals
Load `references/competitive-mapping.md` for landscape templates and `references/market-positioning.md` for positioning frameworks.
**Gate**: Direct competitors analyzed. Gaps and weaknesses identified.
**Phase 3: POSITION** -- Convert intelligence into defensible differentiation.
- Build a positioning map on two dimensions where you can differentiate
- Define: positioning statement, defensible advantages, vulnerable advantages, strategic gaps to exploit
- Set monitoring cadence: monthly (direct competitors), quarterly (full landscape), trigger-based (major moves)
**Gate**: Positioning map built. Differentiation strategy defined. Monitoring cadence set.
---
### Mode: EVALUATION
**Framework**: SCOPE -> EVALUATE -> VERDICT
**Phase 1: SCOPE** -- Define the project and what success looks like.
- Define done before estimating effort ("build an app" is not a project)
- Separate the vision from the MVP -- evaluate the minimum viable version
- Name the binding constraint (time, money, skills, attention)
- Define success criteria, the problem it solves, who benefits, and why now
**Gate**: Project defined with measurable success criteria. MVP scope identified. Binding constraint named.
**Phase 2: EVALUATE** -- Assess feasibility, estimate effort, calculate ROI.
- Feasibility across three dimensions: Technical, Resource, Market (each High/Medium/Low confidence)
- Effort in ranges, not points ("2-5 weeks, most likely 3")
- Include hidden costs: learning curve, integration, testing, documentation (add 20-40%)
- ROI: direct value, indirect value, strategic value vs. build cost, ongoing cost, opportunity cost
Load `references/feasibility-scoring.md` for the three-dimension model and `references/roi-frameworks.md` for estimation templates.
**Gate**: Feasibility assessed. Effort estimated in ranges. ROI calculated with confidence level.
**Phase 3: VERDICT** -- Deliver a clear go/no-go recommendation.
- Verdict: GO, GO WITH CONDITIONS, DEFER, or NO-GO
- Include: summary, key factors, conditions (if conditional), what would change the verdict, recommended next step
- For multiple projects: rank using RICE scoring (Reach * Impact * Confidence / Effort)
**Gate**: Verdict stated with confidence. Conditions specified. Next step identified.
---
## Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| Too many options | 5+ options creating paralysis | Eliminate obviously inferior options first. Get to 2-4 before running full framework. |
| Not enough information | User cannot answer framing questions | Identify 2-3 critical unknowns. Recommend time-boxed research sprint before deciding. |
| Analysis paralysis | Keeps adding criteria or second-guessing | Apply reversibility test. If reversible, recommend best current option with checkpoint. |
| Emotional attachment | User has already decided, wants validation | Name the pattern directly. Ask: stress-test the choice, or genuinely evaluate all options? |
| Comparing apples to oranges | Options at different abstraction levels | Normalize to the capability level. Compare what each option gives for the specific need. |
| Vendor lock-in fear | Over-weights lock-in, under-weights time-to-value | Quantify actual switching cost. Compare concrete switching cost against concrete speed benefit. |
| Build bias (NIH) | Team wants to build because it is more interesting | Apply core competency test: "If this disappeared, would customers notice?" |
| Vanity metrics | Optimizes followers/likes instead of outcomes | Redirect to "one metric that matters" -- what action should the audience take? |
| Scope creep during evaluation | Keeps adding features to project definition | Freeze scope at end of Phase 1. Additional features evaluate as v2. |
| Optimism bias | Effort estimates too low | Apply reference class test. If no similar project, add 50% to pessimistic estimate. |
---
## References
| Reference | When to Load | Content |
|-----------|-------------|---------|
| `references/strategic-frameworks.md` | STRATEGY mode: market entry, competitive dynamics, SWOT, OKR alignment | Porter's Five Forces, SWOT scoring, OKR alignment matrices |
| `references/decision-matrices.md` | STRATEGY mode: structured scoring, comparison, pre-mortem | Weighted decision matrices, ICE/RICE scoring, pre-mortem templates |
| `references/tco-framework.md` | TECHNOLOGY mode: TCO modeling, cost projections, build vs buy scorecard | TCO templates, hidden cost checklists, migration cost models |
| `references/vendor-evaluation.md` | TECHNOLOGY mode: vendor comparison, RFP criteria, integration complexity | Vendor scorecards, RFP criteria, red flag detection, contract checklist |
| `references/audience-segmentation.md` | GROWTH mode: audience analysis, ICP definition, persona development | ICP scoring matrix, persona templates, segmentation frameworks |
| `references/channel-evaluation.md` | GROWTH mode: channel selection, CAC/LTV modeling, content funnel | Channel scoring matrices, CAC/LTV models, funnel stage mapping |
| `references/competitive-mapping.md` | COMPETITIVE mode: landscape mapping, feature comparison, competitor profiling | Landscape map templates, feature matrices, activity tracker |
| `references/market-positioning.md` | COMPETITIVE mode: positioning strategy, differentiation scoring | Positioning maps, differentiation scoring, win/loss frameworks |
| `references/feasibility-scoring.md` | EVALUATION mode: feasibility assessment, risk evaluation, go/no-go | Three-dimension feasibility model, confidence calibration, decision tree |
| `references/roi-frameworks.md` | EVALUATION mode: effort estimation, ROI calculation, project comparison | T-shirt sizing, three-point estimation, risk-adjusted NPV |
references/customer-support.md
# Customer Support
Umbrella skill for customer-facing support workflows: triage incoming tickets, draft calibrated responses, convert resolutions into KB articles, package escalations, and research customer context. Each mode loads its own reference files on demand.
**Scope**: Customer-facing support work. Use professional-communication for internal business formats, csuite for strategic decisions, pr-workflow for code PRs.
---
## Mode Detection
Classify the request into exactly one mode. If the request spans modes, choose the primary and note the secondary.
| Mode | Signal Phrases | Core Output |
|------|---------------|-------------|
| **TRIAGE** | New ticket, categorize, prioritize, route, P1-P4, severity, SLA | Structured triage assessment with priority, routing, initial response |
| **RESPOND** | Draft response, reply to customer, follow up, de-escalate, bad news, decline | Customer-facing message with tone calibration and internal notes |
| **KB** | Knowledge base, document this, FAQ, write article, how-to guide, troubleshooting doc | Publish-ready KB article with metadata and search optimization |
| **ESCALATE** | Escalate, engineering attention, SLA breach, churn risk, multiple customers, leadership | Structured escalation brief with impact assessment and repro steps |
| **RESEARCH** | Look up, investigate, what did we tell them, has this been reported, check history | Research brief with source attribution and confidence scoring |
---
## Reference Loading Table
Load only the references required by the detected mode.
| Mode | Reference |
|------|-----------|
| TRIAGE | `references/customer-support/triage-methodology.md` |
| RESPOND | `references/customer-support/response-drafting.md` |
| KB | `references/customer-support/knowledge-base.md` |
| ESCALATE | `references/customer-support/triage-methodology.md`, `references/customer-support/response-drafting.md` |
| RESEARCH | `references/customer-support/knowledge-base.md` |
| Any mode | `references/customer-support/llm-support-failure-modes.md` (always load -- LLM failure awareness is non-negotiable in support) |
---
## Mode: TRIAGE
**Framework**: PARSE -> CLASSIFY -> ROUTE -> RESPOND
**Phase 1: PARSE** -- Extract the actual problem from the ticket.
- Core problem vs. stated symptom (customers describe symptoms, not root causes)
- Urgency signals: production down, data loss, blocked, multiple users, time-sensitive
- Emotional state: frustrated, confused, matter-of-fact, escalating
- Customer context: account tier, history, previous tickets if available
**Phase 2: CLASSIFY** -- Assign category and priority.
Apply the category taxonomy from `references/customer-support/triage-methodology.md`:
| Category | When |
|----------|------|
| Bug | "It used to work and now it doesn't" |
| How-to | "How do I make it work?" |
| Feature request | "I want it to work differently" |
| Billing | Payment, subscription, invoice, refund |
| Account | Login, permissions, SSO, access |
| Integration | API, webhook, third-party, sync |
| Security | Data exposure, unauthorized access, compliance |
| Performance | Slow, timeout, degraded, unavailable |
Assign priority P1-P4. When in doubt, err higher -- easier to de-escalate than recover from a missed SLA.
| Priority | Criteria | SLA Response |
|----------|----------|-------------|
| P1 Critical | Production down, data loss, security breach, all users | 1 hour |
| P2 High | Major feature broken, no workaround, many users | 4 hours |
| P3 Medium | Partial break, workaround exists, small impact | 1 business day |
| P4 Low | Cosmetic, feature request, general question | 2 business days |
**Phase 3: ROUTE** -- Determine the right team.
| Route to | When |
|----------|------|
| Tier 1 | How-to, known issues with docs, billing inquiries, password resets |
| Tier 2 | Bugs needing investigation, complex config, integration troubleshooting |
| Engineering | Confirmed bugs needing code fixes, infrastructure, performance degradation |
| Product | Feature requests with demand, design decisions, workflow gaps |
| Security | Data access concerns, vulnerability reports, compliance (bypasses tier progression) |
**Phase 4: RESPOND** -- Draft initial response using the category templates from `references/customer-support/triage-methodology.md`.
**Gate**: Triage output includes: category, priority with justification, routing recommendation, suggested initial response, internal notes.
---
## Mode: RESPOND
**Framework**: CONTEXT -> CALIBRATE -> DRAFT -> VERIFY
**Phase 1: CONTEXT** -- Understand the full situation before writing a word.
- Who: customer name, account tier, relationship stage (new/established/frustrated)
- What: situation type (question, issue, escalation, bad news, good news, decline)
- Channel: email, ticket, chat (adjusts length and formality)
- Stakeholder level: end user, manager, executive, technical
- History: previous communications, commitments made, tone of thread
**Phase 2: CALIBRATE** -- Select tone from `references/customer-support/response-drafting.md`.
| Situation | Tone | Key Characteristic |
|-----------|------|--------------------|
| Good news | Celebratory | Forward-looking, enthusiastic |
| Routine update | Professional | Clear, concise, friendly |
| Technical response | Precise | Accurate, patient, structured |
| Delayed delivery | Accountable | Honest, action-oriented |
| Bad news / won't-fix | Candid | Direct, empathetic, alternative-offering |
| Issue / outage | Urgent | Transparent, actionable, reassuring |
| Escalation | Executive | Composed, ownership-taking, plan-presenting |
| Billing | Precise | Factual, resolution-focused |
Adjust by relationship stage:
- **New customer**: More formal, extra context, proactive help
- **Established**: Warm, direct, reference shared history
- **Frustrated**: Extra empathy first, concrete plan, shorter feedback loops
**Phase 3: DRAFT** -- Write the response following the structure:
1. **Acknowledgment** (1-2 sentences) -- show you understand their situation
2. **Core message** (1-3 paragraphs) -- the actual answer, update, or information
3. **Next steps** (1-3 bullets) -- what you will do, what they need to do, when they hear from you
4. **Close** (1 sentence) -- warm, professional, available
**Phase 4: VERIFY** -- Run the quality checks before presenting.
- [ ] Tone matches situation and relationship stage
- [ ] No unauthorized commitments (timelines, features, exceptions)
- [ ] No product roadmap details that shouldn't be external
- [ ] Clear ownership and next steps
- [ ] Appropriate length for channel
- [ ] No corporate jargon, no passive voice to dodge accountability
**Gate**: Draft includes internal notes covering: rationale for tone, facts to verify before sending, risk factors, follow-up actions needed.
---
## Mode: KB
**Framework**: SOURCE -> STRUCTURE -> DRAFT -> OPTIMIZE
**Phase 1: SOURCE** -- Understand what you're documenting.
- Original problem, question, or error
- Resolution, workaround, or answer
- Who this affects (user type, plan level, configuration)
- Frequency: one-off or recurring
- Article type: how-to, troubleshooting, FAQ, known issue, reference
**Phase 2: STRUCTURE** -- Choose the right template from `references/customer-support/knowledge-base.md`.
| Type | Purpose | Structure |
|------|---------|-----------|
| How-to | Step-by-step task completion | Prerequisites -> Steps -> Verify -> Common Issues |
| Troubleshooting | Diagnose and fix a problem | Symptoms -> Cause -> Solution(s) -> Prevention |
| FAQ | Quick answer to common question | Direct Answer -> Details -> Related Questions |
| Known issue | Document a bug with workaround | Status -> Symptoms -> Workaround -> Fix Timeline |
**Phase 3: DRAFT** -- Write the article following formatting standards.
- Title: specific, searchable, uses customer language ("How to configure SSO with Okta" not "SSO Setup")
- Opening sentence: restate the problem in plain language
- Headers for scannability. Numbered lists for sequences. Bullet lists for non-sequential items.
- Code blocks for commands, error messages, configuration values
- Short paragraphs (2-4 sentences max). One idea per section.
**Phase 4: OPTIMIZE** -- Search optimization and metadata.
- Include exact error messages (customers copy-paste into search)
- Use customer language, not internal jargon
- Add common synonyms (delete/remove, dashboard/home page, export/download)
- Tag with product areas matching customer mental models
- Set review date and identify SME for technical verification
**Gate**: Article includes metadata (title, type, category, tags, audience), full content, publishing notes (source, related articles, review needed, suggested review date).
---
## Mode: ESCALATE
**Framework**: CONFIRM -> GATHER -> ASSESS -> PACKAGE
**Phase 1: CONFIRM** -- Verify this warrants escalation.
Escalate when:
- Bug confirmed, needs code fix
- Multiple customers affected (3+ = pattern)
- Production down or data at risk
- SLA breach imminent or occurred
- Customer threatening churn
- Requires access or authority beyond support tier
Handle in support when:
- Documented solution or workaround exists
- Configuration or setup issue you can resolve
- Known limitation with documented alternative
**Phase 2: GATHER** -- Pull together all context.
- Timeline: when it started, how long the customer has waited
- What's been tried: troubleshooting steps and results
- Reproduction steps (for bugs): from clean state, specific values, environment details, frequency, evidence
- Related tickets and pattern detection
**Phase 3: ASSESS** -- Quantify business impact.
| Dimension | Assess |
|-----------|--------|
| Breadth | How many customers/users? Growing? |
| Depth | Blocked vs. inconvenienced? |
| Duration | How long? Getting worse? |
| Revenue | ARR at risk? Deals affected? |
| Contractual | SLA breach? Contractual obligations? |
Severity shorthand:
- **Critical**: Production down, data at risk, security breach. Immediate attention.
- **High**: Major function broken, key customer blocked, SLA at risk. Same day.
- **Medium**: Significant issue with workaround, important but not urgent. This week.
**Phase 4: PACKAGE** -- Structure the escalation brief.
Determine target:
- **Engineering**: confirmed bugs, infrastructure, code changes needed
- **Product**: feature gaps, design decisions, competing priorities
- **Security**: data exposure, vulnerability, compliance (bypasses normal tier progression)
- **Leadership**: high-revenue churn risk, SLA breach on critical account, policy exception needed
Include: severity, target team, impact summary, issue description, what's been tried, reproduction steps, customer communication status, specific ask with deadline, supporting context.
**Gate**: Escalation brief complete. Follow-up cadence set (Critical: every 2h internal / 2-4h customer. High: every 4h / 4-8h. Medium: daily / 1-2 business days).
---
## Mode: RESEARCH
**Framework**: SCOPE -> SEARCH -> SYNTHESIZE -> CAPTURE
**Phase 1: SCOPE** -- Define what you're looking for.
Research types:
- **Customer question**: needs a factual answer
- **Issue investigation**: has this been reported, what's the workaround
- **Account context**: what was previously communicated
- **Topic research**: best practices, general domain knowledge
Clarify: factual vs. contextual, audience (internal vs. customer), scope boundaries.
**Phase 2: SEARCH** -- Work through source tiers systematically.
| Tier | Source Type | Confidence |
|------|------------|------------|
| 1 | Official docs, KB, policies | High |
| 2 | CRM, support tickets, meeting notes | Medium-High |
| 3 | Chat, email, calendar | Medium |
| 4 | Web, forums, third-party docs | Low-Medium |
| 5 | Inference, analogies, best practices | Low -- flag explicitly |
Cross-reference across multiple sources.
**Phase 3: SYNTHESIZE** -- Compile findings with attribution.
- Lead with the bottom-line answer
- Assign confidence level (High / Medium / Low / Unable to Determine)
- Attribute every finding to its source
- Note contradictions explicitly -- present both sides, recommend the conservative answer for customer-facing use
- Identify gaps and unknowns
**Phase 4: CAPTURE** -- Suggest knowledge preservation.
If the research took significant effort, was a common question, or corrected a misunderstanding, offer to create a KB article or FAQ entry. Knowledge capture prevents duplicate research.
**Gate**: Research brief includes: direct answer, confidence level, key findings with source attribution, gaps and unknowns, recommended next steps.
---
## Cross-Mode Patterns
These apply regardless of mode. Internalize them.
**Empathy is not performance.** Acknowledge the customer's situation genuinely. "I understand how frustrating this must be" is fine when they're frustrated. The same phrase when they asked a simple how-to question is patronizing. Match the emotional register. Read `references/customer-support/llm-support-failure-modes.md` on tone mismatch.
**Own it.** Active voice. "We" not "the system." "I'll investigate" not "this will be investigated." Take responsibility where appropriate. Take ownership in all customer-facing communication.
**Specificity over reassurance.** "I'll update you by Friday at 3pm" beats "I'll get back to you soon." Concrete details build trust. Vague reassurance erodes it.
**Close the loop.** Every interaction ends with clear next steps: what you will do, what they need to do, when they hear from you next.
**Confirm feature existence before stating it.** Say 'let me check' when uncertain. A wrong answer that sends a customer down a dead-end path is worse than "let me check and get back to you." See `references/customer-support/llm-support-failure-modes.md`.
**Limit commitments to actions within your authority.** No timeline commitments on behalf of engineering. No policy exceptions without approval. No "we'll definitely build that." The trust cost of a broken promise exceeds the short-term relief of making one.
references/customer-support/knowledge-base.md
# Knowledge Base
Deep reference for converting resolved support tickets into publish-ready KB articles: article type selection, structure templates, search optimization, formatting standards, maintenance cadence, and categorization taxonomy.
---
## When to Create a KB Article
Not every resolved ticket deserves an article. Create one when:
| Signal | Article Type |
|--------|-------------|
| Same question asked 3+ times | FAQ or How-to |
| Ticket resolution required non-obvious steps | Troubleshooting |
| Product behavior surprises customers | Known Issue or FAQ |
| New feature launched without adequate docs | How-to |
| Workaround exists for a known bug | Known Issue |
| Complex setup customers struggle with | How-to |
| Common error message with non-obvious fix | Troubleshooting |
Skip when:
- One-off account-specific issue
- Issue already covered by existing article (update instead)
- Bug will be fixed within days and workaround is simple
- Issue requires internal-only context to understand
---
## Article Type Selection
| Type | Purpose | Best For | Structure |
|------|---------|----------|-----------|
| **How-to** | Step-by-step task completion | Setup, configuration, workflow guidance | Prerequisites -> Steps -> Verify -> Common Issues |
| **Troubleshooting** | Diagnose and fix a problem | Error resolution, unexpected behavior | Symptoms -> Cause -> Solution(s) -> Prevention |
| **FAQ** | Quick answer to common question | Policy, capability, pricing, simple factual | Direct Answer -> Details -> Related Questions |
| **Known Issue** | Document a bug with workaround | Active bugs, limitations, regressions | Status -> Symptoms -> Workaround -> Fix Timeline |
| **Reference** | Technical specification or comparison | API docs, plan comparisons, config options | Overview -> Details (tables) -> Examples |
### Type Disambiguation
| Customer says | Type | Reasoning |
|--------------|------|-----------|
| "How do I...?" | How-to | Task-oriented, needs steps |
| "It's not working" | Troubleshooting | Problem-oriented, needs diagnosis |
| "Can I...?" or "Does it...?" | FAQ | Yes/no answer with context |
| "I'm seeing an error that I heard is a known bug" | Known Issue | Active defect, needs status/workaround |
| "What are the limits?" | Reference | Factual, tabular |
---
## Article Structure Templates
### How-to Article
```markdown
# How to [accomplish task]
[1-2 sentence overview: what this guide covers and when you'd use it]
## Prerequisites
- [What's needed before starting]
- [Required permissions, plan level, etc.]
## Steps
### 1. [Action verb phrase]
[Instruction with specific UI paths: "Go to Settings > Integrations > API Keys"]
[What you should see after completing: "A green confirmation banner appears"]
### 2. [Action verb phrase]
[Instruction]
### 3. [Action verb phrase]
[Instruction]
## Verify It Worked
[How to confirm the task succeeded -- specific observable outcome]
## Common Issues
| Issue | Fix |
|-------|-----|
| [Problem that might occur] | [Resolution] |
| [Another common snag] | [Resolution] |
## Related Articles
- [Link to related how-to or troubleshooting]
```
**Writing rules for how-tos:**
- Start each step with a verb
- Include the full navigation path ("Settings > Integrations > API Keys")
- State what the user should see after each step
- Test steps against the actual product before publishing
- One task per article -- split multi-task guides
### Troubleshooting Article
```markdown
# [Problem description -- what the user sees]
If you're seeing [symptom in customer language], this article
explains how to fix it.
## Symptoms
- [Observable behavior 1]
- [Observable behavior 2]
- [Exact error message if applicable, in code block]
## Cause
[Why this happens -- brief, non-jargon explanation.
Keep it simple even if the root cause is complex.]
## Solution
### Option 1: [Primary fix -- most likely to work]
1. [Step]
2. [Step]
3. [Step]
### Option 2: [Alternative if Option 1 doesn't work]
1. [Step]
2. [Step]
## Prevention
[How to avoid this in the future, if applicable]
## Still Having Issues?
If these steps didn't resolve the problem, contact support
with the following information:
- [What to include in their ticket]
```
**Writing rules for troubleshooting:**
- Lead with symptoms, not causes -- customers search for what they see
- Include exact error messages in code blocks (customers copy-paste into search)
- Provide multiple solutions ordered by likelihood
- Always include "Still having issues?" pointing to support
- Keep customer-facing explanation simple even if root cause is complex
### FAQ Article
```markdown
# [Question in the customer's words]
[Direct answer -- 1-3 sentences. Answer in the first sentence.]
## Details
[Additional context, nuance, or edge cases if needed.
Keep brief -- if this needs a walkthrough, it's a how-to.]
## Related Questions
- [Link to related FAQ]
- [Link to related FAQ]
```
**Writing rules for FAQs:**
- Answer the question in the first sentence
- Keep it concise -- if the answer needs steps, it's a how-to
- Use the customer's language in the title, not internal terminology
- Group related FAQs and cross-link
### Known Issue Article
```markdown
# Known Issue: [Brief description]
**Status:** [Investigating | Workaround Available | Fix In Progress | Resolved]
**Affected:** [Who/what is affected -- plan, feature, browser, etc.]
**Last updated:** [Date]
## Symptoms
[What users experience]
## Workaround
[Steps to work around the issue]
[Or: "No workaround currently available."]
## Fix Timeline
[Expected fix date, or current investigation status]
## Updates
- [Date]: [Update -- most recent first]
- [Date]: [Update]
```
**Writing rules for known issues:**
- Keep status current -- stale known-issue articles erode trust fast
- Update when fix ships, mark as Resolved
- Keep resolved articles live for 30 days (customers still searching old symptoms)
- If no workaround, say so honestly -- don't fabricate one
### Reference Article
```markdown
# [Topic]: [Specific scope]
[1-2 sentence overview of what this reference covers]
## [Section]
| [Column 1] | [Column 2] | [Column 3] |
|------------|------------|------------|
| [Data] | [Data] | [Data] |
## Examples
[Concrete examples showing usage or application]
## Related
- [Links to related reference or how-to articles]
```
---
## Search Optimization
Articles are useless if customers can't find them. Every article must be findable through search using the words a customer would type.
### Title Rules
| Good Title | Bad Title | Why |
|------------|-----------|-----|
| "How to configure SSO with Okta" | "SSO Setup" | Specific, includes the tool name |
| "Fix: Dashboard shows blank page" | "Dashboard Issue" | Includes the symptom |
| "API rate limits and quotas" | "API Information" | Includes specific terms |
| "Error: 'Connection refused' when importing data" | "Import Problems" | Includes exact error message |
| "Why is my export missing rows?" | "Export FAQ" | Matches customer's question |
### Keyword Strategies
- **Exact error messages**: customers copy-paste error text into search
- **Customer language**: "can't log in" not "authentication failure"
- **Common synonyms**: delete/remove, dashboard/home page, export/download
- **Alternate phrasings**: address the issue from different angles in the overview
- **Product area tags**: match how customers think about the product
### Opening Sentence Formulas
| Article Type | Formula |
|-------------|---------|
| How-to | "This guide shows you how to [accomplish X]." |
| Troubleshooting | "If you're seeing [symptom], this article explains how to fix it." |
| FAQ | "[Question in customer's words]? Here's the answer." |
| Known issue | "Some users are experiencing [symptom]. Here's what we know and how to work around it." |
| Reference | "This page documents [topic scope] for [audience]." |
---
## Formatting Standards
### Universal Rules
- **Headers (H2, H3)** for scannable sections
- **Numbered lists** for sequential steps (order matters)
- **Bullet lists** for non-sequential items (order doesn't matter)
- **Bold** for UI element names, key terms, emphasis
- **Code blocks** for commands, API calls, error messages, config values
- **Tables** for comparisons, options, reference data
- **Callouts** for warnings, tips, important caveats
- **Short paragraphs**: 2-4 sentences max
- **One idea per section**: if covering two topics, split
### Metadata Block
Every article includes:
```yaml
Title: [searchable, specific]
Type: [How-to | Troubleshooting | FAQ | Known Issue | Reference]
Category: [Product area]
Tags: [comma-separated searchable terms]
Audience: [All users | Admins | Developers | Specific plan]
Last updated: [Date]
```
---
## Categorization Taxonomy
Organize articles into a hierarchy matching customer mental models:
```
Getting Started
├── Account setup
├── First-time configuration
└── Quick start guides
Features & How-tos
├── [Feature area 1]
├── [Feature area 2]
└── [Feature area 3]
Integrations
├── [Per-integration guides]
└── API reference
Troubleshooting
├── Common errors
├── Performance issues
└── Known issues
Billing & Account
├── Plans and pricing
├── Billing questions
└── Account management
```
### Linking Rules
| From | To | Purpose |
|------|----|---------|
| Troubleshooting | How-to | "For setup instructions, see [Guide]" |
| How-to | Troubleshooting | "If you encounter errors, see [Troubleshooting]" |
| FAQ | Detailed article | "For a full walkthrough, see [Guide]" |
| Known issue | Workaround | Keep problem-to-solution chain short |
Link articles in one direction unless both are genuinely useful entry points. Use relative links within the KB — they survive restructuring.
---
## Maintenance Cadence
Knowledge bases decay without maintenance. Stale or wrong articles are worse than no articles.
| Activity | Frequency | Owner |
|----------|-----------|-------|
| New article review | Before publishing | Peer + SME for technical content |
| Accuracy audit | Quarterly | Support reviews top-traffic articles |
| Stale content check | Monthly | Flag articles not updated in 6+ months |
| Known issue updates | Weekly | Update status on all open known issues |
| Analytics review | Monthly | Low helpfulness ratings, high bounce rates |
| Gap analysis | Quarterly | Top ticket topics without KB articles |
### Article Lifecycle
| State | Meaning | Trigger |
|-------|---------|---------|
| Draft | Written, needs review | New article created |
| Published | Live, available to customers | Review passed |
| Needs Update | Flagged for revision | Product change, feedback, or age |
| Archived | Not live but preserved | No longer relevant |
| Retired | Removed from KB | Content permanently obsolete |
### Update vs. Create New
**Update existing when:**
- Product changed and steps need refreshing
- Article is mostly right but missing a detail
- Customer feedback says a section is confusing
- Better workaround or solution found
**Create new when:**
- New feature or product area needs docs
- Resolved ticket reveals a gap (no article exists)
- Existing article covers too many topics (split it)
- Different audience needs different explanation
---
## Ticket-to-Article Conversion Process
### Step 1: Identify the Candidate
After resolving a ticket, ask:
- Would this resolution help future customers self-serve?
- Has this question come up before (or will it again)?
- Is the workaround non-obvious?
### Step 2: Extract the Essentials
From the ticket thread, extract:
- The original problem in customer language
- The resolution steps (in order, with specifics)
- Edge cases or caveats discovered during resolution
- Environment or configuration details that matter
### Step 3: Generalize
Transform the ticket-specific resolution into a general guide:
- Remove customer-specific account details
- Replace specific values with representative examples
- Add prerequisites that were implicit in the ticket context
- Cover alternate scenarios the ticket didn't encounter
### Step 4: Validate
- Follow the steps yourself (or have someone unfamiliar follow them)
- Verify against the current product version
- Get SME review for technical accuracy
- Check that search terms match customer language
### Step 5: Publish and Link
- Add to the appropriate category
- Link from related articles
- Tag for searchability
- Set a review date (30-90 days depending on product change velocity)
---
## KB Quality Failure Modes
| Failure Mode | Problem | Fix |
|-------------|---------|-----|
| Internal jargon in title | Customers can't find it via search | Rewrite in customer language |
| Steps that skip prerequisites | Customer gets stuck at step 1 | Add explicit prerequisites section |
| No "Still Having Issues?" section | Dead end when article doesn't help | Always include support escalation path |
| Screenshots without text descriptions | Breaks when UI changes, not accessible | Describe UI elements in text, use screenshots as supplement |
| Monster article covering 5 topics | Hard to find, hard to maintain | Split into focused single-topic articles |
| Outdated known issue without status update | Erodes trust in entire KB | Weekly status sweeps on all known issues |
| Written for internal audience | Customer can't follow | Rewrite with zero assumed internal knowledge |
references/customer-support/llm-support-failure-modes.md
# LLM Support Failure Modes
Where LLMs fail in customer support contexts. These are not theoretical risks -- they are observed, recurring failure patterns that produce real customer harm. Load this reference in every support mode.
This file exists because LLM failures in support are categorically different from LLM failures in code or analysis. A wrong code suggestion gets caught by a compiler. A wrong support response gets sent to a customer and damages trust. The feedback loop is slower and the cost is higher.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers support-specific failures only.
---
## Failure Mode 1: Hallucinated Product Features
### What Happens
The LLM invents product capabilities that don't exist, describes features from a competitor as your own, or states that a feature works in a way it doesn't. The customer follows instructions to a dead end, then contacts support again -- now frustrated and distrustful.
### Why LLMs Do This
- Training data contains documentation from many products. Feature boundaries blur.
- LLMs optimize for helpfulness and will fabricate a plausible answer rather than say "I don't know."
- Product-specific knowledge decays rapidly as products ship updates after training cutoff.
### Concrete Examples
| Hallucination | Reality | Damage |
|--------------|---------|--------|
| "You can export to PDF from the dashboard" | PDF export doesn't exist | Customer wastes 20 min looking for a button that isn't there |
| "Enable the SSO option in Settings > Security" | SSO is enterprise-only | Customer on starter plan hits a wall, feels misled |
| "The API supports webhook retries natively" | Retries must be implemented client-side | Customer builds integration on false assumption, discovers at production |
| "Bulk delete is available in the admin panel" | Feature was removed two versions ago | Customer can't find it, questions their own competence |
### Prevention Rules
1. **Never describe a feature you cannot verify exists.** If uncertain, say: "Let me verify whether [feature] is available and get back to you."
2. **Treat feature existence as a factual claim requiring evidence**, not a knowledge retrieval. Check documentation, product, or ask.
3. **State plan/tier restrictions explicitly.** "This feature is available on [plan]" prevents mismatched expectations.
4. **Flag when your knowledge may be outdated.** "I want to confirm this is current" is honest. "Yes, you can do that" when you're guessing is dangerous.
### Detection Signals
- You're describing a multi-step UI workflow from memory without having seen the current product
- You're using phrases like "should be" or "typically" about specific product behavior
- You're describing a feature that sounds similar to what a competitor offers
- You can't point to a documentation URL or product screen to back the claim
---
## Failure Mode 2: Tone Mismatch
### What Happens
The LLM applies the wrong emotional register to the situation. Overly cheerful when the customer is angry. Patronizingly empathetic when they asked a simple question. Casually dismissive when their production system is down.
### Why LLMs Do This
- Default training biases toward relentless positivity and helpfulness
- Lack of genuine emotional intelligence -- the LLM pattern-matches, not empathizes
- "Be empathetic" instructions produce formulaic empathy regardless of context
- LLMs can't distinguish between "mildly curious" and "existentially frustrated"
### The Tone Mismatch Matrix
| Customer State | Wrong Response | Right Response | Why It Matters |
|---------------|---------------|----------------|----------------|
| Angry (production down for 6 hours) | "I'd be happy to help! Let's take a look." | "I understand your production has been down for 6 hours. That's unacceptable and I'm treating this as our top priority." | Happy-to-help tone trivializes a crisis |
| Simple question | "I completely understand how confusing this can be, and I really appreciate your patience..." | "The setting is under Dashboard > Settings > Notifications." | Over-empathizing a simple question is patronizing |
| Frustrated (third time asking) | "Great question! Here's how to..." | "I see this is the third time you've reached out about this. I'm sorry we haven't resolved it yet. Here's what I'm doing differently this time:" | Ignoring history feels like nobody reads tickets |
| Reporting a minor bug | "I am so sorry you're experiencing this. We take every issue extremely seriously..." | "Thanks for reporting this. I've logged it as [priority] and we'll have it fixed in [timeframe]." | Disproportionate gravity for a minor issue is unsettling |
| Executive escalation | "No worries! We'll get this sorted out." | "I've reviewed the situation with [name]. Here's our action plan and timeline." | Casual tone to an executive signals lack of seriousness |
### Prevention Rules
1. **Read the customer's emotional state before writing.** Is this person frustrated, confused, angry, or matter-of-fact? Match accordingly.
2. **Scale empathy to impact.** Production down = full empathy. Typo in docs = thank them and fix it.
3. **Never open with "I'd be happy to help" when the customer is unhappy.** It reads as tone-deaf.
4. **Check the ticket history before responding.** A customer on their third contact about the same issue needs acknowledgment of that history, not a fresh "great question!"
5. **Match formality to stakeholder level.** End users get warm. Executives get composed and data-driven.
---
## Failure Mode 3: Unauthorized Promises
### What Happens
The LLM commits to timelines, features, exceptions, or actions that the support agent cannot authorize. The customer holds the organization to these promises. When they aren't kept, trust damage compounds.
### Why LLMs Do This
- Optimizing for immediate resolution over long-term accuracy
- Training on support transcripts where agents did make such promises
- Conflating "good customer service" with "say yes to everything"
- No model of organizational authority boundaries
### The Promise Spectrum
| Promise Type | Example | Risk Level | Why It's Dangerous |
|-------------|---------|-----------|-------------------|
| Timeline commitment | "This will be fixed by next Tuesday" | High | Engineering timelines are estimates, not promises |
| Feature commitment | "We're building that -- it'll be ready Q3" | Critical | Roadmap changes. Customers plan around your statement. |
| Policy exception | "I'll waive that fee for you" | Medium | Sets precedent, may not be authorized |
| Refund guarantee | "We'll definitely refund the full amount" | Medium | May need approval, may not be policy |
| SLA override | "I'll make sure it's resolved within the hour" | High | Can't guarantee engineering response time |
| Competitive claim | "Our product does everything [competitor] does" | Critical | Almost certainly false in specifics |
### Prevention Rules
1. **Never commit to timelines you don't control.** "Our engineering team is investigating and I'll update you with their timeline" is honest. "This will be fixed by Friday" when you don't know is not.
2. **Never promise features on the roadmap.** "I've shared this feedback with our product team" is safe. "We're planning to build this" creates a contract.
3. **State what you CAN do, not what you wish you could.** "I'm able to [specific action]" keeps commitments within your authority.
4. **Distinguish between "I will" and "I'll request."** "I'll request a refund for you -- the billing team typically processes these within [timeframe]" is honest about the approval chain.
5. **If you must give a timeline, pad it and frame it as an estimate.** "I expect to have an update by [date], but I'll let you know if anything changes."
### Detection Signals
- You're using "will" instead of "expect to" or "plan to" for future actions
- You're promising outcomes that depend on another team
- You're waiving fees or making exceptions without checking policy
- You're stating "we can do that" about a feature you haven't verified
---
## Failure Mode 4: Over-Templated Responses
### What Happens
The response is technically correct but reads like it was generated by a mail merge. Generic acknowledgment, generic empathy, generic next steps. The customer feels like they're talking to a script, not a person.
### Why LLMs Do This
- Template-following is a core LLM strength -- and weakness
- Response templates in training data create strong attractors
- Without specific context to anchor, the LLM defaults to generic patterns
- Instructions to "be professional" often produce corporate-speak
### Generic vs. Specific Examples
| Generic (bad) | Specific (good) | What Changed |
|--------------|----------------|--------------|
| "Thank you for reaching out about this issue." | "Thank you for the detailed report about the dashboard blank-page error." | Named the actual issue |
| "I understand this is frustrating." | "I can see this has been blocking your team's reporting workflow since Monday." | Named the specific impact |
| "Our team is looking into it." | "I've assigned this to our data team, and they're investigating the query timeout that's causing the blank results." | Named the team and the technical lead |
| "I'll follow up soon." | "I'll update you by 3pm ET tomorrow with what we find." | Specific time commitment |
| "Please let us know if you have any questions." | "Would it help if I set up a 15-minute call with our integration specialist to walk through the configuration?" | Specific, actionable offer |
### Prevention Rules
1. **Reference the customer's specific situation in every response.** Use their words, their product area, their specific error.
2. **Replace "this issue" with the actual issue name.** "The CSV export timeout" not "this issue."
3. **Replace "our team" with the actual team or person.** "Sarah on our billing team" not "our team."
4. **Replace "soon" with a specific time.** "By 3pm ET Friday" not "soon."
5. **After drafting, read the response and ask: could this response apply to any other customer with any other issue?** If yes, it's too generic. Add specifics.
---
## Failure Mode 5: Premature Closure
### What Happens
The LLM declares the issue resolved or suggests closing the ticket before the customer has confirmed the fix works. Or it provides a solution without verifying it addresses the actual problem, not just the stated symptom.
### Why LLMs Do This
- Optimizing for resolution speed over resolution quality
- Pattern-matching symptom to known solution without validating assumptions
- Treating "provided an answer" as equivalent to "solved the problem"
- No concept of "did this actually work for you?"
### Premature Closure Patterns
| Pattern | Problem | Correct Approach |
|---------|---------|-----------------|
| "This should resolve your issue" | "Should" != "does" | "Please try these steps and let me know if the issue persists" |
| Closing ticket after sending instructions | Customer may not have tried them yet | Wait for customer confirmation before closing |
| "Let me know if you need anything else" as sign-off on an unresolved issue | Implies the issue is resolved when it isn't | "I'll check back in [timeframe] to see if this resolved the issue" |
| Providing a solution for the symptom, not the root cause | Issue recurs | Ask clarifying questions before prescribing a fix |
| Marking "resolved" when a workaround was given | Customer still has the underlying problem | Workaround tickets stay open until the root fix ships |
### Prevention Rules
1. **Never close a ticket without customer confirmation.** "Did this resolve the issue?" is required before any close action.
2. **Distinguish between "answered" and "resolved."** Providing information is not the same as fixing the problem.
3. **Follow up proactively.** If the customer goes silent after your response, check in -- don't assume silence means success.
4. **Keep workaround tickets open.** A workaround is a bridge, not a resolution.
---
## Failure Mode 6: Context Amnesia
### What Happens
The LLM treats each interaction as isolated. It asks questions the customer already answered. It suggests solutions already tried. It contradicts what a colleague said in the previous response. The customer feels like they're starting over every time.
### Why LLMs Do This
- Context window limitations
- No persistent memory across interactions
- Each response generated independently without thread awareness
- Switching between agents loses accumulated context
### Context Amnesia Patterns
| Pattern | Customer Experience | Prevention |
|---------|-------------------|------------|
| Asking for information already provided | "I already told you my OS in my first message" | Read the full thread before responding |
| Suggesting already-tried solutions | "We tried that yesterday -- it didn't work" | Check what's been attempted before suggesting |
| Contradicting previous agent | "Your colleague said X, now you're saying Y" | Review previous responses before contradicting |
| Losing track of commitments | "You promised to update me Tuesday" | Check for any commitments made in thread history |
| Restarting troubleshooting from scratch | "I've already done all the basic steps" | Acknowledge what's been tried, start from where they left off |
### Prevention Rules
1. **Read the full ticket history before responding.** Every time. No exceptions.
2. **Acknowledge previous interactions.** "I see you spoke with [name] about this on [date]" signals continuity.
3. **Note previous troubleshooting.** "Since you've already tried [X, Y, Z], let's move to [next step]."
4. **Honor previous commitments.** If a colleague promised an update, deliver it or explain the change.
5. **When transferring, include a summary.** The customer should never have to re-explain.
---
## Failure Mode 7: Defensive or Deflective Language
### What Happens
When the LLM receives criticism or is confronted with a product failure, it deflects blame, minimizes the issue, or becomes subtly defensive. This reads as the company not taking responsibility.
### Deflection Patterns
| Deflective (bad) | Ownership (good) |
|-----------------|------------------|
| "This is working as designed" | "I understand this behavior doesn't match your expectation. Let me look into options." |
| "That's a third-party issue" | "Even though [partner] is involved, I'll own coordination until this is resolved." |
| "We haven't seen this issue before" | "I want to investigate this thoroughly. Can you share [details]?" |
| "You may need to check your configuration" | "Let's review your configuration together to see if we can find the issue." |
| "The system was performing maintenance" | "We had maintenance that affected your access. I should have communicated that proactively." |
### Prevention Rules
1. **"We" not "the system."** Take organizational ownership of the experience.
2. **Never blame the customer's configuration first.** Investigate collaboratively.
3. **Own cross-vendor issues.** Customers don't care whose fault it is -- they care who fixes it.
4. **Acknowledge the impact before explaining the cause.** Impact first, root cause second.
---
## Cross-Cutting Detection Rules
Apply these checks to every response before presenting:
| Check | Question | If Yes |
|-------|----------|--------|
| Feature claim | Am I stating a product capability I haven't verified? | Verify or hedge: "Let me confirm" |
| Tone calibration | Does the emotional register match the customer's state? | Re-read their message, adjust |
| Authority boundary | Am I committing to something outside my authority? | Reframe: "I'll request" or "I expect" |
| Specificity | Could this response apply to any customer with any issue? | Add specifics from this ticket |
| Closure readiness | Have they confirmed the fix works? | Don't close until confirmed |
| Thread awareness | Have I read and acknowledged the full history? | Re-read thread, reference prior context |
| Defensiveness | Am I deflecting blame or minimizing impact? | Take ownership, acknowledge impact |
---
## The Meta-Failure: Confidence Without Verification
All seven failure modes share a root cause: the LLM generates confident responses without verifying the claims against reality. Confident tone is not correlated with accuracy. The most dangerous support response is one that sounds authoritative and is wrong.
**The single most important rule**: if you are not certain, say so. "Let me verify and get back to you" is always better than a confident wrong answer. Customers forgive uncertainty. They do not forgive being misled.
references/customer-support/response-drafting.md
# Response Drafting
Deep reference for customer-facing response drafting: tone calibration by situation, relationship-stage adjustments, de-escalation patterns, channel-appropriate formatting, follow-up templates, and quality verification.
---
## Tone Spectrum
Select tone based on the situation type. Every response sits somewhere on this spectrum.
| Situation | Tone | Characteristics | Risk if Wrong |
|-----------|------|----------------|---------------|
| Good news / wins | Celebratory | Enthusiastic, warm, congratulatory, forward-looking | Flat delivery kills the moment |
| Routine update | Professional | Clear, concise, informative, friendly | Too formal reads cold; too casual reads unserious |
| Technical response | Precise | Accurate, detailed, structured, patient | Condescension if you over-explain; confusion if you under-explain |
| Delayed delivery | Accountable | Honest, apologetic, action-oriented, specific | Defensiveness or vagueness destroys trust |
| Bad news | Candid | Direct, empathetic, solution-oriented, respectful | Burying the news in jargon feels dishonest |
| Issue / outage | Urgent | Immediate, transparent, actionable, reassuring | Downplaying severity insults intelligence |
| Escalation | Executive | Composed, ownership-taking, plan-presenting, confident | Panic signals incompetence |
| Billing / account | Precise | Clear, factual, empathetic, resolution-focused | Vagueness about money creates anxiety |
---
## Relationship Stage Calibration
### New Customer (0-3 months)
- More formal and professional tone
- Extra context and explanation (don't assume product knowledge)
- Proactively offer help and resources
- Build trust through reliability and responsiveness
- Include links to relevant documentation
### Established Customer (3+ months)
- Warm and collaborative tone
- Reference shared history and previous conversations
- More direct and efficient -- skip the preamble
- Show awareness of their goals and priorities
- "As you know" or "building on our previous conversation" signals continuity
### Frustrated or Escalated Customer
- Lead with empathy -- acknowledge before solving
- Urgency in response time and language
- Concrete action plan with specific commitments
- Shorter feedback loops ("I'll update you every 2 hours")
- Name the person accountable for resolution
- Offer a call or meeting for severity-appropriate situations
---
## De-escalation Patterns
When a customer is angry, frustrated, or escalating. The goal is not to win an argument. The goal is to move from emotion to resolution.
### The De-escalation Sequence
1. **Acknowledge** -- Name what they're feeling without being patronizing
2. **Validate** -- Confirm their concern is reasonable (even if the specific complaint isn't)
3. **Own** -- Take responsibility where appropriate, never deflect
4. **Plan** -- Present concrete next steps with timeline
5. **Commit** -- State exactly what you will do and by when
### De-escalation Language
| Instead of | Write | Why |
|-----------|-------|-----|
| "I understand your frustration" (generic) | "I can see how [specific impact] would be frustrating for your team" | Specificity proves you actually read their message |
| "Unfortunately, we can't..." | "Here's what we can do: [alternative]" | Lead with the solution, not the limitation |
| "This is working as designed" | "The current behavior is [X]. I can see why you'd expect [Y]. Let me explore options." | Validate their expectation before explaining the gap |
| "Per our policy..." | "Here's what I'm able to do in this situation: [action]" | Policy citations feel bureaucratic |
| "As I mentioned previously..." | [Just restate the information] | Implies they should have read your earlier message |
| "That's not really a bug" | "I see what you're experiencing. Let me dig into why it behaves that way." | Classification debate alienates; investigation helps |
| "You need to..." | "Here's how to resolve this: [steps]" | Directive language feels adversarial under stress |
<!-- no-pair-required: failure modes are self-explanatory failure modes; positive approach is in surrounding context -->
### De-escalation Failure Modes
| Failure Mode | Why It Fails |
|-------------|-------------|
| Matching their energy/anger | Escalates, never de-escalates |
| Excessive apologizing without action | Apology fatigue -- they want fixes, not sorry |
| Explaining the technical root cause first | They don't care about your architecture; they care about their workflow |
| Citing policy to justify the experience | Policy is your constraint, not their problem |
| Dismissing with "works for me" | Invalidates their experience entirely |
| Over-promising to end the conversation | Creates a bigger problem when you can't deliver |
---
## Response Structure
### Standard Structure (all situations)
```
1. Acknowledgment (1-2 sentences)
- Acknowledge what they said, asked, or are experiencing
- Show you understand their specific situation
2. Core Message (1-3 paragraphs)
- Deliver the main information, answer, or update
- Be specific and concrete
- Include relevant details they need
3. Next Steps (1-3 bullets)
- What YOU will do and by when
- What THEY need to do (if anything)
- When they'll hear from you next
4. Closing (1 sentence)
- Warm but professional sign-off
- Reinforce availability
```
### Length by Channel
| Channel | Target Length | Structure Notes |
|---------|-------------|-----------------|
| Chat/IM | 1-4 sentences | Get to the point immediately. Skip formal structure. |
| Support ticket | 1-3 short paragraphs | Structured and scannable. Use bullets freely. |
| Email | 3-5 paragraphs max | Respect their inbox. Headers for long responses. |
| Escalation response | As needed | Well-structured with headers. Thoroughness > brevity. |
| Executive communication | 2-3 paragraphs max | Shorter is better. Lead with the decision or ask. Data-driven. |
---
## Situation-Specific Approaches
### Answering a Product Question
- Lead with the direct answer, then context
- Link to documentation if relevant
- If you don't know: say so, commit to finding out, give a timeline
- Never guess or speculate about product capabilities
### Responding to a Bug Report
- Acknowledge the impact on their work specifically
- State what you know about the issue and its status
- Provide workaround if available
- Set expectations for resolution timeline
- Commit to regular updates
### Handling an Escalation
- Acknowledge severity and their frustration
- Take ownership (no deflecting, no "the engineering team...")
- Present a clear action plan with timeline
- Name the person accountable
- Offer a call if severity warrants it
### Delivering Bad News
- Be direct -- don't bury it in paragraph three
- Explain the reasoning honestly
- Acknowledge their specific impact
- Offer alternatives or mitigation
- Provide a clear path forward
### Declining a Request
- Acknowledge the request and its reasoning
- Be honest about the decision
- Explain why without being dismissive
- Offer alternatives when possible
- Leave the door open for future conversation
### Outage Communication
- Lead with status: what's affected, current state, ETA
- Be transparent about what you know and don't know
- Provide workaround if one exists
- Commit to update cadence
- Address prevention after resolution
---
## Writing Style Rules
### Do
- Active voice: "We'll investigate" not "This will be investigated"
- "I" for personal commitments, "we" for team commitments
- Name people when assigning actions: "Sarah from our engineering team will..."
- Use their terminology, not your internal jargon
- Specific dates and times: "by Friday January 24" not "in a few days"
- Break up long responses with headers or bullets
### Don't
- Corporate jargon: "synergy", "leverage", "paradigm shift", "circle back"
- Deflect blame to other teams, systems, or processes
- Passive voice to avoid ownership: "Mistakes were made"
- Excessive hedging that undermines confidence
- Exclamation marks beyond one per message (if any)
- "Please be advised" or "Kindly note" (bureaucratic distancing)
- "As per my last email" (passive-aggressive)
---
## Follow-up Templates
### Post-Resolution Check-in
```
Hi [Name],
I wanted to follow up on the [issue] we resolved on [date].
Is everything working as expected on your end?
If anything comes up, don't hesitate to reach out.
Best,
[Your name]
```
### After Silence (No Response)
```
Hi [Name],
I sent over [what you sent] on [date] and wanted to make
sure it didn't get lost.
[Brief reminder of what you need or what you offered]
If now isn't a good time, no worries -- let me know when
would work and I'm happy to reconnect then.
Best,
[Your name]
```
### Outage Resolution Notification
```
Hi [Name],
Good news -- the [issue] affecting [service/feature] has
been resolved as of [time].
**What happened:** [Clear, non-technical explanation]
**Root cause:** [Brief explanation]
**What we've done to prevent recurrence:** [Steps taken]
Your team should be able to [resume normal activity] now.
If you notice anything unusual, please let us know immediately.
I'm sorry for the disruption. We take reliability seriously
and have [specific preventive action] in place.
[Your name]
```
### Status Update (No New Information)
```
Hi [Name],
I wanted to give you an update on [issue]. Our team is
still actively investigating and I don't have new information
to share yet.
Here's where things stand:
- [What we've checked so far]
- [What we're looking at next]
- [Expected next milestone]
I'll update you again by [specific time]. If anything changes
before then, I'll let you know immediately.
[Your name]
```
---
## Follow-up Cadence
| Situation | Follow-up Timing |
|-----------|-----------------|
| Unanswered question | 2-3 business days |
| Open critical issue | Every 2-4 hours until resolved |
| Open standard issue | Daily until resolved |
| Post-resolution | 3-5 days to confirm |
| After delivering bad news | 1 week to check sentiment |
| Post-meeting action items | Within 24 hours (notes), then at deadline |
---
## Quality Verification Checklist
Run before presenting any draft:
### Content Checks
- [ ] Directly addresses the customer's actual question or concern
- [ ] Accurate -- no claims about features, timelines, or policies you can't verify
- [ ] No roadmap details that shouldn't be shared externally
- [ ] No internal jargon or team names the customer wouldn't know
### Commitment Checks
- [ ] No unauthorized timeline commitments
- [ ] No policy exceptions you can't approve
- [ ] No promises about features or fixes being built
- [ ] All next steps are achievable and within your authority
### Tone Checks
- [ ] Matches the situation (empathetic for frustration, direct for questions)
- [ ] Matches the relationship stage (formal for new, warm for established)
- [ ] No corporate jargon or bureaucratic distancing
- [ ] Active voice -- clear ownership of actions
- [ ] Appropriate length for the channel
### Structural Checks
- [ ] Clear acknowledgment of their situation
- [ ] Core message is specific and concrete
- [ ] Next steps include who does what by when
- [ ] Professional close with availability signal
---
## Internal Notes Format
Every drafted response includes internal notes (not sent to customer):
```
### Notes (internal -- do not send)
- **Tone rationale:** [Why this tone was chosen]
- **Verify before sending:** [Facts or commitments to confirm]
- **Risk factors:** [Anything sensitive about this response]
- **Follow-up needed:** [Actions after sending]
- **Escalation note:** [If someone else should review first]
```
references/customer-support/triage-methodology.md
# Triage Methodology
Deep reference for ticket triage: category taxonomy, priority scoring, SLA mapping, routing rules, duplicate detection, escalation triggers, and auto-response templates.
---
## Category Taxonomy
Every ticket gets a **primary category** and optionally a **secondary category**. Root cause drives the category, not the symptom described.
| Category | Description | Signal Words |
|----------|-------------|-------------|
| **Bug** | Product behaving incorrectly or unexpectedly | Error, broken, crash, not working, unexpected, wrong, failing |
| **How-to** | Customer needs guidance on using the product | How do I, can I, where is, setting up, configure, help with |
| **Feature request** | Customer wants capability that doesn't exist | Would be great if, wish I could, any plans to, requesting |
| **Billing** | Payment, subscription, invoice, pricing | Charge, invoice, payment, subscription, refund, upgrade, downgrade |
| **Account** | Access, permissions, settings, user management | Login, password, access, permission, SSO, locked out, can't sign in |
| **Integration** | Third-party tools, APIs, webhooks | API, webhook, integration, connect, OAuth, sync, third-party |
| **Security** | Security concerns, data access, compliance | Data breach, unauthorized, compliance, GDPR, SOC 2, vulnerability |
| **Data** | Data quality, migration, import/export | Missing data, export, import, migration, incorrect data, duplicates |
| **Performance** | Speed, reliability, availability | Slow, timeout, latency, down, unavailable, degraded |
### Category Disambiguation Rules
| Symptom | Category | Reasoning |
|---------|----------|-----------|
| Bug AND feature request in same ticket | Bug (primary) | Bugs take priority -- investigate before dismissing |
| Can't log in due to a bug | Bug, not Account | Root cause drives category |
| "It used to work and now it doesn't" | Bug | Regression signal |
| "I want it to work differently" | Feature request | Enhancement signal |
| "How do I make it work?" | How-to | Guidance signal |
| Unclear whether bug or user error | Bug | Err toward investigation over dismissal |
---
## Priority Framework
### P1 -- Critical
**Criteria**: Production system down, data loss/corruption, security breach, all or most users affected.
Indicators:
- Customer cannot use the product at all
- Data being lost, corrupted, or exposed
- Security incident in progress
- Issue worsening or expanding in scope
**SLA**:
- Respond: 1 hour
- Work pattern: continuous until resolved or mitigated
- Update cadence: every 1-2 hours
### P2 -- High
**Criteria**: Major feature broken, significant workflow blocked, many users affected, no workaround.
Indicators:
- Core workflow broken but product partially usable
- Multiple users affected or key account impacted
- Blocking time-sensitive work
- No reasonable workaround
**SLA**:
- Respond: 4 hours
- Work pattern: active investigation same day
- Update cadence: every 4 hours
### P3 -- Medium
**Criteria**: Feature partially broken, workaround available, single user or small team affected.
Indicators:
- Feature not working correctly but workaround exists
- Inconvenient but not blocking critical work
- Single user or small team affected
- Customer not escalating urgently
**SLA**:
- Respond: 1 business day
- Resolution or update: 3 business days
### P4 -- Low
**Criteria**: Minor inconvenience, cosmetic issue, general question, feature request.
Indicators:
- Cosmetic or UI issues not affecting functionality
- Feature requests and enhancement ideas
- General questions or how-to inquiries
- Issues with simple, documented solutions
**SLA**:
- Respond: 2 business days
- Resolution: normal pace
### Priority Bump Triggers
Automatically escalate priority when:
| Trigger | Action |
|---------|--------|
| Customer waiting longer than SLA | Bump one level |
| Multiple customers report same issue | Bump to at least P2 |
| Customer explicitly escalates or mentions executives | Bump one level |
| Workaround stops working | Bump one level |
| Issue expands in scope (more users, data, symptoms) | Reassess from scratch |
| Revenue at risk (churn signal, contract renewal) | Bump to at least P2 |
---
## Routing Rules
| Route to | When | Expected Action |
|----------|------|-----------------|
| **Tier 1** (frontline) | How-to questions, known issues with documented solutions, billing inquiries, password resets | Resolve using KB, templates, standard procedures |
| **Tier 2** (senior) | Bugs requiring investigation, complex configuration, integration troubleshooting, account issues needing elevated access | Deep investigation, advanced troubleshooting, cross-team coordination |
| **Engineering** | Confirmed bugs needing code fixes, infrastructure issues, performance degradation | Root cause analysis, code fix, deploy |
| **Product** | Feature requests with significant demand, design decisions, workflow gaps | Prioritization decision, roadmap consideration |
| **Security** | Data access concerns, vulnerability reports, compliance questions | Immediate assessment, containment if needed |
| **Billing/Finance** | Refund requests, contract disputes, complex billing adjustments | Account adjustment, approval workflow |
### Routing Complexity Signals
Route to Tier 2+ when:
- Issue requires access the frontline agent doesn't have
- Troubleshooting has exceeded 3 back-and-forth exchanges without resolution
- Issue involves custom configuration or non-standard setup
- Customer is a high-value account (enterprise, high ARR)
- Issue crosses product areas or teams
---
## Duplicate Detection Process
Before routing, check for duplicates:
1. **Search by symptom**: similar error messages or descriptions in open tickets
2. **Search by customer**: does this customer have an open ticket for the same issue
3. **Search by product area**: recent tickets in the same feature area
4. **Check known issues**: compare against documented known issues list
### When Duplicate Found
| Scenario | Action |
|----------|--------|
| Exact duplicate from same customer | Merge into existing ticket, notify customer |
| Same issue, different customer | Link tickets, bump priority if pattern emerging (3+ = escalate the pattern) |
| Similar but different root cause | Create new ticket, reference related ticket in notes |
| Resolved duplicate exists | Check if resolution applies, link for context |
---
## Escalation Trigger Matrix
These conditions should trigger escalation regardless of initial triage assessment:
| Signal | Escalation Target | Urgency |
|--------|-------------------|---------|
| "We're considering switching to [competitor]" | Leadership + Account team | High |
| "This is a security/compliance issue" | Security team | Critical |
| "Our CEO / VP / CTO wants to talk to someone" | Leadership | High |
| 3+ customers report identical issue in 24h | Engineering + Leadership | High |
| Customer has been waiting > 2x SLA | Manager review | Medium |
| Data loss or corruption confirmed | Engineering | Critical |
| Customer references contract terms or legal | Legal + Leadership | High |
| Issue blocks go-live or launch | Engineering + Account team | Critical |
---
## Auto-Response Templates by Category
### Bug -- Initial Response
```
Thank you for reporting this. I can see how [specific impact]
would be disruptive for your work.
I've logged this as a [priority] issue and our team is
investigating. [If workaround: "In the meantime, you can
[workaround]."]
I'll update you within [SLA timeframe] with what we find.
```
### How-to -- Initial Response
```
Great question. [Direct answer or link to documentation]
[If multi-step: "Here's how to do this:"]
[Steps or guidance]
Let me know if that helps or if you have follow-up questions.
```
### Feature Request -- Initial Response
```
Thank you for this suggestion -- I can see why [capability]
would be valuable for your workflow.
I've documented this and shared it with our product team.
I can't commit to a specific timeline, but your feedback
directly informs our roadmap priorities.
[If alternative exists: "In the meantime, you might find
[alternative] helpful for achieving something similar."]
```
### Billing -- Initial Response
```
I understand billing issues need prompt attention. Let me
look into this.
[If straightforward: resolution details]
[If complex: "I'm reviewing your account now and will have
an answer for you within [timeframe]."]
```
### Security -- Initial Response
```
Thank you for flagging this -- we take security concerns
seriously and are reviewing this immediately.
I've escalated to our security team for investigation.
We'll follow up within [timeframe] with our findings.
[If action needed: "In the meantime, we recommend
[protective action]."]
```
### Account -- Initial Response
```
I can see you're having trouble accessing your account.
Let me help sort this out.
[If quick fix: resolution steps]
[If investigation needed: "I'm looking into this now and
will update you within [timeframe]."]
[If security-adjacent: "For your account's security, I may
need to verify some details before making changes."]
```
---
## Triage Output Template
Every triage produces this structured output:
```
## Triage: [One-line issue summary]
**Category:** [Primary] / [Secondary if applicable]
**Priority:** [P1-P4] -- [Brief justification]
**Product area:** [Area/team]
### Issue Summary
[2-3 sentence summary of what the customer is experiencing]
### Key Details
- **Customer:** [Name/account if known]
- **Impact:** [Who and what is affected]
- **Workaround:** [Available / Not available / Unknown]
- **Related tickets:** [Links to similar issues if found]
- **Known issue:** [Yes -- link / No / Checking]
### Routing Recommendation
**Route to:** [Team or queue]
**Why:** [Brief reasoning]
### Suggested Initial Response
[Draft first response to the customer]
### Internal Notes
- [Additional context for the agent picking this up]
- [Reproduction hints if it's a bug]
- [Escalation triggers to watch for]
```
---
## Triage Failure Modes
| Failure Mode | Problem | Correct Approach |
|-------------|---------|-----------------|
| Categorizing by symptom instead of root cause | Misroutes ticket, delays resolution | Investigate root cause before categorizing |
| Defaulting everything to P3 | Under-serves critical issues, over-serves low ones | Apply priority criteria honestly, err high when uncertain |
| Skipping duplicate check | Creates redundant work, fragments context | Always search before routing |
| Writing vague internal notes | Next agent wastes time re-investigating | Include what you checked, what you ruled out, what to try next |
| Routing to engineering without repro steps | Engineers send it back, customer waits longer | Reproduce or document what you tried before escalating |
| Dismissing "feature request" that's actually a bug | Customer feels unheard, real issue persists | "It used to work" = bug, not feature request |
references/decision-matrices.md
---
title: Decision Matrices and Scoring Models
domain: strategic-decision
level: 3
skill: strategic-decision
---
# Decision Matrices and Scoring Models
> **Scope**: Structured decision scoring, opportunity evaluation matrices, and pre-mortem analysis for CEO-level strategic decisions. Does NOT cover financial modeling (see strategic-frameworks.md) or operational project decisions.
> **Version range**: Framework-agnostic — applies to any strategic decision with 2-5 options.
> **Generated**: 2026-04-09 — validate weights against your organization's actual priorities.
---
## Overview
Strategic decisions fail most often because they stay qualitative too long. "This feels right" and "my gut says no" are not frameworks — they are ways to avoid accountability. Decision matrices force explicit weight assignments before scoring, which exposes hidden assumptions and surfaces disagreement early. The pre-mortem technique catches the most common failure: decisions optimized for the scenario where everything works.
---
## Weighted Decision Scoring Matrix
Fill in weights (must sum to 100%) and scores (1-10) before discussion. Filling in collaboratively introduces anchoring bias.
| Criterion | Weight | Option A | Option B | Option C |
|-----------|--------|----------|----------|----------|
| Strategic fit | 30% | ___ | ___ | ___ |
| Technical feasibility | 25% | ___ | ___ | ___ |
| Cost (3-year TCO) | 20% | ___ | ___ | ___ |
| Time to value | 15% | ___ | ___ | ___ |
| Reversibility | 10% | ___ | ___ | ___ |
| **Weighted score** | 100% | **calc** | **calc** | **calc** |
**Calculation**: `weighted_score = sum(criterion_weight * option_score)` for each option.
**Worked example** (market entry decision — enter APAC vs. expand EU vs. deepen US):
| Criterion | Weight | APAC | EU | US |
|-----------|--------|------|----|----|
| Strategic fit | 30% | 8 | 7 | 9 |
| Technical feasibility | 25% | 5 | 8 | 9 |
| Cost (3-year TCO) | 20% | 4 | 7 | 8 |
| Time to value | 15% | 6 | 7 | 9 |
| Reversibility | 10% | 4 | 7 | 9 |
| **Weighted score** | 100% | **5.95** | **7.30** | **8.90** |
Result: Deepen US wins on current resources. APAC revisit in 18 months when localization team is hired.
---
## Opportunity Scoring Model (ICE)
Use when comparing multiple opportunities to prioritize which deserves full decision matrix treatment.
| Opportunity | Impact (1-10) | Confidence (1-10) | Ease (1-10) | ICE Score |
|-------------|---------------|-------------------|-------------|-----------|
| ___ | ___ | ___ | ___ | **calc** |
| ___ | ___ | ___ | ___ | **calc** |
| ___ | ___ | ___ | ___ | **calc** |
**ICE Score** = (Impact + Confidence + Ease) / 3
**Impact**: How significant is the upside if this succeeds? (1 = marginal improvement, 10 = transformative)
**Confidence**: How certain are you in the impact estimate? (1 = pure guess, 10 = validated data)
**Ease**: How easy is execution relative to available resources? (1 = requires 2 years and $2M, 10 = one week, in-house)
**Cutoffs**: ICE < 4 = don't analyze further. ICE 4-6 = low priority. ICE 7+ = run full decision matrix.
---
## RICE Scoring for Strategic Opportunity Ranking
More precise than ICE when you have effort data. Use when resource allocation is the core question.
| Opportunity | Reach | Impact | Confidence | Effort (person-months) | RICE Score |
|-------------|-------|--------|------------|----------------------|------------|
| ___ | ___ | ___ | ___ | ___ | **calc** |
| ___ | ___ | ___ | ___ | ___ | **calc** |
**RICE** = (Reach × Impact × Confidence) / Effort
**Reach**: Users/customers impacted per quarter (absolute number or 1-10 scale, consistent across rows)
**Impact**: 0.25 (minimal) / 0.5 (low) / 1 (medium) / 2 (high) / 3 (massive)
**Confidence**: 0.5 (low) / 0.8 (medium) / 1.0 (high)
**Effort**: Person-months to complete
---
## Pre-Mortem Analysis Template
Run this AFTER scoring but BEFORE committing. It catches the most common failure mode: decisions that optimize for success scenarios.
**Setup**: Tell your team "Imagine 12 months have passed. This decision failed spectacularly. What happened?"
```
## Pre-Mortem: [Decision Name]
### Failure Mode 1 (Most Likely)
- What went wrong: ___
- Which assumption was wrong: ___
- Early warning signal we could have detected: ___
- Mitigation to add to the plan: ___
### Failure Mode 2 (Most Catastrophic)
- What went wrong: ___
- Which assumption was wrong: ___
- Early warning signal we could have detected: ___
- Mitigation to add to the plan: ___
### Failure Mode 3 (Most Unexpected)
- What went wrong: ___
- Which assumption was wrong: ___
- Early warning signal we could have detected: ___
- Mitigation to add to the plan: ___
### Decision change
Does any failure mode change the recommended option? [yes/no]
If yes, which option becomes preferred: ___
```
**Real-but-anonymized example**: A SaaS company scored "build partner integrations" as their top ICE opportunity (score: 8.1). Pre-mortem revealed failure mode 2: "Partners deprioritize the integration after signing the agreement, leaving us with dead listings." Mitigation added: co-marketing commitment required before integration goes live. Decision unchanged, but execution plan materially different.
---
## Reversibility Classification
Apply before scoring. Changes how much analysis is warranted.
| Decision | Reversibility | Analysis Depth | Time Box |
|----------|--------------|----------------|----------|
| Hire a VP of Sales | Low — severance + disruption | Full matrix + pre-mortem | 2-4 weeks |
| New pricing tier | Medium — churn risk, but revertable | Full matrix | 1 week |
| A/B test new homepage | High — can roll back in hours | Skip matrix, just do it | 1 day |
| Acquire a company | Irreversible — culture + integration | Full matrix + pre-mortem + external validation | 4-12 weeks |
| New market entry (pilot) | High if scoped correctly | Matrix on go/no-go, light on details | 3-5 days |
**Classification test**: "If this turns out to be wrong, how long does it take to undo and what is the cost?" Under 2 weeks and under $10K: high reversibility. Over 6 months or over $100K: low reversibility.
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Criteria added after scoring
**What it looks like**: Team scores options, then someone says "shouldn't we also consider X?" and adds it — but only because option they preferred scored low.
**Detection**: Track criteria list from start of session. Any addition after first score row is filled = suspect.
**Do instead**: Lock the criteria list before any scoring begins. If a new criterion is raised mid-session, note it for the next decision cycle. Any addition after scoring starts triggers a full restart.
**Fix**: Lock criteria before any scoring. New criteria trigger a full restart.
### Equal weighting as a default
**What it looks like**: 5 criteria, each gets 20% because "everything matters equally."
**Why wrong**: Hides real priorities. If cost and strategic fit genuinely matter equally, that is an unusual organization — say so explicitly. Usually it means the weighter avoided a political conversation about what actually matters.
**Do instead**: Force-rank all criteria before assigning weights. The top criterion must receive at least 2x the weight of the bottom criterion. If you cannot force-rank them, that disagreement is the real conversation to have.
**Fix**: Force-rank criteria first. Top criterion gets at least 2× weight of bottom criterion.
### Scoring without evidence
**What it looks like**: Team assigns scores in 60 seconds per criterion without citing data.
**Why wrong**: Scores become opinions wearing the costume of analysis.
**Do instead**: Require a one-sentence evidence statement for any score above 7 or below 4. The statement must cite a specific data point, past project, or validated assumption. "We believe" does not qualify as evidence.
**Fix**: Require a one-sentence evidence statement for any score above 7 or below 4. "I scored technical feasibility 3 because we have zero experience with Kubernetes and our last infrastructure project overran by 3x" is valid.
### Ignoring the default option
**What it looks like**: Matrix compares Option A vs. Option B but does not include "do nothing."
**Why wrong**: Every decision has a status quo. If the status quo scores higher than the options, that is the most important finding in the analysis.
**Do instead**: Always add "do nothing / stay the course" as an explicit option in the matrix before scoring. If the status quo wins, that result must be surfaced to the decision-maker — it is the most valuable finding the analysis can produce.
**Fix**: Always include "do nothing / stay the course" as an explicit option.
---
## Detection Commands Reference
These are for use in structured workshops to spot rationalization in real time:
```bash
# Check if "do nothing" was included as an explicit option
# (manual review — check option list at top of decision doc)
# Check if criteria weights sum to 100%
python3 -c "weights=[30,25,20,15,10]; print('OK' if sum(weights)==100 else f'ERROR: sums to {sum(weights)}')"
# Check weighted scores for each option
python3 -c "
weights=[0.30,0.25,0.20,0.15,0.10]
option_a=[8,5,4,6,4]
score = sum(w*s for w,s in zip(weights,option_a))
print(f'Option A weighted: {score:.2f}')
"
```
---
## See Also
- `strategic-frameworks.md` — Porter's Five Forces, SWOT scoring, OKR alignment for market context
- `skills/research/decision-helper/SKILL.md` — lightweight decision support for technical choices
references/estimation-techniques.md
---
title: Estimation Techniques — Cone of Uncertainty, Delphi Method, Reference Class Ratios, Planning Fallacy Mitigation
domain: project-evaluation
level: 3
skill: project-evaluation
---
# Estimation Techniques Reference
> **Scope**: Advanced estimation techniques for project evaluation — cone of uncertainty calibration, Delphi aggregation, reference class forecasting with empirical ratios by project type, decomposition patterns, and planning fallacy mitigation. Extends roi-frameworks.md (which covers three-point estimation and T-shirt sizing basics) with deeper methodology and calibration data.
> **Version range**: Framework-agnostic — applicable to software projects, content initiatives, and infrastructure work. Empirical ratios below are from software engineering research; validate against your team's historical actuals.
> **Generated**: 2026-04-09 — planning fallacy research is stable; specific overrun ratios should be replaced with your team's historical data as soon as you have 5+ comparable projects.
---
## Overview
Estimation is a skill that degrades without deliberate practice. Teams that never compare estimates to actuals never improve. Teams that estimate only at the start of projects and never update mid-project confuse commitment with prediction. The techniques here address the two most common estimation failures: the planning fallacy (systematic underestimation of effort and time) and scope uncertainty (not knowing what you're estimating). The cone of uncertainty makes scope uncertainty explicit. Reference class forecasting combats the planning fallacy with data. The Delphi method aggregates expert estimates without anchoring bias. Together they produce estimates that are calibrated — honest about what they know and what they don't.
---
## Cone of Uncertainty
The cone of uncertainty describes how estimation accuracy changes over project lifecycle. From original NASA and software engineering research.
### Cone Calibration by Project Phase
| Phase | Cone Width (high end) | Cone Width (low end) | Typical Accuracy Achievable |
|-------|-----------------------|----------------------|-----------------------------|
| Concept / Initial idea | 4× | 0.25× | "This is a 40-160 hour project" |
| Requirements defined | 2× | 0.5× | "This is a 60-120 hour project" |
| Architecture designed | 1.5× | 0.67× | "This is a 70-105 hour project" |
| Detailed design complete | 1.25× | 0.8× | "This is a 75-95 hour project" |
| Coding underway (50%) | 1.1× | 0.9× | "This is an 81-89 hour project" |
**What this means in practice**: At the concept stage, providing a point estimate (e.g., "this will take 80 hours") is an illusion of precision. The only honest estimate at concept stage is a range: "this could take anywhere from 40 to 320 hours depending on what we find during requirements." Stakeholders who demand point estimates at concept stage are asking for false confidence.
### Cone of Uncertainty Template
```
## Estimation: [Project Name]
## Current project phase: [Concept / Requirements / Architecture / Design / Coding]
### Current phase cone multipliers
High estimate multiplier: ___× (from table above)
Low estimate multiplier: ___× (from table above)
### Base estimate (from three-point or T-shirt sizing in roi-frameworks.md)
Point estimate: ___ hours
### Cone range
High end: ___ hours × ___× = ___ hours
Low end: ___ hours × ___× = ___ hours
### Communication to stakeholders
"Based on current information, this project will take between ___ and ___ hours.
We are at [phase], so the range will narrow to ___ to ___ hours once [next phase deliverable] is complete.
Requesting budget commitment up to the high end until design is complete."
### Phase-gate commitment schedule
| Phase Gate | Date | Commit to | Range Expected |
|-----------|------|-----------|----------------|
| Architecture done | ___ | Budget range | ___×-___× |
| Detailed design done | ___ | Final estimate | ±25% |
| Coding starts | ___ | Hard deadline | ±10% |
```
---
## Reference Class Forecasting with Empirical Ratios
Standard three-point estimation (in roi-frameworks.md) asks estimators to imagine the pessimistic scenario from their own experience. Reference class forecasting bypasses imagination and uses actual historical data from comparable projects.
### Empirical Overrun Ratios by Project Type
From software engineering research (Kahneman/Lovallo, Flyvbjerg, and industry surveys):
| Project Type | Median Planned Hours | Median Actual Hours | Typical Overrun Ratio | Range |
|-------------|---------------------|--------------------|-----------------------|-------|
| New feature, known tech stack | 100 | 130 | **1.3×** | 1.1-1.8× |
| New feature, partially known stack | 100 | 160 | **1.6×** | 1.2-2.5× |
| Integration with external API/vendor | 100 | 200 | **2.0×** | 1.3-4.0× |
| Data migration | 100 | 250 | **2.5×** | 1.5-6.0× |
| Security / compliance implementation | 100 | 220 | **2.2×** | 1.4-5.0× |
| Major refactor / rewrite | 100 | 300 | **3.0×** | 1.5-8.0× |
| New product / greenfield | 100 | 350 | **3.5×** | 2.0-10× |
| Infrastructure migration | 100 | 280 | **2.8×** | 1.5-7.0× |
| UI/UX redesign | 100 | 200 | **2.0×** | 1.3-4.0× |
| Performance optimization | 100 | 180 | **1.8×** | 1.2-4.0× |
**Critical note**: These are industry medians. Your team's actual ratios will differ. Build your own reference class table after completing 5+ comparable projects. Your historical data is more accurate than industry medians for your specific team.
### Reference Class Forecasting Template
```
## Reference Class Forecast: [Project Name]
## Project type: [from table above]
### Step 1: Classify project type
Match to closest project type from table:
Project type: ___
Industry median overrun ratio: ___×
### Step 2: Find comparable past projects (your team's data)
If 3+ comparable projects exist in team history, use these instead of industry median.
| Past Project | Category | Planned Hours | Actual Hours | Overrun Ratio |
|-------------|----------|---------------|--------------|---------------|
| ___ | ___ | ___ | ___ | ___× |
| ___ | ___ | ___ | ___ | ___× |
| ___ | ___ | ___ | ___ | ___× |
| Team median overrun ratio | | | | **___×** |
Use team median if N ≥ 3. Fall back to industry median if N < 3.
### Step 3: Apply reference class adjustment
Bottom-up estimate (from three-point or detailed breakdown): ___ hours
Reference class ratio applied: × ___
Reference class forecast: ___ hours
### Step 4: Inside view adjustment (optional, bounded)
Is this project materially different from the reference class? If so, adjust ratio ±20% max.
Adjustment reason: ___
Adjusted ratio: ___×
Final forecast: ___ hours
### Honest communication
Bottom-up estimate: ___ hours
Reference class forecast: ___ hours
Recommended budget: ___ hours (use reference class forecast, not bottom-up)
```
---
## Delphi Method for Team Estimation
The Delphi method aggregates estimates from multiple people without anchoring. Use when a project involves significant uncertainty and has multiple people who could estimate it.
### Standard Delphi Process
```
## Delphi Estimation Session: [Project Name]
### Prerequisites
- 3-5 estimators with relevant domain knowledge
- Written project description available (no verbal briefing that could anchor)
- Estimators work independently — no pre-session discussion of estimates
### Round 1: Individual estimates (silent)
1. Distribute project description to all estimators
2. Each estimator provides: low estimate, most likely estimate, high estimate, and 1-2 assumptions
3. Collect all estimates before anyone sees others' estimates
Round 1 results:
| Estimator | Low | Most Likely | High | Key Assumption |
|-----------|-----|------------|------|----------------|
| A | ___ | ___ | ___ | ___ |
| B | ___ | ___ | ___ | ___ |
| C | ___ | ___ | ___ | ___ |
### Round 2: Reveal and discuss outliers
1. Share all estimates anonymously
2. Focus discussion ONLY on extreme outliers (highest and lowest)
3. Goal: understand which assumptions caused divergence, not to converge by social pressure
4. No one is required to change their estimate
Discussion record:
- Highest estimate assumed: [what extra scope or risk did this estimator see?]
- Lowest estimate assumed: [what simplification did this estimator assume?]
- Key disagreement: [the specific assumption that caused spread]
### Round 3: Revised estimates (with new information)
Each estimator may revise, or maintain their Round 1 estimate.
Round 3 results:
| Estimator | Revised Most Likely | Change from Round 1 |
|-----------|--------------------|--------------------|
| A | ___ | +/- ___% |
| B | ___ | +/- ___% |
| C | ___ | +/- ___% |
### Aggregation
Method 1 (simple): Average of Round 3 most-likely estimates: ___ hours
Method 2 (PERT weighted): (Low_avg + 4 × ML_avg + High_avg) / 6 = ___ hours
Use method 2 when range is wide (High > 2× Low).
### Residual disagreement
If estimators still diverge by > 50% after Round 3:
- The disagreement IS the finding — the project is not scoped clearly enough to estimate
- Action required: Clarify scope before estimation proceeds
- Do not average through large disagreements; resolve them
```
---
## Decomposition-Based Estimation
When top-down estimation is unreliable (first time doing this type of project), break the project into components and estimate each independently. Research shows decomposed estimates are 30-40% more accurate than holistic estimates.
### Work Breakdown Structure (WBS) Template
```
## WBS Estimate: [Project Name]
### Decomposition rules
- No component should be larger than 40 hours (if it is, decompose further)
- Estimate each component with the person most likely to do the work
- Include explicit non-coding components (design, testing, documentation, review)
### WBS
| Component | Sub-component | Responsible | Low (hrs) | ML (hrs) | High (hrs) | PERT |
|-----------|--------------|-------------|-----------|----------|------------|------|
| Design | UI mockups | ___ | ___ | ___ | ___ | calc |
| Design | API contract | ___ | ___ | ___ | ___ | calc |
| Backend | Data model | ___ | ___ | ___ | ___ | calc |
| Backend | API endpoints | ___ | ___ | ___ | ___ | calc |
| Backend | Integration X | ___ | ___ | ___ | ___ | calc |
| Frontend | Component A | ___ | ___ | ___ | ___ | calc |
| Frontend | Component B | ___ | ___ | ___ | ___ | calc |
| Testing | Unit tests | ___ | ___ | ___ | ___ | calc |
| Testing | Integration tests | ___ | ___ | ___ | ___ | calc |
| Testing | QA pass | ___ | ___ | ___ | ___ | calc |
| Documentation | Runbook | ___ | ___ | ___ | ___ | calc |
| Review | Code review | ___ | ___ | ___ | ___ | calc |
| Review | Security review | ___ | ___ | ___ | ___ | calc |
| Overhead | Coordination | ___ | ___ | ___ | ___ | calc |
| **TOTAL** | | | **___** | **___** | **___** | **___** |
PERT per component: (Low + 4×ML + High) / 6
```
**Overhead budget**: Add 15-20% of total PERT estimate for coordination, unplanned interruptions, and administrative overhead. This is separate from the PERT range.
```python
# PERT calculator for WBS components
components = [
("UI mockups", 4, 8, 16),
("API contract", 2, 4, 8),
("Data model", 4, 8, 20),
# (name, low, most_likely, high)
]
total_pert = 0
for name, lo, ml, hi in components:
pert = (lo + 4*ml + hi) / 6
total_pert += pert
print(f"{name}: {pert:.1f} hours")
print(f"\nTotal PERT: {total_pert:.0f} hours")
print(f"With 15% overhead: {total_pert * 1.15:.0f} hours")
print(f"Reference class check: multiply by your team's overrun ratio")
```
---
## Planning Fallacy Mitigation Techniques
The planning fallacy is systematic. It cannot be eliminated by "trying harder." These techniques structurally counteract it.
### Mitigation Technique Reference
| Technique | When to Apply | Effort Required | Effectiveness |
|-----------|--------------|-----------------|---------------|
| Reference class forecasting | All projects M or larger | Low | High — single highest-impact technique |
| Delphi estimation | Projects with 3+ potential estimators | Medium | High — eliminates anchoring and social convergence |
| WBS decomposition | Any project with unclear scope | High | High — decomposition reduces blind spots |
| Pre-mortem on estimate | Any project before commitment | Low | Medium — catches obvious optimistic assumptions |
| Separate optimistic/pessimistic estimators | Projects with schedule pressure | Low | Medium — explicit adversarial roles |
| History comparison requirement | Before finalizing estimate | Low | High — forces acknowledgement of past performance |
| Commitments after design, not before | Project planning | Low organizational cost, high political cost | Very high — reduces commitment-stage overconfidence |
### Estimation Pre-Mortem Template
Run this on any estimate before it is committed to stakeholders.
```
## Estimation Pre-Mortem: [Project Name]
## Estimate being reviewed: ___ hours / ___ weeks
### Setup
"Assume this project has been delivered and it took 3× longer than estimated.
It is not over budget because of bad luck — the estimate was wrong.
Why was the estimate wrong?"
### Round 1: Reasons estimate might be wrong
| Reason | Probability (Low/Med/High) | Hours Impact | Detection Signal |
|--------|---------------------------|--------------|-----------------|
| Scope expanded during implementation | ___ | +___ hrs | Scope change requests within first 2 weeks |
| Integration complexity higher than expected | ___ | +___ hrs | First integration takes 2× as long as estimated |
| Team member unavailable (illness, competing priority) | ___ | +___ hrs | Any team member goes part-time in first month |
| Key technical assumption proved wrong | ___ | +___ hrs | First proof-of-concept takes longer than 1 day |
| External dependency delayed | ___ | +___ hrs | Dependency misses first checkpoint |
| Quality issues require rework | ___ | +___ hrs | First code review requires >10% rewrite |
### Round 2: Adjustments
Does any High-probability row change the estimate?
- [ ] Yes — revised estimate: ___ hours
- [ ] No — estimate stands with risk mitigations noted
### Risk mitigations added
[List specific actions added to the project plan based on the pre-mortem]
```
---
## Estimation Calibration Tracking
Improve over time by tracking actuals against estimates.
```
## Team Estimation Calibration Log
| Project | Date | Type | Initial Estimate (hrs) | Actual (hrs) | Ratio | Notes |
|---------|------|------|----------------------|--------------|-------|-------|
| ___ | ___ | ___ | ___ | ___ | ___× | ___ |
| ___ | ___ | ___ | ___ | ___ | ___× | ___ |
### Quarterly calibration review
- Average overrun ratio: ___×
- Worst overrun: ___× (project: ___)
- Best accuracy: within ___% (project: ___)
- Estimation technique used for most accurate projects: ___
- Most common source of underestimation: ___
### Calibration update
Apply this ratio to future estimates of the same type:
Team ratio (integration work): ___×
Team ratio (new features): ___×
Team ratio (greenfield): ___×
```
**Target**: After 10 projects, your team's calibration ratio should be between 1.1× and 1.4×. If it is consistently above 1.5×, systematic structural changes are needed (scoping discipline, WBS decomposition requirement, or reference class forecasting mandate).
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Point estimate at concept stage
**What it looks like**: "We'll budget 200 hours for this project." — said before requirements exist.
**Why wrong**: At concept stage, cone of uncertainty is 4×. A 200-hour estimate could legitimately be anywhere from 50 to 800 hours. Committing to 200 hours is not an estimate — it is a wish.
**Do instead**: Provide a range at concept stage: "Based on similar projects, this is 100-400 hours. We need requirements before we can narrow it." If stakeholders demand a point estimate, give them the reference class upper end as the planning figure, not the midpoint.
**Fix**: At concept stage, provide a range. "Based on similar projects, this could be 100-400 hours. We need requirements defined before we can narrow this." If stakeholders insist on a point estimate, provide the reference class upper end as the planning figure.
### Averaging through expert disagreement
**What it looks like**: Three estimators give 40, 200, and 350 hours. Team averages to 197 and moves on.
**Why wrong**: 5× divergence between estimators means they are estimating different projects. Averaging does not resolve the disagreement — it obscures it. The 350-hour estimator saw something the 40-hour estimator did not.
**Do instead**: Before averaging Delphi estimates, require all estimates in the same round to be within 2x of each other. If they are not, stop and facilitate a scope clarification discussion. The divergence is the signal worth investigating.
**Fix**: Before averaging Delphi estimates, require that estimates within the same round are within 2× of each other. If they aren't, facilitate a scope clarification discussion first.
### Historical data ignored because "this project is different"
**What it looks like**: Team has a 2.5× overrun history on integration projects. New integration project is estimated without reference class adjustment because "this integration is simpler."
**Why wrong**: "This one is different" is the most common entry point for planning fallacy. The outside view exists precisely to counteract the inside view's over-confidence in project uniqueness.
**Do instead**: Apply reference class adjustment before finalizing any estimate. Name the 3 most similar past projects and their actual outcomes. Reductions to the historical ratio are capped at 20% and require written justification.
**Fix**: Reference class adjustment is mandatory before finalizing any estimate. If the estimator believes the project warrants an adjustment, they may reduce the reference class ratio by up to 20% — with a written justification. More than 20% reduction requires review by a second estimator.
### Estimation accuracy measured at delivery
**What it looks like**: Team compares initial estimate to final delivery date — and it's always wrong. No intermediate comparison, no component-level tracking.
**Why wrong**: By measuring only at the end, you cannot improve estimation technique. You don't know whether the estimate was wrong because of scope change, wrong assumptions, or poor technique.
**Do instead**: Capture estimates at three points: initial (concept), post-design, and actual. Compare initial-to-post-design separately from post-design-to-actual. Each comparison diagnoses a different failure mode and requires a different correction.
**Fix**: Track estimates at three points: initial estimate, post-design estimate, and actual. Compare initial-to-post-design (scope discovery accuracy) separately from post-design-to-actual (execution accuracy). They require different improvements.
---
## Detection Commands Reference
```bash
# Calculate PERT estimate for a set of components
python3 -c "
components = [
('Component A', 4, 8, 16),
('Component B', 8, 20, 40),
('Component C', 2, 4, 12),
# (name, optimistic, most_likely, pessimistic)
]
total_pert = 0
total_variance = 0
for name, o, m, p in components:
pert = (o + 4*m + p) / 6
sigma = (p - o) / 6
total_pert += pert
total_variance += sigma**2
print(f'{name}: {pert:.1f}h (σ={sigma:.1f})')
import math
sigma_total = math.sqrt(total_variance)
print(f'Total PERT: {total_pert:.0f}h')
print(f'80% confidence range: {total_pert - sigma_total:.0f}-{total_pert + sigma_total:.0f}h')
print(f'95% confidence range: {total_pert - 2*sigma_total:.0f}-{total_pert + 2*sigma_total:.0f}h')
"
# Reference class ratio calculator from your team's history
python3 -c "
history = [
('Project A', 100, 150), # (name, planned, actual)
('Project B', 80, 200),
('Project C', 120, 130),
]
ratios = [actual/planned for _, planned, actual in history]
avg_ratio = sum(ratios) / len(ratios)
print(f'Overrun ratios: {[f\"{r:.2f}x\" for r in ratios]}')
print(f'Average ratio: {avg_ratio:.2f}x')
print(f'For a 100-hour estimate, budget: {100 * avg_ratio:.0f} hours')
"
# Cone of uncertainty range calculator
python3 -c "
phases = {
'concept': (4.0, 0.25),
'requirements': (2.0, 0.5),
'architecture': (1.5, 0.67),
'design': (1.25, 0.8),
'coding_50pct': (1.1, 0.9),
}
estimate = 100 # replace with your estimate
phase = 'requirements' # replace with current phase
high_mult, low_mult = phases[phase]
print(f'Phase: {phase}')
print(f'Point estimate: {estimate}h')
print(f'Range: {estimate * low_mult:.0f}h - {estimate * high_mult:.0f}h')
"
```
---
## See Also
- `roi-frameworks.md` — T-shirt sizing, basic three-point estimation, story points, and ROI calculation
- `feasibility-scoring.md` — confidence-adjusted feasibility assessment using estimation data
- `skills/strategic-decision/references/risk-assessment.md` — probability-weighted outcomes when estimate uncertainty is high
references/feasibility-scoring.md
---
title: Feasibility Scoring — Models, Criteria, Weights, and Go/No-Go Frameworks
domain: project-evaluation
level: 3
skill: project-evaluation
---
# Feasibility Scoring Reference
> **Scope**: Feasibility scoring models with concrete criteria and weights, go/no-go decision frameworks, and risk-adjusted feasibility assessment. Use when evaluating whether a project should start, not when estimating effort (see roi-frameworks.md).
> **Version range**: Framework-agnostic — applies to software projects, content initiatives, business experiments, and infrastructure work.
> **Generated**: 2026-04-09 — validate scoring weights against your organization's actual constraint profile.
---
## Overview
Feasibility scoring fails when it is vague ("feasibility: high") or when it conflates feasibility with desirability. A project can be highly feasible (you can definitely build it) and deeply undesirable (nobody will use it). Separate these dimensions explicitly. The scoring models here produce a structured verdict across three independent dimensions: technical, resource, and market feasibility. Each dimension gets an independent confidence rating so you know where your uncertainty is concentrated.
---
## Three-Dimension Feasibility Scoring Model
Score each dimension independently on a 1-10 scale. Do not average them — each dimension is a separate gate.
### Dimension 1: Technical Feasibility
| Factor | Score (1-10) | Evidence / Notes |
|--------|-------------|-----------------|
| All required technologies exist and are proven | ___ | |
| Team has experience with core technical challenges | ___ | |
| Hardest technical problem has a known solution path | ___ | |
| Integration with existing systems is understood | ___ | |
| No unsolved research/novel engineering required | ___ | |
| Security and compliance approach is clear | ___ | |
| **Technical Feasibility Score (average)** | **___** | |
| **Confidence** | H / M / L | |
**Score interpretation**:
- 8-10: High technical feasibility. No material unknowns.
- 6-7: Medium. 1-2 technical risks exist; mitigations are possible.
- 4-5: Borderline. Significant unknowns; recommend spike/prototype first.
- 1-3: Low feasibility. Novel engineering or missing capabilities required.
### Dimension 2: Resource Feasibility
| Factor | Score (1-10) | Evidence / Notes |
|--------|-------------|-----------------|
| Team has required skills (or can hire/contract quickly) | ___ | |
| Timeline is achievable given current capacity | ___ | |
| Budget is allocated or approvable | ___ | |
| Dependencies are available when needed | ___ | |
| Key team members have capacity (not already overcommitted) | ___ | |
| **Resource Feasibility Score (average)** | **___** | |
| **Confidence** | H / M / L | |
### Dimension 3: Market / Demand Feasibility
| Factor | Score (1-10) | Evidence / Notes |
|--------|-------------|-----------------|
| Target audience has validated demand (not just assumed) | ___ | |
| Distribution channel to reach audience exists | ___ | |
| Timing is right (market ready; not too early, not too late) | ___ | |
| Success criteria can be measured within the timeline | ___ | |
| **Market Feasibility Score (average)** | **___** | |
| **Confidence** | H / M / L | |
---
## Feasibility Summary Table
| Dimension | Score (1-10) | Confidence | Gate Result |
|-----------|-------------|------------|-------------|
| Technical | ___ | H/M/L | PASS / FLAG / BLOCK |
| Resource | ___ | H/M/L | PASS / FLAG / BLOCK |
| Market | ___ | H/M/L | PASS / FLAG / BLOCK |
**Gate thresholds**:
- PASS: Score ≥ 7 with High or Medium confidence
- FLAG: Score 5-6, OR score ≥ 7 with Low confidence (must resolve uncertainty before committing)
- BLOCK: Score < 5 on any dimension — do not proceed until addressed
---
## Go / No-Go Decision Tree
```
START: All three dimensions scored
Is any dimension BLOCK (score < 5)?
├── YES → NO-GO
│ Specify: which dimension failed, what must change to re-evaluate
└── NO → Continue
Is any dimension FLAG?
├── YES → What is the flag type?
│ ├── Score 5-6 → GO WITH CONDITIONS
│ │ Specify: what must be validated before full commit
│ └── Score ≥ 7 but Low confidence → GO WITH CONDITIONS
│ Specify: what information would raise confidence to Medium+
└── NO → Continue
Are all three dimensions PASS?
├── YES → Does ROI justify proceeding?
│ ├── YES → GO
│ └── NO → DEFER (feasible but not worth it now)
└── NO → unreachable (handled above)
```
---
## Verdict Definitions
| Verdict | Criteria | Required Output |
|---------|----------|-----------------|
| **GO** | All 3 dimensions PASS, positive ROI | First action within 48 hours |
| **GO WITH CONDITIONS** | Any FLAG dimension | Named conditions + owner + deadline |
| **DEFER** | All PASS but ROI insufficient vs. alternatives | Trigger condition for re-evaluation |
| **NO-GO** | Any BLOCK dimension | Root cause + what must change |
| **SPIKE FIRST** | Technical dimension FLAG with Low confidence | 1-2 week prototype to resolve unknowns |
---
## Risk-Adjusted Feasibility Model
Use when a dimension has Medium or Low confidence. Adjusts the effective score for risk.
```
## Risk Adjustment Formula
Risk-adjusted score = Raw score × Confidence multiplier
Confidence multipliers:
- High confidence: 1.0 (no adjustment)
- Medium confidence: 0.85 (15% haircut)
- Low confidence: 0.65 (35% haircut)
Example:
Technical score: 7.5 (raw)
Confidence: Low
Risk-adjusted score: 7.5 × 0.65 = 4.9 → changes PASS to BLOCK
```
**Worked example** (new SaaS feature — async video review):
| Dimension | Raw | Confidence | Multiplier | Adjusted | Gate |
|-----------|-----|------------|------------|----------|------|
| Technical | 8.2 | High | 1.0 | 8.2 | PASS |
| Resource | 6.0 | Medium | 0.85 | 5.1 | FLAG |
| Market | 7.5 | Low | 0.65 | 4.9 | BLOCK |
Pre-adjustment verdict: GO. Post-adjustment verdict: NO-GO pending market validation.
Action: Run a 2-week user interview sprint to validate demand before committing engineers. Market score must reach 6.0+ with Medium confidence before re-evaluating.
---
## Rapid Feasibility Check (< 30 minutes)
For low-stakes decisions where full scoring would be overkill.
```
## Rapid Feasibility (5 questions)
1. Can this be built with technology the team already knows?
YES / NO / UNSURE
2. Does the team have the time to do this without dropping other commitments?
YES / NO / UNSURE
3. Is there evidence that people want this (conversations, search data, existing requests)?
YES / NO / UNSURE
4. Can this be completed in the available timeline?
YES / NO / UNSURE
5. If we started this and it failed, what is the worst-case impact?
LOW (embarrassment) / MEDIUM (wasted weeks) / HIGH (significant cost or damage)
Verdict:
- 4-5 YES answers + LOW/MEDIUM worst-case → GO
- 3 YES answers OR 1+ UNSURE on questions 1-2 → SPIKE or CONDITIONS
- 2 or fewer YES answers → NO-GO or DEFER
- HIGH worst-case → full scoring model required before deciding
```
---
## Confidence Level Calibration
Use these anchors to score confidence consistently:
| Confidence | What It Means | Evidence Examples |
|------------|---------------|------------------|
| **High** | You have directly validated this claim | Customer interviews, historical data from similar projects, working prototype |
| **Medium** | You have indirect evidence or analogies | Industry benchmarks, team experience with similar work, positive conversations |
| **Low** | You are reasoning from first principles | No comparable data, novel market, new technology, team has never done this type of work |
**Warning**: Teams consistently rate confidence too high. If you have not spoken directly to potential users in the last 30 days, market feasibility confidence is at most Medium. If the technical approach has never been tested by your team, technical confidence is at most Medium.
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Feasibility conflated with desirability
**What it looks like**: "We can build it, so it's feasible." Team with strong engineering capability scores technical feasibility 9 for everything, ignoring whether the market wants it.
**Why wrong**: Feasibility and desirability are orthogonal. The most common startup failure is a product that was technically feasible and built on time, that nobody wanted.
**Do instead**: Score all three dimensions separately: technical feasibility, market feasibility, and organizational feasibility. A high technical score does not compensate for a market feasibility score below 5. Treat them as independent gates, not an average.
**Fix**: Require three separate dimension scores. A project that is technically feasible but has market feasibility below 5 is a NO-GO, regardless of engineering capability.
### Feasibility assessed at vision scope, not MVP scope
**What it looks like**: "We can't build a real-time collaboration tool — that's too hard." But the MVP is just shared document state with 5-second polling, which is straightforward.
**Why wrong**: Feasibility must be assessed against the specific, scoped version of the project — the MVP — not the aspirational full version.
**Do instead**: Freeze the project definition at MVP scope before scoring begins. Score what is actually being built in the first phase. Re-score when scope changes, because scope changes change feasibility.
**Fix**: Before scoring, confirm the project definition is frozen at MVP scope (from Phase 1 of project-evaluation SKILL.md). Re-score when scope changes.
### High confidence on everything
**What it looks like**: All three dimensions rated High confidence. Team is decisive and agrees.
**Why wrong**: If you have not validated assumptions through direct evidence, high confidence is overconfidence. Teams in alignment often reinforce each other's assumptions rather than challenging them.
**Detection**: Ask "what evidence do we have that this is true?" for each High confidence rating. If the answer is "we believe" or "we've discussed," confidence is actually Medium.
**Do instead**: Treat team agreement as a prompt for scrutiny, not a signal of correctness. For each High confidence rating, name the specific evidence. If no direct evidence exists, lower the rating to Medium and identify the validation step needed to earn High confidence.
---
## See Also
- `roi-frameworks.md` — ROI calculation templates, effort estimation, risk-adjusted NPV
- `skills/project-evaluation/SKILL.md` — full project evaluation workflow with RICE scoring
references/finance.md
# Finance & Accounting
Umbrella skill for finance and accounting workflows. Detects the domain from the user's request, loads the right reference files, and executes. All output is working material for qualified professionals — not financial advice.
**Scope**: Journal entries, reconciliation, variance analysis, financial statements, audit/SOX support, month-end close. Use csuite for strategic finance decisions, data-analysis for ad-hoc analytics.
---
## Mode Detection
Classify into exactly one mode before proceeding.
| Mode | Signal Phrases | Reference to Load |
|------|---------------|-------------------|
| **JOURNAL ENTRY** | book, accrue, accrual, depreciation, prepaid, payroll entry, revenue recognition, deferred revenue, journal entry | `references/finance/journal-entries.md` |
| **RECONCILIATION** | reconcile, bank rec, subledger, GL-to-sub, intercompany, reconciling items, aging | `references/finance/reconciliation.md` |
| **VARIANCE** | variance, flux, budget vs actual, period-over-period, price/volume, waterfall, bridge | `references/finance/variance-analysis.md` |
| **STATEMENTS** | P&L, income statement, balance sheet, cash flow, financial statements, GAAP presentation | `references/finance/financial-statements.md` |
| **AUDIT/SOX** | SOX, control testing, sample selection, workpaper, deficiency, material weakness, ITGC, audit | `references/finance/journal-entries.md` + `references/finance/reconciliation.md` |
| **CLOSE** | month-end close, close calendar, close checklist, close day, hard close, soft close | `references/finance/reconciliation.md` + `references/finance/journal-entries.md` |
Always load `references/finance/llm-finance-failure-modes.md` as a guard rail regardless of mode.
---
## Workflow
### Phase 1: CLASSIFY
1. Detect mode from user request
2. Load the corresponding reference file(s)
3. Load `references/finance/llm-finance-failure-modes.md`
4. Confirm scope with user if ambiguous
### Phase 2: GATHER
Collect the data needed for the task:
| Mode | Required Inputs |
|------|----------------|
| JOURNAL ENTRY | Entry type, period, account codes, amounts or source data |
| RECONCILIATION | GL balance, comparison source (bank statement, subledger, counterparty), period |
| VARIANCE | Current period data, comparison period data, materiality thresholds |
| STATEMENTS | Trial balance or financial data, comparison periods, presentation preferences |
| AUDIT/SOX | Control area, testing period, population data, prior results |
| CLOSE | Close calendar dates, task ownership, current status |
If the user provides partial data, ask for the minimum additional data needed. Do not fabricate account codes, balances, or transaction details.
### Phase 3: EXECUTE
Follow the mode-specific workflow from the loaded reference file. Key constraints apply across all modes:
**Debit/Credit Rules (always enforce):**
| Account Type | Normal Balance | To Increase | To Decrease |
|-------------|---------------|-------------|-------------|
| Asset | Debit | Debit | Credit |
| Liability | Credit | Credit | Debit |
| Equity | Credit | Credit | Debit |
| Revenue | Credit | Credit | Debit |
| Expense | Debit | Debit | Credit |
| Contra Asset | Credit | Credit | Debit |
| Contra Revenue | Debit | Debit | Credit |
**Verification gates (every journal entry):**
- Debits = Credits (balanced entry, no exceptions)
- Account codes come from user data, never invented
- Amounts traced to source calculations
- Period is explicit and correct
- Reversal flag set for accruals
**Verification gates (every reconciliation):**
- Both sides reconcile to the same adjusted balance
- Every reconciling item is categorized (timing, adjustment, investigation)
- Items aged >60 days flagged for escalation
**Verification gates (every variance):**
- Decomposition components sum to total variance (verify arithmetic)
- Materiality threshold applied before narrative generation
- Narratives are causal (why), not circular ("revenue was higher due to higher revenue")
### Phase 4: VALIDATE
Before presenting output:
1. **Arithmetic check**: Verify all calculations. Re-derive totals from components. Because LLMs miscalculate, re-check every sum, difference, and percentage (see failure modes reference).
2. **Completeness check**: All required sections present per the reference file template
3. **Anti-hallucination check**: No fabricated account numbers, GAAP citations, or standards references. If uncertain about a specific GAAP rule, say so.
4. **Consistency check**: Treatment matches prior period (if prior period data provided)
### Phase 5: DELIVER
Present output in the format specified by the reference file for the mode. Include:
- The working artifact (entry, reconciliation, analysis, statement)
- Supporting calculations
- Items flagged for professional review
- Suggested next steps
---
## LLM Failure Modes in Finance
See `references/finance/llm-finance-failure-modes.md` for the complete failure mode catalog (calculation errors, fabricated standards, materiality misapplication, period errors, account code fabrication). Universal failure modes (hallucination, overconfidence, generic output) are in `skills/shared-patterns/llm-domain-failure-modes-base.md`.
---
## Failure Modes by Mode
### Journal Entry Failure Modes
| Failure Mode | Why It Fails |
|-------------|-------------|
| Unbalanced entry presented as complete | Violates fundamental accounting equation |
| Missing reversal flag on accruals | Creates double-counting in the next period |
| Round-number estimates without calculation basis | Signals fabrication, fails audit |
| Booking to "Miscellaneous Expense" | Lacks specificity for variance analysis and audit trail |
| Same person prepares and approves | Violates segregation of duties |
### Reconciliation Failure Modes
| Failure Mode | Why It Fails |
|-------------|-------------|
| Forcing the rec to balance by plugging a number | Hides real differences that may indicate errors or fraud |
| Carrying items forward indefinitely without investigation | Stale items may represent losses or control failures |
| Reconciling to an unverified source | Both sides must come from authoritative sources |
| Skipping intercompany for "immaterial" entities | Intercompany must eliminate to zero in consolidation |
### Variance Analysis Failure Modes
| Failure Mode | Why It Fails |
|-------------|-------------|
| Circular narrative ("revenue is higher because revenue increased") | No causal explanation |
| "Timing" without specifying what shifted and when it normalizes | Uninvestigable claim |
| "Various small items" for a material variance | Must decompose until below materiality |
| Ignoring offsetting variances | Net favorable can hide serious unfavorable components |
| Missing outlook (one-time vs recurring) | Variance without context is useless for forecasting |
### Financial Statement Failure Modes
| Failure Mode | Why It Fails |
|-------------|-------------|
| Assets not equal to liabilities + equity | Balance sheet does not balance |
| Mixing functional and nature expense classification | GAAP requires consistency within a statement |
| Omitting non-cash items in cash flow reconciliation | Indirect method requires all non-cash adjustments |
| Current/non-current misclassification | Debt maturing within 12 months must be current |
references/finance/financial-statements.md
# Financial Statements Reference
Working reference for income statement, balance sheet, and cash flow statement generation with GAAP presentation rules and period-end adjustments.
---
## Income Statement (ASC 220 / IAS 1)
### Multi-Column Format
```
INCOME STATEMENT
Period: [Description]
(in thousands)
Current Prior Var ($) Var (%) Budget Bud Var
-------- -------- -------- ------- -------- --------
REVENUE
Product revenue $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
Service revenue $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
Other revenue $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
-------- -------- -------- -------- --------
TOTAL REVENUE $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
COST OF REVENUE $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
-------- -------- -------- -------- --------
GROSS PROFIT $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
Gross Margin XX.X% XX.X%
OPERATING EXPENSES
Research & development $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
Sales & marketing $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
General & administrative $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
-------- -------- -------- -------- --------
TOTAL OPERATING EXPENSES $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
OPERATING INCOME (LOSS) $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
Operating Margin XX.X% XX.X%
OTHER INCOME (EXPENSE)
Interest income $XX,XXX $XX,XXX $X,XXX X.X%
Interest expense ($XX,XXX) ($XX,XXX) $X,XXX X.X%
Other, net $XX,XXX $XX,XXX $X,XXX X.X%
-------- -------- --------
INCOME BEFORE TAXES $XX,XXX $XX,XXX $X,XXX X.X%
Income tax expense $XX,XXX $XX,XXX $X,XXX X.X%
-------- -------- --------
NET INCOME (LOSS) $XX,XXX $XX,XXX $X,XXX X.X% $XX,XXX $X,XXX
Net Margin XX.X% XX.X%
```
### GAAP Presentation Rules — Income Statement
| Rule | Requirement |
|------|-------------|
| Expense classification | By function (COGS, R&D, S&M, G&A) or by nature. Function is standard for US companies. Must be consistent |
| Nature disclosure | If classified by function, disclose depreciation, amortization, and employee benefit costs by nature in notes |
| Operating vs non-operating | Present separately |
| Income tax | Separate line |
| Extraordinary items | Prohibited (both US GAAP and IFRS) |
| Discontinued operations | Separate, net of tax |
| Revenue disaggregation (ASC 606) | Disaggregate by nature, amount, timing, and uncertainty factors |
| Stock-based compensation | Classify within functional expense categories; disclose total SBC in notes |
| Restructuring charges | Present separately if material, or within OpEx with note disclosure |
| Non-GAAP measures | If presented, clearly label and reconcile to GAAP |
### Key Metrics
```
Revenue growth (%) X.X%
Gross margin (%) XX.X%
Operating margin (%) XX.X%
Net margin (%) XX.X%
OpEx as % of revenue XX.X%
Effective tax rate (%) XX.X%
```
---
## Balance Sheet (ASC 210 / IAS 1)
### Standard Format
```
BALANCE SHEET
As of [Date]
(in thousands)
ASSETS
Current Assets
Cash and cash equivalents $XX,XXX
Short-term investments $XX,XXX
Accounts receivable, net $XX,XXX
Inventory $XX,XXX
Prepaid expenses and other current assets $XX,XXX
Total Current Assets $XX,XXX
Non-Current Assets
Property and equipment, net $XX,XXX
Operating lease right-of-use assets $XX,XXX
Goodwill $XX,XXX
Intangible assets, net $XX,XXX
Long-term investments $XX,XXX
Other non-current assets $XX,XXX
Total Non-Current Assets $XX,XXX
TOTAL ASSETS $XX,XXX
LIABILITIES AND STOCKHOLDERS' EQUITY
Current Liabilities
Accounts payable $XX,XXX
Accrued liabilities $XX,XXX
Deferred revenue, current $XX,XXX
Current portion of long-term debt $XX,XXX
Operating lease liabilities, current $XX,XXX
Other current liabilities $XX,XXX
Total Current Liabilities $XX,XXX
Non-Current Liabilities
Long-term debt $XX,XXX
Deferred revenue, non-current $XX,XXX
Operating lease liabilities, non-current $XX,XXX
Other non-current liabilities $XX,XXX
Total Non-Current Liabilities $XX,XXX
Total Liabilities $XX,XXX
Stockholders' Equity
Common stock $XX,XXX
Additional paid-in capital $XX,XXX
Retained earnings (accumulated deficit) $XX,XXX
Accumulated other comprehensive income (loss)$XX,XXX
Treasury stock ($XX,XXX)
Total Stockholders' Equity $XX,XXX
TOTAL LIABILITIES AND STOCKHOLDERS' EQUITY $XX,XXX
```
### GAAP Presentation Rules — Balance Sheet
| Rule | Requirement |
|------|-------------|
| Current vs non-current | Distinguish explicitly |
| Current definition | Realized, consumed, or settled within 12 months (or operating cycle if longer) |
| Asset ordering | Most liquid first (US standard) |
| Accounts receivable | Net of allowance for credit losses (ASC 326) |
| PP&E | Net of accumulated depreciation |
| Goodwill | Not amortized — annual impairment test (ASC 350) |
| Leases (ASC 842) | Recognize ROU assets and lease liabilities for both operating and finance leases |
| Debt reclassification | Long-term debt maturing within 12 months reclassified to current |
| Fundamental equation | Assets = Liabilities + Stockholders' Equity (must balance) |
---
## Cash Flow Statement (ASC 230 / IAS 7)
### Indirect Method Format
```
CASH FLOWS FROM OPERATING ACTIVITIES
Net income (loss) $XX,XXX
Adjustments to reconcile to net cash from operations:
Depreciation and amortization $XX,XXX
Stock-based compensation $XX,XXX
Amortization of debt issuance costs $XX,XXX
Deferred income taxes $XX,XXX
Loss (gain) on disposal of assets $XX,XXX
Impairment charges $XX,XXX
Other non-cash items $XX,XXX
Changes in operating assets and liabilities:
Accounts receivable ($XX,XXX)
Inventory ($XX,XXX)
Prepaid expenses and other assets ($XX,XXX)
Accounts payable $XX,XXX
Accrued liabilities $XX,XXX
Deferred revenue $XX,XXX
Other liabilities $XX,XXX
Net Cash Provided by (Used in) Operating Activities $XX,XXX
CASH FLOWS FROM INVESTING ACTIVITIES
Purchases of property and equipment ($XX,XXX)
Purchases of investments ($XX,XXX)
Proceeds from sale/maturity of investments $XX,XXX
Acquisitions, net of cash acquired ($XX,XXX)
Other investing activities $XX,XXX
Net Cash Provided by (Used in) Investing Activities ($XX,XXX)
CASH FLOWS FROM FINANCING ACTIVITIES
Proceeds from issuance of debt $XX,XXX
Repayment of debt ($XX,XXX)
Proceeds from issuance of common stock $XX,XXX
Repurchases of common stock ($XX,XXX)
Dividends paid ($XX,XXX)
Payment of debt issuance costs ($XX,XXX)
Other financing activities $XX,XXX
Net Cash Provided by (Used in) Financing Activities ($XX,XXX)
Effect of exchange rate changes on cash $XX,XXX
Net Increase (Decrease) in Cash $XX,XXX
Cash and cash equivalents, beginning of period $XX,XXX
Cash and cash equivalents, end of period $XX,XXX
```
### GAAP Presentation Rules — Cash Flow
| Rule | Requirement |
|------|-------------|
| Method | Indirect most common (start with net income, adjust). Direct permitted but rare |
| Required disclosures | Interest paid and income taxes paid (face or notes) |
| Non-cash activities | Disclosed separately (lease assets, stock-for-acquisition, etc.) |
| Cash equivalents | Original maturity ≤ 3 months, highly liquid |
| Operating changes sign convention | Increase in asset = use of cash (negative). Increase in liability = source of cash (positive) |
| Verification | Beginning cash + net change = ending cash. Cross-check to balance sheet |
### Working Capital Changes — Sign Convention
| Change | Cash Flow Effect | Sign |
|--------|-----------------|------|
| AR increases | Cash used (sold but not collected) | Negative |
| AR decreases | Cash provided (collected) | Positive |
| Inventory increases | Cash used (purchased) | Negative |
| Inventory decreases | Cash provided (sold from existing) | Positive |
| AP increases | Cash provided (purchased but not paid) | Positive |
| AP decreases | Cash used (paid down) | Negative |
| Deferred revenue increases | Cash provided (collected in advance) | Positive |
| Deferred revenue decreases | Cash used (recognized previously collected) | Negative |
---
## Period-End Adjustments
### Required Adjustments
| Adjustment | Description | Accounts Affected |
|-----------|-------------|-------------------|
| Accruals | Expenses incurred, not paid | Expense / Accrued liabilities |
| Deferrals | Prepaid expenses, deferred revenue | Expense or Revenue / Prepaid or Deferred |
| Depreciation/amortization | Periodic allocation from schedules | D&A expense / Accumulated D&A |
| Bad debt provision | Adjust allowance per aging and loss rates | Bad debt expense / Allowance for credit losses |
| Inventory adjustments | Write-downs for obsolete/impaired | COGS or loss / Inventory |
| FX revaluation | Revalue foreign currency monetary items | FX gain/loss / Asset or Liability |
| Tax provision | Current and deferred income tax | Tax expense / Tax payable, Deferred tax |
| Fair value | Mark-to-market investments, derivatives | Gain/loss or OCI / Asset or Liability |
### Reclassifications
| Reclassification | Purpose |
|-----------------|---------|
| Current/non-current | Reclassify debt maturing within 12 months to current |
| Contra account netting | Net allowances against gross balances |
| Intercompany elimination | Eliminate IC balances in consolidation |
| Discontinued operations | Reclassify to separate line, net of tax |
| Equity method | Record share of investee income/loss |
| Segment | Ensure correct operating segment classification |
---
## Materiality Thresholds for Variance Investigation
| Line Item Size | Dollar Threshold | Percentage Threshold |
|---------------|-----------------|---------------------|
| > $10M | $500K | 5% |
| $1M - $10M | $100K | 10% |
| < $1M | $50K | 15% |
Investigate when either threshold is exceeded. Flag for professional review.
---
## Comparison Periods
| Report Type | Comparison Columns |
|------------|-------------------|
| Monthly | Prior month, prior year same month, budget |
| Quarterly | Prior quarter, prior year same quarter, budget |
| Annual | Prior year, budget |
| YTD | Prior year YTD, budget YTD |
Always include both dollar and percentage variances. Express margin changes in basis points (1 bp = 0.01%).
references/finance/journal-entries.md
# Journal Entries Reference
Working reference for preparing journal entries with proper debits, credits, supporting documentation, and review workflows.
---
## Debit/Credit Rules
| Account Type | Normal Balance | Debit Does | Credit Does |
|-------------|---------------|------------|-------------|
| Asset | Debit | Increase | Decrease |
| Contra Asset (e.g., Accum. Depreciation) | Credit | Decrease | Increase |
| Liability | Credit | Decrease | Increase |
| Equity | Credit | Decrease | Increase |
| Revenue | Credit | Decrease | Increase |
| Contra Revenue (e.g., Sales Returns) | Debit | Increase | Decrease |
| Expense | Debit | Increase | Decrease |
**Fundamental constraint**: Every journal entry must balance. Sum of debits = sum of credits. No exceptions.
---
## Standard Accrual Categories
### 1. Accounts Payable Accruals
Goods/services received but not yet invoiced at period end.
**Entry:**
- Debit: Expense account (or asset if capitalizable)
- Credit: Accrued liabilities
**Sources for amount:**
- Open POs with confirmed receipts
- Contracts with services rendered, not yet billed
- Recurring vendor arrangements (utilities, professional services, subscriptions)
- Employee expense reports submitted but unprocessed
**Controls:**
- Auto-reverse in the following period
- Consistent estimation methodology period over period
- Document basis: PO amount, contract terms, or historical run-rate
- Track actual vs accrual to calibrate future estimates
### 2. Fixed Asset Depreciation
Periodic depreciation/amortization for tangible and intangible assets.
**Entry:**
- Debit: Depreciation/amortization expense (by department/cost center)
- Credit: Accumulated depreciation/amortization
**Methods:**
| Method | Formula | Use Case |
|--------|---------|----------|
| Straight-line | (Cost - Salvage) / Useful life | Default for financial reporting |
| Declining balance | Rate x Net book value | Accelerated; tax reporting |
| Units of production | (Cost - Salvage) x (Actual units / Total expected units) | Usage-based assets |
**Controls:**
- Source from fixed asset register or depreciation schedule
- Verify new additions have correct useful life and method
- Check for disposals or impairments requiring write-off
- Track book vs tax depreciation separately
### 3. Prepaid Expense Amortization
Amortize prepaid expenses over their benefit period.
**Entry:**
- Debit: Expense account (insurance, software, rent, etc.)
- Credit: Prepaid expense
**Common categories:**
| Category | Typical Term | Amortization Basis |
|----------|-------------|-------------------|
| Insurance premiums | 12 months | Straight-line monthly |
| Software licenses | 12-36 months | Per contract term |
| Prepaid rent | Per lease term | Monthly |
| Maintenance contracts | Per contract | Straight-line |
| Conference deposits | Event date | Expense at event or upon forfeiture |
**Controls:**
- Maintain amortization schedule with start/end dates and monthly amounts
- Review for immaterial items that should be expensed immediately
- Check for cancelled contracts requiring accelerated amortization
- Add new prepaids to schedule promptly
### 4. Payroll Accruals
Accrue compensation and related costs earned but not yet paid.
**Entries:**
| Component | Debit | Credit |
|-----------|-------|--------|
| Salary accrual (partial pay period) | Salary expense by dept | Accrued payroll |
| Bonus accrual | Bonus expense by dept | Accrued bonus |
| Benefits | Benefits expense | Accrued benefits |
| Employer payroll taxes | Payroll tax expense | Accrued payroll taxes |
**Calculation basis:**
- Salary: Working days in period / total days in pay period x gross pay
- Bonus: Plan terms (target x performance factor x proration)
- Benefits: Employer share of health, retirement match, PTO liability
- Taxes: FICA (6.2% SS up to wage base + 1.45% Medicare), FUTA, SUTA
### 5. Revenue Recognition (ASC 606)
Five-step framework:
| Step | Action |
|------|--------|
| 1. Identify the contract | Agreement with commercial substance, identifiable rights/obligations, payment terms, approval |
| 2. Identify performance obligations | Distinct goods/services promised. Distinct = customer can benefit independently AND separately identifiable |
| 3. Determine transaction price | Fixed + variable consideration (constraint: include only amounts not subject to significant reversal) |
| 4. Allocate to obligations | Standalone selling price for each obligation. Methods: adjusted market, expected cost + margin, residual |
| 5. Recognize upon satisfaction | Point in time (control transfers) or over time (customer simultaneously receives/consumes benefit) |
**Common entries:**
| Scenario | Debit | Credit |
|----------|-------|--------|
| Recognize deferred revenue | Deferred revenue | Revenue |
| Recognize with new receivable | Accounts receivable | Revenue |
| Receive payment in advance | Cash / AR | Deferred revenue |
---
## Entry Format Template
```
Journal Entry: [Type] — [Period]
Prepared by: [Name]
Date: [Period end date]
Reversal: [Yes/No — if yes, reversal date]
| Line | Account Code | Account Name | Debit | Credit | Dept | Memo |
|------|-------------|--------------|-------|--------|------|------|
| 1 | [Code] | [Name] | X,XXX | | [D] | [Detail] |
| 2 | [Code] | [Name] | | X,XXX | [D] | [Detail] |
| | | **Total** | X,XXX | X,XXX | | |
Supporting Detail:
- Calculation basis and assumptions
- Reference to source documentation
- Comparison to prior period entry (if available)
```
---
## Supporting Documentation Requirements
Every entry must include:
| Requirement | Purpose |
|------------|---------|
| Entry description/memo | Audit trail — what and why |
| Calculation support | How amounts were derived |
| Source documents | PO numbers, invoice numbers, contract refs, payroll register |
| Period | Accounting period the entry applies to |
| Preparer ID | Who prepared and when |
| Approval evidence | Review per authorization matrix |
| Reversal indicator | Whether entry auto-reverses and reversal date |
---
## Approval Matrix
| Entry Type | Threshold | Approver |
|-----------|-----------|----------|
| Standard recurring | Any | Accounting manager |
| Non-recurring manual | < $50K | Accounting manager |
| Non-recurring manual | $50K-$250K | Controller |
| Non-recurring manual | > $250K | CFO / VP Finance |
| Top-side / consolidation | Any | Controller+ |
| Out-of-period adjustments | Any | Controller+ |
*Thresholds are illustrative. Set based on organization's materiality.*
---
## Review Checklist
Before approving any entry:
- [ ] Debits = credits (balanced)
- [ ] Correct accounting period
- [ ] Account codes valid and appropriate
- [ ] Amounts mathematically accurate with supporting calculations
- [ ] Description clear, specific, audit-sufficient
- [ ] Department/cost center coding correct
- [ ] Consistent with prior period treatment
- [ ] Auto-reversal set for accruals
- [ ] Supporting documentation complete and referenced
- [ ] Within preparer's authority level
- [ ] No duplicate of existing entry
- [ ] Unusual amounts explained
---
## Common Errors
| Error | Detection | Risk |
|-------|-----------|------|
| Unbalanced entry | Sum check: debits != credits | Misstatement |
| Wrong period | Entry date vs period end | Cut-off error |
| Wrong sign | Debit as credit or vice versa | Double the intended impact |
| Duplicate | Same transaction recorded twice | Overstatement |
| Wrong account | Similar account codes transposed | Misclassification |
| Missing reversal | Accrual not set to auto-reverse | Double-counting next period |
| Stale accrual | Recurring accrual not updated for changed circumstances | Inaccurate estimate |
| Round-number estimate | $100,000 exactly without calculation basis | Audit flag for fabrication |
| Wrong FX rate | Foreign currency at incorrect rate or date | Translation error |
| Missing intercompany elimination | One-sided intercompany entry | Consolidation error |
| Capitalization error | Expense capitalized or capital item expensed | Asset/expense misstatement |
| Cut-off error | Recorded in wrong period based on delivery/service date | Revenue/expense timing |
---
## Intercompany Journal Entries
When booking intercompany transactions:
1. Both entities must record their side of the transaction
2. Use designated intercompany accounts (receivable/payable pairs)
3. Amounts must match exactly (same currency, same date, same FX rate)
4. Elimination entries must zero out the intercompany balances in consolidation
5. Document the business purpose for each intercompany transaction
**Common intercompany types:**
- Management fees / shared services allocations
- Intercompany sales of goods or services
- Intercompany loans and interest
- Cost recharges and royalties
- Dividend declarations between entities
references/finance/llm-finance-failure-modes.md
# LLM Finance Failure Modes
Where LLMs fail in finance work. Load this reference for every finance task as a guard rail. Each failure mode includes the pattern, why it happens, and the mitigation.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers finance-specific failures only.
---
## Category 1: Arithmetic Errors
LLMs do not compute. They predict tokens that look like computation. Every number in LLM-generated financial output must be verified.
### Wrong Sums
**Pattern**: Three line items ($127K, $89K, $214K) totaled as $420K. Actual: $430K.
**Why**: The model generates a plausible-looking total rather than computing one. The error is usually small enough to look correct on casual inspection.
**Mitigation**: Re-derive every total from its components. Do not trust any sum, subtotal, or grand total without explicit verification. When possible, use a script or calculator for arithmetic.
### Percentage Errors
**Pattern**: $28K variance on $500K base reported as "4.6%". Actual: 5.6%.
**Why**: The model may divide by the wrong base, transpose digits, or round incorrectly.
**Mitigation**: Recompute every percentage: variance / base x 100. Verify the base is correct (budget, prior period, actual — whichever was specified). State the formula used.
### Sign Errors
**Pattern**: Favorable revenue variance presented as unfavorable, or vice versa.
**Why**: The model confuses the direction convention. For revenue, actual > budget = favorable. For expenses, actual > budget = unfavorable. The model sometimes applies the wrong convention.
**Mitigation**: Apply the sign convention table explicitly:
| Line Type | Actual > Comparison | Actual < Comparison |
|-----------|-------------------|-------------------|
| Revenue / Income | Favorable (+) | Unfavorable (-) |
| Expense / COGS | Unfavorable (-) | Favorable (+) |
### Rounding Cascade
**Pattern**: Individual items rounded, then summed. Rounded sum differs from sum-then-round.
**Why**: Premature rounding introduces cumulative error. Three items at $33.33K each: rounded individually = $33K + $33K + $33K = $99K. Actual sum = $100K.
**Mitigation**: Carry full precision through calculations. Round only at the final presentation step.
### Debit/Credit Imbalance
**Pattern**: Entry presented as "balanced" with $150K debit and $145K credit.
**Why**: The model generates the entry line by line and does not verify the sum constraint at the end.
**Mitigation**: After generating any journal entry, verify: sum(all debits) == sum(all credits). This is an absolute constraint — no tolerance, no rounding exception.
---
## Category 2: Fabricated Standards and Citations
LLMs generate confident-sounding references to accounting standards that do not exist.
### Invented ASC/IFRS Paragraph Numbers
**Pattern**: "Per ASC 842-30-55-12, the lessee must..." when no such paragraph exists, or the paragraph says something different.
**Why**: The model has seen many ASC citations in training data and generates plausible-looking ones. The topic-level code (e.g., ASC 842 for leases) is usually correct. The sub-paragraph is often wrong.
**Mitigation**:
- Cite at the topic level only: "Per ASC 842 (Leases)..."
- Never cite specific paragraph numbers unless verified against the actual codification
- When uncertain: "Consult the relevant ASC guidance for [topic]"
- Flag all standard citations for professional verification
### Misapplied Rules
**Pattern**: Applying lease accounting (ASC 842) rules to a service contract, or revenue recognition (ASC 606) rules to an investment.
**Why**: The model pattern-matches on surface features (contract with payments over time) without verifying scope applicability.
**Mitigation**: State the standard being applied and its scope. Let the professional verify that the standard applies to the specific transaction.
### Conflated GAAP and IFRS
**Pattern**: Mixing US GAAP and IFRS rules in the same analysis without distinguishing them.
**Why**: Training data includes both frameworks. The model does not maintain framework boundaries.
**Mitigation**:
- Default to US GAAP unless the user specifies IFRS
- State which framework is being applied
- Never blend rules from different frameworks without explicit callout
### Superseded Guidance
**Pattern**: Referencing pre-ASC 606 revenue recognition rules (e.g., SAB 101/104) or pre-ASC 842 lease rules (FAS 13).
**Why**: Training data includes historical guidance that has been superseded.
**Mitigation**: Flag when guidance may have been superseded. Key supersession dates:
| Old Standard | Replaced By | Effective |
|-------------|------------|-----------|
| SAB 101/104, SOP 97-2 | ASC 606 (Revenue) | 2018 |
| FAS 13 | ASC 842 (Leases) | 2019 |
| FAS 5 (loss contingencies) | ASC 326 (CECL for credit losses) | 2020 |
| FAS 141R | ASC 805 (Business Combinations) | 2009 |
---
## Category 3: Hallucinated Account Codes and Chart of Accounts
### Fabricated Account Numbers
**Pattern**: LLM generates "4100 — Product Revenue" or "6200 — Salaries" when no chart of accounts was provided.
**Why**: These are common account codes in training data. The model generates plausible defaults rather than acknowledging it lacks the information.
**Mitigation**:
- Never generate account codes unless provided by the user
- Use descriptive placeholders: `[Revenue Account]`, `[Salary Expense Account]`
- If the user provides a partial chart, use only the codes given and flag gaps
### Assumed Chart Structure
**Pattern**: LLM assumes a specific numbering convention (1xxx = Assets, 2xxx = Liabilities, etc.) that may not match the user's system.
**Why**: This convention is common but not universal. Many organizations use different numbering schemes.
**Mitigation**: Ask for the user's chart of accounts or use descriptive names rather than assumed codes.
---
## Category 4: Materiality and Judgment Errors
### No Materiality Threshold Applied
**Pattern**: Investigating a $500 variance on a $50M revenue line. Generating 200-word narratives for immaterial items.
**Why**: The model treats all variances as equally important. It has no inherent sense of materiality.
**Mitigation**: Apply materiality thresholds before investigation. If no thresholds are provided, ask for them. Suggest defaults based on organization size.
### Wrong Materiality Benchmark
**Pattern**: Using net income as the materiality base for a balance sheet assertion, or total assets as the base for an income statement item.
**Why**: The model does not consistently match the benchmark to the financial statement element.
**Mitigation**:
| Financial Statement | Typical Benchmark |
|--------------------|-------------------|
| Income statement | Revenue or pre-tax income |
| Balance sheet | Total assets or equity |
| Cash flow | Cash from operations |
### Inconsistent Materiality Application
**Pattern**: Investigating a 3% variance on one line while ignoring an 8% variance on another of similar magnitude.
**Why**: The model processes items sequentially and may lose track of the threshold applied to earlier items.
**Mitigation**: Apply thresholds to ALL line items in a single pass. Present results in a table showing which items exceed thresholds and which do not.
---
## Category 5: Period and Temporal Errors
### Wrong Period Allocation
**Pattern**: December services booked as January expense because the invoice date is in January.
**Why**: The model may confuse invoice date with service date. Under accrual accounting, the expense period is when the service was received, not when it was invoiced or paid.
**Mitigation**: Always determine the service/delivery date, not the invoice/payment date. Accrue in the period the obligation was incurred.
### Ignoring Accrual Basis
**Pattern**: Recording expense only when cash is paid. Recognizing revenue only when cash is received.
**Why**: Cash basis is simpler and the model defaults to the simpler approach.
**Mitigation**: Unless explicitly told otherwise, assume accrual basis. Revenue recognized when earned (performance obligation satisfied). Expense recognized when incurred (goods/services received).
### Partial Period Errors
**Pattern**: Full-month depreciation on an asset placed in service on the 15th.
**Why**: The model defaults to the simplest calculation rather than applying the entity's partial-period convention.
**Mitigation**: Ask for the entity's convention: half-month, mid-quarter, exact days. Apply it consistently.
### Incorrect Comparison Periods
**Pattern**: Comparing Q4 actual to Q3 actual and calling it "year-over-year."
**Why**: Terminology confusion. YoY = same period in the prior year, not prior sequential period.
**Mitigation**: Use precise labels:
| Term | Meaning |
|------|---------|
| YoY | Same period, prior year (Q4 2025 vs Q4 2024) |
| QoQ | Sequential quarter (Q4 vs Q3) |
| MoM | Sequential month (Dec vs Nov) |
| YTD | Year-to-date (Jan through current month) |
---
## Category 6: Structural Accounting Errors
### Balance Sheet Does Not Balance
**Pattern**: LLM generates a balance sheet where Assets != Liabilities + Equity.
**Why**: The model generates each section independently and does not verify the fundamental equation.
**Mitigation**: After generating any balance sheet, verify: Total Assets = Total Liabilities + Total Stockholders' Equity. This is an absolute constraint.
### Cash Flow Does Not Reconcile
**Pattern**: Beginning cash + net change != ending cash. Or ending cash != balance sheet cash.
**Why**: Same issue — sections generated independently without cross-verification.
**Mitigation**: Verify:
1. Beginning cash + net cash change = ending cash
2. Ending cash = balance sheet cash and cash equivalents
### Intercompany Does Not Eliminate
**Pattern**: Intercompany receivables and payables that do not net to zero in consolidation.
**Why**: The model may create one side of an intercompany entry without the corresponding elimination.
**Mitigation**: For every intercompany balance, verify that the offsetting balance exists in the counterparty entity and that the elimination entry zeroes both out.
---
## Category 7: Overconfidence and Missing Caveats
### Presenting Estimates as Facts
**Pattern**: "The accrual should be $47,382" without stating assumptions or ranges.
**Why**: The model presents a single point estimate with false precision rather than acknowledging uncertainty.
**Mitigation**: For any estimated amount, state:
- Calculation methodology
- Key assumptions
- Sensitivity to assumption changes
- Whether the estimate requires professional review
### Omitting "Consult Professional" Disclaimers
**Pattern**: Providing definitive tax advice, legal interpretations, or audit opinions.
**Why**: The model optimizes for helpfulness over appropriate caution.
**Mitigation**: All financial output is working material for professionals. Never present output as:
- Tax advice
- Legal interpretation
- Audit opinion
- Authoritative GAAP interpretation
Always include: "Review by qualified financial professional required before [posting/filing/reporting]."
---
## Quick Reference: Verification Checklist
Run after every finance output:
| Check | Action |
|-------|--------|
| All sums verified | Re-derive totals from components |
| All percentages verified | Recompute: value / base x 100 |
| Signs correct | Apply favorable/unfavorable convention |
| Debits = credits | Sum check on every journal entry |
| A = L + E | Verify on every balance sheet |
| Cash reconciles | Beginning + change = ending |
| No fabricated codes | Account codes from user data only |
| No fabricated standards | ASC/IFRS citations at topic level only |
| Materiality applied | Thresholds set before investigation |
| Period correct | Expense in period incurred, revenue in period earned |
| Caveats present | Professional review disclaimer included |
references/finance/reconciliation.md
# Reconciliation Reference
Working reference for account reconciliation: GL-to-subledger, bank reconciliations, intercompany, aging analysis, and reconciling item classification.
---
## Reconciliation Types
### GL-to-Subledger Reconciliation
Compare the general ledger control account balance to the detailed subledger balance.
**Common accounts:**
| Control Account | Subledger Source |
|----------------|-----------------|
| Accounts receivable | AR aging report |
| Accounts payable | AP aging report |
| Fixed assets | Fixed asset register |
| Inventory | Inventory valuation report |
| Prepaid expenses | Prepaid amortization schedule |
| Accrued liabilities | Accrual detail schedules |
**Process:**
1. Pull GL balance for the control account at period end
2. Pull subledger trial balance at the same date
3. Compare totals — should match if posting is real-time
4. Investigate differences
**Common causes of GL-to-sub differences:**
| Cause | Direction | Resolution |
|-------|-----------|------------|
| Manual JE to control account, not reflected in subledger | GL differs from sub | Post corresponding subledger entry or reclassify |
| Subledger transactions not interfaced to GL | Sub differs from GL | Run interface or post manual JE |
| Batch posting timing | Either | Wait for batch completion; document as timing |
| Reclassification in GL without subledger adjustment | GL differs from sub | Post subledger reclassification |
| System interface error / failed posting | Either | Diagnose and reprocess; escalate to IT |
### Bank Reconciliation
Compare GL cash balance to bank statement balance.
**Process:**
1. Obtain bank statement balance at period end
2. Pull GL cash account balance at same date
3. Identify outstanding checks (issued, not cleared)
4. Identify deposits in transit (recorded in GL, not credited by bank)
5. Identify bank charges/interest/adjustments not yet in GL
6. Reconcile both sides to adjusted balance
**Standard format:**
```
Balance per bank statement: $XX,XXX
Add: Deposits in transit $X,XXX
Less: Outstanding checks ($X,XXX)
Add/Less: Bank errors $X,XXX
Adjusted bank balance: $XX,XXX
Balance per general ledger: $XX,XXX
Add: Interest/credits not recorded $X,XXX
Less: Bank fees not recorded ($X,XXX)
Add/Less: GL errors $X,XXX
Adjusted GL balance: $XX,XXX
Difference: $0.00
```
**Constraint**: Adjusted bank balance must equal adjusted GL balance. A nonzero difference means the reconciliation is incomplete.
### Intercompany Reconciliation
Reconcile balances between related entities to ensure elimination to zero on consolidation.
**Process:**
1. Pull IC receivable/payable balances for each entity pair
2. Compare Entity A's receivable from B to Entity B's payable to A
3. Identify and resolve differences
4. Confirm all IC transactions recorded on both sides
5. Verify elimination entries are correct
**Common causes of IC differences:**
| Cause | Resolution |
|-------|------------|
| Timing: one entity recorded, other has not | Confirm and post on the late side |
| Different FX rates used by each entity | Agree on rate source and date; adjust |
| Misclassification (IC vs third-party) | Reclassify to correct account |
| Disputed amounts / unapplied payments | Resolve dispute; apply payment |
| Different cut-off practices across entities | Standardize cut-off procedures |
---
## Reconciling Item Classification
### Category 1: Timing Differences
Items that will self-clear without action within the normal processing cycle (1-5 business days).
| Item | Description | Action |
|------|------------|--------|
| Outstanding checks | Issued and in GL, pending bank clearance | Monitor |
| Deposits in transit | In GL, pending bank credit | Monitor |
| In-transit transactions | Posted in one system, pending interface | Monitor |
| Pending approvals | Awaiting approval to post | Monitor |
No adjusting entry needed.
### Category 2: Adjustments Required
Items requiring a journal entry to correct.
| Item | Description | Action |
|------|------------|--------|
| Unrecorded bank charges | Fees, wire charges, returned item fees | Book JE |
| Unrecorded interest | Interest income or expense | Book JE |
| Recording errors | Wrong amount, wrong account, duplicate | Correcting JE |
| Missing entries | Transaction in one system, no counterpart | Book missing entry |
| Classification errors | Correct amount, wrong account | Reclassify |
### Category 3: Requires Investigation
Items that cannot be immediately explained.
| Item | Description | Action |
|------|------------|--------|
| Unidentified differences | No obvious cause | Root cause analysis |
| Disputed items | Contested between parties | Escalate to resolution |
| Aged outstanding items | Beyond expected clearance window | Supervisor review |
| Recurring unexplained differences | Same type each period | Process investigation |
---
## Aging Analysis
### Age Buckets and Escalation
| Age | Status | Action |
|-----|--------|--------|
| 0-30 days | Current | Monitor — within normal processing cycle |
| 31-60 days | Aging | Investigate — why has item not cleared? |
| 61-90 days | Overdue | Escalate to supervisor, document investigation |
| 90+ days | Stale | Escalate to management — potential write-off or adjustment |
### Aging Report Format
```
| Item # | Description | Amount | Date Originated | Age (Days) | Category | Status | Owner |
|--------|-------------|--------|-----------------|------------|----------|--------|-------|
```
### Trending Analysis
Track reconciling item totals over time:
- Compare total outstanding items to prior period
- Flag if total reconciling items exceed materiality threshold
- Flag if item count is growing period over period
- Identify recurring items (may indicate underlying process failure)
---
## Escalation Thresholds
| Trigger | Example Threshold | Escalation |
|---------|-------------------|------------|
| Individual item amount | > $10K | Supervisor review |
| Individual item amount | > $50K | Controller review |
| Total reconciling items | > $100K | Controller review |
| Item age | > 60 days | Supervisor follow-up |
| Item age | > 90 days | Controller / management |
| Unreconciled difference | Any amount | Cannot close — must resolve or document |
| Growing trend | 3+ consecutive periods | Process improvement investigation |
*Set thresholds based on organization's materiality level and risk appetite.*
---
## Reconciliation Best Practices
| Practice | Standard |
|----------|----------|
| Timeliness | Complete within close calendar (typically T+3 to T+5) |
| Completeness | All balance sheet accounts on defined frequency (monthly for material, quarterly for immaterial) |
| Documentation | Preparer, reviewer, date, clear explanation of all reconciling items |
| Segregation | Reconciler is not the transaction processor for that account |
| Follow-through | Track open items to resolution — never carry forward indefinitely |
| Root cause | For recurring items, fix the underlying process |
| Standardization | Consistent templates and procedures across all accounts |
| Retention | Per organization's document retention policy |
---
## Reconciliation Template
```
ACCOUNT RECONCILIATION
Account: [Code] — [Name]
Period: [Month/Year]
Prepared by: [Name] Date: [Date]
Reviewed by: [Name] Date: [Date]
GL Balance: $XX,XXX
Subledger/Source Balance: $XX,XXX
Difference: $X,XXX
Reconciling Items:
| # | Description | Amount | Category | Age | Status |
|---|-------------|--------|----------|-----|--------|
| 1 | [Detail] | $X,XXX | Timing | XX | Open |
| 2 | [Detail] | $X,XXX | Adj Req | XX | JE #XX |
Total Reconciling Items: $X,XXX
Adjusted Difference: $0.00
Prior Period Comparison:
- Total reconciling items last period: $X,XXX
- Change: $X,XXX ([direction])
- Items carried forward from prior period: [count]
```
---
## Month-End Close Integration
Reconciliation fits into the close calendar as follows:
| Close Day | Reconciliation Activity |
|-----------|------------------------|
| T+1 | Bank reconciliation (with final bank statement) |
| T+2 | AR and AP subledger reconciliations |
| T+3 | All balance sheet reconciliations, intercompany recs |
| T+3 | Post adjusting JEs identified during reconciliation |
| T+4 | Management review of reconciliation results |
### Close Task Dependencies
```
Level 1 (T+1): Cash entries, bank statement retrieval
↓
Level 2 (T+2): Bank rec, AR/AP subledger recs
↓
Level 3 (T+3): All balance sheet recs, IC rec, adjusting entries
↓
Level 4 (T+4): Preliminary trial balance, draft financials
↓
Level 5 (T+5): Management review, hard close, period lock
```
### Close Metrics
| Metric | Target |
|--------|--------|
| Close duration (period end to hard close) | Reduce over time |
| Adjusting entries after soft close | Minimize |
| Late tasks | Zero |
| Reconciliation exceptions | Reduce over time |
| Post-close corrections | Zero |
---
## Continuous Reconciliation
For organizations targeting a 3-day close, shift reconciliation work into the month:
| Traditional | Continuous |
|------------|-----------|
| Reconcile all items at month-end | Reconcile daily/weekly; month-end is final verification |
| Start from scratch each period | Carry forward the prior rec, update incrementally |
| Investigate items at close | Investigate items as they arise |
| All adjusting JEs at close | Post adjustments throughout the month |
Benefits: shorter close, earlier detection of errors, reduced month-end stress.
references/finance/variance-analysis.md
# Variance Analysis Reference
Working reference for variance decomposition, materiality thresholds, narrative generation, and waterfall analysis.
---
## Decomposition Techniques
### 1. Price / Volume Decomposition
The foundational technique. Applies to any metric expressible as Price x Volume.
**Two-way decomposition:**
```
Total Variance = Actual - Budget (or Prior)
Volume Effect = (Actual Volume - Budget Volume) x Budget Price
Price Effect = (Actual Price - Budget Price) x Actual Volume
Verification: Volume Effect + Price Effect = Total Variance
```
**Three-way decomposition (isolating mix):**
```
Volume Effect = (Actual Volume - Budget Volume) x Budget Price x Budget Mix
Price Effect = (Actual Price - Budget Price) x Budget Volume x Actual Mix
Mix Effect = Budget Price x Budget Volume x (Actual Mix - Budget Mix)
Verification: Volume + Price + Mix = Total Variance
```
**Worked example — Revenue variance:**
| | Units | Price | Revenue |
|---|------|-------|---------|
| Budget | 10,000 | $50 | $500,000 |
| Actual | 11,000 | $48 | $528,000 |
| Variance | +1,000 | -$2 | +$28,000 |
Decomposition:
- Volume effect: +1,000 x $50 = +$50,000 favorable (sold more)
- Price effect: -$2 x 11,000 = -$22,000 unfavorable (lower ASP)
- Net: +$28,000 favorable
- Verify: $50,000 + (-$22,000) = $28,000
### 2. Rate / Mix Decomposition
Used when analyzing blended rates across segments with different unit economics.
```
Rate Effect = Sum_i (Actual Volume_i x (Actual Rate_i - Budget Rate_i))
Mix Effect = Sum_i (Budget Rate_i x (Actual Volume_i - Expected Volume_i at Budget Mix))
```
**Worked example — Gross margin:**
- Product A: 60% margin, Product B: 40% margin
- Budget mix: 50% A, 50% B → Blended margin 50%
- Actual mix: 40% A, 60% B → Blended margin 48%
- Mix effect: 2 percentage points of margin compression
### 3. Headcount / Compensation Decomposition
For payroll and people-cost variances.
```
Total Comp Variance = Actual - Budget Compensation
Components:
1. Headcount variance = (Actual HC - Budget HC) x Budget Avg Comp
2. Rate variance = (Actual Avg Comp - Budget Avg Comp) x Budget HC
3. Mix variance = Level/department mix shift
4. Timing variance = Hiring earlier/later than plan (partial-period)
5. Attrition impact = Unplanned departure savings (net of backfill costs)
```
### 4. Spend Category Decomposition
For operating expenses where price/volume does not apply.
```
Total OpEx Variance = Actual - Budget
Categories:
1. Headcount-driven (salaries, benefits, payroll taxes, recruiting)
2. Volume-driven (hosting, transaction fees, commissions, shipping)
3. Discretionary (travel, events, professional services, marketing programs)
4. Contractual/fixed (rent, insurance, software licenses, subscriptions)
5. One-time (severance, legal settlements, write-offs, project costs)
6. Timing/phasing (spend shifted between periods vs plan)
```
---
## Materiality Thresholds
### Setting Thresholds
Base thresholds on:
| Factor | Approach |
|--------|----------|
| Financial statement materiality | 1-5% of key benchmark (revenue, total assets, net income) |
| Line item size | Larger items warrant lower % thresholds |
| Volatility | More volatile items may need higher thresholds to filter noise |
| Decision relevance | What level would change a management decision? |
### Threshold Framework
| Comparison | Dollar Threshold | % Threshold | Trigger |
|-----------|-----------------|-------------|---------|
| Actual vs Budget | Org-specific | 10% | Either exceeded |
| Actual vs Prior Period | Org-specific | 15% | Either exceeded |
| Actual vs Forecast | Org-specific | 5% | Either exceeded |
| Sequential (MoM) | Org-specific | 20% | Either exceeded |
*Dollar thresholds: typically 0.5%-1% of revenue for income statement items.*
### Investigation Priority
When multiple variances exceed thresholds, prioritize:
| Priority | Criterion |
|----------|-----------|
| 1 | Largest absolute dollar variance |
| 2 | Largest percentage variance |
| 3 | Unexpected direction (opposite to trend) |
| 4 | New variance (was on track, now off) |
| 5 | Cumulative/trending (growing each period) |
---
## Narrative Generation
### Structure
```
[Line Item]: [Favorable/Unfavorable] variance of $[amount] ([percentage]%)
vs [comparison basis] for [period]
Driver: [Primary driver description]
[2-3 sentences: business reason, quantified contributing factors]
Outlook: [One-time / Expected to continue / Improving / Deteriorating]
Action: [None / Monitor / Investigate / Update forecast]
```
### Quality Checklist
Every narrative must be:
| Criterion | Test |
|-----------|------|
| Specific | Names the actual driver, not "higher than expected" |
| Quantified | Dollar and percentage impact of each driver |
| Causal | Explains WHY, not just WHAT |
| Forward-looking | States whether variance continues, normalizes, or worsens |
| Actionable | Identifies required follow-up or decision |
| Concise | 2-4 sentences. No filler |
### Failure Modes (detect and fix these)
| Failure Mode | Why It Fails | Fix |
|-------------|-------------|-----|
| "Revenue was higher due to higher revenue" | Circular — no explanation | Name the driver: volume, pricing, customer, product |
| "Expenses were elevated" | Vague — which expenses? Why? | Specify line item, amount, and cause |
| "Timing" (standalone) | Unverifiable without detail | State what was early/late and when it normalizes |
| "One-time" (standalone) | Unnamed one-time item is suspicious | Name the specific event and amount |
| "Various small items" for material variance | Insufficient decomposition | Break down until each component is below materiality |
| Only largest driver mentioned | Hides offsetting items | Show all material contributing factors |
---
## Waterfall / Bridge Analysis
### Data Structure
```
Starting value: [Base / Budget / Prior period]
Drivers: [Signed amounts for each contributing factor]
Ending value: [Actual / Current period]
Verification: Starting + Sum(Drivers) = Ending
```
### Text Waterfall Format
```
WATERFALL: [Metric] — [Period] [Comparison]
[Comparison] Amount $XX,XXXK
|
|--[+] [Driver 1] +$X,XXXK
|--[+] [Driver 2] +$X,XXXK
|--[-] [Driver 3] -$X,XXXK
|--[-] [Driver 4] -$X,XXXK
|--[+] [Driver 5] +$XXXK
|
[Current Period] Amount $XX,XXXK
Net Variance: +$X,XXXK (+X.X% favorable)
```
### Bridge Reconciliation Table
```
| Driver | Amount | % of Variance | Cumulative |
|--------|--------|---------------|------------|
| [Driver 1] | +$XXK | XX% | +$XXK |
| [Driver 2] | +$XXK | XX% | +$XXK |
| [Driver 3] | -$XXK | -XX% | +$XXK |
| **Total** | **+$XXK** | **100%** | |
```
*Individual driver percentages can exceed 100% when offsetting items exist.*
### Waterfall Best Practices
| Practice | Rationale |
|----------|-----------|
| Order by magnitude (largest positive to largest negative) or by logical business sequence | Readability |
| Limit to 5-8 drivers | Aggregate smaller items into "Other" |
| Verify reconciliation | Start + drivers = end (always) |
| Color code when visual | Green = favorable, red = unfavorable |
| Label each bar | Amount + brief description |
| Include total variance bar | Summary view |
---
## Budget vs Actual vs Forecast
### Three-Way Comparison
```
| Metric | Budget | Forecast | Actual | Bud Var ($) | Bud Var (%) | Fcast Var ($) | Fcast Var (%) |
|--------|--------|----------|--------|-------------|-------------|---------------|---------------|
```
### When to Use Each Comparison
| Comparison | Use Case |
|-----------|----------|
| Actual vs Budget | Annual performance, compensation, board reporting |
| Actual vs Forecast | Operational management, emerging issues |
| Forecast vs Budget | How expectations changed since planning |
| Actual vs Prior Period | Trend analysis, seasonality, new business lines |
| Actual vs Prior Year | YoY growth, seasonality-adjusted comparison |
### Forecast Accuracy Tracking
```
Forecast Accuracy = 1 - |Actual - Forecast| / |Actual|
MAPE = Average of |Actual - Forecast| / |Actual| across periods
```
### Variance Trending Interpretation
| Pattern | Interpretation |
|---------|---------------|
| Consistently favorable | Budget may be too conservative (sandbagging) |
| Consistently unfavorable | Budget too aggressive or execution issues |
| Growing unfavorable | Deteriorating performance or unrealistic targets |
| Shrinking variance | Forecast accuracy improving (normal intra-year pattern) |
| Volatile | Unpredictable business or poor forecasting methodology |
---
## Favorable vs Unfavorable Convention
| Line Item Type | Actual > Budget | Actual < Budget |
|---------------|----------------|----------------|
| Revenue | Favorable | Unfavorable |
| COGS / Expense | Unfavorable | Favorable |
| Gross Profit | Favorable | Unfavorable |
| Net Income | Favorable | Unfavorable |
**Sign convention**: Express favorable as positive, unfavorable as negative. Always label direction explicitly — do not rely on sign alone.
---
## Common Calculation Verification
After producing any variance analysis, verify:
| Check | Method |
|-------|--------|
| Decomposition sums to total | Volume + Price + Mix = Total Variance |
| Percentages are correct | Variance / Base x 100 (not variance / actual) |
| Bridge reconciles | Start + all drivers = End |
| Signs are consistent | Favorable = positive for income items, negative for expense items |
| Basis is stated | "vs Budget" or "vs Prior Year" — never ambiguous |
| Margin changes in basis points | 1 pp = 100 bps. State "X.X pp" or "XXX bps" |
references/hr.md
# HR — People Operations
Umbrella skill for all people operations: recruiting, performance management, compensation, offer drafting, interview design, onboarding, org planning, people analytics, and policy lookup. Each mode loads its own references on demand.
**Scope**: Decisions and artifacts involving people's careers, compensation, and organizational structure. Use csuite for strategic business decisions, data-analysis for general analytics, professional-communication for non-HR business writing.
---
## Mode Detection
Classify the request into exactly one mode. If the request spans modes, choose the primary and note the secondary.
| Mode | Signal Phrases | References |
|------|---------------|------------|
| **RECRUITING** | Pipeline, candidates, sourcing, screening, hiring status, time to fill | `references/hr/recruiting.md` |
| **PERFORMANCE** | Review, self-assessment, calibration, feedback, rating, promotion case | `references/hr/performance-management.md` |
| **COMPENSATION** | Pay, salary, equity, comp bands, benchmarking, offer competitive, retention risk | `references/hr/compensation.md` |
| **OFFER** | Draft offer, offer letter, comp package, signing bonus, start date | `references/hr/compensation.md` |
| **INTERVIEW** | Interview plan, questions, scorecard, evaluation rubric, debrief | `references/hr/recruiting.md` |
| **ONBOARDING** | New hire, first week, 30/60/90, onboarding checklist, buddy | `references/hr/recruiting.md` |
| **ORG-PLANNING** | Headcount, reorg, team structure, span of control, org design | `references/hr/org-planning.md` |
| **PEOPLE-ANALYTICS** | Attrition, headcount report, diversity metrics, org health, flight risk | `references/hr/org-planning.md` |
| **POLICY** | PTO, benefits, leave, expenses, handbook, remote work policy | (no reference — use user-provided policy docs) |
**Always load**: `references/hr/llm-hr-failure-modes.md` — applies to every mode.
---
## Sensitivity Guardrails
HR content touches people's careers, livelihoods, and legal rights. These guardrails are non-negotiable.
| Rule | Rationale |
|------|-----------|
| Source compensation data from user-provided or public databases | Invented market rates cause real pay decisions. State "I don't have current market data" when you don't. |
| Focus recommendations on skills, behaviors, outcomes | "Hire more [group]" or "this candidate fits the culture" introduces bias. Focus on skills, behaviors, outcomes. |
| Include legal review disclaimer on all binding language | Offer letters, policy interpretations, and termination language need legal review. Always state this. |
| Ask for jurisdiction before advising on compliance | Employment law varies by country, state, city. Ask for jurisdiction before advising on compliance, leave, or termination. |
| Minimize PII retention; use role/level identifiers | Names, salaries, SSNs, demographics — minimize retention. Use role/level when names aren't needed. |
| Flag when output needs legal review | Offer letters, PIPs, termination docs, policy changes, accommodation decisions — always flag. |
---
## Workflow
### Mode: RECRUITING
**Framework**: DEFINE → PIPELINE → EVALUATE
**Phase 1: DEFINE** — Clarify role requirements and pipeline structure.
- Define role: title, level, team, location, hiring manager
- Establish pipeline stages: Sourced → Screen → Interview → Debrief → Offer → Accepted
- Set target metrics: time-to-fill, pipeline velocity, source mix
- Design job posting — check against `references/hr/llm-hr-failure-modes.md` for biased language (because gendered/exclusionary language reduces qualified applicant pools by 10-40%)
**Gate**: Role defined. Pipeline stages agreed. Posting reviewed for bias.
**Phase 2: PIPELINE** — Track and manage candidates through stages.
- Report pipeline health: candidates per stage, days in stage, conversion rates
- Flag bottlenecks: stages with >2x average dwell time
- Track source effectiveness: which channels produce hires, not just applicants
**Gate**: Pipeline metrics current. Bottlenecks identified.
**Phase 3: EVALUATE** — Structure interviews and decisions.
- Generate interview plan: 4-6 competencies, behavioral questions per competency, scoring rubric (1-4 scale with anchors)
- Assign panel: map interviewers to competencies, ensure diverse perspectives
- Produce debrief template: structured format, evidence-based, no "gut feel" fields
- Score candidates against rubric, not against each other (because comparative scoring amplifies similarity bias)
**Gate**: Interview kit complete. Debrief structured. Decision evidence-based.
**Onboarding sub-mode** (after offer acceptance):
- Pre-start checklist: accounts, equipment, buddy assignment, welcome email
- Day 1 schedule: orientation, IT setup, team intros, expectations
- Week 1 plan: compliance training, documentation, shadowing, first task
- 30/60/90-day goals: measurable, role-specific milestones
---
### Mode: PERFORMANCE
**Framework**: STRUCTURE → WRITE → CALIBRATE
**Phase 1: STRUCTURE** — Select review type and load template.
| Type | Use When |
|------|----------|
| Self-assessment | Employee writing their own review |
| Manager review | Manager writing review for direct report |
| Calibration prep | Preparing rating distributions for calibration meeting |
**Gate**: Review type selected. Template loaded.
**Phase 2: WRITE** — Generate review content with behavioral specificity.
Self-assessment:
- Key accomplishments: situation, contribution, impact (measurable)
- Goals review: status + evidence per goal
- Growth areas and challenges
- Next-period goals (specific, measurable)
Manager review:
- Overall rating with 2-3 sentence summary
- Strengths with specific behavioral examples (not personality traits)
- Development areas with actionable guidance (not vague directives)
- Goal achievement ratings with observations
- Development plan: skill → current level → target → actions
- Compensation recommendation with justification
**Constraint**: Describe observable behavior with specific examples ("documentation was incomplete on 3 of 5 deliverables") — because personality feedback triggers defensiveness and has no actionable path. See `references/hr/llm-hr-failure-modes.md`.
**Gate**: Review content complete. All feedback behavior-based. Development areas actionable.
**Phase 3: CALIBRATE** — Prepare rating distribution and promotion cases.
- Team overview: employee, role, level, tenure, proposed rating
- Distribution check against targets: Exceeds (~15-20%), Meets (~60-70%), Below (~10-15%)
- Discussion points: borderline cases, role changes, first-at-level reviews
- Promotion candidates: current level, proposed level, evidence of next-level performance
- Compensation actions: promotions, equity refreshes, market adjustments, retention grants
**Constraint**: Present rating targets as guidelines, with flexibility for team context (because forced ranking creates perverse incentives and has been abandoned by most organizations).
**Gate**: Distribution documented. Promotion cases evidence-based. Compensation actions justified.
---
### Mode: COMPENSATION
**Framework**: BENCHMARK → ANALYZE → RECOMMEND
**Phase 1: BENCHMARK** — Establish market data context.
- Identify components: base salary, equity, bonus (target + signing), benefits
- Key variables: role, level, location, company stage, industry
- Data sources: user-provided data, public salary databases, uploaded CSVs
- **Never invent percentile numbers** — if you don't have data, say so explicitly and offer to analyze user-provided data
**Gate**: Components identified. Variables specified. Data sources declared.
**Phase 2: ANALYZE** — Score against market and internal equity.
- Percentile bands: 25th, 50th, 75th, 90th for each component
- Band placement: where each employee falls within their band (below/at/above)
- Internal equity: same-role comparisons, compression detection, tenure-pay correlation
- Outlier detection: employees significantly above/below band midpoints
- Retention risk: below-band + high performer = flight risk
**Constraint**: Always state data vintage and source limitations. "Based on 2024 Levels.fyi data" not "the market rate is $X" — because stale data presented as current causes underpayment or overpayment.
**Gate**: Analysis complete. Sources cited. Limitations stated.
**Phase 3: RECOMMEND** — Deliver actionable compensation recommendations.
- Adjustment recommendations with priority ranking
- Budget impact modeling: total cost of recommended changes
- Equity refresh guidance: vesting cliffs, refresh cadence, retention timing
- Offer structuring: base/equity/bonus mix by company stage and candidate preference
**Gate**: Recommendations prioritized. Budget impact calculated. Sources documented.
**Offer drafting sub-mode:**
- Assemble package: base, equity (shares + vesting schedule + valuation method), bonus (target + signing), benefits summary
- Draft offer letter text — include disclaimer: "This draft requires legal review before sending"
- Negotiation guidance for hiring manager: flexibility ranges, non-monetary levers, walk-away points
- Flag compliance requirements by jurisdiction (at-will language, non-compete enforceability, benefits mandates)
---
### Mode: ORG-PLANNING
**Framework**: ASSESS → MODEL → PLAN
**Phase 1: ASSESS** — Map current organizational state.
| Metric | Healthy Range | Warning Sign |
|--------|---------------|--------------|
| Span of control | 5-8 direct reports | <3 (too narrow) or >12 (too wide) |
| Management layers | 4-6 per 500 people | Excess layers = slow decisions |
| IC-to-manager ratio | 6:1 to 10:1 | <4:1 = top-heavy |
| Team size | 5-9 people | <4 = fragile, >12 = unmanageable |
| Single points of failure | 0 | Any = structural risk |
**Gate**: Current state mapped. Structural issues identified.
**Phase 2: MODEL** — Design target state and transition.
- Headcount modeling: role, level, location, cost, timeline, hiring sequence
- Org design options: functional, product, matrix, pod — with trade-offs per option
- Capacity planning: current capacity vs. planned work, gap analysis
- Sequencing: which hires unlock the most capacity or reduce the most risk
**Constraint**: Never recommend org changes based on individuals ("move Alice because she's difficult") — structure around roles and capabilities (because person-dependent org design creates fragility and masks management problems).
**Gate**: Target state modeled. Sequencing justified. Budget estimated.
**Phase 3: PLAN** — Convert to executable hiring roadmap.
- Phased hiring plan: Q1/Q2/Q3/Q4 with roles, cost, dependencies
- Reporting line changes with communication plan
- Risk mitigation: what happens if key hires take longer than planned
- Success metrics: time to productivity, team health scores, delivery velocity
**Gate**: Roadmap executable. Risks mitigated. Success metrics defined.
**People analytics sub-mode:**
- Headcount reports: by team, location, level, tenure — snapshot and trend
- Attrition analysis: voluntary/involuntary, regrettable/non-regrettable, by team, trend lines
- Diversity metrics: representation by level/team/function, pipeline diversity, promotion rate parity, pay equity
- Engagement indicators: survey scores, eNPS, participation rates, theme analysis
- Flight risk modeling: below-band compensation + low engagement + tenure inflection points
---
### Mode: POLICY
**Framework**: FIND → EXPLAIN → CAVEAT
1. Search user-provided policy documents or handbook
2. Answer in plain language — no legalese
3. Quote specific policy language with source citation
4. Note exceptions and special cases
5. For legal/compliance questions: "Consult HR or legal directly for your specific situation"
**Constraint**: Answer only from user-provided policy documents. State when no source is available — do not guess (because fabricated policy guidance creates liability).
---
## Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| No compensation data | User asks "what should we pay" without data | State limitation explicitly. Offer to analyze user-provided data or recommend public sources (Levels.fyi, Glassdoor, Radford). |
| Biased language in output | LLM generates gendered, ageist, or exclusionary phrasing | Run output against `references/hr/llm-hr-failure-modes.md` bias checklist. Rewrite flagged phrases. |
| Jurisdiction unknown | Legal advice requested without location | Ask for jurisdiction before proceeding. Never default to US employment law. |
| Confidentiality scope unclear | User shares individual comp/performance data | Confirm intended audience. Remind that HR data is need-to-know. Minimize PII in outputs. |
| Template vs. real data confusion | User treats template placeholders as recommendations | Label all templates explicitly: "[PLACEHOLDER — replace with actual data]". |
| Policy not found | User asks about policy with no handbook provided | State clearly: no policy source available. Do not fabricate. Suggest uploading handbook. |
---
## References
| Reference | When to Load | Content |
|-----------|-------------|---------|
| `references/hr/recruiting.md` | RECRUITING, INTERVIEW, ONBOARDING modes | Pipeline stages, velocity metrics, interview frameworks, evaluation rubrics, onboarding checklists |
| `references/hr/performance-management.md` | PERFORMANCE mode | Review structure, calibration methodology, feedback patterns, development planning |
| `references/hr/compensation.md` | COMPENSATION, OFFER modes | Market benchmarking, internal equity analysis, offer structuring, equity modeling |
| `references/hr/org-planning.md` | ORG-PLANNING, PEOPLE-ANALYTICS modes | Headcount modeling, org design principles, capacity planning, people analytics |
| `references/hr/llm-hr-failure-modes.md` | **Every mode** | Bias detection, fabrication risks, compliance gaps, inappropriate language patterns |
references/hr/compensation.md
# Compensation Reference
Market benchmarking, internal equity analysis, offer structuring, and equity modeling.
---
## Total Compensation Components
| Component | Description | Variability | Key Considerations |
|-----------|-------------|-------------|-------------------|
| **Base salary** | Fixed annual cash compensation | Low — changes annually | Primary anchor. Affects bonus %, equity expectations, benefits calculations |
| **Equity** | RSUs, stock options, restricted stock | High — value fluctuates | Vesting schedule, grant cadence, dilution, tax treatment vary widely |
| **Target bonus** | Annual performance-based cash | Medium — 0-200% payout | Percentage of base. Individual + company performance multipliers |
| **Signing bonus** | One-time cash at hire | One-time | Often clawback if departure <12 months. Bridges gap when base is constrained. |
| **Benefits** | Health, dental, vision, 401k match, HSA/FSA | Low — plan-level | Hard to quantify per-person. $15-25K value for US tech companies. |
| **Perks** | WFH stipend, meals, commuter, wellness | Low | Marginal comp. Rarely a deciding factor. |
### Total Comp Calculation
```
Total First-Year Comp = Base + (Equity Year 1 Value) + (Target Bonus * Expected Payout) + Signing Bonus
Total Annual Comp (steady state) = Base + (Annual Equity Value) + (Target Bonus * Expected Payout)
```
**Year 1 is anomalous.** Signing bonus inflates it. Equity cliff (25% vesting at month 12) means zero equity for the first year at many companies. Always present both Year 1 and steady-state.
---
## Market Benchmarking
### Data Sources
| Source | Strengths | Limitations | Freshness |
|--------|-----------|-------------|-----------|
| Levels.fyi | Individual-reported, strong tech coverage, includes equity | Self-selection bias, US-heavy | Rolling, ~6 months lag |
| Glassdoor | Broad coverage across industries | Less granular on equity/bonus | Rolling |
| Radford (Aon) | Survey-based, rigorous methodology, enterprise coverage | Expensive, annual release, 12-18 month lag | Annual |
| Pave/Carta | Real-time, equity-focused, startup coverage | Biased toward funded startups | Real-time |
| Mercer/WTW | Global coverage, job architecture framework | Expensive, enterprise-focused | Annual |
| H1B salary data | Public, granular, actual paid salaries | Only H1B employees, legal minimums may apply | Rolling, ~6 month lag |
### Benchmarking Methodology
**Step 1: Job matching.** Match internal roles to benchmark roles based on responsibilities and scope, not title (because title inflation makes title-based matching unreliable).
**Step 2: Level calibration.** Map internal levels to benchmark levels. Common mapping:
| Internal Level | Typical Years | Scope | Benchmark Equivalent |
|---------------|---------------|-------|---------------------|
| Junior/L1-L2 | 0-2 | Task-level execution with guidance | Entry/Junior |
| Mid/L3 | 2-5 | Independent execution on defined work | Intermediate |
| Senior/L4 | 5-8 | Owns projects end-to-end, mentors | Senior |
| Staff/L5 | 8-12 | Cross-team technical leadership | Staff/Principal |
| Principal/L6 | 12+ | Org-wide technical direction | Distinguished |
**Step 3: Location adjustment.** Apply geographic differential. Typical adjustments relative to SF Bay Area:
| Location Tier | Adjustment | Examples |
|--------------|------------|---------|
| Tier 1 (HCOL) | 100% (baseline) | SF, NYC, Seattle |
| Tier 2 (MCOL) | 85-95% | Austin, Denver, Boston, LA |
| Tier 3 (LCOL) | 70-85% | Midwest, Southeast, smaller metros |
| International — Tier 1 | 80-100% | London, Zurich, Tel Aviv |
| International — Tier 2 | 50-80% | Berlin, Amsterdam, Toronto, Sydney |
| International — Tier 3 | 30-60% | Eastern Europe, Southeast Asia, Latin America |
**Step 4: Company stage adjustment.** Equity/cash mix shifts dramatically by stage:
| Stage | Base Percentile Target | Equity | Bonus | Rationale |
|-------|----------------------|--------|-------|-----------|
| Pre-seed/Seed | 25th-50th | Heavy (0.5-2%) | Rare | Cash constrained, equity is the draw |
| Series A-B | 50th-65th | Moderate (0.1-0.5%) | Sometimes | Balancing cash with equity upside |
| Series C+ / Growth | 60th-75th | Moderate (RSUs typical) | Common | Competitive cash + meaningful equity |
| Public | 50th-75th | RSUs, refresh grants | Standard | Full comp package, liquid equity |
| Enterprise/non-tech | 50th-75th | Limited or none | Standard | Cash-heavy, pension/benefits |
### Percentile Bands
Standard reporting format:
| Percentile | Meaning | When to Target |
|-----------|---------|----------------|
| 25th | Below market | Constrained budget. Risk of attrition for strong performers. |
| 50th (median) | At market | Standard target for most roles. "Competitive." |
| 75th | Above market | For critical/hard-to-fill roles or retention situations |
| 90th | Top of market | Exceptional talent, niche skills, competitive bidding situations |
---
## Internal Equity Analysis
### Band Structure
| Term | Definition |
|------|-----------|
| **Band minimum** | Floor for the level. No one should be paid below this for the role/level. |
| **Band midpoint** | Target market rate. New hires at expected experience should land here. |
| **Band maximum** | Ceiling for the level. Above this = promotion or re-leveling needed. |
| **Compa-ratio** | Employee pay / band midpoint. 1.0 = at midpoint. <0.8 = significantly below. >1.2 = above band. |
| **Range penetration** | (Pay - Min) / (Max - Min). Shows position within the band. |
### Equity Audit Checklist
| Check | Method | Flag When |
|-------|--------|-----------|
| Band outliers | Compa-ratio for each employee | <0.85 or >1.15 |
| Compression | Senior vs. junior pay gap within same team | Gap <10% between levels |
| Inversion | Junior employee paid more than senior in same role | Any occurrence |
| Tenure-pay disconnect | Tenure vs. compa-ratio scatter plot | Long tenure + low compa-ratio |
| Demographic parity | Compa-ratio by demographic group at same level/role | Statistical gap >3% at same level |
| New hire premium | New hire comp vs. existing employee comp at same level | New hires >10% above existing |
### Compression and Inversion
**Compression**: Pay gap between levels narrows to the point where promotion provides minimal pay increase. Causes: market-rate new hires at higher rates, insufficient adjustment for promotions, cost-of-living raises that don't keep pace with market movement.
**Fix**: Regular market adjustment cycles (annual minimum), promotion-triggered band placement review, new hire offer calibration against existing team.
**Inversion**: Junior employee compensated higher than senior peer in same role family. Causes: hot market hiring, location arbitrage changes, acquisition-inherited comp.
**Fix**: Immediate review of senior employee's compensation. Inversion is never acceptable as a steady state.
---
## Offer Structuring
### Package Assembly Checklist
| Component | Inputs Needed | Typical Flexibility |
|-----------|---------------|-------------------|
| Base salary | Role, level, location, band, budget | ±5-10% from target |
| Equity | Stage, grant pool, vesting schedule, valuation method | Moderate — pool-constrained |
| Signing bonus | Competing offer gap, relocation, clawback terms | High — one-time cost |
| Target bonus | Role level, company plan | Low — typically standardized |
| Start date | Candidate availability, team readiness | Moderate |
| Title | Internal leveling, candidate expectations | Low — must match internal framework |
| Location | Office requirement, remote policy, tax implications | Varies by company policy |
### Equity Grant Design
| Parameter | Common Structures |
|-----------|------------------|
| **Vesting schedule** | 4-year with 1-year cliff (standard), 4-year monthly (increasingly common), 3-year annual (uncommon) |
| **Grant type** | RSUs (public/late-stage), ISOs (startup, tax-advantaged), NSOs (startup, above ISO limit) |
| **Refresh cadence** | Annual (top companies), ad-hoc (most companies), none (concerning for retention) |
| **Valuation method** | 409A (private), market price (public), projected (speculative — flag this) |
### Negotiation Framework for Hiring Managers
| Lever | When to Use | Watch For |
|-------|-------------|-----------|
| Base increase | Candidate cites competing offer or market data | Internal equity impact — raising one creates comparison |
| Equity increase | Candidate values long-term upside | Pool constraints, dilution limits |
| Signing bonus | Bridge gap without raising base (no recurring cost) | Clawback terms, candidate sees through short-term fix |
| Title upgrade | Candidate has title sensitivity | Must match internal leveling. Never give a title the person can't perform at. |
| Start date flexibility | Candidate needs time for transition | Team's hiring urgency |
| Remote work | Candidate prefers flexibility | Must align with policy. Don't make one-off exceptions. |
| Relocation package | Candidate must relocate | One-time cost, specify exactly what's covered |
### Offer Letter Compliance
| Jurisdiction Issue | What to Check |
|-------------------|---------------|
| At-will employment | Required in most US states. Some states (Montana) differ. |
| Salary transparency | NYC, CO, CA, WA, and growing list require salary ranges in postings and offers |
| Non-compete | Unenforceable in CA, limited in many states, FTC rule pending |
| Benefits enrollment | Enrollment windows, waiting periods, state-mandated minimums |
| Equity tax treatment | ISOs vs. NSOs, 83(b) elections, RSU tax at vesting — always recommend tax advisor |
| International | Employment law varies dramatically. Local counsel required. |
**Every offer letter draft must include the disclaimer: "This draft requires legal review before sending to the candidate."**
---
## Equity Modeling
### RSU Modeling Template
| Parameter | Input | Notes |
|-----------|-------|-------|
| Grant size | [X] shares/units | |
| Current share price | $[X] | 409A for private, market for public |
| Vesting schedule | [X]-year, [cliff?], [monthly/quarterly/annual] | |
| Annual refresh | [X] shares/year (if applicable) | |
| Projected growth rate | [X]% annually (if modeling upside) | Always label as projection, never as guarantee |
### Option Modeling
| Parameter | Description |
|-----------|-------------|
| Strike price | Price at which the option can be exercised (409A value at grant) |
| Current FMV | Fair market value at time of analysis |
| Spread | FMV - Strike = current paper value per share |
| Shares granted | Total option shares |
| Vested shares | Shares exercisable today |
| Cost to exercise | Vested shares * strike price |
| Paper value | Vested shares * spread |
| Tax implications | ISO: AMT on exercise, capital gains on sale. NSO: ordinary income on exercise. |
### Equity Communication Rules
| Do | Don't |
|-----|-------|
| Present equity as a range of outcomes | Present a single "expected" value |
| Label projections explicitly | Imply growth rates are guaranteed |
| Explain dilution risk | Ignore dilution in projections |
| Recommend a tax advisor | Provide tax advice |
| Explain vesting schedule clearly | Bury cliff or forfeiture terms |
| Show paper value AND liquidity constraints | Show only paper value |
---
## Retention Risk Assessment
### Flight Risk Indicators
| Factor | Risk Signal | Weight |
|--------|------------|--------|
| Compa-ratio | <0.85 | High |
| Time since last adjustment | >18 months | Medium |
| Performance rating | Exceeds + below-market comp | High (regrettable attrition risk) |
| Tenure | 2-3 years (vesting cliff), 18 months (common transition point) | Medium |
| Manager relationship | Low manager effectiveness scores | Medium |
| Market demand | Hot market for their skill set | High |
| Equity cliff | Approaching vesting cliff with no refresh | High |
### Retention Intervention Timing
| Risk Level | Action | Timeline |
|-----------|--------|----------|
| Watch | Monitor, flag for next comp cycle | Next quarter |
| Elevated | Manager 1:1 to assess satisfaction, off-cycle adjustment consideration | Within 2 weeks |
| Critical | Retention package: equity refresh + base adjustment + career conversation | Immediately |
**Retention packages after a counter-offer have 50% failure rate within 12 months.** Proactive retention before the resignation is 3x more effective than reactive counter-offers.
references/hr/llm-hr-failure-modes.md
# LLM HR Failure Modes
Where LLMs fail in HR contexts. Every mode in the HR skill must account for these failure modes. This reference loads with every HR task.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers HR-specific failures only.
---
## Failure Mode Categories
| Category | Risk Level | Prevalence | Detection Difficulty |
|----------|-----------|------------|---------------------|
| **Bias in language** | High | Very common | Medium — requires pattern awareness |
| **Data fabrication** | Critical | Common | Low — easy to detect if you check |
| **Compliance gaps** | Critical | Common | High — requires jurisdictional knowledge |
| **Inappropriate phrasing** | High | Very common | Medium — context-dependent |
| **False confidence** | High | Universal | High — LLMs present uncertainty as certainty |
| **Template over-reliance** | Medium | Common | Low — output feels generic |
---
## 1. Bias in Language
### Gendered Language in Job Postings
LLMs default to language patterns that discourage diverse applicants.
| Biased Term | Problem | Neutral Alternative |
|-------------|---------|-------------------|
| "Rockstar" / "Ninja" / "Guru" | Masculine-coded. Discourages women and non-binary applicants from applying. | "Expert", "Specialist", "Experienced" |
| "Aggressive" / "Dominant" | Masculine-coded. Signals adversarial culture. | "Ambitious", "Results-oriented", "Driven" |
| "Nurturing" / "Supportive" | Feminine-coded. Limits role perception. | Describe the behavior: "Mentors junior engineers" |
| "Culture fit" | Undefined proxy for "like us." Enables homogeneity. | "Values alignment" with specific, defined values |
| "Young and dynamic team" | Age discrimination signal. Illegal in many jurisdictions. | "Collaborative team" — describe the work style, not demographics |
| "Native English speaker" | National origin discrimination. Illegal. | "Fluent English" or "Strong written and verbal English" |
| "Must be willing to work long hours" | Discriminates against caregivers, disabilities. | "Ability to meet project deadlines" — focus on outcome |
| "Digital native" | Age proxy. | "Proficient with [specific tools]" |
### Inflated Requirements
LLMs generate wish lists, not requirements.
| Pattern | Problem | Fix |
|---------|---------|-----|
| "10+ years experience" for mid-level roles | Excludes career changers, fast learners. Years ≠ skill. | State the actual skill needed: "Can design distributed systems independently" |
| "CS degree required" for skills-based roles | Excludes bootcamp grads, self-taught engineers. Degree ≠ competence. | "CS degree or equivalent practical experience" |
| "Expert in [8 technologies]" | No one is expert in 8 things. Discourages qualified candidates. | "Primary expertise in 1-2: [list]. Familiarity with others is a plus." |
| "Must have startup experience" | Excludes enterprise talent who may be excellent. | "Comfortable with ambiguity and changing priorities" |
### Performance Review Bias
LLMs replicate common performance review biases.
| Bias | How LLM Manifests It | Detection | Mitigation |
|------|----------------------|-----------|------------|
| **Halo effect** | One strong accomplishment colors entire review as "Exceeds" | Check: are all competency ratings high, or just one? | Score competencies independently |
| **Similarity bias** | More positive language for people "like" the reviewer | Check: compare adjective intensity across reviews | Use SBI framework, not adjectives |
| **Recency bias** | Review content clusters around recent events | Check: are examples spread across the full review period? | Require examples from each quarter |
| **Attribution bias** | Success attributed to individuals, failure to circumstances | Check: does the same review have "she achieved" and "the project failed"? | Consistent attribution framing |
| **Gender-coded feedback** | Women: "collaborative, helpful, supportive." Men: "technical, visionary, strategic." | Check: swap names and re-read. Does the feedback still fit? | Focus on behaviors and outcomes, not traits |
---
## 2. Data Fabrication
### Compensation Data
**This is the highest-risk failure mode.** LLMs will generate plausible-sounding salary numbers that are entirely invented.
| What LLM Does | Why It's Dangerous | Prevention |
|---------------|-------------------|-----------|
| Generates specific percentile values ($185K base, $50K equity) | These numbers inform real pay decisions. Wrong numbers cause underpayment or overpayment. | Always state: "These are estimates based on training data, not current market data." |
| Cites "Glassdoor" or "Levels.fyi" with specific numbers | Numbers may be fabricated. Citation adds false credibility. | Never cite a specific number from a source without actually querying that source. |
| Presents ranges without stating data vintage | 2022 data applied to 2025 decisions. Market moves 5-15% annually. | Always state data vintage: "Based on training data through [date]. Verify with current sources." |
| Provides location-adjusted figures | Adjustment percentages may be invented or outdated | State the adjustment methodology explicitly. Better: let the user provide local data. |
**Rule: If you don't have real data, say so. Never present generated numbers as market data.**
### People Analytics
| Fabrication Risk | Example | Prevention |
|-----------------|---------|-----------|
| Generating benchmark rates | "Industry average attrition is 13.2%" | State source explicitly. If no source, say "typical ranges are X-Y% based on general HR knowledge." |
| Inventing correlation | "Our data shows a strong correlation between..." | Only state correlations present in user-provided data. Never infer from training data. |
| Hallucinating survey results | "Based on your engagement survey..." | Only reference data the user actually provided. |
---
## 3. Compliance Gaps
### Jurisdiction Sensitivity
Employment law varies by country, state, and city. LLMs default to US norms or generate jurisdiction-agnostic advice that may be illegal in the user's location.
| Topic | Jurisdiction Risk | LLM Failure Mode |
|-------|------------------|-------------------|
| **At-will employment** | US concept. Most of the world has termination protections. | Generates "at-will" language for international offers |
| **Non-compete clauses** | Unenforceable in CA. Varying enforceability elsewhere. FTC rule pending. | Includes non-competes without jurisdiction check |
| **Notice periods** | US: often none. Europe: 1-3 months. Some countries: 6 months. | Generates US-style immediate termination language |
| **Salary transparency** | Required in NYC, CO, CA, WA, and growing list. | Omits salary range from job posting |
| **Leave policies** | Parental leave: US (none federal), Europe (extensive), varies globally | Generates US-centric leave language |
| **Discrimination protections** | Protected classes vary by jurisdiction | Generates advice legal in one jurisdiction, illegal in another |
| **Data privacy** | GDPR (EU), CCPA (CA), varying global privacy laws | Suggests collecting or storing PII without privacy consideration |
**Rule: Always ask for jurisdiction before providing employment law guidance. Never default to US law.**
### Offer Letter Compliance
| Issue | LLM Failure | Correct Approach |
|-------|-------------|-----------------|
| Missing required language | Omits at-will disclaimer, arbitration clause, EEOC statement | Use company's legal-approved template as the base |
| Implied contract | "We expect you to be here for at least 2 years" creates implied contract | Avoid duration implications unless it's a fixed-term contract |
| Benefit promises | "You will receive full benefits" when there's a waiting period | Specify enrollment dates, waiting periods, plan options |
| Equity promises | "Your options will be worth $X" | "Subject to board approval. Value is not guaranteed." |
### PIP and Termination
| Risk | LLM Failure | Correct Approach |
|------|-------------|-----------------|
| Discriminatory timing | PIP issued right after employee announces pregnancy/disability | Establish documented performance concerns well before protected event |
| Vague success criteria | "Improve your work quality" — unmeasurable, undefendable | Specific, measurable, time-bound criteria |
| Retaliation appearance | PIP after employee filed HR complaint | Ensure HR/legal review for any performance action near protected activity |
| Inconsistent application | PIP for one employee but not another with same performance | Document policy: same standards applied to all employees |
---
## 4. Inappropriate Phrasing
### Performance Review Language
| Inappropriate | Why | Better |
|-------------|-----|--------|
| "She is very emotional" | Gendered, personality-based, not behavioral | "Expressed frustration during the sprint review, which disrupted the discussion" |
| "He's not a culture fit" | Vague, potentially discriminatory | "Did not follow the team's code review process in 4 of 6 sprints" |
| "Natural leader" | Implies some people aren't natural leaders (coded) | "Organized the incident response, assigned roles, and led the post-mortem" |
| "Aggressive" (about a woman) | Disproportionately applied to women. Same behavior in men = "assertive." | "Advocated strongly for the architectural change in the design review" |
| "Surprisingly technical" | Implies expectation of low competence (often gendered/racial) | "Demonstrated strong technical depth in the system design review" |
| "Doesn't speak up enough" | Penalizes introverts. Communication ≠ volume. | "When [person] shares input, it's high quality. Finding more venues for contribution would increase impact." |
| "Works too hard" | Not a development area. Masks burnout concern. | "I've noticed 60+ hour weeks. Let's discuss workload sustainability." |
### Offer Letter Language
| Inappropriate | Why | Better |
|-------------|-----|--------|
| "We're like a family" | Creates implied obligation, blurs professional boundaries | "We're a collaborative team that supports each other" |
| "Fast-paced, high-intensity" | Signals unsustainable work expectations | "We ship frequently and iterate quickly" |
| "Unlimited PTO" without guidance | In practice people take less PTO. Companies save on accrual liability. | State minimum expected PTO usage: "We expect you to take at least X days" |
### Org Planning Language
| Inappropriate | Why | Better |
|-------------|-----|--------|
| "Move Alice because she's difficult" | Person-based org design, masks management failure | "Restructure the team around the authentication domain, not current reporting lines" |
| "Low performers in this team" | Labels people without evidence | "Three employees have been below target on sprint delivery for 2+ quarters" |
| "Right-sizing" | Euphemism for layoffs. Everyone sees through it. | "Headcount reduction of X positions due to [specific business reason]" |
| "Redundant roles" (without evidence) | Assumes without analysis | "These two roles have 80% task overlap based on job analysis" |
---
## 5. False Confidence
### Patterns
| Pattern | Example | Harm | Fix |
|---------|---------|------|-----|
| Stating opinions as facts | "The market rate for this role is $X" | User takes fabricated data as truth | "Based on general industry knowledge, ranges typically fall between $X-$Y. Verify with current market data." |
| Omitting uncertainty | "This org structure will improve velocity by 20%" | User expects guaranteed outcome | "Based on industry benchmarks, teams with 5-8 direct reports tend to have higher velocity. Measure after 90 days." |
| Overconfident legal advice | "Non-competes are unenforceable" | Varies by jurisdiction. Bad advice. | "Non-compete enforceability varies significantly by jurisdiction. Consult legal counsel for your specific location." |
| Definitive policy interpretation | "Your PTO policy means you can take 3 weeks" | Policy language may be ambiguous | "Based on the policy text you shared, it appears to allow X. Confirm with your HR team." |
---
## 6. Template Over-Reliance
### Problem
LLMs generate templates that look professional but carry no domain insight. A performance review template with "[insert specific example]" placeholders is not a performance review — it's a word processor.
### Detection
| Signal | What It Means |
|--------|-------------|
| Output is >80% boilerplate | Template, not analysis. No domain knowledge applied. |
| All examples are generic | "Improved team productivity" — which team, what productivity, by how much? |
| Structure is perfect, content is empty | The skeleton is the easy part. Filling it with insight is the job. |
| Same template regardless of role/level/context | Junior IC review looks identical to VP review |
### Mitigation
- Always ask clarifying questions before generating HR content
- Incorporate user-provided context into every section (not just the placeholders)
- Flag sections where the user needs to supply specific data
- Prefer fewer sections with real content over many sections with boilerplate
---
## Pre-Output Checklist
Run every HR output through these checks before delivering.
| # | Check | Pass Criteria |
|---|-------|--------------|
| 1 | Bias scan | No gendered, ageist, or exclusionary language |
| 2 | Data sourcing | All numbers sourced or flagged as estimates |
| 3 | Jurisdiction | Location-dependent advice identifies jurisdiction or asks |
| 4 | Legal disclaimer | Documents requiring legal review are flagged |
| 5 | PII minimization | Names/salaries used only when necessary, not casual |
| 6 | Behavioral language | All feedback describes behavior, not personality |
| 7 | Template quality | Content is specific to the user's context, not generic |
| 8 | Confidence calibration | Uncertainty stated explicitly, not hidden |
| 9 | Compliance flags | Offer letters, PIPs, termination language flagged for legal review |
| 10 | Source vintage | Data freshness stated where applicable |
references/hr/org-planning.md
# Org Planning Reference
Headcount modeling, org design principles, capacity planning, and people analytics.
---
## Org Design Principles
### Design Variables
| Variable | Options | Trade-offs |
|----------|---------|------------|
| **Structure type** | Functional, Product, Matrix, Pod | See structure comparison below |
| **Span of control** | Narrow (3-5) vs. Wide (8-12) | Narrow = more managers, more layers. Wide = less overhead, more autonomy required. |
| **Centralization** | Centralized vs. Distributed | Centralized = consistency, efficiency. Distributed = speed, local adaptation. |
| **Reporting depth** | Flat (2-3 layers) vs. Deep (5+ layers) | Flat = faster decisions, harder to manage scale. Deep = more career levels, slower. |
### Structure Comparison
| Structure | Best For | Strengths | Weaknesses |
|-----------|----------|-----------|------------|
| **Functional** | Stable domains, shared expertise | Deep specialization, clear career paths, resource efficiency | Cross-functional coordination is slow, siloed priorities |
| **Product** | Customer-facing delivery, speed | End-to-end ownership, fast decisions, customer-aligned | Duplicate expertise, inconsistent standards, people isolation |
| **Matrix** | Complex orgs needing both depth and delivery | Balances specialization with delivery | Dual reporting confusion, political, requires mature management |
| **Pod/Squad** | Cross-functional delivery at small scale | Autonomy, speed, clear mission | Hard to scale, skill development challenges, resource allocation |
### Healthy Org Benchmarks
| Metric | Healthy Range | Warning | Critical |
|--------|---------------|---------|----------|
| Span of control | 5-8 direct reports | <4 or >10 | <3 (vanity titles) or >12 (neglect) |
| Management layers | 1 per 50-100 people | >1 per 30 | >1 per 20 |
| IC-to-manager ratio | 6:1 to 10:1 | <5:1 | <3:1 (top-heavy, excess overhead) |
| Team size | 5-9 people | <4 or >10 | <3 (fragile) or >12 (unmanageable) |
| Single points of failure | 0 critical | 1-2 | >3 (bus factor risk) |
| Skip-level ratio | Every IC has skip-level access | Some don't | No skip-levels happening |
### Org Design Failure Modes
| Failure Mode | Symptom | Root Cause | Fix |
|-------------|---------|-----------|-----|
| **Empire building** | Managers optimizing headcount over outcome | Headcount = status in org culture | Measure outcomes per person, not team size |
| **Reporting line as reward** | Star performer gets direct reports as "promotion" | No IC growth track | Build a staff/principal IC ladder with real scope |
| **Shadow org** | Informal decision-making bypasses the official chart | Org structure doesn't match actual information flow | Redesign around how work actually flows |
| **Reorg churn** | Annual restructures that never settle | Treating structure as the solution to execution problems | Fix execution, then assess if structure change is still needed |
| **Single point of failure** | One person holds all context for a critical system | Under-investment in documentation and knowledge sharing | Immediate pairing, documentation sprint, cross-training |
| **Absentee manager** | Manager has too many reports, too many meetings, no 1:1 time | Span too wide, manager also an IC | Reduce span or remove IC responsibilities |
---
## Headcount Modeling
### Planning Framework
| Dimension | Questions | Output |
|-----------|-----------|--------|
| **Demand** | What work needs to happen? What's the project/product roadmap? | Required capacity by function |
| **Supply** | Who do we have? What can they deliver? What's the attrition forecast? | Current capacity with risk adjustment |
| **Gap** | Where does demand exceed supply? Which gaps are most critical? | Prioritized hiring list |
| **Sequence** | Which hires unlock the most capacity or reduce the most risk? | Quarterly hiring roadmap |
| **Budget** | What does this cost? What are the trade-offs? | Cost model with scenarios |
### Headcount Request Template
| Field | Content | Example |
|-------|---------|---------|
| Role | Title, level, function | Senior Backend Engineer, L4 |
| Justification | Why this role, why now | "Auth service rewrite blocked on capacity. No existing team member has IdP integration experience." |
| Impact if unfilled | What happens without this hire | "Auth rewrite slips to Q3. SOC2 compliance timeline at risk." |
| Alternative considered | Can existing team absorb? Contractor? Defer? | "Contractor evaluated — context ramp time makes FTE more efficient for 6+ month project." |
| Cost | Fully loaded annual cost | $280K (base $180K + equity $60K + benefits $40K) |
| Timeline | When needed, how long to fill | Start by April 1. Expect 45-day fill time. Req must open by Feb 15. |
| Hiring manager | Who owns the search | [Name] |
### Cost Modeling
| Cost Component | Typical Multiplier | Notes |
|---------------|-------------------|-------|
| Base salary | 1.0x | Cash compensation |
| Benefits | 0.2-0.3x base | Health, dental, vision, retirement, insurance |
| Equity | 0.1-0.5x base | Varies dramatically by stage |
| Payroll taxes | 0.08-0.12x base | FICA, state taxes, unemployment |
| Equipment/onboarding | $3-8K one-time | Laptop, monitor, software licenses |
| Recruiting | 0.15-0.25x first-year base (agency) or $5-15K (in-house) | Per hire |
| **Fully loaded multiplier** | **1.3-1.6x base** | Rule of thumb for total cost |
### Scenario Planning
Model three scenarios for every headcount plan:
| Scenario | Assumption | Impact |
|----------|-----------|--------|
| **Optimistic** | All hires filled on time, no attrition, budget approved in full | Full roadmap delivery |
| **Baseline** | 80% of hires filled, standard attrition (10-15%), budget 85% approved | Adjusted roadmap with prioritization |
| **Conservative** | 60% of hires filled, elevated attrition (20%), budget constrained | Critical path only, defer non-essential |
Always present baseline as the plan and conservative as the contingency.
---
## Capacity Planning
### Capacity Calculation
```
Available capacity = (Headcount) * (Working days) * (Productivity factor)
Productivity factor = 1.0 - (meetings + overhead + PTO + onboarding ramp)
```
Typical productivity factors:
| Role Type | Productivity Factor | Why |
|-----------|-------------------|-----|
| IC (execution-heavy) | 0.65-0.75 | Meetings, overhead, context switching |
| IC (new hire, first 90 days) | 0.25-0.50 | Onboarding ramp |
| Manager | 0.20-0.40 for individual contribution | Most time goes to people management, meetings, planning |
| Tech lead (hybrid) | 0.40-0.55 | Split between execution and coordination |
### Gap Analysis
| Gap Type | Signal | Resolution Options |
|----------|--------|-------------------|
| **Skill gap** | Team lacks expertise for planned work | Hire, train existing, contractor, defer project |
| **Capacity gap** | Team has skills but not enough people | Hire, prioritize ruthlessly, extend timelines |
| **Coverage gap** | Single point of failure for critical function | Cross-train, document, pair, hire backup |
| **Leadership gap** | Team needs a manager or tech lead and doesn't have one | Promote internal (preferred), hire external |
---
## People Analytics
### Headcount Reporting
| Dimension | Cuts | Frequency |
|-----------|------|-----------|
| By team/department | All org units | Monthly |
| By location | Office, remote, geo | Monthly |
| By level | IC levels, management levels | Quarterly |
| By tenure | Bands: <1yr, 1-2yr, 2-5yr, 5+ yr | Quarterly |
| By function | Engineering, product, design, ops, etc. | Monthly |
### Attrition Analysis
| Metric | Formula | Healthy Range |
|--------|---------|---------------|
| Overall attrition | (Departures in period / Average headcount) * 12/months | 10-15% annual |
| Voluntary attrition | Voluntary departures only | 8-12% annual |
| Regrettable attrition | Voluntary departures of high performers | <5% annual |
| Involuntary attrition | Terminations, layoffs | <5% (steady state) |
| New hire attrition (first year) | New hires who leave within 12 months | <15% |
### Attrition Root Cause Categories
| Category | Indicators | Intervention |
|----------|-----------|-------------|
| **Compensation** | Below-market comp, no recent adjustment, competitor offers | Market adjustment, retention package |
| **Manager** | Exit interviews cite manager, low manager scores | Manager coaching, reassignment, skip-level program |
| **Growth** | Exit interviews cite lack of growth, stale role | Career pathing, stretch assignments, lateral moves |
| **Culture** | Exit interviews cite culture, belonging issues | Team health assessment, inclusion initiatives |
| **Burnout** | High hours, no PTO usage, on-call burden | Workload audit, mandatory PTO, on-call rotation fix |
| **Personal** | Relocation, career change, family | Limited intervention — wish them well |
### Diversity Metrics
| Metric | What to Measure | Granularity |
|--------|----------------|-------------|
| Representation | Headcount by demographic group | By level, team, function |
| Hiring pipeline | Applicant, screen, interview, offer, accept by group | By role, source |
| Promotion rate | Promotions per group / eligible population per group | By level transition |
| Attrition rate | Departures per group / headcount per group | By voluntary/involuntary |
| Pay equity | Compa-ratio by demographic group at same level/role | By role family + level |
| Inclusion scores | Survey responses by demographic group | By team |
**Reporting constraints:**
- Minimum group size of 5 before reporting demographics (below that, individuals are identifiable)
- Present as rates, not raw counts (counts without context are misleading)
- Compare like-for-like: same role, same level, same location
- Trends over time matter more than snapshots
### Engagement Metrics
| Metric | Source | Healthy Range |
|--------|--------|---------------|
| eNPS (Employee Net Promoter Score) | Survey: "How likely to recommend as a workplace?" | 10-50 (>30 = strong) |
| Engagement index | Composite of survey questions on motivation, satisfaction, commitment | >70% favorable |
| Participation rate | Survey response rate | >75% |
| Manager effectiveness | Survey questions on manager quality | >75% favorable |
| Growth perception | "I have opportunities to grow here" | >65% favorable |
### Flight Risk Model
Composite risk score based on multiple factors:
| Factor | Weight | Data Source |
|--------|--------|------------|
| Compa-ratio < 0.85 | 25% | Comp data |
| Time since last promotion > 2 years | 20% | HRIS |
| Low engagement score | 20% | Survey |
| Manager effectiveness score (low) | 15% | Survey |
| Approaching vesting cliff | 10% | Equity data |
| Tenure at risk inflection (18mo, 3yr) | 10% | HRIS |
**Risk levels:**
- Low (0-30): Monitor in regular cycle
- Medium (31-60): Proactive manager conversation
- High (61-100): Immediate retention intervention
---
## Reorg Planning
### Decision Framework: When to Reorg
| Situation | Reorg Appropriate? | Why |
|-----------|-------------------|-----|
| Work is flowing fine but "could be better" | No | Churn cost > marginal improvement |
| Two teams do overlapping work with conflicts | Yes | Clear duplication and coordination failure |
| Team is too large for one manager | Partial — split, don't reorg | Targeted change, minimal disruption |
| New product/market requires cross-functional team | Maybe — pod first | Try a pod/squad before restructuring |
| Post-acquisition integration | Yes | Merge teams, eliminate redundancy, align cultures |
| Manager departure creates vacuum | Partial — backfill or merge | Don't reorg around one person's departure |
### Reorg Communication Plan
| Phase | Timeline | Audience | Content |
|-------|----------|----------|---------|
| **Pre-announcement** | 1-2 weeks before | Senior leaders | Changes, rationale, their role in communication |
| **Announcement** | Day 0 | All affected employees | What's changing, why, what it means for each person |
| **1:1 conversations** | Day 0-3 | Each affected individual | Personal impact, new manager, new team, questions |
| **Team kickoff** | Week 1 | New teams | Mission, priorities, working agreements |
| **30-day check-in** | Day 30 | All affected | How's it going, what needs adjustment |
### Reorg Failure Modes
| Failure Mode | Risk | Prevention |
|-------------|------|-----------|
| Announcing before 1:1s | People learn about their new manager in an all-hands | Individual conversations first, then group announcement |
| Reorg without clear "why" | Perceived as political, creates cynicism | State the problem the reorg solves. If you can't, don't reorg. |
| Changing everything at once | Disorientation, productivity cliff | Change structure OR process OR tools — not all three simultaneously |
| No success metrics | Can't tell if the reorg worked | Define measurable outcomes before executing |
| Ignoring informal networks | Reorg severs critical informal communication channels | Map informal networks, preserve key connections intentionally |
references/hr/performance-management.md
# Performance Management Reference
Review structures, calibration methodology, feedback patterns, and development planning.
---
## Review Types and Templates
### Self-Assessment
Purpose: Employee documents their own accomplishments, growth, and goals. Forces reflection. Provides manager with evidence they may not have observed directly.
#### Structure
| Section | Content | Quality Gate |
|---------|---------|-------------|
| Key accomplishments | Top 3-5 achievements with situation, contribution, impact | Each has measurable impact, not just activity |
| Goals review | Status per goal from last period with evidence | Evidence is specific (metrics, artifacts, dates), not claims |
| Growth areas | New skills, expanded scope, leadership moments | Concrete examples, not self-assessment platitudes |
| Challenges | What was difficult, what you'd do differently | Honest reflection, not disguised brags |
| Next-period goals | 3-5 specific, measurable goals | SMART criteria: Specific, Measurable, Achievable, Relevant, Time-bound |
| Manager feedback | How can your manager better support you | Actionable requests, not generic "more feedback" |
#### Accomplishment Writing Formula
**Weak**: "Worked on the migration project."
**Strong**: "Led the database migration from PostgreSQL to CockroachDB, reducing query latency p99 from 450ms to 120ms and eliminating 3 weekly on-call pages."
Pattern: **[Action verb] + [what you did specifically] + [quantified impact]**
| Component | Questions to Answer |
|-----------|-------------------|
| Action verb | Led, designed, built, shipped, reduced, eliminated, launched, established |
| What specifically | Not the project name — your contribution to it |
| Impact | Revenue, cost, time saved, reliability, team efficiency, customer satisfaction |
| Scope | Team, org, company — how broad was the effect? |
---
### Manager Review
Purpose: Document performance assessment, provide development guidance, justify compensation recommendations.
#### Structure
| Section | Content | Constraint |
|---------|---------|-----------|
| Overall rating | Exceeds / Meets / Below Expectations | One rating. No "Meets+" hedging. |
| Performance summary | 2-3 sentence overall assessment | Covers scope, impact, and growth trajectory |
| Key strengths | 2-3 strengths with behavioral examples | Behaviors, not traits. "Documented runbooks" not "is organized." |
| Development areas | 2-3 areas with actionable guidance | Specific next steps, not vague directives |
| Goal achievement | Rating per goal with observations | Evidence-based, references specific work product |
| Development plan | Skill → current → target → actions | Actions are concrete: courses, stretch assignments, mentorship |
| Comp recommendation | Promotion / equity / adjustment / none | Justified by performance evidence, not tenure or likability |
#### Rating Scale
| Rating | Definition | Signal |
|--------|-----------|--------|
| **Exceeds Expectations** | Consistently performs above level expectations. Demonstrates next-level behaviors. High impact, broad scope. | Ready for promotion discussion or significant comp action |
| **Meets Expectations** | Solid performance at expected level. Delivers reliably. Growing in role. | On track. Standard comp progression. |
| **Below Expectations** | Not meeting expectations for current level. Specific improvement needed. | Requires development plan. May lead to PIP if sustained. |
**Rating failure modes:**
- Rating inflation: everyone gets "Exceeds." Loses meaning. Undermines calibration.
- Recency bias: rating based on last 4 weeks, not full period. Keep running notes.
- Halo/horns effect: one strong/weak area colors the entire review. Score dimensions independently.
- Central tendency: everyone gets "Meets." Avoids difficult conversations at the cost of rewarding mediocrity.
---
### Feedback Writing Patterns
#### The SBI Framework
**Situation → Behavior → Impact.** The gold standard for feedback specificity.
| Component | What It Is | Example |
|-----------|-----------|---------|
| **Situation** | When and where it happened | "During the Q3 architecture review..." |
| **Behavior** | What the person did (observable) | "...you presented three options with trade-off analysis and cost projections..." |
| **Impact** | Effect on team, project, outcomes | "...which enabled the VP to make a decision in the meeting instead of scheduling a follow-up." |
#### Positive Feedback Patterns
| Pattern | Example | Why It Works |
|---------|---------|-------------|
| SBI with reinforcement | "During [situation], you [behavior], which [impact]. Keep doing this." | Links behavior to outcome. Person knows what to repeat. |
| Growth recognition | "Six months ago, your design docs lacked trade-off analysis. This quarter, every doc included it. That improvement is visible." | Acknowledges trajectory, not just current state. |
| Scope expansion | "You've been handling incidents for your team. You also wrote the runbook that reduced MTTR for the whole org. That's next-level impact." | Recognizes broader contribution. |
#### Constructive Feedback Patterns
| Pattern | Example | Why It Works |
|---------|---------|-------------|
| SBI with forward-looking | "During [situation], [behavior] led to [impact]. Next time, try [alternative]." | Specific, not personal. Provides a concrete alternative. |
| Gap identification | "You're strong at execution. The gap I see is in proactive communication — stakeholders are sometimes surprised by delays." | Names the specific gap between current and expected. |
| Frequency + evidence | "In the last quarter, 3 of 5 project updates were late by more than a day. Let's discuss what's causing this." | Data-driven. Not "you're always late." |
#### Feedback Failure Modes
| Failure Mode | Problem | Better |
|-------------|---------|--------|
| "Great job!" | No signal. The person learns nothing about what to repeat. | Name the specific behavior and its impact. |
| "You need to communicate better." | Vague. Communicate what, to whom, when, how? | "Send project status updates to stakeholders every Friday by EOD." |
| "You're not a team player." | Personality judgment. Defensive response guaranteed. | "In the last sprint, you merged 3 PRs without requesting review. Code review is a team expectation." |
| Feedback sandwich | "Good job, but [real feedback], and good job again." Transparent. Dilutes the message. | Lead with the feedback directly. Adults can handle it. |
| "You should be more like [other person]." | Demoralizing comparison. | Describe the target behavior without naming someone else. |
| Annual surprise | Saving feedback for the review. | Feedback within 48 hours of the event. Reviews should contain no surprises. |
---
## Calibration Methodology
### Purpose
Ensure rating consistency across managers, teams, and the organization. Prevent inflation, deflation, and bias.
### Pre-Calibration Preparation
| Step | Action | Output |
|------|--------|--------|
| 1 | Manager submits proposed ratings for all direct reports | Rating spreadsheet |
| 2 | HRBP aggregates ratings by team, level, org | Distribution summary |
| 3 | Compare distribution to targets | Variance analysis |
| 4 | Identify discussion candidates | Borderline list |
### Distribution Targets
| Rating | Target Range | Red Flag |
|--------|-------------|----------|
| Exceeds Expectations | 15-20% | >30% = inflation, <10% = suppression |
| Meets Expectations | 60-70% | <50% = bimodal split |
| Below Expectations | 10-15% | 0% sustained = avoidance of difficult conversations |
**These are guidelines, not quotas.** If a team genuinely has 30% exceeds performers, the calibration discussion should surface the evidence, not force the numbers down. Forced distribution destroys trust.
### Calibration Meeting Protocol
1. **Manager presents cases** — 2 minutes per employee. Focus on top/bottom of distribution.
2. **Cross-manager challenge** — "What specific evidence supports Exceeds?" Require behavioral examples.
3. **Bias checks** — HRBP monitors for pattern biases: tenure bias (long tenure = higher rating), recency bias, similarity bias.
4. **Level-norming** — Compare across managers: "Is Manager A's 'Meets' the same as Manager B's?"
5. **Adjustment recording** — Document every rating change and the evidence that drove it.
### Promotion Readiness Assessment
| Criterion | Evidence Required |
|-----------|-------------------|
| Sustained performance | 2+ review cycles at current level demonstrating consistent delivery |
| Next-level behaviors | Already performing at least 2-3 responsibilities of the target level |
| Scope expansion | Impact extends beyond immediate team/project |
| Independence | Operates with decreasing guidance over time |
| Peer recognition | Respected by peers as operating at a higher level |
**Promotion failure modes:**
- Tenure-based promotion: "They've been here 3 years, they deserve it." Time is not evidence.
- Promise-based promotion: "Promote them and they'll step up." Evidence first, promotion second.
- Retention promotion: "If we don't promote them, they'll leave." Solves the wrong problem — likely a comp issue.
- Single-achievement promotion: One great project ≠ sustained next-level performance.
---
## Development Planning
### Development Plan Structure
| Field | Description | Example |
|-------|-------------|---------|
| Skill | The specific competency to develop | "Technical design documentation" |
| Current level | Where they are now (behavioral description) | "Writes implementation docs but not trade-off analysis" |
| Target level | Where they should be (behavioral description) | "Writes design docs with alternatives, trade-offs, and decision rationale" |
| Actions | Concrete steps to close the gap | "1. Review 3 exemplar design docs. 2. Write design doc for Project X with mentor review. 3. Present design at architecture review." |
| Timeline | When to reassess | "Review at next 1:1 in 6 weeks" |
| Support needed | Manager/org resources required | "Pair with Staff Engineer Y on first design doc" |
### Development Action Types
| Type | Best For | Examples |
|------|----------|---------|
| **Stretch assignments** | Building new skills through real work | Lead a project, own a workstream, mentor a junior |
| **Training/courses** | Foundational knowledge gaps | Conference, online course, certification |
| **Mentorship** | Navigating organizational complexity, career growth | Pair with senior leader, cross-functional mentor |
| **Exposure** | Building visibility and breadth | Present at all-hands, join cross-functional initiative, shadow senior leader |
| **Feedback loops** | Improving specific behaviors | Weekly 1:1 check-ins on the development area with specific observations |
---
## Performance Improvement Plans (PIPs)
### When to Use
- Employee has received clear feedback on performance gaps
- Gaps have persisted despite coaching (documented in 1:1 notes)
- Performance is below expectations for current level
- This is not a first conversation — surprise PIPs are failures of management, not of the employee
### PIP Structure
| Component | Content |
|-----------|---------|
| Performance gaps | Specific, behavioral, documented. "Missed 3 of 5 sprint deadlines in Q3" not "underperforming." |
| Expected level | Clear description of what meeting expectations looks like |
| Success criteria | Measurable outcomes. "Complete all sprint commitments for 4 consecutive sprints." |
| Support provided | Training, mentorship, reduced scope, regular check-ins |
| Timeline | Typically 30-60 days. Long enough to demonstrate sustained improvement. |
| Check-in cadence | Weekly minimum. Document progress at each check-in. |
| Outcomes | Meets expectations (PIP closed), partial improvement (extend or adjust), does not meet (separation) |
### PIP Guardrails
- **Legal review required** before issuing. Always.
- **Document everything.** Every conversation, every observation, every check-in.
- **Genuine path to success.** If the PIP is designed so the employee can't pass, it's not a PIP — it's a slow termination. Those destroy trust across the team.
- **Confidential.** Between the employee, their manager, and HR. Not team knowledge.
- **No PIP for skills the employee was never expected to have.** Hire for skill X, PIP for skill Y = management failure.
references/hr/recruiting.md
# Recruiting Reference
Pipeline management, interview design, evaluation methodology, and onboarding frameworks.
---
## Pipeline Architecture
### Standard Stages
| Stage | Owner | Duration Target | Key Actions | Exit Criteria |
|-------|-------|-----------------|-------------|---------------|
| **Sourced** | Recruiter | 0-3 days | Identify, personalize outreach, initial contact | Response received (positive or negative) |
| **Screen** | Recruiter | 3-5 days | Phone/video screen, basic fit assessment | Meets minimum qualifications + culture signal |
| **Interview** | Hiring panel | 5-10 days | Structured competency interviews (2-4 rounds) | All competencies scored, panel aligned |
| **Debrief** | Hiring manager | 1-2 days | Calibrate feedback, make hire/no-hire decision | Clear decision with documented reasoning |
| **Offer** | Recruiter + HM | 2-5 days | Package assembly, approval chain, extension | Offer extended with deadline |
| **Accepted** | Recruiter | 1-7 days | Negotiation (if needed), verbal/written acceptance | Signed offer letter |
| **Onboarding** | Manager + HR | Pre-start → 90 days | Equipment, accounts, buddy, 30/60/90 plan | Employee productive and integrated |
### Stage Transition Rules
- Never skip Screen → Interview. Every candidate gets the same structured evaluation path.
- Debrief happens within 48 hours of final interview. Delayed debriefs degrade recall and produce vague feedback.
- Offer approval chain completes before verbal offer. Verbal offers without approval create legal exposure.
---
## Pipeline Metrics
### Velocity Metrics
| Metric | Definition | Healthy Range | Warning Threshold |
|--------|------------|---------------|-------------------|
| Time to fill | Days from req open to offer accepted | 30-45 days (IC), 45-60 (manager) | >60 days IC, >90 days manager |
| Days in stage | Average time per pipeline stage | 3-7 per stage | Any stage >14 days |
| Pipeline velocity | Candidates moved per week across all stages | Depends on volume | Declining trend over 3+ weeks |
| Offer-to-accept | Days from offer extended to signed | 3-7 days | >14 days |
### Conversion Metrics
| Transition | Target Conversion | Signal When Low |
|------------|-------------------|-----------------|
| Sourced → Screen | 30-50% | Poor targeting or weak outreach |
| Screen → Interview | 40-60% | Screen criteria too loose or too tight |
| Interview → Offer | 20-40% | Bar miscalibration, weak pipeline top |
| Offer → Accepted | 70-90% | Comp not competitive, slow process, poor candidate experience |
| Overall funnel | 5-15% sourced to accepted | Healthy range for most roles |
### Source Effectiveness
Track per channel: volume sourced, conversion to hire, cost per hire, time to fill, quality of hire (90-day retention + performance).
| Source Type | Typical Strengths | Watch For |
|-------------|-------------------|-----------|
| Referrals | Highest conversion, fastest, best retention | Homogeneity risk — referral networks mirror existing team demographics |
| Inbound (job boards) | Volume | Low conversion, high noise |
| Outbound (sourcing) | Passive talent access | Low response rates, high recruiter time cost |
| Agencies | Speed for hard-to-fill roles | High cost (15-25% of first-year comp), misaligned incentives |
| University/intern pipeline | Long-term talent development | Long ramp time, seasonal |
---
## Interview Design
### Competency Framework
For each role, define 4-6 competencies. Each competency gets dedicated interview time.
| Component | Description | Example |
|-----------|-------------|---------|
| **Competency** | Observable skill or behavior the role requires | "System design" or "cross-team collaboration" |
| **Behavioral questions** | Past behavior predicts future behavior. "Tell me about a time..." | "Tell me about a time you had to design a system under tight constraints." |
| **Situational questions** | Hypothetical scenarios revealing approach. "How would you..." | "How would you handle a production outage you'd never seen before?" |
| **Follow-up probes** | Dig deeper on specifics. Force detail beyond rehearsed answers. | "What specifically did you do vs. what the team did?" |
| **Scoring rubric** | 1-4 scale with behavioral anchors per level | See rubric table below |
### Scoring Rubric Template
| Score | Label | Behavioral Anchor |
|-------|-------|-------------------|
| 1 | Does not meet | Cannot demonstrate the competency. No relevant examples. Concerning signals. |
| 2 | Partially meets | Some evidence but inconsistent. Needs significant development. Limited scope examples. |
| 3 | Meets expectations | Clear evidence of competency at expected level. Multiple relevant examples. Appropriate scope. |
| 4 | Exceeds | Strong evidence beyond expected level. Complex examples. Demonstrates teaching/leading in this area. |
**Scoring rules:**
- Score each competency independently. Do not let one strong/weak area contaminate others.
- Score against the rubric, not against other candidates (comparative scoring amplifies similarity bias).
- Score immediately after the interview. Do not wait for the debrief (anchor bias increases with delay).
- Record specific evidence for each score, not just the number.
### Interview Panel Design
| Principle | Implementation |
|-----------|---------------|
| Cover all competencies | Map each interviewer to 1-2 competencies. No competency unassessed. |
| Diverse perspectives | Panel varies by seniority, function, background. Minimum 3 interviewers. |
| No duplicate coverage | Two interviewers on the same competency wastes candidate time and creates anchoring. |
| Trained interviewers | Everyone on the panel has completed structured interviewing training. |
| Consistent questions | Same question set per competency across all candidates for the same role. |
### Question Design Failure Modes
| Failure Mode | Problem | Better Approach |
|-------------|---------|-----------------|
| "What's your greatest weakness?" | Rehearsed answers. Zero signal. | "Tell me about a project that didn't go well and what you learned." |
| Brain teasers | Test puzzle familiarity, not job skills | Use role-relevant problem-solving scenarios |
| "Where do you see yourself in 5 years?" | Penalizes honesty. Rewards performance. | "What kind of work energizes you most?" |
| Leading questions | "You'd agree that X is important, right?" | Open-ended: "How do you think about X?" |
| Illegal questions | Age, family status, religion, disability, national origin | Never ask. Train interviewers on prohibited topics. |
| Culture fit (vague) | Proxy for "like me" — homogeneity bias | Define specific values + behaviors. Evaluate those. |
---
## Debrief Structure
### Debrief Protocol
1. **Independent scoring first.** All interviewers submit scores before the debrief meeting. No anchoring.
2. **Structured discussion.** Walk through each competency. Interviewer presents evidence, then score.
3. **Evidence-based.** "I scored 3 on system design because [specific example from interview]" — not "I liked them."
4. **Disagree on evidence, not feelings.** When scores differ, discuss the evidence. What did one interviewer see that another didn't?
5. **Decision framework.** Strong hire / Hire / No hire / Strong no hire. Require at least one "strong hire" and no "strong no hire" to proceed.
### Debrief Template
```
## Candidate Debrief: [Name] — [Role]
Date: [Date] | Panel: [Names]
### Competency Scores
| Competency | Interviewer | Score | Key Evidence |
|------------|-------------|-------|-------------|
| [Comp 1] | [Name] | [1-4] | [Specific example] |
| [Comp 2] | [Name] | [1-4] | [Specific example] |
### Strengths (evidence-based)
- [Strength with specific example]
### Concerns (evidence-based)
- [Concern with specific example]
### Decision: [Strong hire / Hire / No hire / Strong no hire]
### Reasoning: [2-3 sentences]
```
---
## Onboarding Framework
### Pre-Start Checklist (Before Day 1)
| Category | Task | Owner | Timeline |
|----------|------|-------|----------|
| Communication | Send welcome email (start date, time, logistics, dress code) | Recruiter/Manager | 1 week before |
| Accounts | Set up email, Slack, core tools for role | IT | 3 days before |
| Equipment | Order/configure laptop, monitor, peripherals | IT | 1 week before |
| Access | Grant repo/system access appropriate for role | Manager/IT | Day 1 |
| People | Assign onboarding buddy (not manager, peer-level) | Manager | 1 week before |
| Calendar | Add to team standup, recurring meetings, socials | Manager | 3 days before |
| Space | Prepare desk / ship remote setup | Office ops | 3 days before |
| Documentation | Prepare reading list, team wiki links, architecture docs | Manager | 1 week before |
### Day 1 Schedule Template
| Time | Activity | With | Goal |
|------|----------|------|------|
| 9:00 | Welcome, logistics, building/remote tour | Manager | Comfort, orientation |
| 10:00 | IT setup, tool walkthrough | IT / Buddy | Functional access |
| 11:00 | Team introductions (individual, not group) | Team members | Personal connection |
| 12:00 | Welcome lunch | Manager + 2-3 team | Informal relationship building |
| 1:30 | Company context: mission, values, how we work | Manager | Alignment |
| 3:00 | Role expectations, 30/60/90 goals, first project | Manager | Clarity on success criteria |
| 4:00 | Free exploration: read docs, poke around tools | Self | Self-directed onboarding |
### 30/60/90-Day Framework
| Period | Focus | Success Looks Like |
|--------|-------|-------------------|
| **Days 1-30** | Learn | Understands team, codebase/domain, processes. Completes 1-2 small tasks. Has asked many questions. |
| **Days 31-60** | Contribute | Delivers independently on scoped work. Participates in reviews. Identifies one improvement. |
| **Days 61-90** | Own | Owns a workstream. Contributes to planning. Can onboard the next new hire. |
### Onboarding Failure Modes
| Failure Mode | Problem | Fix |
|-------------|---------|-----|
| Information firehose Day 1 | Overwhelming, nothing retained | Spread learning over 2 weeks. Day 1 = setup + relationships. |
| No buddy assigned | New hire bothers manager for everything or stays silent | Assign a peer buddy. Not the manager. Someone approachable. |
| Unclear first task | New hire sits idle, feels useless | Have a real (small, shippable) task ready for week 1. |
| No check-ins | Problems fester for weeks | Daily check-in week 1, then weekly through day 90. |
| Sink or swim | "Smart people figure it out" — they do, slowly and resentfully | Structure is not hand-holding. It's efficiency. |
---
## Job Posting Quality Checklist
Run every job posting through these checks before publishing.
| Check | What to Look For | Action |
|-------|-----------------|--------|
| Gendered language | "Rockstar", "ninja", "aggressive", "dominant" | Replace with neutral terms. Use tools like Textio or manual review. |
| Unnecessary requirements | "10 years experience" for a mid-level role, degree requirements for skills-based roles | Separate must-have from nice-to-have. Cut inflated requirements. |
| Exclusionary phrasing | "Culture fit", "young and dynamic team", "native English speaker" | Remove. Replace with specific behavioral requirements. |
| Jargon overload | Acronyms, internal terminology | Plain language. A candidate outside your company should understand every requirement. |
| Benefits buried | Salary range and benefits not visible | Lead with or clearly include compensation range. Many jurisdictions require it. |
| Passive voice | "Responsibilities will include..." | Active voice: "You will..." |
references/legal.md
# Legal Workflows
Analysis support for in-house legal teams. Contract review, compliance checks, NDA triage, risk assessment, legal writing, and response generation.
**Disclaimer**: Analysis support, not legal advice. Review by qualified counsel required.
---
## Mode Detection
Classify the request into one mode before proceeding. If the request spans modes, choose the primary and note the secondary.
| Mode | Signal Phrases | Core Output |
|------|---------------|-------------|
| **CONTRACT** | review contract, clause analysis, redline, playbook, negotiate | Clause-by-clause analysis with GREEN/YELLOW/RED flags and redline suggestions |
| **COMPLIANCE** | compliance check, GDPR, HIPAA, CCPA, SOX, regulation, data protection, DSGVO, GoBD, TDDDG, eIDAS, AI Act, NIS2, KRITIS, Grundschutz, TISAX, DORA | Applicable regulations, requirements checklist, risk areas, approvals needed |
| **NDA** | NDA, triage NDA, non-disclosure, confidentiality agreement | GREEN/YELLOW/RED classification with screening checklist |
| **RISK** | legal risk, risk assessment, exposure, severity, escalation | Severity x Likelihood matrix score with escalation path |
| **WRITING** | legal brief, memo, legal response, draft response, template | Structured legal document in appropriate format |
| **VENDOR** | vendor check, vendor status, agreement status, what's signed | Agreement inventory, gap analysis, upcoming deadlines |
---
## Reference Loading Table
Load only the references required by the detected mode.
| Mode | References to Load |
|------|-------------------|
| CONTRACT | `references/legal/contract-review.md` |
| COMPLIANCE | `references/legal/compliance-frameworks.md`, `references/legal/german-business-compliance.md` |
| NDA | `references/legal/nda-triage.md` |
| RISK | `references/legal/risk-assessment.md` |
| WRITING | `references/legal/legal-writing.md` |
| VENDOR | `references/legal/contract-review.md` (for gap analysis context) |
Always load `references/legal/llm-legal-failure-modes.md` for every mode. LLM failure awareness is non-negotiable in legal work.
---
## Mode: CONTRACT
**Framework**: INTAKE -> ANALYZE -> FLAG -> REDLINE -> STRATEGIZE
**Phase 1: INTAKE** -- Accept the contract and gather context.
- Accept contract as file, pasted text, or URL reference
- Determine: which side the user is on (vendor/customer/licensor/licensee/partner), deadline, focus areas, deal context (size, strategic importance, existing relationship)
- If user provides partial context, proceed and note assumptions
**Phase 2: ANALYZE** -- Clause-by-clause review.
Load `references/legal/contract-review.md` for the full clause analysis methodology.
1. Identify contract type (SaaS, services, license, procurement, partnership)
2. Read entire contract before flagging -- clauses interact (uncapped indemnity may be mitigated by broad LOL)
3. Analyze each material clause against playbook or market-standard positions
4. Cover at minimum: LOL, indemnification, IP, data protection, confidentiality, reps/warranties, term/termination, governing law, insurance, assignment, force majeure, payment
**Gate**: Every material clause analyzed. No clause reviewed in isolation.
**Phase 3: FLAG** -- Classify deviations.
| Flag | Meaning | Action |
|------|---------|--------|
| **GREEN** | At or better than standard. Minor commercially reasonable variation. | Note for awareness. No negotiation. |
| **YELLOW** | Outside standard but within negotiable range. Common in market. | Generate redline + fallback + business impact. |
| **RED** | Outside acceptable range. Material risk. Escalation trigger. | Explain risk. Provide market-standard alternative. Recommend escalation. |
**Phase 4: REDLINE** -- Generate specific alternative language for YELLOW and RED items.
Each redline includes: current language (exact quote), proposed language, rationale (suitable for counterparty), priority (must-have / should-have / nice-to-have), fallback position.
**Phase 5: STRATEGIZE** -- Negotiation strategy.
- Tier 1 (deal breakers): uncapped liability, missing DPA for regulated data, IP jeopardizing core assets, regulatory conflicts
- Tier 2 (strong preferences): LOL adjustments, indemnification scope, termination flexibility, audit rights
- Tier 3 (concession candidates): preferred governing law, notice periods, minor definitions, insurance certificates
Lead with Tier 1. Trade Tier 3 to secure Tier 2. Escalate before making any Tier 1 concession.
**Gate**: Top 3 issues identified. Negotiation priority established. Concession candidates named.
**Output format**:
```
## Contract Review Summary
**Document**: [name] | **Parties**: [names] | **Side**: [role] | **Basis**: [Playbook/Generic]
## Key Findings
[Top 3-5 issues with severity flags]
## Clause-by-Clause Analysis
### [Clause] -- [GREEN/YELLOW/RED]
**Contract says**: ... | **Standard**: ... | **Deviation**: ... | **Impact**: ... | **Redline**: ...
## Negotiation Strategy
[Priorities, concessions, approach]
```
---
## Mode: COMPLIANCE
**Framework**: SCOPE -> MAP -> ASSESS -> RECOMMEND
**Phase 1: SCOPE** -- Understand the proposed action.
- What is being done (feature launch, data processing, marketing campaign, new vendor)
- What data is involved (personal data categories, sensitive data, regulated data)
- Which geographies (determines applicable regulations)
- Who is affected (customers, employees, partners, public)
**Phase 2: MAP** -- Identify applicable regulations.
Load `references/legal/compliance-frameworks.md` for regulation-specific requirements.
Map all potentially applicable frameworks. Check for overlapping or conflicting requirements across jurisdictions.
**Phase 3: ASSESS** -- Check each requirement.
| Requirement | Status | Action Needed |
|-------------|--------|---------------|
| [Requirement] | Met / Not Met / Unknown | [Specific action] |
For each risk area, assess severity and mitigation path.
**Phase 4: RECOMMEND** -- Prioritized action list with approvals needed.
**Gate**: All applicable regulations identified. Requirements checked. Approvals mapped.
**Output**: Quick assessment (Proceed / Proceed with conditions / Requires further review), applicable regulations table, requirements checklist, risk areas, recommended actions, approvals needed.
---
## Mode: NDA
**Framework**: ACCEPT -> SCREEN -> CLASSIFY -> REPORT
Load `references/legal/nda-triage.md` for the full screening checklist and common deviations catalog.
**Phase 1: ACCEPT** -- Accept NDA in any format.
**Phase 2: SCREEN** -- Systematic evaluation against 10 screening criteria.
Agreement structure, definition scope, receiving party obligations, standard carveouts (public knowledge, prior possession, independent development, third-party receipt, legal compulsion), permitted disclosures, term/duration, return/destruction, remedies, problematic provisions (non-solicit, non-compete, exclusivity, standstill, residuals, IP assignment).
**Phase 3: CLASSIFY**
| Classification | Criteria | Routing |
|----------------|----------|---------|
| **GREEN** | All criteria pass. Standard mutual, all carveouts, reasonable term, no prohibited provisions. | Standard delegation. Same-day approval. |
| **YELLOW** | Minor deviations: broader definition, longer term, missing one carveout, narrow residuals, non-preferred jurisdiction. | Counsel review. 1-2 business days. |
| **RED** | Wrong type, missing critical carveouts, non-solicit/non-compete, perpetual term, broad residuals, hidden IP assignment, liquidated damages. | Full legal review. Do not sign. 3-5 business days. |
**Phase 4: REPORT** -- Structured triage report with specific issues, risks, and suggested fixes.
**Gate**: Every screening criterion evaluated. Classification justified.
---
## Mode: RISK
**Framework**: IDENTIFY -> SCORE -> CLASSIFY -> DOCUMENT
Load `references/legal/risk-assessment.md` for the full severity/likelihood matrix and documentation standards.
**Phase 1: IDENTIFY** -- Define the risk clearly with background and context.
**Phase 2: SCORE** -- Apply Severity (1-5) x Likelihood (1-5) matrix.
- Severity: Negligible (1) through Critical (5), calibrated to financial exposure as % of deal/contract value
- Likelihood: Remote (1) through Almost Certain (5), calibrated to precedent and triggering events
- Risk Score = Severity x Likelihood
**Phase 3: CLASSIFY**
| Score | Level | Color | Escalation |
|-------|-------|-------|------------|
| 1-4 | Low | GREEN | Accept. Monitor quarterly. |
| 5-9 | Medium | YELLOW | Mitigate. Assign owner. Monthly review. |
| 10-15 | High | ORANGE | Senior counsel. Outside counsel if needed. Weekly review. |
| 16-25 | Critical | RED | GC/C-suite/Board. Outside counsel. Litigation hold if applicable. Daily review. |
**Phase 4: DOCUMENT** -- Risk memo with contributing factors, mitigating factors, mitigation options, recommended approach, residual risk, monitoring plan.
**Gate**: Both severity and likelihood ratings justified with specific rationale. Score calculated. Escalation path defined.
---
## Mode: WRITING
**Framework**: CLASSIFY -> DRAFT -> REVIEW
Load `references/legal/legal-writing.md` for format templates and escalation triggers.
**Phase 1: CLASSIFY** -- Determine document type.
| Type | Use Case |
|------|----------|
| Legal memo | Internal analysis of a legal question |
| Legal brief | Summary of issue, law, and recommendation |
| Legal response | Templated response to common inquiries (DSR, litigation hold, vendor question, NDA request, subpoena) |
| Meeting brief | Pre-meeting context, talking points, action items |
| Incident brief | Rapid brief for developing situations (breach, litigation threat, regulatory inquiry) |
**Phase 2: DRAFT** -- Generate document in the appropriate format.
For legal responses: check escalation triggers before generating. If any trigger fires (regulatory inquiry, potential litigation, criminal exposure, media attention, multiple jurisdictions), stop and recommend escalation instead of a templated response.
**Phase 3: REVIEW** -- Present draft for user review. Note any assumptions, gaps, or areas needing counsel input.
**Gate**: Document type correctly identified. Escalation triggers checked. All required elements present.
---
## Mode: VENDOR
**Framework**: IDENTIFY -> INVENTORY -> GAP ANALYSIS -> REPORT
**Phase 1: IDENTIFY** -- Accept vendor name. Handle variations (legal name vs. trade name, abbreviations, parent/subsidiary).
**Phase 2: INVENTORY** -- Search for all agreements with the vendor. For each agreement found, capture: type (NDA/MSA/SOW/DPA/SLA), status (active/expired/in-negotiation), effective date, expiration date, auto-renewal details, key terms.
**Phase 3: GAP ANALYSIS** -- Identify what exists vs. what should exist.
Required agreements by relationship type:
- **Data processor**: NDA + MSA + DPA + SOW minimum
- **SaaS vendor**: NDA + MSA/SaaS Agreement + DPA (if personal data) + SLA
- **Professional services**: NDA + MSA + SOW(s)
- **Hardware/commodity**: NDA + Purchase Agreement
Flag: agreements expired but with surviving obligations, approaching expirations (90-day window), DPA gaps when vendor handles personal data.
**Phase 4: REPORT** -- Consolidated status report with gap analysis and upcoming actions.
**Gate**: All available sources checked. Gaps identified. Approaching deadlines flagged.
---
## LLM Failure Mode Awareness
Legal analysis is a high-risk domain for LLM failures. Load `references/legal/llm-legal-failure-modes.md` and apply these guards on every mode:
| Failure Mode | Guard |
|-------------|-------|
| Fabricated case law or citations | Verify all citations with user before including. State "verify this citation" when referencing specific law. |
| Invented regulatory requirements | Distinguish between "this regulation requires X" (high confidence, well-known) and "check whether this applies in your jurisdiction" (lower confidence). |
| Jurisdiction confusion | Always ask which jurisdiction applies. Ask which jurisdiction applies. State which jurisdiction the analysis covers. |
| Overconfident analysis | Use calibrated language: "likely," "typically," "in most jurisdictions" rather than absolutes. |
| Missing clause interactions | Read entire contract before analyzing individual clauses. Clauses interact. |
| Stale legal knowledge | Training data has a cutoff. Recommend counsel verify current regulatory state, especially for recently enacted or amended laws. |
---
## Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| No contract provided | User asks for review without document | Prompt for document in any format |
| Ambiguous jurisdiction | Multi-jurisdiction deal | Ask user to specify primary jurisdiction. Note differences. |
| No playbook configured | First use, no organizational standards | Proceed with market-standard positions. Note clearly. |
| Contract too long (50+ pages) | Large agreement | Offer to focus on most material sections first, then complete review |
| Conflicting regulations | Cross-border requirements clash | Flag conflicts explicitly. Do not pick a winner. Recommend counsel. |
| Template needed for unknown type | No template for the inquiry type | Help user create a template following the creation guide in references |
references/legal/compliance-frameworks.md
# Compliance Frameworks
Regulation-specific requirements, checklist gates, and jurisdiction mapping for compliance checks. Covers major privacy, financial, and health data frameworks.
---
## Framework Quick Reference
| Framework | Jurisdiction | Applies When | Key Differentiator |
|-----------|-------------|-------------|-------------------|
| GDPR | EU/EEA + any org processing EU residents' data | Processing personal data of EU individuals | Broadest scope, highest fines (4% global revenue) |
| CCPA/CPRA | California | Business meets revenue/data thresholds + CA residents | Opt-out model, "sale" broadly defined |
| HIPAA | US | Covered entities + business associates handling PHI | Criminal penalties, BAA required |
| SOX | US | Public companies + accounting firms | Internal controls over financial reporting |
| PCI DSS | Global | Any entity storing/processing/transmitting cardholder data | Self-assessment or QSA audit |
| SOC 2 | Global | Service organizations (trust services criteria) | Type I (point-in-time) vs. Type II (period) |
---
## GDPR (General Data Protection Regulation)
### Scope Determination
Applies if ANY of these are true:
- Organization is established in the EU/EEA
- Organization offers goods/services to individuals in the EU (even if org is outside EU)
- Organization monitors behavior of individuals in the EU
### Obligation Checklist
#### Lawful Basis (Article 6)
- [ ] Lawful basis identified and documented for EACH processing activity
- [ ] Basis is one of: consent, contract, legal obligation, vital interest, public task, legitimate interest
- [ ] If consent: freely given, specific, informed, unambiguous, withdrawable
- [ ] If legitimate interest: legitimate interest assessment (LIA) conducted and documented
- [ ] If special category data (Article 9): additional condition identified (explicit consent, employment, vital interest, etc.)
#### Data Subject Rights
| Right | Response Deadline | Extension | Key Requirements |
|-------|------------------|-----------|-----------------|
| Access (Art. 15) | 30 days | +60 days with notice | Copy of data + supplementary information |
| Rectification (Art. 16) | 30 days | +60 days with notice | Correct inaccurate data without undue delay |
| Erasure (Art. 17) | 30 days | +60 days with notice | Delete unless exemption applies (legal obligation, public interest, legal claims) |
| Restriction (Art. 18) | 30 days | +60 days with notice | Mark data, cease processing except storage |
| Portability (Art. 20) | 30 days | +60 days with notice | Structured, commonly used, machine-readable format |
| Objection (Art. 21) | Without undue delay | N/A | Must stop unless compelling legitimate grounds |
#### Breach Notification
- [ ] Process in place to detect breaches
- [ ] Supervisory authority notification within 72 hours of awareness
- [ ] Data subject notification "without undue delay" if high risk
- [ ] Breach register maintained (all breaches, not just notifiable ones)
- [ ] Template notifications prepared for common breach types
#### International Transfers
| Transfer Mechanism | When to Use | Current Status |
|-------------------|-------------|----------------|
| Adequacy decision | Transferring to a country with EU adequacy | Check current list -- subject to change |
| Standard Contractual Clauses (SCCs) | Most common mechanism for non-adequate countries | Must use June 2021 version. Correct module required (C2P, C2C, P2P, P2C). |
| Binding Corporate Rules (BCRs) | Intra-group transfers | Requires supervisory authority approval. 12-18 month process. |
| Derogations (Art. 49) | Limited, specific situations | Explicit consent, contract necessity, important public interest. Not for systematic transfers. |
Transfer Impact Assessment required for SCCs to non-adequate countries:
- [ ] Laws of the destination country assessed
- [ ] Supplementary measures identified (encryption, pseudonymization, split processing)
- [ ] Assessment documented and reviewed periodically
#### Records of Processing (Article 30)
- [ ] Records maintained for all processing activities
- [ ] Each record includes: purposes, data categories, recipients, transfers, retention, security measures
- [ ] Records available to supervisory authority on request
#### Data Protection Impact Assessment (DPIA)
Required when processing is "likely to result in a high risk":
- [ ] Systematic and extensive profiling with significant effects
- [ ] Large-scale processing of special categories or criminal data
- [ ] Large-scale systematic monitoring of public areas
- [ ] New technologies with potential high-risk impact
- [ ] Automated decision-making with legal or significant effects
DPIA must include: systematic description, necessity/proportionality, risk assessment, mitigation measures.
---
## CCPA / CPRA (California)
### Scope Determination
Applies if business meets ANY threshold:
- Annual gross revenue > $25 million
- Buys, sells, or shares personal information of 100,000+ California residents/households/devices
- Derives 50%+ of annual revenue from selling or sharing California residents' personal information
### Obligation Checklist
#### Consumer Rights
| Right | Response Deadline | Extension | Notes |
|-------|------------------|-----------|-------|
| Right to Know | 45 calendar days | +45 days with notice | Acknowledge within 10 business days |
| Right to Delete | 45 calendar days | +45 days with notice | Must direct service providers to delete too |
| Right to Opt-Out of Sale/Sharing | 15 business days | None | "Do Not Sell or Share My Personal Information" link required |
| Right to Correct | 45 calendar days | +45 days with notice | CPRA addition |
| Right to Limit Sensitive PI Use | 15 business days | None | CPRA addition. "Limit the Use of My Sensitive Personal Information" link. |
| Non-Discrimination | N/A | N/A | Cannot penalize consumers exercising rights |
#### Service Provider Requirements
- [ ] Written contract with service providers restricts PI use to specified business purpose
- [ ] Service providers certify they understand and will comply with CCPA/CPRA
- [ ] Service providers must cooperate with consumer rights requests
- [ ] Downstream service providers bound by same restrictions
#### Privacy Notice Requirements
- [ ] Notice at or before collection
- [ ] Categories of PI collected
- [ ] Purposes for each category
- [ ] Whether PI is sold or shared
- [ ] Retention period for each category
- [ ] Consumer rights and how to exercise them
- [ ] Updated at least annually
### CCPA vs. GDPR Key Differences
| Dimension | GDPR | CCPA/CPRA |
|-----------|------|-----------|
| Consent model | Opt-in (consent before processing) | Opt-out (can process until consumer opts out) |
| Scope of "personal data" | Any information relating to identified/identifiable person | Broader: includes household and device data |
| Right to delete exceptions | More limited | Broader exceptions (transaction completion, security, legal obligation, internal use) |
| Private right of action | Generally no (except through supervisory authority) | Yes, for data breaches (statutory damages $100-$750 per consumer per incident) |
| Enforcement body | Supervisory authorities (e.g., ICO, CNIL) | California Privacy Protection Agency (CPPA) + AG |
---
## HIPAA (Health Insurance Portability and Accountability Act)
### Scope Determination
Applies to:
- **Covered entities**: Health plans, healthcare clearinghouses, healthcare providers who transmit health information electronically
- **Business associates**: Any entity that creates, receives, maintains, or transmits PHI on behalf of a covered entity
### Obligation Checklist
#### Administrative Safeguards
- [ ] Security officer designated
- [ ] Risk analysis conducted (initial + periodic)
- [ ] Risk management plan implemented
- [ ] Workforce training on policies and procedures
- [ ] Sanctions policy for violations
- [ ] Information system activity review procedures
- [ ] Contingency plan (backup, recovery, emergency operations)
- [ ] Business Associate Agreements (BAAs) with all BAs
#### Technical Safeguards
- [ ] Access controls (unique user IDs, emergency access, automatic logoff, encryption)
- [ ] Audit controls (hardware, software, procedural mechanisms to record access)
- [ ] Integrity controls (mechanisms to authenticate ePHI, protect from improper alteration/destruction)
- [ ] Transmission security (encryption for ePHI in transit)
- [ ] Authentication (verify identity of persons seeking access)
#### Physical Safeguards
- [ ] Facility access controls
- [ ] Workstation use policies
- [ ] Workstation security
- [ ] Device and media controls (disposal, re-use, accountability, backup)
#### Breach Notification
| Notification | Deadline | Details |
|-------------|----------|---------|
| To individuals | Without unreasonable delay, max 60 days from discovery | Written notice by first-class mail or email (if consented) |
| To HHS | 60 days (if 500+ individuals); annual (if <500) | Via HHS breach reporting portal |
| To media | Without unreasonable delay, max 60 days | If breach affects 500+ residents of a state/jurisdiction |
#### BAA Requirements
Every Business Associate Agreement must include:
- [ ] Permitted uses and disclosures of PHI
- [ ] Prohibition on further use/disclosure beyond contract
- [ ] Appropriate safeguards requirement
- [ ] Breach reporting obligation
- [ ] Subcontractor same-obligation flow-down
- [ ] Access to PHI for individual rights requests
- [ ] Return or destruction of PHI on termination
- [ ] HHS audit/investigation cooperation
---
## SOX (Sarbanes-Oxley Act)
### Scope Determination
Applies to: US public companies and their accounting firms. Section 404 (internal controls) is the primary in-house legal touchpoint.
### Key Compliance Areas
#### Section 302 -- Corporate Responsibility for Financial Reports
- [ ] CEO and CFO certify quarterly and annual reports
- [ ] Certify: reviewed the report, no material misstatements, financial statements fairly present financial condition
- [ ] Certify: responsible for internal controls, evaluated within 90 days, disclosed changes
#### Section 404 -- Internal Controls over Financial Reporting
- [ ] Management assessment of internal controls effectiveness
- [ ] External auditor attestation on management's assessment
- [ ] Material weaknesses in controls identified and disclosed
- [ ] Remediation plans for identified weaknesses
#### Section 802 -- Criminal Penalties for Document Destruction
- [ ] Document retention policies in place
- [ ] Litigation hold procedures documented and enforced
- [ ] Prohibition on destruction of documents related to federal investigation
- [ ] 20-year maximum imprisonment for willful destruction
#### Whistleblower Protections (Section 806)
- [ ] Anti-retaliation policy in place
- [ ] Reporting channels available (anonymous option)
- [ ] Investigation procedures for reports
- [ ] No discharge, demotion, suspension, threats, harassment for reporting
---
## PCI DSS (Payment Card Industry Data Security Standard)
### Scope Determination
Applies to: any entity that stores, processes, or transmits cardholder data (CHD) or sensitive authentication data (SAD).
### 12 Requirements Summary
| Requirement | Area | Key Controls |
|-------------|------|-------------|
| 1 | Network security | Install and maintain network security controls |
| 2 | Default security | Apply secure configurations to all system components |
| 3 | Data protection | Protect stored account data (encryption, masking, hashing) |
| 4 | Transmission | Protect CHD with strong cryptography during transmission |
| 5 | Malware | Protect from malicious software |
| 6 | Software | Develop and maintain secure systems and software |
| 7 | Access control | Restrict access by business need to know |
| 8 | Authentication | Identify users and authenticate access |
| 9 | Physical | Restrict physical access to CHD |
| 10 | Logging | Log and monitor all access to system components and CHD |
| 11 | Testing | Test security of systems and networks regularly |
| 12 | Policies | Support information security with organizational policies and programs |
---
## Additional Privacy Frameworks (Quick Reference)
| Framework | Jurisdiction | Key Differentiators | Enforcement |
|-----------|-------------|--------------------|----|
| LGPD | Brazil | Similar to GDPR. DPO required. ANPD enforcement. | Fines up to 2% of revenue in Brazil (R$50M cap) |
| POPIA | South Africa | Information Regulator oversight. Required registration. | Up to R10M fine or 10 years imprisonment |
| PIPEDA | Canada (federal) | Consent-based. OPC oversight. Being modernized. | Complaints-driven, limited fines currently |
| PDPA | Singapore | Do Not Call registry. Mandatory breach notification. | Up to S$1M fine |
| PIPL | China | Strict cross-border rules. Data localization. CAC oversight. | Up to 5% of annual revenue |
| UK GDPR | United Kingdom | Post-Brexit UK version. ICO oversight. | Up to GBP 17.5M or 4% of global turnover |
| Privacy Act 1988 | Australia | 13 Australian Privacy Principles. Notifiable Data Breaches scheme. | Up to A$50M per violation |
---
## Cross-Framework Compliance Matrix
When multiple frameworks apply, use this matrix to identify the strictest requirement:
| Obligation | GDPR | CCPA/CPRA | HIPAA | Strictest |
|-----------|------|-----------|-------|-----------|
| Consent model | Opt-in | Opt-out | Authorization | GDPR (opt-in) |
| Breach notification (authority) | 72 hours | N/A | 60 days | GDPR (72 hours) |
| Breach notification (individual) | Without undue delay | N/A | 60 days | GDPR |
| Right to delete | Yes (with exceptions) | Yes (broader exceptions) | Limited | GDPR (fewer exceptions) |
| Data retention limits | Purpose limitation | Stated in notice | 6 years (HIPAA records) | Varies by data type |
| Cross-border transfers | Restricted (SCCs, etc.) | No specific restriction | BAA required | GDPR |
| DPO/Privacy officer | Required in certain cases | Not required | Security officer required | Depends on org |
| Fines (maximum) | 4% global revenue | $7,500 per violation | $1.9M per violation category | GDPR |
**Rule**: When frameworks overlap, apply the strictest requirement unless doing so would violate another framework. Document conflicts and escalate to counsel.
---
## Compliance Check Output Template
```
## Compliance Check: [Initiative]
### Quick Assessment
[Proceed / Proceed with conditions / Requires further review]
### Applicable Frameworks
| Framework | Relevance | Key Requirements |
|-----------|-----------|-----------------|
| [Framework] | [How it applies] | [What to do] |
### Requirements Checklist
| # | Requirement | Framework | Status | Action Needed |
|---|-------------|-----------|--------|---------------|
| 1 | [Req] | [Source] | [Met/Not Met/Unknown] | [Action] |
### Risk Areas
| Risk | Severity | Framework | Mitigation |
|------|----------|-----------|------------|
| [Risk] | [H/M/L] | [Source] | [How to address] |
### Approvals Needed
| Approver | Why | Framework | Status |
|----------|-----|-----------|--------|
| [Role] | [Reason] | [Source] | [Pending] |
### Recommended Actions (Priority Order)
1. [Action] -- [Deadline if applicable]
```
references/legal/contract-review.md
# Contract Review: Clause-by-Clause Methodology
Full analysis methodology for contract review. Covers every material clause category with key review points, common issues, and LLM-specific failure modes.
---
## Pre-Analysis Protocol
Before analyzing any clause in isolation:
1. **Read the entire contract.** Clauses interact. An uncapped indemnity may be partially mitigated by a broad limitation of liability. A narrow confidentiality definition may be expanded by a separate DPA.
2. **Identify the contract type.** SaaS, professional services, license, partnership, procurement. The type determines which clauses are most material.
3. **Determine the user's side.** Vendor, customer, licensor, licensee, partner. This fundamentally changes the analysis -- limitation of liability protections favor different parties depending on position.
4. **Note the governing law.** Jurisdiction affects enforceability of specific clauses (e.g., non-compete enforceability, consequential damages exclusions, indemnification scope).
---
## Clause Categories
### 1. Limitation of Liability (LOL)
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Cap amount | Fixed dollar, multiple of fees, or uncapped |
| Cap symmetry | Mutual or different per party |
| Cap carveouts | What liabilities are excluded from the cap |
| Consequential damages | Excluded? Mutual exclusion? |
| Consequential carveouts | What escapes the consequential damages exclusion |
| Cap period | Per-claim, per-year, or aggregate |
**Common issues:**
- Cap at a fraction of fees paid (e.g., "fees paid in prior 3 months" on a low-value contract) -- effective cap may be negligibly small
- Asymmetric carveouts favoring the drafter -- one party's breaches are capped while the other's are not
- Broad carveouts that eliminate the cap ("any breach of Section X" where X covers most obligations)
- No consequential damages exclusion for one party
- Cap referencing "fees paid" rather than "fees payable" -- disadvantages the receiving party early in a multi-year deal
**LLM failure modes:**
- Analyzing LOL without cross-referencing indemnification clause -- an indemnity carved out from the cap effectively creates uncapped liability
- Missing that "super cap" carveouts (IP infringement, confidentiality breach, data breach) are common and often at a higher multiple (2-3x fees)
- Failing to note that consequential damages exclusions may be unenforceable in certain jurisdictions or for certain claim types (e.g., willful misconduct)
- Assuming US-style LOL structure when reviewing contracts governed by civil law jurisdictions
**GREEN**: Cap at 12+ months of fees (or higher), mutual, standard carveouts, mutual consequential damages exclusion.
**YELLOW**: Cap at 6-12 months of fees, minor asymmetry, 1-2 additional carveouts.
**RED**: No LOL clause, uncapped liability, cap below 6 months, broad asymmetric carveouts.
---
### 2. Indemnification
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Mutual vs. unilateral | Both parties indemnify, or only one |
| Triggers | IP infringement, data breach, bodily injury, breach of reps |
| Cap | Subject to LOL cap, separate cap, or uncapped |
| Procedure | Notice requirements, right to control defense, right to settle |
| Mitigation | Indemnitee obligation to mitigate |
| LOL relationship | How indemnification interacts with the liability cap |
**Common issues:**
- Unilateral indemnification for IP infringement when both parties contribute IP
- "Any breach" indemnification -- converts the entire agreement into uncapped liability
- No right to control defense of claims
- Indefinite survival of indemnification obligations
- No mitigation obligation on the indemnitee
- Indemnification for third-party claims only vs. direct claims -- scope matters
**LLM failure modes:**
- Failing to check whether indemnification obligations are carved out from the LOL cap -- this is the most common miss
- Not recognizing that "indemnify and hold harmless" may have different legal meanings in some jurisdictions
- Missing the interplay between indemnification scope and insurance requirements
- Overstating indemnification risks without noting that practical enforcement requires third-party claims in most formulations
**GREEN**: Mutual for core risks (IP, data breach), capped, standard procedures.
**YELLOW**: Unilateral IP indemnification (common market position), cap at LOL cap level.
**RED**: Uncapped, "any breach" scope, no defense control, indefinite survival.
---
### 3. Intellectual Property
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Pre-existing IP | Each party retains their own |
| Developed IP | Ownership of IP created during engagement |
| Work-for-hire | Scope of work-for-hire provisions |
| License grants | Scope, exclusivity, territory, sublicensing |
| Open source | Obligations and restrictions |
| Feedback clauses | Grants on suggestions or improvements |
**Common issues:**
- Broad IP assignment capturing the customer's pre-existing IP
- Work-for-hire extending beyond deliverables to tools, methodologies, or frameworks
- Unrestricted feedback clauses granting perpetual, irrevocable licenses on any suggestion
- License scope broader than the business relationship requires
- No open source disclosure or compliance obligations
- Assignment of "improvements" or "derivative works" without defining these terms
**LLM failure modes:**
- Confusing IP ownership with IP licensing -- assignment transfers ownership permanently; a license grants usage rights
- Not flagging that work-for-hire has specific legal requirements (US Copyright Act Section 101) and may not apply to all work products
- Missing that "background IP" / "pre-existing IP" definitions are critical -- if undefined, disputes arise about what each party brought to the relationship
- Failing to note that IP provisions may be unenforceable without adequate consideration
**GREEN**: Pre-existing IP retained, developed IP ownership appropriate for deal structure, reasonable license scope.
**YELLOW**: Broad feedback clause, license scope wider than needed, work-for-hire scope slightly broad.
**RED**: Pre-existing IP assignment, unrestricted work-for-hire, no IP provisions at all.
---
### 4. Data Protection
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| DPA requirement | Is a DPA needed? Is one attached? |
| Controller/processor | Correct classification of roles |
| Sub-processors | Rights and notification obligations |
| Breach notification | Timeline (72 hours for GDPR) |
| Cross-border transfers | SCCs, adequacy decisions, BCRs |
| Deletion/return | Obligations on termination |
| Security | Requirements and audit rights |
| Purpose limitation | Processing limited to stated purposes |
**DPA review checklist (GDPR Article 28):**
- Subject matter, duration, nature, purpose clearly defined
- Type of personal data and categories of data subjects specified
- Processor processes only on documented instructions
- Confidentiality commitments for personnel
- Appropriate technical and organizational security measures (Article 32)
- Sub-processor requirements: written authorization, notification of changes, same obligations flow down
- Data subject rights assistance
- Breach notification without undue delay (24-48 hours to enable 72-hour regulatory deadline)
- Deletion or return on termination
- Audit rights (SOC 2 Type II + right to audit upon cause is standard compromise)
**Common issues:**
- No DPA when personal data is processed
- Blanket sub-processor authorization without notification
- Breach notification timeline exceeding regulatory requirements
- No cross-border transfer protections for international data flows
- Inadequate deletion provisions (no timeline, no certification)
- Outdated SCCs (must use June 2021 EU SCCs)
- No data processing locations specified
**LLM failure modes:**
- Assuming GDPR applies universally -- must check which regulations apply based on data subjects' locations and organization's presence
- Confusing data controller and data processor roles -- misclassification changes the entire obligation structure
- Not recognizing that SCC modules matter (C2P, C2C, P2P, P2C) -- wrong module invalidates the transfer mechanism
- Stating specific breach notification timelines without noting jurisdiction variations
- Missing that DPA liability must align with (not conflict with) the main services agreement
**GREEN**: DPA attached, correct roles, standard sub-processor provisions, compliant transfer mechanisms, audit rights.
**YELLOW**: DPA present but missing one element, breach timeline slightly long, general sub-processor authorization with notification.
**RED**: No DPA for personal data processing, no transfer protections, blanket sub-processor authorization, no audit rights.
---
### 5. Confidentiality
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Scope | Definition of confidential information |
| Term | Duration of obligations |
| Standard carveouts | Public knowledge, prior possession, independent development, third-party receipt, legal compulsion |
| Return/destruction | Obligations and timeline on termination |
| Permitted disclosures | Employees, contractors, advisors, affiliates |
| Residuals | Whether residuals clause exists and its scope |
**Common issues:**
- Overbroad definition capturing all information regardless of marking
- Missing independent development carveout (creates risk that internal work is claimed as derived)
- Perpetual obligations without trade secret justification
- No retention exception for legal/compliance backups
- Broad residuals clause effectively granting a license to use confidential information
**LLM failure modes:**
- Not distinguishing between confidentiality provisions in standalone NDAs vs. embedded in commercial agreements (different materiality)
- Missing that "residuals" clauses are contentious and often a deal point
- Failing to cross-reference confidentiality with the separate DPA (if any) for data-specific obligations
**GREEN**: Reasonable scope, standard carveouts, 2-5 year term, return/destruction with retention exception.
**YELLOW**: Broader scope, 5-7 year term, missing one carveout, narrow residuals.
**RED**: Overbroad definition, perpetual term, missing critical carveouts, broad residuals.
---
### 6. Representations and Warranties
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Scope | What is warranted (authority, non-infringement, functionality, compliance) |
| Disclaimers | "As-is" disclaimers, exclusion of implied warranties |
| Survival | How long warranties survive after termination |
| Remedy | Exclusive remedy for warranty breach vs. general remedies |
**Common issues:**
- No warranty of non-infringement from the provider
- "As-is" disclaimer on services that should have functionality warranties
- Warranty survival too short to discover issues
- No remedy specified for warranty breach (defaults to general remedies, which may be more favorable)
**LLM failure modes:**
- Failing to note that warranty disclaimers may be unenforceable for certain types of warranties in certain jurisdictions (e.g., implied warranties of merchantability under UCC may require conspicuous disclaimer)
- Not cross-referencing warranties with indemnification -- a warranty of non-infringement typically pairs with an IP indemnification obligation
**GREEN**: Mutual authority reps, appropriate functionality warranties, reasonable disclaimers, 12+ month survival.
**YELLOW**: Limited warranties, short survival, broad disclaimers on services.
**RED**: No warranties, warranty period shorter than a billing cycle, hidden "as-is" for core deliverables.
---
### 7. Term and Termination
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Initial term | Length and appropriateness for deal |
| Renewal | Auto-renewal terms, notice period, renewal term length |
| Termination for convenience | Available? Notice period? Early termination fees? |
| Termination for cause | Cure period? What constitutes cause? |
| Effects | Data return, transition assistance, survival clauses |
| Wind-down | Period and obligations |
**Common issues:**
- Long initial terms with no termination for convenience
- Auto-renewal with short notice windows (30 days for annual renewal is too short)
- No cure period for termination for cause
- Inadequate transition assistance provisions
- Survival clauses that effectively extend the agreement indefinitely
- Early termination fees calculated on remaining term rather than actual damages
**LLM failure modes:**
- Not flagging the interaction between auto-renewal and notice periods -- a 30-day notice window on a 3-year renewal effectively locks in for 3 years if the window is missed
- Missing that "termination for cause" definitions vary -- "material breach" is standard but "any breach" lowers the threshold dramatically
- Failing to check what happens to paid-but-unused fees on termination
**GREEN**: Reasonable term, 90+ day renewal notice, termination for convenience, 30-day cure, adequate transition.
**YELLOW**: 60-day renewal notice, limited termination for convenience, 15-day cure.
**RED**: No termination for convenience, auto-renewal with <30-day notice, no cure period, no transition assistance.
---
### 8. Governing Law and Dispute Resolution
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Choice of law | Governing jurisdiction |
| Dispute mechanism | Litigation, arbitration, mediation first |
| Venue | Court or arbitration seat |
| Arbitration rules | Which rules, number of arbitrators, seat |
| Jury waiver | Present? Mutual? |
| Class action waiver | Present? Enforceable? |
| Prevailing party fees | Attorney's fees provision |
**Common issues:**
- Unfavorable jurisdiction (unusual, remote, or hostile venue)
- Mandatory arbitration with rules favoring the drafter (e.g., drafter chooses arbitrator)
- Jury waiver without corresponding protections
- No escalation process before formal dispute resolution
- Governing law and venue in different jurisdictions (creates conflicts)
**LLM failure modes:**
- Assuming all jurisdictions are equivalent -- Delaware, New York, California, and England are standard commercial jurisdictions; others may present enforcement challenges
- Not recognizing that arbitration vs. litigation choice has significant cost, speed, confidentiality, and appeal implications
- Stating that a particular governing law is "unfavorable" without asking what the user's preferred jurisdiction is
- Missing that choice of law may not be enforceable for certain claims (e.g., employment claims, consumer protection)
**GREEN**: Well-established commercial jurisdiction, consistent governing law and venue, reasonable dispute mechanism.
**YELLOW**: Acceptable but non-preferred jurisdiction, mandatory arbitration with standard rules.
**RED**: Problematic jurisdiction, mandatory arbitration with drafter-favorable rules, governing law and venue misaligned.
---
### 9. Insurance
**Key elements to review:**
| Element | What to Check |
|---------|--------------|
| Coverage types | CGL, E&O/professional liability, cyber, workers' comp |
| Minimums | Dollar amounts appropriate for deal size and risk |
| Evidence | Certificate of insurance requirements |
| Additional insured | Whether the counterparty must be named |
| Notification | Obligation to notify of policy changes or cancellation |
**Common issues:**
- No insurance requirements when the deal involves material risk
- Minimums too low for the contract value
- No cyber insurance when data is being processed
- No requirement to maintain coverage for the term of the agreement
**GREEN**: Appropriate coverage types, reasonable minimums, standard COI requirements.
**YELLOW**: Missing one coverage type, minimums below ideal but present.
**RED**: No insurance clause for a material engagement, or minimums grossly inadequate.
---
### 10. Assignment
**Key elements to review:**
- Consent requirements for assignment
- Change of control provisions
- Exceptions (affiliates, restructuring)
- Whether consent can be unreasonably withheld
**Common issues:**
- No assignment clause (defaults to applicable law, which varies)
- Unilateral assignment right for one party only
- Change of control exception that allows assignment to a competitor
---
### 11. Force Majeure
**Key elements to review:**
- Scope of qualifying events
- Notification requirements
- Duration threshold before termination rights trigger
- Whether pandemic/epidemic is included
- Whether the clause covers performance obligations only or payment obligations too
---
### 12. Payment Terms
**Key elements to review:**
- Net terms (Net 30 standard; Net 60+ disadvantages the payee)
- Late payment interest and fees
- Tax responsibilities
- Price escalation mechanisms
- Currency and exchange rate risk
- Right to suspend services for non-payment
---
## Redline Generation Standards
When generating redlines for YELLOW and RED deviations:
```
**Clause**: [Section reference and clause name]
**Current language**: "[exact quote]"
**Proposed redline**: "[specific alternative language]"
**Rationale**: [1-2 sentences, suitable for counterparty]
**Priority**: [Must-have / Should-have / Nice-to-have]
**Fallback**: [Alternative position if primary redline rejected]
```
**Rules:**
- Be specific -- provide exact language ready to insert, not vague guidance
- Be balanced -- firm on critical points, commercially reasonable on others
- Provide fallback positions for YELLOW items
- Prioritize -- not all redlines are equal
- Consider the relationship -- new vendor vs. strategic partner vs. commodity supplier
---
## Negotiation Priority Framework
**Tier 1 -- Must-Haves (Deal Breakers):**
- Uncapped or materially insufficient liability protections
- Missing data protection requirements for regulated data
- IP provisions jeopardizing core assets
- Terms conflicting with regulatory obligations
**Tier 2 -- Should-Haves (Strong Preferences):**
- LOL cap adjustments within range
- Indemnification scope and mutuality
- Termination flexibility
- Audit and compliance rights
**Tier 3 -- Nice-to-Haves (Concession Candidates):**
- Preferred governing law (if alternative is acceptable)
- Notice period preferences
- Minor definitional improvements
- Insurance certificate requirements
**Strategy**: Lead with Tier 1. Trade Tier 3 concessions to secure Tier 2 wins. Never concede Tier 1 without escalation.
---
## Holistic Assessment Checklist
After individual clause analysis, assess the contract holistically:
- [ ] Overall risk allocation balanced or appropriately skewed for the deal structure?
- [ ] Liability protections consistent across clauses (no clause undermining another)?
- [ ] Indemnification and LOL interact correctly?
- [ ] Data protection provisions adequate for the data involved?
- [ ] Termination rights provide adequate exit options?
- [ ] Survival clauses reasonable in scope and duration?
- [ ] Commercial terms (payment, pricing, SLA) align with business expectations?
references/legal/german-business-compliance.md
# German & EU Business Compliance Frameworks
Regulation-specific requirements for German business operations, digital services, electronic identity, and AI systems.
---
## Framework Quick Reference
| Framework | Jurisdiction | Applies When | Key Differentiator |
|-----------|-------------|-------------|-------------------|
| GoBD | Germany | Tax-relevant digital records | Immutability, audit trails, Verfahrensdokumentation |
| HGB §§238-261 | Germany | All German merchants | Legal foundation for bookkeeping; GoBD operationalizes |
| TDDDG | Germany | All digital service providers | Cookie consent, ePrivacy (was TTDSG, renamed May 2024) |
| eIDAS 2.0 | EU | Any org accepting ID verification | Digital wallets by end 2026, open source requirement |
| EU AI Act | EU | Any org deploying AI systems | Risk-based tiers, mandatory logging, EUR 35M fines |
---
## GoBD (Grundsätze zur ordnungsmäßigen Führung und Aufbewahrung von Büchern)
Federal Ministry of Finance regulation for digital bookkeeping and data retention.
### 2025 Changes
- **E-invoicing mandate (Jan 1, 2025 receiving; Jan 1, 2027 issuing >EUR 800K)**: XML component of ZUGFeRD is legally relevant for archiving
- **Retention reduced**: 10 → 8 years for invoices/receipts (Bürokratieentlastungsgesetz IV, Jan 1, 2025)
- **PDF storage**: If PDF is generatable from XML, separate PDF storage no longer required
### Core Principles
Integrity, Authenticity, Accessibility, Immutability, Traceability.
### Verfahrensdokumentation (Mandatory Since 2015)
Required for every IT-supported tax-relevant process. Four components:
1. General description
2. User documentation
3. Technical system documentation
4. Operations documentation
### Code-Level Checks
| ID | Check | GoBD Ref |
|----|-------|----------|
| GB-01 | Audit trail on all tax-relevant data (user ID, timestamp, old/new value) | Immutability |
| GB-02 | WORM-like write protection for archived records | Retention |
| GB-03 | Verfahrensdokumentation covers all IT-supported tax processes | §151 |
| GB-04 | E-invoices archived in original machine-readable XML format | 2025 Amendment |
| GB-05 | Change management documented for system updates affecting tax data | Verfahrensdoku |
| GB-06 | Data retention configured for 8-year minimum (post-Dec 2024 records) | BEG IV |
| GB-07 | Role-based access control on tax-relevant data with integrity checks | Access/Integrity |
---
## HGB §§238-261 (Handelsgesetzbuch Third Book)
Commercial bookkeeping and accounting requirements for German merchants.
### Key Sections
| Section | Requirement |
|---------|------------|
| §238 | Merchants must keep books recording all business transactions |
| §239 | Clear records, no blank spaces, no erasures |
| §243-245 | Annual financial statements: balance sheet + P&L |
| §257 | Retention: 10y financial statements, 8y invoices (BEG IV) |
**HGB = the WHAT; GoBD = the HOW for electronic systems.**
### Code-Level Checks
| ID | Check | HGB Ref |
|----|-------|---------|
| HG-01 | Double-entry bookkeeping enforced in accounting system | §238 |
| HG-02 | Retention periods configured (10y statements, 8y invoices) | §257 + BEG IV |
| HG-03 | Records immutable after finalization | §239 |
| HG-04 | Financial statements generated with complete balance sheet + P&L | §243-245 |
| HG-05 | Digital records maintain audit-ready accessibility | §257 |
---
## TDDDG (Telekommunikation-Digitale-Dienste-Datenschutz-Gesetz)
Germany's ePrivacy implementation. Renamed TTDSG → TDDDG May 13, 2024 (aligned with EU Digital Services Act).
### Section 25: Cookie/Tracking Consent
- Consent required before storing/accessing info on user terminal
- Exceptions: transmission of communications, strictly necessary for requested service
- "Accept" and "Decline" with equal prominence
- Pre-filled not permitted
- Purpose, duration, third-party access disclosed
- GDPR-compliant consent standard
**Penalties**: Max EUR 300,000 per violation.
### Code-Level Checks
| ID | Check | TDDDG Ref |
|----|-------|-----------|
| TD-01 | Cookie consent banner with equal-prominence Accept/Decline buttons | §25(1) |
| TD-02 | No non-essential cookies set before explicit consent | §25(1) |
| TD-03 | Consent not pre-filled or pre-selected | §25(1) |
| TD-04 | Purpose, duration, third-party access disclosed before consent | §25(1) |
| TD-05 | Strictly-necessary exception properly scoped (no analytics in exception) | §25(2) |
| TD-06 | Consent management follows TDDDG ordinance standards | Consent Ordinance |
---
## eIDAS 2.0 (Regulation 2024/1183)
European Digital Identity Framework. Wallets by end 2026. Regulated industries must accept wallets by December 2027.
### New Trust Services
- QEAA (Qualified Electronic Attestation of Attributes)
- Electronic Archiving
- Electronic Ledgers
- EUDI Wallet
### Developer Requirements
- Wallet components must be open source
- Zero-tracking/profiling by design
- Privacy dashboard
- Selective disclosure
- Local data storage
### Code-Level Checks
| ID | Check | eIDAS Ref |
|----|-------|-----------|
| EI-01 | Electronic signatures implement correct eIDAS level (simple/advanced/qualified) | Art. 25-26 |
| EI-02 | QEAA verification against qualified trust service provider | Art. 45 |
| EI-03 | EUDI Wallet integration accepts wallet-presented credentials | Art. 5a |
| EI-04 | Privacy-by-design: no tracking/profiling of wallet usage | Art. 5a |
| EI-05 | Selective disclosure: request only necessary attributes | Art. 5a |
| EI-06 | Open source wallet components publicly available | Art. 5a |
---
## EU AI Act (Regulation 2024/1689)
Risk-based: Unacceptable (banned Feb 2025) → High-Risk (Aug 2026) → Limited (transparency) → Minimal.
### Timeline
| Date | Milestone |
|------|-----------|
| Feb 2, 2025 | Prohibited practices + AI literacy obligations |
| Aug 2, 2025 | GPAI obligations |
| Aug 2, 2026 | High-risk full compliance |
| Aug 2, 2027 | Annex I regulated products |
### Developer-Facing Requirements (Articles 9-15)
| Article | Requirement |
|---------|------------|
| Art. 9 | Lifecycle risk identification, residual risk documentation, continuous monitoring |
| Art. 11 | Technical documentation (9 Annex IV categories incl. architecture, ADRs, model cards, training data, testing). Retained 10 years. |
| Art. 12 | Automatic logging as architectural requirement (not bolt-on). Schema: user, timestamp, spec version, model ID, input, output, reviewer, tests. Retention: 6 months minimum. |
| Art. 13 | Transparency -- disclose GPAI model, spec, human review, modifications |
| Art. 14 | Human oversight -- override/halt mechanisms, automation bias awareness |
| Art. 15 | Accuracy & robustness testing with documented metrics |
### High-Risk Triggers for Dev Teams
- AI as safety component in regulated products (medical, automotive)
- AI evaluating/screening/monitoring PEOPLE (accidental trigger: dev performance dashboards, AI-based PR routing)
- AI in critical infrastructure
### GPAI Open Source Exception
Exempt from documentation unless systemic risk.
### Penalties
| Tier | Fine | Applies To |
|------|------|-----------|
| Prohibited practices | EUR 35M / 7% global revenue | Art. 5 violations |
| High-risk non-compliance | EUR 15M / 3% global revenue | Arts. 9-15 violations |
| Incorrect information | EUR 7.5M / 1% global revenue | Misleading declarations |
### Code-Level Checks
| ID | Check | AI Act Ref |
|----|-------|-----------|
| AI-01 | AI system risk classification documented | Art. 6 |
| AI-02 | Technical documentation with 9 Annex IV categories | Art. 11 |
| AI-03 | Automatic logging architecture with provenance schema | Art. 12 |
| AI-04 | Human oversight mechanism (override/halt capability) | Art. 14 |
| AI-05 | Accuracy and robustness testing with documented metrics | Art. 15 |
| AI-06 | Risk management system with lifecycle coverage | Art. 9 |
| AI-07 | Transparency disclosure for GPAI model usage | Art. 13 |
| AI-08 | Version control for all AI artifacts (code, weights, configs, datasets, docs) | Art. 11 |
| AI-09 | Prohibited AI practices not implemented (social scoring, real-time biometric mass surveillance, etc.) | Art. 5 |
| AI-10 | AI literacy training documented for staff | Art. 4 |
| AI-11 | CI/CD provenance tracking (model ID, spec version, human review) per artifact | Art. 12 |
---
## Cross-Framework Compliance Matrix (German Business)
| Pair | Relationship |
|------|-------------|
| GoBD ↔ HGB | GoBD operationalizes HGB §§238-261 for electronic systems |
| TDDDG ↔ GDPR | TDDDG is German ePrivacy; GDPR consent standards apply |
| AI Act ↔ GDPR | AI Act references GDPR for personal data in AI training |
| eIDAS ↔ DORA | Financial entities must accept EUDI Wallets by 2027 |
**Rule**: When frameworks overlap, apply the strictest requirement unless doing so would violate another framework. Document conflicts and escalate to counsel.
---
## Cross-References
These reviewer-system references cover adjacent German/EU compliance domains:
- `agents/reviewer-system/references/german-it-security.md` — BSI Grundschutz, KRITIS, NIS2UmsuCG
- `agents/reviewer-system/references/financial-resilience-de-eu.md` — DORA, KWG/MaRisk (eIDAS<>DORA interaction noted in matrix above)
- `agents/reviewer-system/references/industry-specific-compliance.md` — TISAX automotive compliance
- `agents/reviewer-system/references/sovereign-cloud-data-residency.md` — BSI C5, BDSG/DSGVO, data residency
- `agents/reviewer-system/references/compliance-checklists.md` — GDPR, SOC 2, PCI-DSS, HIPAA
---
## Compliance Check Output Template
```
## Compliance Check: [Initiative]
### Quick Assessment
[Proceed / Proceed with conditions / Requires further review]
### Applicable Frameworks
| Framework | Relevance | Key Requirements |
|-----------|-----------|-----------------|
| [Framework] | [How it applies] | [What to do] |
### Requirements Checklist
| # | Requirement | Framework | Status | Action Needed |
|---|-------------|-----------|--------|---------------|
| 1 | [Req] | [Source] | [Met/Not Met/Unknown] | [Action] |
### Risk Areas
| Risk | Severity | Framework | Mitigation |
|------|----------|-----------|------------|
| [Risk] | [H/M/L] | [Source] | [How to address] |
### Approvals Needed
| Approver | Why | Framework | Status |
|----------|-----|-----------|--------|
| [Role] | [Reason] | [Source] | [Pending] |
### Recommended Actions (Priority Order)
1. [Action] -- [Deadline if applicable]
```
references/legal/legal-writing.md
# Legal Writing
Formats, templates, and methodology for legal documents: memos, briefs, responses, meeting briefings, incident briefs. Includes escalation triggers for templated responses and template creation guide.
---
## Document Types
| Type | Purpose | Audience | Formality |
|------|---------|----------|-----------|
| Legal memo | Internal analysis of a legal question | Legal team, stakeholders | High |
| Legal brief | Summary of issue, law, and recommendation | Legal team, executives | High |
| Legal response | Templated response to common inquiry | External parties, regulators, internal teams | Medium-High |
| Meeting brief | Pre-meeting context and talking points | Meeting attendees | Medium |
| Incident brief | Rapid situation assessment | Response team, leadership | Medium-High |
---
## Legal Memo Format
```
## Legal Memorandum
**To**: [recipient]
**From**: [author]
**Date**: [date]
**Re**: [subject]
**Privileged**: [Yes/No -- mark "ATTORNEY-CLIENT PRIVILEGED / WORK PRODUCT" if applicable]
### Question Presented
[Specific legal question in one sentence. Frame precisely.]
### Short Answer
[Direct answer in 2-3 sentences. State the conclusion first.]
### Facts
[Relevant facts only. Chronological or topical organization. Distinguish established facts from assumptions.]
### Analysis
[Apply law to facts. Address counterarguments. Identify ambiguities.]
#### [Sub-issue 1]
[Analysis with supporting reasoning]
#### [Sub-issue 2]
[Analysis with supporting reasoning]
### Conclusion
[Restate answer with key reasoning. Note confidence level.]
### Recommended Action
1. [Specific action with owner and timeline]
2. [Specific action with owner and timeline]
### Open Questions
[What remains unknown. What additional information would change the analysis.]
```
**Writing standards:**
- State the conclusion first, then support it
- Distinguish "the law requires X" from "the law likely requires X" from "this is unsettled"
- Flag jurisdiction-specific variations explicitly
- Note training data limitations for recent legal developments
- Never present LLM analysis as authoritative legal opinion
---
## Legal Brief Format
```
## Legal Brief: [Topic]
**Date**: [date]
**Prepared for**: [recipient/audience]
**Privileged**: [Yes/No]
### Executive Summary
[2-3 sentence summary. Conclusion first.]
### Background
[Context and history. What led to this issue.]
### Issue
[Precise statement of the legal question or situation]
### Analysis
[Structured analysis. Use headings for sub-issues.]
### Risk Assessment
| Risk | Severity | Likelihood | Mitigation |
|------|----------|------------|------------|
| [Risk] | [H/M/L] | [H/M/L] | [How to address] |
### Recommendation
[Clear recommendation with rationale and confidence level]
### Next Steps
1. [Action -- Owner -- Deadline]
```
---
## Legal Response Templates
### Response Categories
| Category | Sub-types | Key Elements |
|----------|-----------|-------------|
| Data Subject Request (DSR) | Acknowledgment, verification, fulfillment, partial denial, full denial, extension | Applicable regulation, timeline, verification requirements, rights information |
| Litigation Hold | Initial notice, reminder, modification, release | Matter reference, preservation obligations, scope, prohibition on spoliation |
| Vendor Question | Contract status, amendment request, compliance certification, audit response | Agreement reference, specific response, caveats, next steps |
| NDA Request | Send standard form, accept with markup, decline, renew | Purpose, standard terms, execution instructions, timeline |
| Privacy Inquiry | Cookie/tracking, policy questions, data sharing, cross-border | Privacy notice reference, specific answers, privacy team contact |
| Subpoena Response | Acknowledgment, objection, extension request, compliance cover | Case reference, objections, preservation confirmation, privilege log |
| Insurance Notification | Initial claim, supplemental info, reservation of rights response | Policy number, matter description, timeline, coverage request |
### DSR Response Templates
**Acknowledgment:**
```
Subject: Your Data [Access/Deletion/Correction] Request -- Reference [ID]
Dear [Name],
We received your request dated [date] to [access/delete/correct] your personal
data under [applicable regulation].
We will respond substantively by [deadline: 30 days GDPR / 45 days CCPA].
To verify your identity, please provide: [verification requirements].
If you have questions, contact [privacy team contact].
[Rights information: right to lodge complaint with supervisory authority]
```
**Extension notification:**
```
Subject: Update on Your Data Request -- Reference [ID]
Dear [Name],
We are writing regarding your [request type] request dated [date].
Due to [complexity of request / number of requests], we require additional time
to respond. Under [regulation], we are extending our response period by
[60 days GDPR / 45 days CCPA].
We will provide our substantive response by [new deadline].
[Contact information]
```
### Litigation Hold Template
```
Subject: LEGAL HOLD NOTICE -- [Matter Name] -- Action Required
PRIVILEGED AND CONFIDENTIAL
ATTORNEY-CLIENT COMMUNICATION
Dear [Custodian Name],
You are receiving this notice because you may possess documents, communications,
or data relevant to the matter referenced above.
PRESERVATION OBLIGATION:
Effective immediately, you must preserve all documents and electronically stored
information (ESI) related to:
- Subject matter: [scope]
- Date range: [start date] to present
- Document types: [email, chat, files, voicemail, text messages, etc.]
- Systems: [list specific systems]
DO NOT delete, destroy, modify, or discard any potentially relevant materials.
This includes suspending any automatic deletion or archival processes.
WHAT TO PRESERVE:
- All emails, attachments, and calendar entries related to the subject matter
- All chat messages (Slack, Teams, etc.) related to the subject matter
- All documents, spreadsheets, presentations, and notes
- All drafts, even if a final version exists
- All voicemails and recorded calls
- All text messages and instant messages
WHAT TO DO:
1. Read this notice carefully
2. Identify all locations where relevant materials may exist
3. Take steps to preserve those materials
4. Acknowledge receipt by [deadline]
5. Contact [legal contact] with any questions
IMPORTANT: Do not discuss the substance of this hold or the related matter
with anyone outside the legal team.
Please acknowledge receipt by responding to this email by [date].
[Legal contact information]
```
### Hold Release Template
```
Subject: RELEASE OF LEGAL HOLD -- [Matter Name]
PRIVILEGED AND CONFIDENTIAL
Dear [Custodian Name],
The legal hold issued on [original date] regarding [Matter Name] is hereby released.
You may resume normal document retention and deletion practices for materials
that were subject to this hold.
Note: This release applies only to the hold identified above. If you are subject
to any other legal holds, those obligations remain in effect.
[Legal contact information]
```
---
## Escalation Triggers
Before generating ANY templated legal response, check these triggers. If any fire, recommend escalation to qualified counsel instead of generating a template.
### Universal Triggers (All Categories)
| Trigger | Why It Matters |
|---------|---------------|
| Potential litigation or regulatory investigation | Templated response could create admissions or waive rights |
| From regulator, government agency, or law enforcement | Requires individualized, counsel-reviewed response |
| Could create binding legal commitment or waiver | Template may not account for specific implications |
| Potential criminal liability | Requires specialized counsel |
| Media attention involved or likely | Response becomes public statement |
| Unprecedented situation | No template can account for novel facts |
| Multiple jurisdictions with conflicting requirements | Template designed for single jurisdiction |
| Involves executive leadership or board members | Heightened sensitivity and scrutiny |
### Category-Specific Triggers
**DSR:**
- Minor's data or request from/on behalf of a minor
- From regulatory authority (not individual)
- Data subject to litigation hold
- Current/former employee with active dispute or HR matter
- Unusually broad scope (possible fishing expedition)
- Special category data (health, biometric, genetic)
**Litigation Hold:**
- Potential criminal liability
- Unclear, disputed, or overbroad preservation scope
- Prior holds for same/related matter exist
- Conflicts with regulatory deletion requirements (GDPR right to erasure vs. hold)
- Custodian objects to hold scope
**Vendor:**
- Dispute or potential breach
- Vendor threatening litigation or termination
- Involves regulatory compliance (not just contract terms)
- Could create binding commitment or waiver
- Could affect ongoing negotiation
**NDA Request:**
- Counterparty is a competitor
- Government classified information
- Potential M&A transaction
- Unusual subject matter (AI training data, biometric data)
**Subpoena:**
- ALWAYS requires counsel review (templates are starting points only)
- Privilege issues identified
- Third-party data involved
- Cross-border production
- Unreasonable timeline
### When Trigger Fires
1. **Stop** -- Do not generate templated response
2. **Alert** -- Inform user which trigger was detected
3. **Explain** -- Why this matters and what could go wrong with a template
4. **Recommend** -- Appropriate escalation path (senior counsel, outside counsel, specific team)
5. **Offer** -- Draft marked "DRAFT -- FOR COUNSEL REVIEW ONLY" rather than final response
---
## Meeting Brief Format
```
## Meeting Brief
### Meeting Details
**Meeting**: [title] | **Date**: [date/time/timezone] | **Duration**: [duration]
**Location**: [physical/video] | **Your Role**: [advisor/presenter/negotiator/observer]
### Participants
| Name | Organization | Role | Key Interests | Notes |
|------|-------------|------|---------------|-------|
| [name] | [org] | [role] | [what they care about] | [context] |
### Agenda / Expected Topics
1. [Topic] -- [brief context]
### Background and Context
[2-3 paragraphs: relevant history, current state, why this meeting is happening]
### Open Issues
| Issue | Status | Owner | Priority |
|-------|--------|-------|----------|
| [issue] | [status] | [who] | [H/M/L] |
### Legal Considerations
[Risks, compliance issues, privilege concerns relevant to meeting topics]
### Talking Points
1. [Point with supporting context]
### Questions to Raise
- [Question] -- [why it matters]
### Decisions Needed
- [Decision] -- [options and recommendation]
### Red Lines / Non-Negotiables
[Positions that cannot be conceded, if negotiation meeting]
### Preparation Gaps
[Information not found, questions for user]
```
### Meeting-Type Preparation Guide
| Meeting Type | Additional Sections |
|-------------|-------------------|
| Deal review | Deal summary, contract status, approval requirements, counterparty dynamics |
| Board/Committee | Legal department update, risk highlights, regulatory update, pending approvals, litigation summary |
| Vendor call | Agreement status, open issues, performance metrics, negotiation objectives |
| Regulatory | Regulatory body context, matter history, compliance posture, privilege considerations |
| Litigation/Dispute | Case status, recent developments, strategy, settlement parameters |
---
## Incident Brief Format
```
## Incident Brief: [Topic]
**Prepared**: [timestamp]
**Classification**: [severity if determinable]
### Situation Summary
[What is known]
### Timeline
[Chronological events]
### Immediate Legal Considerations
- Regulatory notification deadlines (e.g., 72 hours GDPR)
- Preservation obligations
- Privilege concerns
### Relevant Agreements
[Contracts, insurance, indemnification provisions implicated]
### Recommended Immediate Actions
1. [Most urgent]
2. [Second priority]
### Information Gaps
[What is not yet known]
```
**Incident brief rules:**
- Speed over completeness. Produce quickly with available information.
- Flag litigation hold obligations immediately
- Mark as privileged if appropriate
- If data breach: flag applicable notification deadlines
- Recommend outside counsel if matter is significant
---
## Template Creation Guide
When no template exists for an inquiry type:
### 1. Define Use Case
- Inquiry type and frequency
- Typical audience and urgency
### 2. Identify Required Elements
- Legal requirements for this response type
- Organizational policies that govern it
### 3. Define Variables
- Use clear names: `{{requester_name}}`, `{{response_deadline}}`, `{{matter_reference}}`
- Distinguish what changes per use from what stays constant
### 4. Draft Template
- Clear, professional language
- All legally required elements
- Subject line template for email use
### 5. Define Escalation Triggers
- Specific situations where this template should NOT be used
### 6. Add Metadata
- Version, last reviewed date, author, approver
- Follow-up actions checklist
### Template Metadata Format
```
## Template: [Name]
**Category**: [type] | **Version**: [n] | **Last Reviewed**: [date]
**Approved By**: [name]
### Use When
- [Condition]
### Do NOT Use When (Escalation Triggers)
- [Trigger]
### Variables
| Variable | Description | Example |
|----------|-------------|---------|
| {{var}} | [what it is] | [example] |
### Body
[Template text with {{variables}}]
### Follow-Up Actions
1. [Post-send action]
```
---
## Action Item Tracking
After meetings or incident responses, capture action items:
```
## Action Items -- [Context] -- [Date]
| # | Action | Owner | Deadline | Priority | Type | Status |
|---|--------|-------|----------|----------|------|--------|
| 1 | [specific task] | [name] | [date] | [H/M/L] | [Legal/Business/External] | Open |
```
**Rules:**
- Be specific ("Send redline of Section 4.2" not "Follow up on contract")
- One owner per item (not a team)
- Specific date, not "soon" or "ASAP"
- Note dependencies on other actions or external input
- Distinguish: legal team actions, business team actions, external actions, follow-up meetings
**Review cadence:**
- High priority: daily until completed
- Medium priority: weekly
- Low priority: monthly
- Overdue: escalate to owner and manager
references/legal/llm-legal-failure-modes.md
# LLM Legal Failure Modes
Where LLMs fail in legal analysis. Specific failure patterns, detection methods, and mitigation guards. Load this reference for every legal workflow mode.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers legal-specific failures only.
---
## Why Legal Is a High-Risk Domain for LLMs
Legal analysis has properties that amplify LLM weaknesses:
1. **Precision matters.** A single word ("shall" vs. "may", "and" vs. "or") changes legal meaning. LLMs optimize for plausible text, not precise legal language.
2. **Authority matters.** Legal analysis requires citing actual statutes, regulations, and case law. LLMs fabricate citations that look real but are not.
3. **Jurisdiction matters.** The same legal question has different answers in different jurisdictions. LLMs blend jurisdictions without flagging it.
4. **Currency matters.** Laws change. Training data has a cutoff. LLMs do not know what they do not know about recent changes.
5. **Consequences are asymmetric.** A plausible-sounding wrong answer can cause material harm. In most domains, a plausible wrong answer is merely unhelpful.
---
## Failure Mode Catalog
### 1. Fabricated Case Law and Citations
**What happens:** LLM generates realistic-looking case citations, statute numbers, or regulatory references that do not exist. Case names follow real naming conventions. Statute formats look correct. But the citations are invented.
**Detection signals:**
- Case name follows "[Name] v. [Name]" pattern but does not appear in legal databases
- Statute citation has correct format but wrong section number
- Regulation cited does not exist or was repealed
- LLM provides citation with unusual confidence (no hedging language)
**Mitigation:**
- Mark every case citation, statute number, and regulatory reference as "requires verification" and instruct the user to confirm with primary sources
- Use hedging: "regulations such as..." or "provisions similar to..." rather than citing specific sections
- For well-known, foundational statutes (GDPR, CCPA, HIPAA), cite the act name and general provision (e.g., "GDPR Article 28") but avoid citing specific subsections unless highly confident
- When the user needs specific citations, recommend legal research tools (Westlaw, Lexis, regulatory authority websites)
- Always include: "Verify all citations with authoritative sources before relying on them"
**Severity: CRITICAL.** Fabricated citations in legal filings have led to sanctions against attorneys. This is not a theoretical risk.
---
### 2. Jurisdiction Confusion
**What happens:** LLM blends legal principles from different jurisdictions without flagging the mix. Applies US law to UK contracts, EU regulations to Australian situations, California requirements to Texas entities. Often presents blended analysis as if it were universally applicable.
**Detection signals:**
- Analysis uses terms from one jurisdiction while discussing another (e.g., "reasonable person" standard in a civil law jurisdiction)
- Regulatory requirements cited are from a different jurisdiction than the one governing the contract
- Analysis assumes common law principles in a civil law jurisdiction (or vice versa)
- No jurisdiction specified but analysis implicitly assumes one
**Mitigation:**
- Always ask which jurisdiction governs before analyzing
- State explicitly which jurisdiction the analysis covers at the top of every output
- When multiple jurisdictions apply, analyze separately and note conflicts
- Flag when analysis may not apply to other jurisdictions: "This analysis applies to [jurisdiction]. Requirements may differ in other jurisdictions."
- Never default to US law -- ask first
- For cross-border matters, note that each element may be governed by different law
**Severity: HIGH.** Wrong jurisdiction analysis is worse than no analysis -- it creates false confidence.
---
### 3. Overconfident Analysis
**What happens:** LLM presents uncertain legal conclusions with the same confidence as well-established ones. Does not distinguish between settled law, evolving standards, and genuinely unsettled questions. Uses definitive language ("this violates..." "this is enforceable...") when the correct answer is "it depends" or "this is unsettled."
**Detection signals:**
- Absolute language without hedging: "This clause is unenforceable" (should be "this clause may be unenforceable in [jurisdiction] because...")
- No acknowledgment of counterarguments or alternative interpretations
- Presenting majority position as unanimous
- No mention of factual dependencies that could change the analysis
- Stating legal conclusions without identifying the governing standard or test
**Mitigation:**
- Use calibrated language consistently:
| Instead of | Use |
|-----------|-----|
| "This is illegal" | "This likely violates [specific provision] in [jurisdiction], though enforcement depends on [factors]" |
| "This is enforceable" | "Courts in [jurisdiction] have generally enforced similar provisions, subject to [conditions]" |
| "This is standard" | "This is common in [market/industry], though terms vary" |
| "You must do X" | "Under [regulation] in [jurisdiction], organizations meeting [criteria] are required to [obligation]" |
- Always note when an area of law is:
- Settled (high confidence)
- Evolving (moderate confidence, note recent developments)
- Unsettled (low confidence, recommend counsel)
- Present both sides when reasonable arguments exist on each
**Severity: HIGH.** Overconfidence in legal analysis leads to uninformed decision-making.
---
### 4. Invented Regulatory Requirements
**What happens:** LLM fabricates specific regulatory requirements that sound plausible but do not exist. May invent filing deadlines, notification requirements, approval thresholds, or compliance obligations. Often combines real requirements from different frameworks into a fictional composite.
**Detection signals:**
- Specific numeric thresholds or deadlines that cannot be verified
- Requirements attributed to a regulation that does not actually contain them
- Composite requirements mixing elements from different frameworks
- Requirements that are "too specific to be wrong" but are not from the cited source
**Mitigation:**
- For well-known requirements (GDPR 72-hour breach notification, CCPA 45-day response), cite confidently
- For specific thresholds, deadlines, or filing requirements: state the general obligation and recommend verifying specific requirements with the regulatory authority or counsel
- Distinguish between: "this regulation requires [well-known general obligation]" and "verify the specific thresholds and deadlines with current regulatory guidance"
- Never invent specific dollar amounts for fines or penalties unless they are well-established maximums (e.g., GDPR 4% of global turnover)
- When in doubt, describe the type of obligation without fabricating specifics
**Severity: HIGH.** Acting on invented requirements wastes resources. Missing real requirements creates liability.
---
### 5. Missing Clause Interactions
**What happens:** LLM analyzes each contract clause in isolation, missing critical interactions between clauses. An indemnification clause that is "uncapped" may be effectively limited by a well-crafted LOL with narrow carveouts. A confidentiality clause may be expanded by a separate DPA. LLM flags issues that are already mitigated elsewhere, or misses issues created by clause combinations.
**Detection signals:**
- Flagging a risk in one clause without checking if another clause mitigates it
- Not cross-referencing indemnification with LOL
- Analyzing confidentiality without reference to the DPA (if any)
- Missing that survival clauses extend obligations beyond the term
- Not noting that definitions in one section affect terms used throughout
**Mitigation:**
- Read the entire contract before analyzing individual clauses (explicit instruction in contract review workflow)
- After individual clause analysis, perform a holistic cross-reference check:
- Does LOL adequately cap indemnification? Are indemnification carveouts consistent with LOL carveouts?
- Do confidentiality provisions and DPA provisions align? Conflict?
- Do termination provisions account for survival clauses?
- Are defined terms used consistently throughout?
- Flag clause interactions explicitly in the analysis
- Note when one clause mitigates a risk flagged in another
**Severity: MEDIUM-HIGH.** Missing interactions leads to either false alarms (flagging mitigated risks) or missed risks (not seeing how clauses combine to create exposure).
---
### 6. Stale Legal Knowledge
**What happens:** LLM's training data has a cutoff date. Laws enacted, amended, or repealed after that date are unknown to the model. LLM may confidently apply repealed provisions, miss new requirements, or be unaware of significant court decisions.
**Detection signals:**
- References to laws that may have been amended (check for recent privacy laws, which change frequently)
- Analysis based on frameworks that have been superseded (e.g., old EU SCCs)
- No mention of well-known recent developments in the area
- Confidently stating current status of rapidly evolving areas (AI regulation, cryptocurrency, social media)
**Areas of highest staleness risk:**
- Privacy and data protection (new state laws in US, EU developments, international frameworks)
- AI regulation (rapidly evolving globally)
- Cryptocurrency and digital assets
- Cross-border data transfers (adequacy decisions change)
- Employment law (remote work, gig economy, non-compete restrictions)
- ESG and sustainability reporting
**Mitigation:**
- Include a standard caveat for any regulatory analysis: "This analysis is based on information available as of [training cutoff]. Verify current regulatory requirements with authoritative sources."
- For rapidly evolving areas, explicitly flag that the law may have changed and recommend checking with counsel or regulatory authority websites
- Never state "this is current law" -- state "as of [knowledge cutoff], the applicable framework is..."
- Recommend specific authoritative sources for current information (regulatory authority websites, not secondary sources)
**Severity: MEDIUM-HIGH.** Applying repealed law or missing new requirements can cause concrete harm.
---
### 7. Template Overfitting
**What happens:** LLM applies a standard template or framework to a situation that does not fit. Uses boilerplate language when the situation requires specific, tailored analysis. Generates a response that looks professional but does not address the actual question.
**Detection signals:**
- Response follows a clear template but does not address the specific facts
- Boilerplate language that could apply to any situation
- Missing engagement with the unusual or difficult aspects of the question
- Response is generic where specific analysis was requested
**Mitigation:**
- After generating any templated response, check: does this address the specific facts and question?
- For unusual situations, note explicitly that standard templates may not apply
- When escalation triggers fire, do not force-fit a template -- escalate
- Include the specific facts in the analysis, not just the framework
**Severity: MEDIUM.** Template overfitting produces responses that look helpful but are not.
---
### 8. False Equivalence in Risk Assessment
**What happens:** LLM treats all risks as roughly equivalent, compressing scores toward the middle (3-4 range) and failing to differentiate between genuinely critical risks and minor concerns. May also assign extreme scores without justification.
**Detection signals:**
- All risks scored in the same narrow range
- No risk scored 1 or 5 on either dimension
- Critical risks (active litigation, data breach) scored the same as minor contract deviations
- No differentiation in urgency or escalation recommendations
**Mitigation:**
- Use the full 1-5 scale. A routine NDA with a known counterparty is a 1/1 risk. Active litigation with significant exposure is a 5/4.
- Require specific justification for each severity and likelihood score
- Cross-check: would a reasonable in-house counsel agree with this score?
- Compare relative scores: is this risk really the same severity as [other scored risk]?
**Severity: MEDIUM.** Score compression reduces the value of risk assessment by failing to prioritize.
---
### 9. Confidentiality and Privilege Blindness
**What happens:** LLM does not inherently understand attorney-client privilege, work product protection, or confidentiality obligations. May suggest actions that would waive privilege, recommend sharing privileged information, or fail to flag privilege considerations.
**Detection signals:**
- Recommending sharing legal analysis with opposing party or in non-privileged settings
- Not flagging that a document should be marked as privileged
- Suggesting inclusion of privileged content in non-privileged communications
- Not noting privilege considerations when recommending who to consult or share with
**Mitigation:**
- When generating legal analysis: note whether the output should be treated as privileged
- When recommending sharing or distribution: flag privilege considerations
- For incident briefs and litigation-adjacent matters: include privilege marking guidance
- Never recommend waiving privilege without explicitly flagging the implications
**Severity: MEDIUM.** Inadvertent privilege waiver can be costly and is often irreversible.
---
### 10. Precedent Bias
**What happens:** LLM defaults to common-law, US-centric legal reasoning even when analyzing situations governed by other legal systems. Assumes adversarial system, precedent-based reasoning, and common-law contract interpretation when civil law, statutory interpretation, or other frameworks may apply.
**Detection signals:**
- Analysis structured around "case law" in civil law jurisdictions
- Assuming freedom of contract principles in jurisdictions with mandatory terms
- Applying UCC concepts to non-US commercial transactions
- Using "reasonable person" standard where different standards apply
**Mitigation:**
- Identify the legal system (common law, civil law, mixed) before analyzing
- For civil law jurisdictions: focus on code provisions, not case precedent
- For non-US contracts: do not assume UCC, Restatement, or other US frameworks apply
- Ask about governing law before defaulting to any legal tradition
**Severity: MEDIUM.** Applying wrong legal tradition produces analysis that is internally consistent but fundamentally inapplicable.
---
## Guard Protocol Summary
Apply these guards to every legal analysis output:
| # | Guard | Implementation |
|---|-------|---------------|
| 1 | Jurisdiction stated | Every analysis names its governing jurisdiction at the top |
| 2 | Citations hedged | No citation presented as authoritative without "verify with authoritative sources" |
| 3 | Calibrated language | No absolutes. "Likely," "typically," "in most jurisdictions" unless truly certain |
| 4 | Staleness caveat | Note training cutoff for any regulatory analysis |
| 5 | Cross-reference check | After clause analysis, check for interactions between clauses |
| 6 | Escalation check | Before generating responses, check escalation triggers |
| 7 | Privilege awareness | Flag when output should be privileged or when privilege considerations arise |
| 8 | Full-scale scoring | Risk scores use full 1-5 range with specific justification |
| 9 | Specific engagement | Output addresses the specific facts, not just the general framework |
| 10 | Disclaimer present | "Analysis support, not legal advice. Review by qualified counsel required." |
references/legal/nda-triage.md
# NDA Triage
Rapid NDA pre-screening methodology. GREEN/YELLOW/RED classification with a 10-criterion screening checklist and common deviations catalog.
---
## Screening Checklist
### Criterion 1: Agreement Structure
| Check | Pass | Flag | Fail |
|-------|------|------|------|
| Type identified (Mutual / Unilateral disclosing / Unilateral receiving) | Type clear and appropriate for context | Type not ideal but workable | Type wrong for the relationship |
| Appropriate for business context | Mutual for exploratory, unilateral for one-way disclosure | Slight mismatch but manageable | Fundamentally inappropriate |
| Standalone agreement | Yes, standalone NDA | Confidentiality section in larger agreement (note broader context) | Not actually an NDA (contains commercial terms, exclusivity) |
**If the document is not actually an NDA** (labeled as NDA but contains substantive commercial terms): flag immediately as RED, recommend full contract review.
---
### Criterion 2: Definition of Confidential Information
| Check | Pass | Flag | Fail |
|-------|------|------|------|
| Reasonable scope | Limited to non-public info disclosed for stated purpose | Broader than preferred but not unreasonable | "All information of any kind whether or not marked" |
| Marking requirements | None, or marking within 30 days of oral disclosure | Immediate marking required for oral disclosures | Impractical marking requirements |
| Exclusions present | All standard exclusions defined | Missing one non-critical exclusion | Missing critical exclusions or none defined |
| No problematic inclusions | Clean | Minor concern | Defines publicly available or independently developed info as confidential |
---
### Criterion 3: Obligations of Receiving Party
| Check | Pass | Flag | Fail |
|-------|------|------|------|
| Standard of care | Reasonable care or same as own confidential info | Slightly higher but workable | Strict liability or impractical standard |
| Use restriction | Limited to stated purpose | Purpose broader than expected | No use restriction or "any purpose" |
| Disclosure restriction | Need-to-know basis, bound by similar obligations | Slightly broader but controlled | No meaningful restriction |
| No onerous obligations | Standard obligations | Minor unusual requirements | Impractical requirements (encrypt all communications, physical logs) |
---
### Criterion 4: Standard Carveouts
All five must be present. Missing any critical carveout affects classification.
| Carveout | Required | Impact if Missing |
|----------|----------|-------------------|
| **Public knowledge** | Information publicly available through no fault of receiving party | YELLOW -- add in redline |
| **Prior possession** | Information already known before disclosure | YELLOW -- add in redline |
| **Independent development** | Information independently developed without reference to confidential info | RED if missing -- creates risk that internal work is claimed as derived |
| **Third-party receipt** | Information rightfully received from unrestricted third party | YELLOW -- add in redline |
| **Legal compulsion** | Right to disclose when required by law/regulation/legal process (with notice where permitted) | RED if missing -- could prevent compliance with legal obligations |
---
### Criterion 5: Permitted Disclosures
| Recipient Category | Standard | Flag If |
|--------------------|----------|---------|
| Employees | Need-to-know employees | Not explicitly permitted |
| Contractors/Advisors | Under similar confidentiality obligations | Not permitted or restricted |
| Affiliates | If needed for business purpose | Prohibited when needed |
| Legal/Regulatory | As required by law | Not addressed (legal compulsion carveout may cover) |
---
### Criterion 6: Term and Duration
| Element | GREEN | YELLOW | RED |
|---------|-------|--------|-----|
| Agreement term | 1-3 years | 3-5 years | 5+ years without justification |
| Confidentiality survival | 2-5 years from termination | 5-7 years | Perpetual (unless trade secret carveout) |
| Trade secret protection | As long as info remains a trade secret | Slightly broader | Perpetual for all information regardless of trade secret status |
---
### Criterion 7: Return and Destruction
| Check | Pass | Flag | Fail |
|-------|------|------|------|
| Obligation triggered | On termination or upon request | Only on termination (no upon-request) | No return/destruction obligation |
| Scope | Return or destroy all copies | Reasonable scope | Overbroad or unclear |
| Retention exception | Allows retention for legal/compliance/backup | No explicit exception (likely implied) | Requires destruction of all copies including legal/compliance backups |
| Certification | Certification of destruction acceptable | Sworn affidavit required | Notarized affidavit or third-party attestation |
---
### Criterion 8: Remedies
| Check | Pass | Flag | Fail |
|-------|------|------|------|
| Injunctive relief | Acknowledgment of irreparable harm, equitable relief appropriate | Slightly broader but standard | Automatic injunction without showing harm |
| Damages | No pre-determined damages | Minor provisions | Liquidated damages in NDA |
| Symmetry | Mutual remedies (in mutual NDA) | Minor asymmetry | Substantially one-sided |
---
### Criterion 9: Problematic Provisions
Any of these present should trigger at minimum YELLOW, and most trigger RED.
| Provision | Classification if Present | Standard Position |
|-----------|--------------------------|-------------------|
| **Non-solicitation of employees** | RED | Does not belong in NDA. Delete entirely. If counterparty insists: limit to targeted solicitation (not general recruitment), 12-month term. |
| **Non-compete** | RED | Does not belong in NDA. Delete entirely. |
| **Exclusivity** | RED | Should not restrict either party from similar discussions with others |
| **Standstill** | RED (unless M&A context) | Inappropriate outside M&A. In M&A: should be time-limited with clear terms. |
| **Broad residuals clause** | RED | Effectively creates license to use confidential info. Resist entirely. |
| **Narrow residuals clause** | YELLOW | If limited to: unaided memory of authorized individuals, excludes trade secrets and patentable info, does not grant IP license. |
| **IP assignment or license** | RED | NDA should not grant any IP rights |
| **Audit rights** | YELLOW-RED | Unusual in standard NDAs. If present, must have reasonable scope and notice. |
---
### Criterion 10: Governing Law and Jurisdiction
| Check | GREEN | YELLOW | RED |
|-------|-------|--------|-----|
| Jurisdiction | Well-established commercial jurisdiction | Acceptable but non-preferred | Unfavorable or unusual jurisdiction |
| Consistency | Governing law and jurisdiction in same/related jurisdictions | Minor mismatch | Different countries for law vs. venue |
| Dispute mechanism | Litigation (standard for NDA disputes) | Mediation-first with litigation fallback | Mandatory arbitration with drafter-favorable rules |
---
## Classification Rules
### GREEN -- Standard Approval
**ALL of the following must be true:**
- Mutual NDA (or unilateral in appropriate direction)
- All five standard carveouts present
- Term within standard range (1-3 year agreement, 2-5 year survival)
- No non-solicitation, non-compete, or exclusivity provisions
- No residuals clause, or residuals narrowly scoped
- Reasonable governing law jurisdiction
- Standard remedies (no liquidated damages)
- Permitted disclosures include employees, contractors, advisors
- Return/destruction includes retention exception
- Definition of confidential information reasonably scoped
**Routing**: Approve via standard delegation of authority. Same-day.
---
### YELLOW -- Counsel Review
**One or more present, but NDA is not fundamentally problematic:**
- Definition broader than preferred but not unreasonable
- Term longer than standard but within market range (5-year agreement, 7-year survival)
- Missing one standard carveout that could be added with minor redline
- Narrow residuals clause (unaided memory, excludes trade secrets)
- Governing law in acceptable but non-preferred jurisdiction
- Minor asymmetry in mutual NDA
- Marking requirements present but workable
- Return/destruction lacks explicit retention exception
- Unusual but non-harmful provisions (obligation to notify of potential breach)
**Routing**: Flag specific issues for counsel review. Single review pass expected. 1-2 business days.
---
### RED -- Full Legal Review
**One or more present:**
- Unilateral when mutual is required (or wrong direction)
- Missing critical carveouts (independent development or legal compulsion)
- Non-solicitation or non-compete provisions
- Exclusivity or standstill provisions without appropriate context
- Unreasonable term (10+ years, or perpetual without trade secret justification)
- Overbroad definition capturing public or independently developed information
- Broad residuals clause creating effective license
- IP assignment or license grant
- Liquidated damages or penalty provisions
- Audit rights without reasonable scope or notice
- Highly unfavorable jurisdiction with mandatory arbitration
- Document is not actually an NDA (contains substantive commercial terms)
**Routing**: Full legal review. Do not sign. Requires negotiation, counterproposal with standard form, or rejection. 3-5 business days.
---
## Common Deviations Catalog
### Deviation: Overbroad Confidential Information Definition
**Frequency**: Very common
**Risk**: Captures information that should not be confidential, creates compliance burden
**Standard redline**: Narrow to information marked or identified as confidential, or that a reasonable person would understand to be confidential given nature and circumstances.
**Fallback**: Accept broader definition but add robust exclusions.
---
### Deviation: Missing Independent Development Carveout
**Frequency**: Common
**Risk**: Could create claims that internally-developed products were derived from counterparty's confidential information
**Standard redline**: Add: "Information independently developed by or for the Receiving Party without use of or reference to the Disclosing Party's Confidential Information."
**Fallback**: None. This carveout is essential. Escalate if counterparty refuses.
---
### Deviation: Non-Solicitation of Employees
**Frequency**: Moderate (often embedded by counterparty legal as boilerplate)
**Risk**: Restricts hiring, may be unenforceable in some jurisdictions, creates unnecessary liability
**Standard redline**: Delete entirely.
**Fallback**: If counterparty insists: limit to targeted solicitation (not general recruitment or response to job postings), cap at 12 months, limit to employees directly involved in the engagement.
---
### Deviation: Broad Residuals Clause
**Frequency**: Moderate (more common in technology NDAs)
**Risk**: Effectively grants a license to use confidential information for any purpose, undermining the NDA's core protection
**Standard redline**: Delete entirely.
**Fallback**: If required: (a) limited to general ideas, concepts, know-how retained in unaided memory of authorized individuals, (b) explicitly excludes trade secrets and patentable information, (c) does not create any IP license, (d) does not override any other obligation in the agreement.
---
### Deviation: Perpetual Confidentiality Obligation
**Frequency**: Common
**Risk**: Creates indefinite compliance burden, may be unenforceable
**Standard redline**: Replace with defined term (2-5 years from disclosure or termination). Offer trade secret carveout: "provided that Confidential Information constituting trade secrets shall be protected for so long as such information retains trade secret status."
**Fallback**: Accept 7-year term with trade secret carveout.
---
### Deviation: Mandatory Arbitration
**Frequency**: Moderate
**Risk**: Limits remedies (no injunctive relief without court involvement), may be more expensive for NDA disputes (which are typically straightforward)
**Standard redline**: Replace with litigation in agreed jurisdiction. Preserve right to seek injunctive relief in any court of competent jurisdiction.
**Fallback**: Accept arbitration with: (a) right to seek injunctive relief in court, (b) neutral arbitration rules (e.g., AAA, ICC, JAMS), (c) arbitration seat in reasonable location, (d) single arbitrator for disputes under threshold.
---
### Deviation: No Legal Compulsion Carveout
**Frequency**: Uncommon but serious when missing
**Risk**: Could prevent compliance with legal obligations (subpoena, regulatory inquiry, court order)
**Standard redline**: Add: "The Receiving Party may disclose Confidential Information to the extent required by applicable law, regulation, or legal process, provided that the Receiving Party gives prompt written notice to the Disclosing Party (to the extent legally permitted) to allow the Disclosing Party to seek a protective order or other appropriate remedy."
**Fallback**: None. This carveout is essential. Escalate if counterparty refuses.
---
## Triage Report Template
```
## NDA Triage Report
**Classification**: [GREEN / YELLOW / RED]
**Parties**: [party names]
**Type**: [Mutual / Unilateral (disclosing) / Unilateral (receiving)]
**Term**: [agreement duration] | **Survival**: [confidentiality duration]
**Governing Law**: [jurisdiction]
**Review Basis**: [Playbook / Default Standards]
## Screening Results
| # | Criterion | Status | Notes |
|---|-----------|--------|-------|
| 1 | Agreement Structure | [PASS/FLAG/FAIL] | [details] |
| 2 | Definition Scope | [PASS/FLAG/FAIL] | [details] |
| 3 | Receiving Party Obligations | [PASS/FLAG/FAIL] | [details] |
| 4 | Standard Carveouts | [PASS/FLAG/FAIL] | [which present/missing] |
| 5 | Permitted Disclosures | [PASS/FLAG/FAIL] | [details] |
| 6 | Term and Duration | [PASS/FLAG/FAIL] | [details] |
| 7 | Return and Destruction | [PASS/FLAG/FAIL] | [details] |
| 8 | Remedies | [PASS/FLAG/FAIL] | [details] |
| 9 | Problematic Provisions | [PASS/FLAG/FAIL] | [list any found] |
| 10 | Governing Law | [PASS/FLAG/FAIL] | [details] |
## Issues Found
### [Issue -- YELLOW/RED]
**What**: [description]
**Risk**: [what could go wrong]
**Suggested Fix**: [specific redline language]
**Fallback**: [alternative position]
## Recommendation
[Approve / Send for review with notes / Reject and counter]
## Routing
| Classification | Action | Timeline |
|---|---|---|
| GREEN | Approve per delegation of authority | Same day |
| YELLOW | Counsel review with flagged issues | 1-2 business days |
| RED | Full review, negotiation, or counterproposal | 3-5 business days |
```
references/legal/risk-assessment.md
# Legal Risk Assessment Framework
Structured methodology for identifying, scoring, classifying, and documenting legal risks. Severity x Likelihood matrix with escalation criteria and business impact scoring.
---
## Risk Scoring Model
### Severity Scale
| Level | Label | Description | Financial Exposure | Operational Impact |
|-------|-------|-------------|-------------------|-------------------|
| 1 | **Negligible** | Minor inconvenience, no material impact | < 0.1% of contract/deal value | None. Normal operations. |
| 2 | **Low** | Limited impact, minor disruption | < 1% of relevant value | Minor disruption, easily absorbed |
| 3 | **Moderate** | Meaningful impact, noticeable disruption | 1-5% of relevant value | Noticeable disruption, potential limited public attention |
| 4 | **High** | Significant impact, regulatory scrutiny likely | 5-25% of relevant value | Significant disruption, likely public attention, potential regulatory action |
| 5 | **Critical** | Severe impact, fundamental business threat | > 25% of relevant value | Business disruption, reputational damage, regulatory action likely, potential personal liability |
### Likelihood Scale
| Level | Label | Description | Precedent | Triggering Events |
|-------|-------|-------------|-----------|-------------------|
| 1 | **Remote** | Highly unlikely | No known precedent in similar situations | Would require exceptional circumstances |
| 2 | **Unlikely** | Could occur but not expected | Limited precedent | Would require specific triggering events |
| 3 | **Possible** | May occur | Some precedent exists | Triggering events are foreseeable |
| 4 | **Likely** | Probably will occur | Clear precedent | Triggering events are common |
| 5 | **Almost Certain** | Expected to occur | Strong precedent or pattern | Triggering events are present or imminent |
### Risk Matrix
```
LIKELIHOOD
Remote Unlikely Possible Likely Almost Certain
(1) (2) (3) (4) (5)
SEVERITY
Critical (5) | 5 | 10 | 15 | 20 | 25 |
High (4) | 4 | 8 | 12 | 16 | 20 |
Moderate (3) | 3 | 6 | 9 | 12 | 15 |
Low (2) | 2 | 4 | 6 | 8 | 10 |
Negligible(1) | 1 | 2 | 3 | 4 | 5 |
```
**Risk Score = Severity x Likelihood**
---
## Risk Classification and Response
### GREEN -- Low Risk (Score 1-4)
**Characteristics:**
- Minor issues unlikely to materialize
- Standard business risks within normal parameters
- Well-understood risks with established mitigations
**Response protocol:**
- Accept and proceed with standard controls
- Document in risk register
- Monitor quarterly or annually
- No escalation required
**Examples:**
- Vendor contract with minor deviation in non-critical area
- Routine NDA with known counterparty in standard jurisdiction
- Administrative compliance task with clear deadline and owner
---
### YELLOW -- Medium Risk (Score 5-9)
**Characteristics:**
- Moderate issues that could materialize under foreseeable circumstances
- Warrant attention but not immediate action
- Established management precedent exists
**Response protocol:**
- Implement specific controls or negotiate to reduce exposure
- Assign a single owner responsible for monitoring and mitigation
- Review monthly or as trigger events occur
- Document risk, mitigations, and rationale thoroughly
- Brief relevant business stakeholders
- Define trigger events that would elevate the risk level
**Examples:**
- Contract with LOL cap below standard but within negotiable range
- Vendor processing personal data in jurisdiction without clear adequacy determination
- Regulatory development that may affect business activity in medium term
- IP provision broader than preferred but common in market
---
### ORANGE -- High Risk (Score 10-15)
**Characteristics:**
- Significant issues with meaningful probability of materializing
- Could result in substantial financial, operational, or reputational impact
- Requires senior attention and dedicated mitigation
**Response protocol:**
- Escalate to senior counsel (head of legal or designated senior)
- Develop specific, actionable mitigation plan
- Brief business leadership
- Review weekly or at defined milestones
- Consider outside counsel engagement
- Full risk memo with analysis, options, recommendations
- Define contingency plan (what if risk materializes?)
**Examples:**
- Contract with uncapped indemnification in material area
- Data processing activity potentially violating regulatory requirements
- Threatened litigation from significant counterparty
- IP infringement allegation with colorable basis
- Regulatory inquiry or audit request
---
### RED -- Critical Risk (Score 16-25)
**Characteristics:**
- Severe issues likely or certain to materialize
- Could fundamentally impact the business, officers, or stakeholders
- Requires immediate executive attention
**Response protocol:**
- Immediate escalation to General Counsel, C-suite, Board as appropriate
- Engage outside counsel immediately
- Establish dedicated response team with clear roles
- Consider insurance notification
- Activate crisis management protocols if reputational risk involved
- Implement litigation hold if legal proceedings possible
- Daily or more frequent review until resolved or reduced
- Include in board risk reporting
- Make any required regulatory notifications
**Examples:**
- Active litigation with significant exposure
- Data breach affecting regulated personal data
- Regulatory enforcement action
- Material contract breach (by or against the organization)
- Government investigation
- Credible IP infringement claim against core product/service
---
## Escalation Decision Tree
```
Is there active litigation or government investigation?
├── YES → RED. Engage outside counsel immediately.
└── NO → Continue.
Could the risk result in criminal liability?
├── YES → RED. Engage criminal defense counsel.
└── NO → Continue.
Does the risk involve regulated personal data (breach, non-compliance)?
├── YES → Is there a notification deadline?
│ ├── YES → ORANGE minimum. Check deadline. May be RED.
│ └── NO → Score normally but minimum YELLOW.
└── NO → Continue.
Is the financial exposure > 25% of relevant deal/contract value?
├── YES → Score normally but minimum ORANGE.
└── NO → Continue.
Apply standard Severity x Likelihood scoring.
```
---
## When to Engage Outside Counsel
### Mandatory
| Trigger | Why |
|---------|-----|
| Active litigation | Defense or prosecution requires litigation counsel |
| Government investigation | Regulatory, law enforcement, or agency inquiry |
| Criminal exposure | Potential criminal liability for org or personnel |
| Securities issues | Could affect disclosures or filings |
| Board-level matters | Requires board notification or approval |
### Strongly Recommended
| Trigger | Why |
|---------|-----|
| Novel legal issues | First impression or unsettled law |
| Jurisdictional complexity | Unfamiliar jurisdiction or conflicting requirements |
| Material financial exposure | Exceeds organizational risk tolerance |
| Specialized expertise needed | Antitrust, FCPA, patent prosecution, etc. |
| New regulations | Materially affect business, require compliance program |
| M&A transactions | Due diligence, deal structuring, regulatory approvals |
### Consider
| Trigger | Why |
|---------|-----|
| Complex contract disputes | Significant disagreements with material counterparties |
| Employment matters | Discrimination, harassment, wrongful termination, whistleblower |
| Data incidents | Potential breaches triggering notification obligations |
| IP disputes | Infringement allegations involving material products |
| Insurance coverage disputes | Disagreements over coverage for material claims |
---
## Business Impact Scoring
Beyond the risk matrix, assess business impact across five dimensions:
| Dimension | Questions to Answer | Score (1-5) |
|-----------|-------------------|-------------|
| **Financial** | Direct costs? Potential damages? Insurance coverage? | __ |
| **Operational** | Business disruption? Process changes needed? Timeline impact? | __ |
| **Reputational** | Public attention likely? Customer trust affected? Media risk? | __ |
| **Regulatory** | Fines possible? License/certification at risk? Ongoing scrutiny? | __ |
| **Strategic** | Affects competitive position? Partnership implications? Market access? | __ |
**Composite Business Impact = Average of dimension scores**
Use to prioritize when multiple risks compete for attention. Two risks with the same matrix score may have very different business impact profiles.
---
## Risk Assessment Documentation
### Risk Assessment Memo Format
```
## Legal Risk Assessment
**Date**: [date]
**Assessor**: [name]
**Matter**: [description]
**Privileged**: [Yes/No]
### 1. Risk Description
[Clear, concise description]
### 2. Background and Context
[Relevant facts, history, business context]
### 3. Risk Analysis
**Severity**: [1-5] -- [Label]
[Rationale: potential financial exposure, operational impact, reputational considerations]
**Likelihood**: [1-5] -- [Label]
[Rationale: precedent, triggering events, current conditions]
**Risk Score**: [Score] -- [GREEN/YELLOW/ORANGE/RED]
### 4. Business Impact
| Dimension | Score | Rationale |
|-----------|-------|-----------|
| Financial | [1-5] | [Why] |
| Operational | [1-5] | [Why] |
| Reputational | [1-5] | [Why] |
| Regulatory | [1-5] | [Why] |
| Strategic | [1-5] | [Why] |
### 5. Contributing Factors
[What increases the risk]
### 6. Mitigating Factors
[What decreases the risk or limits exposure]
### 7. Mitigation Options
| Option | Effectiveness | Cost/Effort | Recommended? |
|--------|-------------|-------------|-------------|
| [Option] | [H/M/L] | [H/M/L] | [Yes/No] |
### 8. Recommended Approach
[Specific recommendation with rationale]
### 9. Residual Risk
[Expected risk level after mitigation]
### 10. Monitoring Plan
[How and how often monitored. Trigger events for re-assessment.]
### 11. Next Steps
1. [Action -- Owner -- Deadline]
2. [Action -- Owner -- Deadline]
```
### Risk Register Entry Format
| Field | Content |
|-------|---------|
| Risk ID | Unique identifier |
| Date Identified | [date] |
| Description | Brief description |
| Category | Contract / Regulatory / Litigation / IP / Data Privacy / Employment / Corporate |
| Severity | [1-5] with label |
| Likelihood | [1-5] with label |
| Risk Score | [calculated] |
| Risk Level | GREEN / YELLOW / ORANGE / RED |
| Business Impact | [composite score] |
| Owner | Person responsible |
| Mitigations | Current controls |
| Status | Open / Mitigated / Accepted / Closed |
| Review Date | Next scheduled review |
| Notes | Additional context |
---
## Common Scoring Errors
| Error | Problem | Correction |
|-------|---------|------------|
| Anchoring on worst case | Setting severity to 5 because the worst possible outcome is catastrophic, ignoring that the worst case is extremely unlikely | Score severity based on the most likely adverse outcome, not the theoretical worst case |
| Conflating severity and likelihood | High-severity risk automatically scored as high-likelihood | Assess independently. A catastrophic risk can be remote. |
| Recency bias | Recent incident drives likelihood up beyond what data supports | Use base rates and precedent, not recent events alone |
| Ignoring mitigations | Scoring raw risk without accounting for controls already in place | Score residual risk (after existing controls), not inherent risk |
| Range compression | All risks scored 3-4 (moderate-high) to avoid extremes | Use the full scale. Scores of 1 and 5 exist for a reason. |
| Single-dimension focus | Scoring on financial exposure only, ignoring operational/reputational | Use the five-dimension business impact model |
---
## Risk Review Cadence
| Risk Level | Review Frequency | Escalation Check |
|------------|-----------------|-----------------|
| GREEN | Quarterly or annually | At each review |
| YELLOW | Monthly or at trigger events | At each review + when conditions change |
| ORANGE | Weekly or at milestones | At each review + when new information arrives |
| RED | Daily or more frequently | Continuous until resolved or reduced |
**Trigger events requiring immediate re-assessment regardless of cadence:**
- New litigation filed or threatened
- Regulatory inquiry or investigation initiated
- Data breach discovered
- Material contract breach (by either party)
- Significant business change (M&A, restructuring, new market entry)
- Change in applicable law or regulation
- Insurance coverage change
references/market-positioning.md
---
title: Market Positioning — Positioning Maps, Differentiation Scoring, Win/Loss Analysis
domain: competitive-intel
level: 3
skill: competitive-intel
---
# Market Positioning Reference
> **Scope**: Positioning map templates, differentiation scoring frameworks, and win/loss analysis for competitive intelligence. Use when determining how to differentiate from specific competitors and where to position in a market.
> **Version range**: Framework-agnostic — applies to SaaS, content, services, and product markets.
> **Generated**: 2026-04-09 — positioning strategy should be refreshed quarterly; this framework is stable.
---
## Overview
Positioning is not marketing copy. "The leading platform for..." is not a position — it is a claim. A position is a specific place in the customer's mind relative to alternatives: "When I need X, I use them because Y, unlike Z which gives me W." Positioning strategy finds a defensible location in the competitive landscape that you can own before competitors can replicate it. The three tools here — positioning maps, differentiation scoring, and win/loss analysis — build that location from evidence rather than aspiration.
---
## Positioning Statement Template
The standard Geoffrey Moore template, made concrete.
```
## Positioning Statement
For [target customer segment]
who [have a specific problem or need]
[product/content name] is a [category]
that [key differentiator — one concrete, specific thing]
unlike [primary competitor]
which [competitor's specific limitation].
```
**Worked examples**:
Bad (generic):
> "For developers who want to learn Go, GoLearn is a platform that provides high-quality tutorials, unlike other resources which are low quality."
Good (specific):
> "For platform engineers at series A-C startups who need to debug Kubernetes production incidents, PlatformNotes is a technical blog that publishes post-mortem analyses with actual root-cause investigation, unlike the Official Kubernetes docs which explain what to do but not why failures occur."
**Test**: Can your target reader fill in the blank: "I read [you] when ___"? If they can, positioning is working.
---
## Positioning Map Templates
### Template 1: 2×2 Quadrant Map
Choose axes that genuinely separate players in your market.
```
## Positioning Map: [Category]
## Updated: [YYYY-MM-DD]
High [Y-axis]
|
| [Aspirational A] [US target position]
|
| [Direct Competitor B]
|
| [Adjacent C]
| [Legacy D]
|
+--------------------------------> High [X-axis]
Low [Y-axis]
Low [X-axis]
### Player coordinates
| Player | X (1-10) | Y (1-10) | Notes |
|--------|----------|----------|-------|
| Our current | ___ | ___ | |
| Our target | ___ | ___ | 12-month goal |
| Competitor A | ___ | ___ | |
| Competitor B | ___ | ___ | |
```
### Template 2: Value Curve (Blue Ocean Canvas)
Compare offerings on 5-8 factors. Reveals where you are competing on the same terms vs. where you offer something unique.
| Factor | Competitor A | Competitor B | Industry Avg | Us |
|--------|-------------|-------------|--------------|-----|
| Price | High | Medium | Medium | Low |
| Depth of content | High | Medium | Low | High |
| Publication frequency | Low | High | Medium | Medium |
| Community | None | Large | Small | Medium |
| Interactivity (Q&A, office hours) | None | None | None | **High** |
| Industry-specific focus | Broad | Broad | Broad | **Narrow** |
**Interpretation**: Where your line is distinct from competitors, you have differentiation. Where it follows the industry average, you are competing on commodity terms — and losing.
---
## Differentiation Scoring Framework
Score your differentiation on each dimension to determine what is defensible vs. vulnerable.
| Differentiation Dimension | Our Advantage (1-10) | Ease for Competitor to Copy (1-10) | Defensibility Score |
|---------------------------|---------------------|-----------------------------------|---------------------|
| Brand/voice/trust | ___ | ___ | ___ - ___ = ___ |
| Proprietary data/research | ___ | ___ | ___ - ___ = ___ |
| Community/network effects | ___ | ___ | ___ - ___ = ___ |
| Technical depth/expertise | ___ | ___ | ___ - ___ = ___ |
| Speed/responsiveness | ___ | ___ | ___ - ___ = ___ |
| Price | ___ | ___ | ___ - ___ = ___ |
| Breadth of coverage | ___ | ___ | ___ - ___ = ___ |
**Defensibility Score** = Our Advantage − Ease to Copy. Higher = more defensible.
| Score Range | Classification | Strategy |
|-------------|---------------|---------|
| 7-10 | Core moat — protect and amplify | Double down, use in all positioning |
| 4-6 | Competitive advantage — maintain | Defend, monitor for competitor improvement |
| 1-3 | Temporary advantage | Do not build strategy on this; find moat |
| Negative | Vulnerability | Address or concede this dimension |
**Worked example** (technical content creator):
| Dimension | Our Advantage | Ease to Copy | Defensibility |
|-----------|--------------|--------------|---------------|
| Authentic voice (10 years same author) | 9 | 1 | **8** — core moat |
| Proprietary production incident data | 8 | 2 | **6** — strong |
| Community Slack (2,000 members) | 7 | 4 | **3** — defend |
| SEO (2,000 pages indexed) | 6 | 5 | **1** — temporary |
| Price (free) | 10 | 9 | **1** — not a moat |
Conclusion: Build strategy on voice + proprietary data. The community is valuable but replicable. Price and SEO are not differentiators.
---
## Win/Loss Analysis Framework
Collect data on decisions where you won or lost. For content: subscribe vs. bounce. For product: bought vs. churned.
### Win Analysis Template
```
## Win Record: [Date/Cohort]
### Why they chose us (ask or infer)
- Factor 1: ___
- Factor 2: ___
- Factor 3: ___
### What they tried before (alternatives considered)
- Alternative 1: ___ — why rejected: ___
- Alternative 2: ___ — why rejected: ___
### Trigger (what made them choose now)
- ___
### Segment match
- Does this winner match our ICP? [Yes / Partially / No]
- If no: why did they engage and is this a segment worth expanding to?
```
### Loss Analysis Template
```
## Loss Record: [Date/Cohort]
### Why they did NOT choose us (ask or infer)
- Factor 1: ___
- Factor 2: ___
### Who they chose instead
- Competitor: ___
- Why: ___
### What we would need to offer to win this segment
- ___
### Should we compete for this segment?
- [Yes — add to roadmap / No — outside our ICP / Maybe — investigate further]
```
### Win/Loss Pattern Table
After 10+ records, aggregate:
| Decision Factor | % Won when present | % Lost when present | Action |
|-----------------|-------------------|---------------------|--------|
| [Factor A] | 80% | 20% | Amplify in positioning |
| [Factor B] | 40% | 60% | Not a differentiator |
| [Factor C] | 10% | 90% | Address or stop competing on it |
| [Factor D] | 95% | 5% | Core moat — protect |
---
## Positioning Validation Checklist
Before finalizing positioning:
```
[ ] Can you name 3 real competitors and state clearly why a customer would choose you over each?
[ ] Does your positioning statement pass the "for who" test? (specific segment, not "developers")
[ ] Does your differentiation score show at least 2 defensibility scores of 5+?
[ ] Have you validated positioning with 3 real customers/readers? ("Why do you read/use us?")
[ ] Is your positioning distinct on the value curve from industry average on at least 2 factors?
[ ] Does your positioning avoid competing on price as primary differentiator? (price = race to bottom)
[ ] Is your target position on the positioning map currently unoccupied or weakly occupied?
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Positioning by negation only
**What it looks like**: "We're not like those enterprise tools. We're different."
**Why wrong**: Negation tells the customer what you are not. It does not tell them what you are. The brain cannot hold "not X" — it has to form a positive association.
**Do instead**: Lead with a positive claim in the form "We are the [specific thing] for [specific people] who need [specific outcome]." Use competitor references only as clarifiers after the positive identity is established.
**Fix**: Every positioning statement must have a positive claim: "We are the [specific thing] for [specific people] who need [specific outcome]." The competitor reference is a clarifier, not the definition.
### Claiming differentiation on commodity dimensions
**What it looks like**: "We differentiate on quality, service, and reliability."
**Why wrong**: Every competitor claims quality, service, and reliability. These are table stakes, not differentiators. Scoring these high in the differentiation matrix with low ease-to-copy is usually wishful thinking.
**Detection**: Run the defensibility scoring. Any dimension where "ease for competitor to copy" is above 7 is not a strategic differentiator.
**Do instead**: Run the defensibility scoring on every claimed differentiator. Retire any dimension where ease-to-copy scores above 7. What remains after that filter is your actual differentiation surface.
### Win/loss analysis from memory instead of data
**What it looks like**: "We think we lose to Competitor A because of price." No records, no patterns, just team consensus.
**Why wrong**: Teams systematically attribute losses to price because it is the most face-saving reason. Actual win/loss data frequently reveals positioning and capability gaps are more important than price.
**Do instead**: Capture win/loss data at the moment of decision using forced-choice surveys (for products: cancel surveys; for content: exit surveys). Analyze only after 10 or more records. Treat team consensus as a hypothesis to test, not a conclusion.
**Fix**: Collect win/loss data at the point of decision. For content: exit surveys. For products: cancel surveys with forced-choice options. Analyze at 10+ records before concluding.
---
## Detection Commands Reference
```bash
# Check Google Search Console for branded vs. non-branded traffic split
# (Manual: Search Console > Performance > filter "type: web" > compare queries with/without brand name)
# Find what competitors rank for that you don't (requires ahrefs/semrush or similar)
# Manual: search "[competitor] vs [your brand]" — see what comparison content exists
# Monitor competitor brand mentions on Reddit
curl -s "https://www.reddit.com/search.json?q={competitor-name}&sort=new&limit=10" | \
python3 -c "import json,sys; d=json.load(sys.stdin); \
[print(p['data']['title'], p['data']['created_utc']) for p in d['data']['children']]"
```
---
## See Also
- `competitive-mapping.md` — competitive landscape mapping, feature comparison matrices
- `skills/strategic-decision/references/strategic-frameworks.md` — SWOT scoring and Porter's Five Forces
- `skills/growth-strategy/references/audience-segmentation.md` — ICP scoring for segment selection
references/migration-planning.md
---
title: Migration Planning — Cost Estimation, Phased Rollout, Rollback Plans, Data Migration Patterns
domain: build-vs-buy
level: 3
skill: build-vs-buy
---
# Migration Planning Reference
> **Scope**: Migration cost estimation frameworks, phased rollout templates, rollback plan design, and data migration patterns for build-vs-buy transitions. Use when switching from an existing solution — whether migrating FROM a custom-built system to a vendor, or FROM one vendor to another. Covers the gap between "we decided to switch" and "we are fully switched."
> **Version range**: Framework-agnostic — patterns apply to database migrations, SaaS replacements, and internal system cutover equally.
> **Generated**: 2026-04-09 — validate migration effort estimates against your specific data volume and integration count before using.
---
## Overview
Migration cost is the most underestimated dimension of build-vs-buy decisions. The TCO framework (tco-framework.md) includes a migration cost line item, but that line is typically a single number arrived at quickly. Migrations fail most often not because the new system is bad, but because migration planning underestimated the data complexity, left no rollback path, and used a big-bang cutover that put the whole business at risk simultaneously. This reference gives migration the depth it deserves.
---
## Migration Cost Estimation Framework
Fill this out BEFORE finalizing the build-vs-buy decision. Migration cost is a critical input to the TCO comparison.
### Phase 1: Inventory Assessment
```
## Migration Inventory
### Data Volume
- Total records to migrate: ___
- Data size (GB/TB): ___
- Age range of data: ___ years
- Data quality assessment: [Clean / Mixed / Poor — check for nulls, duplicates, format inconsistencies]
### Schema Complexity
- Number of tables/entities to migrate: ___
- Many-to-many relationships requiring transformation: ___
- Denormalized columns requiring normalization (or vice versa): ___
- Custom data types not supported by target: ___
### Integration Count
- APIs consuming data from current system: ___
- APIs producing data to current system: ___
- Webhooks/events being sent or received: ___
- Batch jobs or ETL pipelines touching current system: ___
### User Impact
- Estimated users affected by cutover: ___
- Acceptable downtime window: [None / <1 hour / <4 hours / Scheduled maintenance]
- Geographic distribution (time zone complexity): ___
```
### Phase 2: Effort Estimation by Component
| Migration Component | Effort Level | Hours Estimate | Dependencies |
|---------------------|-------------|----------------|--------------|
| Schema mapping document | Low | 8-16 hrs | Inventory complete |
| Data transformation scripts | Low/Med/High | ___ hrs | Schema mapping |
| Historical data migration (one-time) | Med/High | ___ hrs | Transform scripts |
| Delta sync (catch up during cutover) | Med/High | ___ hrs | Historical complete |
| API client updates (per integration) | Low per integration | ___ × ___ hrs | New system ready |
| Webhook reconfiguration | Low | 4-8 hrs | New system ready |
| Auth/credential rotation | Low | 4-8 hrs | All integrations |
| User communication and training | Low/Med | 8-40 hrs | Launch date set |
| Parallel run validation | Med | 16-40 hrs | Both systems live |
| Data reconciliation audit | Med | 8-24 hrs | Migration complete |
| **Total** | | **___ hrs** | |
**Complexity multipliers**:
- Data quality is Poor: multiply transform estimate × 2.5
- Zero downtime required: multiply cutover estimate × 2
- 10+ integrations: add 20% to total (coordination overhead)
- No migration precedent in team: multiply total × 1.5 (learning curve)
---
## Phased Rollout Template
Use for any migration affecting more than 10 users or involving more than 50K records. Big-bang cutover is the most common migration failure pattern.
### Phase Gate Structure
```
## Migration Rollout Plan: [System Name]
### Phase 0: Foundation (weeks 1-2)
- [ ] New system provisioned in production environment
- [ ] Data transformation scripts written and reviewed
- [ ] Rollback plan documented and tested in staging
- [ ] Monitoring alerts configured for new system
- [ ] Communication plan drafted
Gate criteria: Staging migration completed successfully. Rollback drill completed in under 15 minutes.
---
### Phase 1: Shadow Mode (weeks 3-4)
New system receives all writes but is NOT the system of record.
- [ ] New system receiving writes (via dual-write or event forwarding)
- [ ] Data reconciliation running daily: [old system vs. new system comparison]
- [ ] Divergence rate below threshold: ___% acceptable drift
- [ ] No reads redirected yet
Gate criteria: Zero data divergence for 5 consecutive business days. Performance benchmarks met.
---
### Phase 2: Pilot Group (weeks 5-6)
Small group of users reads from and writes to new system. Old system remains for everyone else.
- [ ] Pilot group selected: [internal team / power users / specific account]
- [ ] Pilot users notified and trained
- [ ] Support channel for pilot feedback established
- [ ] Escalation path to rollback pilot users only (not full revert)
Gate criteria: Pilot group satisfaction ≥ ___ (survey). Zero data loss incidents. No blocking bugs.
---
### Phase 3: Progressive Rollout (weeks 7-10)
Expand in cohorts. Use percentage-based or account-based expansion.
| Cohort | % of Users | Start Date | Gate Metric |
|--------|-----------|------------|-------------|
| Cohort 1 | 10% | ___ | Error rate < 0.1%, 48 hours stable |
| Cohort 2 | 25% | ___ | Error rate < 0.1%, 72 hours stable |
| Cohort 3 | 50% | ___ | Zero P1 incidents, 1 week stable |
| Cohort 4 | 100% | ___ | All previous gates passed |
Gate criteria per cohort: Error rate below threshold for minimum stable period.
---
### Phase 4: Cutover and Decommission (weeks 11-14)
Old system moved to read-only, then decommissioned.
- [ ] Old system made read-only (writes disabled)
- [ ] Final data reconciliation audit run
- [ ] All integrations confirmed pointing to new system
- [ ] Old system retained in backup mode for [30 / 60 / 90 days]
- [ ] Decommission scheduled: ___
Gate criteria: Full week of operation with 100% of users on new system, zero rollback requests, data reconciliation showing 100% match.
```
---
## Rollback Plan Design
Every migration phase must have a rollback procedure. Design rollback BEFORE starting migration.
### Rollback Decision Matrix
| Trigger | Rollback Scope | Time to Execute | Authorization Required |
|---------|---------------|----------------|----------------------|
| Data corruption detected | Full rollback | < 30 minutes | On-call engineer |
| Error rate > 1% for > 15 minutes | Cohort rollback | < 15 minutes | On-call engineer |
| P1 incident lasting > 1 hour | Full rollback | < 30 minutes | Engineering lead |
| User data loss confirmed | Full rollback + incident | < 15 minutes | Any engineer |
| Performance degradation > 3× baseline | Cohort rollback | < 15 minutes | On-call engineer |
### Rollback Procedure Template
```
## Rollback Procedure: [Migration Phase N]
### Pre-conditions for rollback
- [ ] [Old system is still running and current]
- [ ] [Reverse sync scripts are ready and tested]
- [ ] [Communication template is drafted]
### Rollback Steps (must complete in < ___ minutes)
1. [ ] Declare rollback decision in incident channel (#incident or #migration)
2. [ ] Switch traffic back to old system:
[specific command or feature flag change]
3. [ ] Verify old system is serving traffic:
[curl or health check command]
4. [ ] Disable writes to new system:
[specific command]
5. [ ] Run delta sync (new system → old system) for any writes that occurred:
[command or script path]
6. [ ] Verify data consistency:
[reconciliation script command]
7. [ ] Notify affected users:
[template: "We've temporarily reverted [feature] due to [issue]. Your data is intact."]
8. [ ] Post-rollback audit (within 24 hours):
- Root cause of rollback trigger
- Data integrity confirmation
- Next steps for re-attempt
### Rollback drill schedule
Run this procedure in staging: [quarterly / before each phase gate]
Last drill date: ___
Drill completion time: ___ minutes (target: < ___ minutes)
```
**Non-negotiable rollback rules**:
1. Rollback procedures must be tested in staging before production cutover
2. Any engineer on the team must be able to execute the rollback, not just the migration lead
3. Rollback must be possible without database admin access (use application-level controls)
4. If rollback time exceeds the acceptable downtime window, the architecture must change before proceeding
---
## Data Migration Patterns
### Pattern 1: Extract-Transform-Load (ETL) — Offline Migration
Use when: Historical data needs format transformation. Downtime is acceptable. Volume < 100GB.
```
## ETL Migration Script Structure
# Step 1: Extract from source
extract_query = """
SELECT id, user_id, created_at, metadata_json
FROM old_table
WHERE migrated = FALSE
LIMIT {batch_size}
"""
# Step 2: Transform
def transform_row(row):
return {
'external_id': str(row['id']),
'owner': row['user_id'],
'created_at': row['created_at'].isoformat(),
# Flatten JSON: old_table.metadata_json.settings → new_table.settings
'settings': json.loads(row['metadata_json']).get('settings', {}),
}
# Step 3: Load in batches (never row-by-row in production)
def load_batch(transformed_rows):
# Use bulk insert API or COPY command — not INSERT per row
new_system_api.bulk_create(transformed_rows)
mark_source_as_migrated(transformed_rows)
# Idempotency: mark migrated=TRUE in source after successful load
# Resume: re-run will skip migrated=TRUE rows automatically
```
**Checklist**:
```
[ ] Script is idempotent — safe to re-run without duplicating data
[ ] Processes data in configurable batches (default: 1,000 records)
[ ] Writes migration log with timestamp, batch_id, count, and errors
[ ] Handles NULL values in every column explicitly
[ ] Has a dry-run mode that transforms but does not load
[ ] Error handling: logs failed rows, does not abort entire batch on single failure
```
### Pattern 2: Change Data Capture (CDC) — Zero-Downtime Migration
Use when: Zero downtime required. Volume > 100GB. Migration window is weeks.
```
## CDC Migration Phases
Phase A: Historical load
- Run ETL (Pattern 1) on all data older than T₀ (cutoff timestamp)
- Mark T₀ explicitly: migration_start_timestamp = NOW()
Phase B: Change capture (parallel operation)
- Enable CDC on source: capture all INSERT/UPDATE/DELETE after T₀
- Forward changes to new system via event queue or direct replication
- Both systems diverge until cutover but CDC closes the gap continuously
Phase C: Delta reconciliation
- Stop writes to old system (maintenance window, or use DB-level read-only flag)
- Drain the CDC queue completely (time estimate: queue_size / event_throughput)
- Run final reconciliation: count matches, spot-check 1% of rows
Phase D: Cutover
- Switch application traffic to new system
- Old system enters read-only mode
- Total downtime = drain time + validation time (typically 1-15 minutes)
```
### Pattern 3: Dual-Write — Low-Risk Incremental Migration
Use when: Cannot accept rollback. New and old systems must stay in sync during transition.
```
## Dual-Write Architecture
# Application layer writes to BOTH systems simultaneously
def save_record(data):
old_result = old_system.save(data) # Primary (source of truth)
try:
new_result = new_system.save(data) # Shadow (not yet authoritative)
except Exception as e:
log_error('dual-write-shadow-failure', e, data)
# Do NOT fail the request — old system is primary
return old_result
compare_results(old_result, new_result) # Log divergences
return old_result # Return old system result during shadow phase
# When ready to flip:
def save_record_post_cutover(data):
new_result = new_system.save(data) # New system is now primary
try:
old_result = old_system.save(data) # Old system in shadow mode
except Exception as e:
log_error('dual-write-legacy-failure', e, data)
return new_result # Return new system result
```
**Dual-write checklist**:
```
[ ] Failures in shadow system DO NOT fail the primary request
[ ] Divergences are logged to monitoring, not silently ignored
[ ] Shadow system has separate error budget (does not pollute primary alerts)
[ ] Cutover is a single feature flag change, not a code deployment
[ ] Old system write path can be disabled independently of read path
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Big-bang cutover
**What it looks like**: All users moved from old to new system in a single weekend deployment. No pilot, no phased rollout, no shadow period.
**Why wrong**: Big-bang cutover concentrates all migration risk into a single point. If anything goes wrong — and something always does — the entire user base is affected simultaneously and the rollback affects the entire user base.
**Detection**: If the migration plan says "go live on [date]" with no mention of cohorts, pilot groups, or shadow periods, it is a big-bang plan.
**Do instead**: Structure every migration affecting more than 10 users as a phased rollout: pilot group first, then progressive cohorts. Each cohort is a checkpoint where you can halt, observe, and correct before proceeding.
**Fix**: Any migration affecting more than 10 users or more than 30 days of data requires at minimum a pilot group phase before full rollout.
### Rollback plan written after cutover
**What it looks like**: Migration proceeds, and then someone asks "how do we roll back if this fails?" The answer is "we'd figure it out."
**Why wrong**: Under incident pressure, ad-hoc rollback takes 4-10× longer than a planned rollback and introduces additional errors. The migration should not start until rollback has been tested.
**Do instead**: Treat rollback as a phase gate. Write and drill the rollback procedure for Phase 1 in staging before Phase 1 begins. The question "how do we roll this back?" must have a documented, tested answer before any production traffic moves.
**Fix**: Rollback is a phase gate — migration to Phase 1 cannot proceed until rollback procedure for Phase 1 has been documented AND drilled in staging.
### Migrating data quality problems forward
**What it looks like**: Source database has 15% null values in required fields, duplicate records, and inconsistent date formats. Migration copies all of this into the new system verbatim.
**Why wrong**: Migrating dirty data validates bad data in the new system and causes downstream failures that are blamed on the new system, not the migration.
**Do instead**: Run a data quality assessment before any migration work begins. For every failing field, define the transformation rule before writing a single migration script. Data cleaning is part of the transform step, not a post-migration cleanup task.
**Fix**: Before migration begins, run a data quality assessment (null counts, duplicate counts, format inconsistency counts). For every failing field, define the transformation rule. Data cleaning happens in the transform step, not as an afterthought.
### Testing migration on wrong data volume
**What it looks like**: Migration scripts tested on a 1,000-row staging dataset. Production has 80 million rows. Performance characteristics are completely different.
**Why wrong**: Migration performance is non-linear. Scripts that complete in 2 minutes on 1K rows may take 8 hours on 80M rows due to index contention, network I/O, and API rate limits.
**Do instead**: Test migration on an anonymized 5% sample of production data in staging. Measure actual duration, then extrapolate. If 5% takes 30 minutes, plan for 10+ hours in production and schedule the cutover window accordingly.
**Fix**: Migrate a random 5% sample of production data (anonymized) in a staging environment and extrapolate timing. If 5% takes 30 minutes, production will take 10+ hours — plan accordingly.
---
## Detection Commands Reference
```bash
# Assess source data quality before migration
python3 -c "
import sqlite3 # replace with your DB connector
conn = sqlite3.connect('source.db')
cursor = conn.cursor()
# Null count per column (replace table_name and column list)
cursor.execute('''
SELECT
COUNT(*) as total,
SUM(CASE WHEN user_id IS NULL THEN 1 ELSE 0 END) as null_user_id,
SUM(CASE WHEN email IS NULL OR email = '' THEN 1 ELSE 0 END) as null_email
FROM users
''')
print(cursor.fetchone())
conn.close()
"
# Estimate migration duration from batch test
python3 -c "
import time
BATCH_SIZE = 1000
TOTAL_RECORDS = 5_000_000 # replace with actual count
# Time a sample batch (replace with actual migration function)
start = time.time()
# migrate_batch(sample_rows) # your function here
elapsed = time.time() - start
batches = TOTAL_RECORDS / BATCH_SIZE
total_seconds = elapsed * batches
print(f'Estimated total time: {total_seconds/3600:.1f} hours')
print(f'Estimated completion: {total_seconds/60:.0f} minutes')
"
# Check dual-write divergence rate
# (Assuming divergences are logged to a file or monitoring system)
grep "dual-write-shadow-failure" /var/log/app.log | \
python3 -c "
import sys
lines = sys.stdin.readlines()
print(f'Shadow failures in log: {len(lines)}')
"
```
---
## See Also
- `tco-framework.md` — TCO calculation that includes migration cost as a line item
- `vendor-evaluation.md` — data portability and export capability checklist
- `skills/project-evaluation/references/feasibility-scoring.md` — technical feasibility assessment for migration complexity
references/operations.md
# Operations
Umbrella skill for business operations: vendor management, runbooks, process documentation, risk assessment, capacity planning, change management, compliance tracking, status reporting, and process optimization. Each mode loads its own reference files on demand.
**Scope**: Operational workflows that keep the business running. Use csuite for strategic decisions, finance for budgeting/forecasting, and hr for people operations.
---
## Mode Detection
Classify the request into exactly one mode. If it spans multiple, choose the primary and note the secondary.
| Mode | Signal Phrases | Reference |
|------|---------------|-----------|
| **RUNBOOK** | runbook, procedure, on-call, playbook, step-by-step, ops task | `references/operations/runbook-authoring.md` |
| **RISK** | risk assessment, risk register, what could go wrong, risk matrix | `references/operations/risk-assessment.md` |
| **VENDOR** | vendor review, vendor evaluation, contract review, procurement | `references/operations/vendor-management.md` |
| **PROCESS** | process doc, SOP, RACI, workflow documentation, process map | `references/operations/process-documentation.md` |
| **CHANGE** | change request, change management, CAB, rollout, deployment change | `references/operations/change-management.md` |
| **CAPACITY** | capacity plan, resource allocation, utilization, headcount planning | `references/operations/process-documentation.md` |
| **COMPLIANCE** | compliance, audit prep, SOC 2, ISO 27001, GDPR, regulatory | `references/operations/risk-assessment.md` |
| **STATUS** | status report, weekly update, project health, KPIs | (no deep reference needed) |
| **OPTIMIZE** | process improvement, bottleneck, streamline, too many steps | `references/operations/process-documentation.md` |
Always load `references/operations/llm-ops-failure-modes.md` regardless of mode. It contains the failure patterns that apply across all operations work.
---
## Instructions
### Mode: RUNBOOK
**Framework**: SCOPE -> AUTHOR -> VERIFY
**Phase 1: SCOPE** -- Define what the runbook covers.
- Name the task, its frequency, and who runs it
- List prerequisites: access, tools, credentials, prior state
- Identify the trigger: scheduled, event-driven, or manual invocation
- Ask: "If a new hire ran this at 3am during an incident, what would they need?"
**Gate**: Task named. Prerequisites listed. Trigger defined.
**Phase 2: AUTHOR** -- Write the procedure with painful specificity.
Load `references/operations/runbook-authoring.md`.
Critical rules:
- Every step has: exact command/action, expected result, failure handling
- "Run the script" is NOT a step. `python sync.py --prod --dry-run` from `/opt/ops/` as `deploy-user` IS a step
- Include verification after every state-changing step
- Rollback procedure for the entire runbook AND per-step rollback where applicable
- Escalation paths with names, contact methods, and when-to-escalate triggers
| Step Component | Required | Example |
|---------------|----------|---------|
| Action | Yes | `kubectl rollout restart deployment/api -n production` |
| Expected result | Yes | "Pods restart within 60s. `kubectl get pods` shows 3/3 Running." |
| Failure handling | Yes | "If pods stay in CrashLoopBackOff >2min, proceed to Rollback." |
| Verification | Yes | `curl -s https://api.example.com/health | jq .status` returns `"ok"` |
| Rollback | Per-step | `kubectl rollout undo deployment/api -n production` |
**Gate**: Every step has all five components. Rollback procedure exists. Escalation path defined.
**Phase 3: VERIFY** -- Validate the runbook is actually usable.
- Walk through the runbook as if you have never seen the system
- Flag any step that requires unstated knowledge
- Confirm the troubleshooting table covers symptoms from each step's failure mode
- Check: could someone follow this at 3am with no prior context?
**Gate**: All steps self-contained. No implicit knowledge. Troubleshooting table complete.
---
### Mode: RISK
**Framework**: IDENTIFY -> ASSESS -> MITIGATE
**Phase 1: IDENTIFY** -- Enumerate risks systematically by category.
Load `references/operations/risk-assessment.md`.
| Category | What to Look For |
|----------|-----------------|
| Operational | Process failures, staffing gaps, system outages, single points of failure |
| Financial | Budget overruns, vendor cost increases, revenue impact, currency exposure |
| Compliance | Regulatory violations, audit findings, policy breaches, certification gaps |
| Strategic | Market changes, competitive threats, technology shifts, dependency risks |
| Reputational | Customer impact, public perception, partner relationships, data incidents |
| Security | Data breaches, access control failures, third-party vulnerabilities |
- Extend risk identification beyond obvious items. Ask: "What kills us if it happens, even if it seems unlikely?"
- Separate risks from issues. A risk might happen. An issue already has.
**Gate**: Risks enumerated across all applicable categories. Each risk has a clear description.
**Phase 2: ASSESS** -- Score each risk on probability and impact.
Apply the probability x impact matrix:
| | Low Impact | Medium Impact | High Impact |
|---|-----------|---------------|-------------|
| **High Probability** | Medium | High | Critical |
| **Medium Probability** | Low | Medium | High |
| **Low Probability** | Low | Low | Medium |
For each risk:
- Probability: base on evidence, not optimism. "It hasn't happened yet" is not "low probability"
- Impact: quantify in dollars, hours, or affected users where possible
- Risk level: derived from matrix, not gut feel
**Gate**: Every risk scored. No unquantified "High" without supporting rationale.
**Phase 3: MITIGATE** -- Plan mitigations and track residual risk.
For each High/Critical risk:
- Mitigation action (specific, not "monitor the situation")
- Owner (named person, not "the team")
- Timeline (date, not "soon")
- Residual risk after mitigation
- Acceptance criteria: what makes the residual risk acceptable?
**Gate**: All High/Critical risks have mitigations with owners and dates. Residual risk documented.
---
### Mode: VENDOR
**Framework**: EVALUATE -> SCORE -> RECOMMEND
**Phase 1: EVALUATE** -- Gather structured information.
Load `references/operations/vendor-management.md`.
Required inputs:
- Vendor name and what they provide
- Context: new evaluation, renewal, or comparison
- Available data: proposal, contract, performance history, pricing
Due diligence checklist (minimum):
- Financial stability signals
- Security posture (SOC 2, penetration testing, incident history)
- Customer references (ask for churned customers too)
- Contract terms: auto-renewal, termination notice, price escalation clauses
- Data portability: export format, migration support, data deletion
**Gate**: Due diligence complete. Contract terms reviewed. Red flags documented.
**Phase 2: SCORE** -- Apply the vendor scorecard.
| Dimension | Weight | Score (1-10) |
|-----------|--------|-------------|
| Functional fit | 5x | |
| Total cost of ownership | 4x | |
| Integration complexity | 4x | |
| Support quality | 3x | |
| Security/compliance | 3x | |
| Data portability | 3x | |
| Company stability | 2x | |
| Contract flexibility | 2x | |
TCO must include: license, implementation, training, support, ongoing maintenance, exit costs. License price alone is not TCO.
**Gate**: All dimensions scored with rationale. TCO calculated for Year 1 and Year 3.
**Phase 3: RECOMMEND** -- Deliver verdict with negotiation points.
- Verdict: Proceed / Negotiate / Pass
- Key strengths and concerns
- Negotiation leverage points
- Contract review triggers (price escalation >X%, SLA misses, acquisition)
- Performance monitoring plan post-contract
**Gate**: Recommendation stated. Negotiation points listed. Monitoring plan defined.
---
### Mode: PROCESS
**Framework**: MAP -> DOCUMENT -> OPTIMIZE
**Phase 1: MAP** -- Capture how the process actually works today.
Load `references/operations/process-documentation.md`.
- Map the real process, not the idealized version. "We're supposed to do X but actually we do Y" is the most valuable input.
- Capture every step, decision point, handoff, and exception
- Identify: who does it (role, not name), what triggers it, what it produces, how long it takes
- Document exceptions and edge cases — these are where processes actually break
**Gate**: Current state mapped. All steps, handoffs, and exceptions documented.
**Phase 2: DOCUMENT** -- Produce the SOP.
Structure:
- Purpose and scope
- RACI matrix (Responsible, Accountable, Consulted, Informed for each step)
- Step-by-step procedure with inputs, actions, outputs, and exception handling
- Metrics: what to measure, target values, how to measure
RACI rules:
- Exactly one A (Accountable) per step. Multiple As = no one is accountable.
- R (Responsible) is who does the work. A is who owns the outcome.
- If a step has no R, it doesn't get done. If it has no A, no one notices when it doesn't.
**Gate**: SOP complete. RACI has exactly one A per step. Exceptions documented.
**Phase 3: OPTIMIZE** -- Identify improvement opportunities.
- Identify waste: waiting, rework, unnecessary handoffs, over-processing, manual work that could be automated
- Bottleneck analysis: which step constrains throughput?
- Recommendations: eliminate, automate, parallelize, simplify
- Before/after comparison with estimated impact
**Gate**: Bottlenecks identified. Recommendations specific and measurable.
---
### Mode: CHANGE
**Framework**: ASSESS -> PLAN -> EXECUTE
**Phase 1: ASSESS** -- Define the change and its impact.
Load `references/operations/change-management.md`.
- What is changing and why (business justification, not just "improvement")
- Who is affected: users, systems, processes, teams
- Impact assessment by area (users, systems, processes, cost) rated High/Medium/Low
- Risk assessment: what could go wrong, likelihood, mitigation
- Resistance forecast: who will resist and why
**Gate**: Change defined. Impact assessed by area. Risks identified.
**Phase 2: PLAN** -- Build the implementation and communication plan.
| Plan Component | Required Elements |
|---------------|-------------------|
| Implementation | Steps, owners, timeline, dependencies |
| Communication | Audience, message, channel, timing |
| Training | What skills needed, delivery method, timeline |
| Rollback | Trigger criteria, steps, verification |
| Approval | Who approves, role, current status |
Communication rules:
- Explain WHY before WHAT
- Communicate early. Surprises create resistance; previews create buy-in.
- Acknowledge what is being lost, not just what is being gained
- "Everyone" is not a stakeholder. "200 users in the billing team" is.
**Gate**: All plan components complete. Rollback plan has trigger criteria. Approvals identified.
**Phase 3: EXECUTE** -- Monitor adoption and sustain.
- Track adoption metrics
- Address resistance (specific actions, not "manage change")
- Reinforce new behaviors
- Document lessons learned
- Define success criteria and measurement timeline
**Gate**: Adoption metrics defined. Lessons captured. Success criteria measurable.
---
### Mode: CAPACITY
**Framework**: INVENTORY -> FORECAST -> DECIDE
**Phase 1: INVENTORY** -- Map current capacity.
- List team members, roles, and current allocation
- Calculate actual available hours (subtract meetings, PTO, on-call, admin)
- Map current work to people
| Role Type | Target Utilization | Rationale |
|-----------|-------------------|-----------|
| IC / Specialist | 75-80% | Buffer for reactive work and growth |
| Manager | 60-70% | Management overhead, 1:1s, meetings |
| On-call / Support | 50-60% | Interrupt-driven work is unpredictable |
**Gate**: Current capacity mapped. Utilization calculated. Overallocations identified.
**Phase 2: FORECAST** -- Model upcoming demand.
- List upcoming projects with resource requirements and timelines
- Identify skill bottlenecks (the constraint is usually a specific skill, not generic headcount)
- Model scenarios: do nothing, hire X, deprioritize Y, contract Z
**Gate**: Demand mapped. Bottlenecks identified. Scenarios modeled.
**Phase 3: DECIDE** -- Recommend action.
- Compare scenarios: outcome, cost, timeline, risk
- Recommend: hire, contract, reprioritize, delay, or split
- Preserve buffer capacity for unplanned work (target 75-80%). 100% means no buffer for surprises.
**Gate**: Recommendation stated. Trade-offs explicit. Buffer preserved.
---
### Mode: COMPLIANCE
**Framework**: MAP -> GAP -> REMEDIATE
**Phase 1: MAP** -- Identify applicable frameworks and current state.
| Framework | Focus | Key Requirements |
|-----------|-------|-----------------|
| SOC 2 | Service organizations | Security, availability, processing integrity, confidentiality, privacy |
| ISO 27001 | Information security | Risk assessment, security controls, continuous improvement |
| GDPR | Data privacy (EU) | Consent, data rights, breach notification, DPO |
| HIPAA | Healthcare (US) | PHI protection, access controls, audit trails |
| PCI DSS | Payment card data | Encryption, access control, vulnerability management |
- Inventory controls mapped to framework requirements
- Document control owners and evidence locations
**Gate**: Frameworks identified. Controls inventoried.
**Phase 2: GAP** -- Find what is missing or deficient.
Load `references/operations/risk-assessment.md` for risk-based prioritization of gaps.
- Requirements vs. current state for each control
- Evidence gaps: what evidence is needed but not collected
- Priority: compliance risk x remediation effort
**Gate**: Gaps identified and prioritized. Evidence gaps documented.
**Phase 3: REMEDIATE** -- Plan and track closure.
- Remediation plan with owner, timeline, and verification method
- Audit calendar with evidence collection deadlines
- Monitoring: control effectiveness tracking
**Gate**: Remediation plan complete. Audit calendar set. Owners assigned.
---
### Mode: STATUS
**Framework**: GATHER -> SYNTHESIZE -> DELIVER
Produce a status report covering:
| Section | Content |
|---------|---------|
| Executive Summary | 3-4 sentences. What is on track, what needs attention, key wins. |
| Overall Status | On Track / At Risk / Off Track with justification |
| Key Metrics | KPI, target, actual, trend, status |
| Accomplishments | What got done this period |
| In Progress | Item, owner, status, ETA |
| Risks and Issues | Risk, impact, mitigation, owner |
| Decisions Needed | Decision, context, deadline, recommendation |
| Next Priorities | Top 3 for next period |
Rules:
- Lead with the headline. Leadership reads the first 3 lines.
- Be honest about risks. Surfacing issues early builds trust. Surprises erode it.
- For each decision needed, provide context AND a recommendation. Include a recommendation with every decision request.
---
### Mode: OPTIMIZE
**Framework**: MAP -> ANALYZE -> REDESIGN
**Phase 1: MAP** -- Document current state with timing.
Load `references/operations/process-documentation.md`.
- Every step, decision point, and handoff with time estimates
- Identify: who, what, how long, what triggers, what blocks
**Phase 2: ANALYZE** -- Identify waste.
| Waste Type | What to Look For |
|-----------|-----------------|
| Waiting | Time in queues, waiting for approvals |
| Rework | Steps that fail and repeat |
| Handoffs | Each handoff = potential failure/delay point |
| Over-processing | Steps that add no value |
| Manual work | Tasks that could be automated |
**Phase 3: REDESIGN** -- Propose improvements.
- Eliminate unnecessary steps
- Automate where deterministic
- Reduce handoffs
- Parallelize independent steps
- Before/after comparison with: time saved, error rate reduction, cost savings
---
## Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| Vague runbook steps | LLM defaults to abstract language | Force each step through the 5-component template. Reject steps without exact commands. |
| Underestimated risk | Optimism bias in probability scoring | Challenge every "Low" probability. Ask: "What evidence supports Low, not Medium?" |
| Generic process docs | Template fill without reality check | Ask how the process actually works today, not how it should work. |
| Missing rollback | Assumed success path only | Require rollback before marking any change/runbook complete. |
| Scorecard inflation | All vendors score 7+ | Force relative scoring. At least one dimension per vendor must be below 5. |
| RACI with multiple As | Accountability diffusion | Enforce exactly one A per step. Multiple As = no one accountable. |
| Compliance checkbox theater | Controls documented but not tested | Require evidence of control effectiveness, not just existence. |
---
## Reference Loading Table
| Mode | Reference | Content |
|------|-----------|---------|
| RUNBOOK | `references/operations/runbook-authoring.md` | Step structure, verification checklists, rollback procedures, escalation paths |
| RISK, COMPLIANCE | `references/operations/risk-assessment.md` | Probability x impact matrix, risk categories, mitigation planning, residual risk tracking |
| VENDOR | `references/operations/vendor-management.md` | Vendor scorecard, due diligence checklist, contract review triggers, performance monitoring |
| PROCESS, CAPACITY, OPTIMIZE | `references/operations/process-documentation.md` | Process mapping, RACI matrices, bottleneck analysis, optimization methodology |
| CHANGE | `references/operations/change-management.md` | Change request workflows, impact assessment, stakeholder communication, rollback criteria |
| ALL | `references/operations/llm-ops-failure-modes.md` | LLM failure patterns in operations: vague procedures, underestimated risks, generic templates |
references/operations/change-management.md
---
title: Change Management — Request Workflows, Impact Assessment, Communication, Rollback
domain: operations
level: 3
skill: operations
---
# Change Management Reference
> **Scope**: Operational change management for system, process, and organizational changes. Covers change request workflows, impact assessment matrices, stakeholder communication planning, and rollback criteria definition. Use when proposing changes that need approval, documenting impact, planning communication, or defining rollback procedures.
> **Generated**: 2026-05-05 — Change management frameworks should be reviewed annually or when organizational change volume significantly shifts.
---
## Overview
Change management fails at three points. First: impact is underestimated because the assessment only considers the primary system, not downstream dependencies. Second: communication is an afterthought — people learn about the change when it breaks their workflow. Third: rollback is an optimistic sentence instead of a tested procedure.
Every change request is a contract: "Here is what will happen, who it affects, how we will tell them, and what we do if it goes wrong." If any section is vague, the contract is incomplete.
---
## Change Request Workflow
### Change Classification
Classify the change before selecting the workflow. Over-classifying wastes time. Under-classifying misses risks.
| Class | Criteria | Approval Path | Lead Time |
|-------|----------|---------------|-----------|
| **Standard** | Pre-approved, low-risk, routine | Auto-approved per policy | Same day |
| **Normal** | Moderate risk, planned, reversible | Manager + change owner | 5+ business days |
| **Major** | High risk, broad impact, or irreversible | Change Advisory Board (CAB) | 10+ business days |
| **Emergency** | Unplanned, required to restore service | Post-hoc approval within 48h | Immediate |
### Classification Decision Tree
```
Is this change pre-approved and routine?
Yes -> STANDARD (auto-approve)
No -> Does it affect production systems?
No -> NORMAL (manager approval)
Yes -> Is it reversible within 1 hour?
Yes -> How many users affected?
<50 -> NORMAL
50+ -> MAJOR
No -> MAJOR
Is this an emergency to restore service?
Yes -> EMERGENCY (implement now, approve within 48h)
```
### Change Request Template
```
## Change Request: [CR-XXXX] [Title]
### Metadata
| Field | Value |
|-------|-------|
| Requester | [Name] |
| Date submitted | [YYYY-MM-DD] |
| Classification | Standard / Normal / Major / Emergency |
| Priority | Critical / High / Medium / Low |
| Status | Draft / Submitted / Approved / In Progress / Completed / Rolled Back |
| Implementation window | [Date + time range] |
### Description
**What is changing**: [Specific technical or process change]
**Why**: [Business justification — the problem being solved or opportunity being captured]
**What happens if we don't**: [Risk of inaction — establishes urgency]
### Impact Assessment
[See Impact Assessment section below]
### Risk Assessment
[See Risk Assessment reference — load references/operations/risk-assessment.md]
### Implementation Plan
[See Implementation Plan section below]
### Communication Plan
[See Stakeholder Communication section below]
### Rollback Plan
[See Rollback Criteria section below]
### Approvals
| Approver | Role | Decision | Date |
|----------|------|----------|------|
| [Name] | [Role] | Pending / Approved / Rejected | |
```
---
## Impact Assessment
### Impact Dimensions
Assess each dimension independently. Rate High / Medium / Low / None.
| Dimension | High | Medium | Low | None |
|-----------|------|--------|-----|------|
| **Users** | >100 users change daily workflow | 20-100 users, minor workflow change | <20 users, minimal change | No user-facing change |
| **Systems** | Multiple production systems modified | Single production system modified | Non-production system only | No system changes |
| **Processes** | Core business process changes | Supporting process changes | Documentation-only changes | No process changes |
| **Data** | Schema change, data migration, format change | New data fields, optional changes | Read-only access changes | No data impact |
| **Security** | Access control, authentication, encryption changes | Permission scope changes | Audit logging changes | No security impact |
| **Cost** | >$50K budget impact | $5K-$50K budget impact | <$5K budget impact | No budget impact |
| **Compliance** | Regulatory or audit-relevant changes | Policy updates | Documentation updates | No compliance impact |
### Impact Assessment Template
```
## Impact Assessment: [CR-XXXX]
| Area | Rating | Details | Affected Parties |
|------|--------|---------|-----------------|
| Users | [H/M/L/N] | [Specific impact description] | [Who specifically] |
| Systems | [H/M/L/N] | [Which systems, how] | [System owners] |
| Processes | [H/M/L/N] | [Which processes change] | [Process owners] |
| Data | [H/M/L/N] | [What data changes] | [Data owners] |
| Security | [H/M/L/N] | [What controls affected] | [Security team] |
| Cost | [H/M/L/N] | [Budget impact] | [Budget owner] |
| Compliance | [H/M/L/N] | [Regulatory implications] | [Compliance team] |
### Dependencies
| This change depends on | Status |
|----------------------|--------|
| [Prerequisite change/system/team] | [Ready / Not ready / Blocked] |
### Downstream Effects
| What depends on this change | Impact if delayed |
|----------------------------|-------------------|
| [Downstream system/process/team] | [Consequence] |
```
### Dependency Mapping
Changes rarely exist in isolation. Map what this change depends on and what depends on this change.
```
[Upstream Change A] ----\
\
[Upstream Change B] -------> [THIS CHANGE] -------> [Downstream System X]
/ \
[Prerequisite C] -------/ \---> [Downstream Process Y]
```
Questions to ask:
- What must be complete BEFORE this change?
- What breaks if this change is delayed?
- What other changes are happening in the same window?
- Which teams need to coordinate?
---
## Stakeholder Communication
### Communication Planning Matrix
For each stakeholder group, define what, when, how, and who.
| Audience | What They Need to Know | Channel | Timing | Owner | Message Type |
|----------|----------------------|---------|--------|-------|-------------|
| **Directly affected users** | What changes, what to do differently, where to get help | Email + team meeting | 2 weeks before | Change owner | Detailed instructions |
| **Indirectly affected teams** | What is changing, potential impact on their work | Email + Slack | 1 week before | Change owner | Summary + FAQ |
| **Leadership/sponsors** | Status, risks, go/no-go decision | Status report | Per approval cycle | Change owner | Executive summary |
| **Support/helpdesk** | What to expect, how to troubleshoot, escalation path | Training session + runbook | 1 week before | Support lead | Operational guide |
| **External customers** | What changes for them, timeline, support contact | Email / in-app notification | Per SLA/contract | Comms team | Customer-facing notice |
### Communication Principles
1. **Explain WHY before WHAT.** People accept change better when they understand the reason. "We are migrating databases" meets resistance. "The current database cannot handle peak load, which caused the outage last month. We are migrating to a system that handles 10x the traffic" gets buy-in.
2. **Communicate early.** Surprises create resistance. Previews create buy-in. Even incomplete information ("We are planning a change to X. Details coming next week.") is better than silence followed by disruption.
3. **Acknowledge what is being lost.** Every change removes something familiar. Saying "this is an improvement" while ignoring that people's workflows are disrupted is tone-deaf. "We know the current system is familiar, and this change will require adjusting your daily routine for a few weeks" builds trust.
4. **Be specific about impact.** "Everyone will be affected" is not a communication. "The 200 users in the billing team will need to learn the new approval workflow. Training sessions are scheduled for June 3-5." is a communication.
5. **Provide a path for questions.** Announce and disappear = anxiety. Announce with a FAQ, a Slack channel, a drop-in session, and a named contact = manageable change.
### Communication Timeline Template
```
## Communication Timeline: [CR-XXXX]
| When | What | Audience | Channel | Owner | Status |
|------|------|----------|---------|-------|--------|
| T-30 days | Preview announcement | All stakeholders | Email | [Name] | |
| T-14 days | Detailed notification | Affected users | Email + meeting | [Name] | |
| T-7 days | Training session | Affected users | Workshop | [Name] | |
| T-3 days | Final reminder | Affected users | Slack + email | [Name] | |
| T-0 (go-live) | Go-live announcement | All | Slack | [Name] | |
| T+1 day | Status update | All stakeholders | Email | [Name] | |
| T+7 days | Feedback survey | Affected users | Survey | [Name] | |
| T+30 days | Adoption report | Leadership | Status report | [Name] | |
```
### Resistance Management
| Resistance Pattern | Cause | Response |
|-------------------|-------|----------|
| "This was fine before" | Comfort with status quo | Quantify the problem the change solves. Show data. |
| "Nobody asked us" | Lack of involvement | Include resistors in design review. Give them influence. |
| "This is more work" | Change increases short-term effort | Acknowledge the transition cost. Show long-term benefit. Provide support. |
| "It won't work here" | Past failed changes eroded trust | Address specific concerns. Start with pilot group. Show early results. |
| Passive non-adoption | Silent disagreement | Track adoption metrics. Follow up individually with non-adopters. |
| Active undermining | Deep disagreement or threatened role | Address privately. Understand root concern. Escalate if persistent. |
---
## Implementation Planning
### Implementation Plan Template
```
## Implementation Plan: [CR-XXXX]
### Pre-Implementation
| Step | Owner | When | Dependencies | Verification |
|------|-------|------|--------------|-------------|
| [Backup/snapshot] | [Name] | T-1h | None | Backup confirmed |
| [Stakeholder notification] | [Name] | T-1h | None | Sent |
| [Monitoring dashboards open] | [Name] | T-15min | None | URLs loaded |
### Implementation Steps
| # | Step | Owner | Expected Duration | Expected Result | Failure Action |
|---|------|-------|-------------------|-----------------|----------------|
| 1 | [Action] | [Name] | [Xmin] | [Observable result] | [What to do] |
| 2 | [Action] | [Name] | [Xmin] | [Observable result] | [What to do] |
| 3 | [Action] | [Name] | [Xmin] | [Observable result] | [What to do] |
### Post-Implementation
| Step | Owner | When | Verification |
|------|-------|------|-------------|
| [Verify service health] | [Name] | T+15min | [Health check passes] |
| [Monitor for 30 minutes] | [Name] | T+30min | [No new alerts] |
| [Notify stakeholders: complete] | [Name] | T+30min | [Message sent] |
| [Close maintenance window] | [Name] | T+1h | [Window closed] |
```
### Go/No-Go Checklist
Run before implementation begins. All items must pass.
```
## Go/No-Go Checklist: [CR-XXXX]
### Go Criteria (all must be YES)
- [ ] All approvals obtained
- [ ] Implementation team available and briefed
- [ ] Rollback procedure reviewed and understood
- [ ] Backups/snapshots taken and verified
- [ ] Monitoring and alerting configured
- [ ] Communication sent to affected parties
- [ ] Maintenance window confirmed (if applicable)
- [ ] No conflicting changes in the same window
- [ ] Support/on-call team briefed on the change
### No-Go Criteria (any one triggers postpone)
- [ ] Outstanding unresolved risks rated High or Critical
- [ ] Key implementation team member unavailable
- [ ] Dependent upstream change not yet complete
- [ ] Production incident in progress
- [ ] Insufficient time remaining in maintenance window
```
---
## Rollback Criteria and Procedures
### When to Roll Back
Define trigger criteria BEFORE implementation. During an incident is too late to debate.
| Trigger Type | Example | Action |
|-------------|---------|--------|
| **Hard trigger** (automatic rollback) | Error rate >5% for 5 minutes | Roll back immediately. No discussion. |
| **Soft trigger** (evaluate, then decide) | Performance degraded 20% but stable | Assess. Roll back if not improving within 30 minutes. |
| **Time-based** | Change not verified within maintenance window | Roll back. Reschedule. |
| **User-reported** | >10 user reports of the same issue | Assess severity. Roll back if user-facing impact confirmed. |
### Rollback Plan Template
```
## Rollback Plan: [CR-XXXX]
### Decision Authority
- **Who can authorize rollback**: [Name/Role]
- **Automatic rollback triggers**: [List hard triggers]
- **Escalation if unsure**: [Contact]
### Rollback Steps
| # | Step | Command/Action | Expected Result | Duration |
|---|------|---------------|-----------------|----------|
| 1 | [Undo step N] | [Exact action] | [Expected state] | [Xmin] |
| 2 | [Undo step N-1] | [Exact action] | [Expected state] | [Xmin] |
| 3 | [Verify restoration] | [Check command] | [Normal state confirmed] | [Xmin] |
### Total Rollback Time: [X minutes]
### Post-Rollback
- [ ] Verify system returned to pre-change state
- [ ] Notify stakeholders: "Change [CR-XXXX] rolled back. [Reason]. No data loss."
- [ ] Create incident record if applicable
- [ ] Schedule post-mortem within 48 hours
- [ ] Update change request status to "Rolled Back"
### Non-Reversible Components
| Component | Why Non-Reversible | Compensating Action |
|-----------|-------------------|---------------------|
| [Component] | [Reason] | [What to do instead] |
```
### Rollback Testing
Rollback procedures must be tested before the change. An untested rollback is a hope, not a plan.
| Test Method | When to Use |
|------------|------------|
| Full rehearsal in staging | Major changes with complex rollback |
| Tabletop walkthrough | Normal changes with straightforward rollback |
| Automated rollback test | Standard changes with scripted rollback |
| Documented from prior experience | Emergency changes (tested post-hoc if time permits) |
---
## Change Management Failure Modes
| Failure Mode | Symptom | Fix |
|-------------|---------|-----|
| Classification avoidance | Everything is "Standard" to skip approval | Audit classification decisions. Auto-escalate when impact assessment contradicts classification. |
| Communication as afterthought | Users discover changes by encountering them | Communication plan is a required section. No approval without it. |
| Vague rollback | "Undo the changes" | Rollback must have exact steps, expected results, and tested procedure. |
| Ignoring downstream | Only assessed impact on primary system | Dependency mapping is required. Ask: "What depends on this?" |
| Emergency abuse | "Emergency" classification to bypass approval | Track emergency frequency. If >20% of changes are emergency, the process or the system has problems. |
| No post-implementation review | Change completed, nobody checks if it worked | Require verification within 24 hours. Track success rate. |
| Change collision | Multiple changes in same window, failure attribution impossible | Maintain change calendar. No overlapping changes on same system. |
| Resistance dismissed | "People will get used to it" | Resistance is signal. Track adoption. Follow up with non-adopters. Adjust. |
references/operations/llm-ops-failure-modes.md
---
title: LLM Operations Failure Modes — Where LLMs Fail in Operations Work
domain: operations
level: 3
skill: operations
---
# LLM Operations Failure Modes
> **Scope**: Catalog of predictable LLM failure modes when generating operational content — runbooks, risk assessments, process documentation, change requests, vendor evaluations, and capacity plans. Each failure mode includes detection criteria and countermeasures. Load this reference for ALL operations modes. It is the quality guard for everything this skill produces.
> **Generated**: 2026-05-05 — LLM capabilities evolve. Re-validate failure modes when underlying models change.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers operations-specific failures only.
---
## Overview
LLMs produce plausible-looking operational content that fails in production. The failure is insidious because the output reads well — proper formatting, complete sections, professional tone — while containing gaps that only become visible when someone tries to follow the runbook at 3am, assess actual risk under pressure, or execute the change plan.
This reference catalogs where LLMs predictably fail in operations work, how to detect each failure, and what to do about it. Every operations mode should be checked against this list before delivery.
---
## Failure Mode 1: Vague Procedures
### The Problem
LLMs default to abstract language. "Deploy the changes" instead of the exact command. "Verify the system" instead of the specific check. "Update the configuration" instead of which file, which key, which value.
This is the #1 failure in runbook generation. The LLM produces a document that looks complete but cannot be followed by someone who does not already know the procedure.
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| Missing paths | "Edit the config file" | Does the step specify which file? Full path? |
| Missing commands | "Run the migration" | Is the exact command present with all flags? |
| Missing users | "SSH to the server" | Which server? As which user? With which key? |
| Missing output | "Check the logs" | Which log? What to grep for? What does success look like? |
| Weasel verbs | "Ensure", "Verify", "Confirm" without specifying HOW | Is there a concrete action after every verification verb? |
| Implicit knowledge | "The usual process" | Would a new hire understand this without asking someone? |
| Assumed context | "As before" / "Same as above" | Is each step self-contained? |
### Countermeasures
1. **Five-component enforcement**: Every step must have Action, Expected Result, Failure Handling, Verification, and Rollback. Reject steps missing any component.
2. **New-hire test**: Read each step as if you have never seen the system. If any step requires knowledge not present in the document, it fails.
3. **Command audit**: Every step that involves running a command must include the full command with path, flags, user context, and expected output.
4. **Verb specificity check**: Flag "ensure", "verify", "confirm", "check", "validate" that are not followed by a concrete how-to within the same sentence.
---
## Failure Mode 2: Underestimated Risk Ratings
### The Problem
LLMs have a systematic bias toward rating risks as "Low" or "Medium" when evidence is ambiguous. This manifests as:
- "Low probability" without evidence (absence of incidents ≠ low probability)
- Downplaying impact by using qualitative descriptions instead of quantifiable measures
- Rating everything "Medium" to avoid seeming alarmist (the "everything is Medium" trap)
- Conflating "unlikely" with "impossible"
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| Unsupported "Low" | "Probability: Low" with no rationale | Is there evidence justifying Low, or is it default? |
| Unquantified impact | "Impact: High" | Is impact expressed in dollars, hours, users, or SLA? |
| Everything Medium | 5+ risks all rated Medium | Is there genuine differentiation, or is Medium the default? |
| Missing worst case | Risk only describes likely outcome | What happens in the worst realistic case? |
| Optimism framing | "Should not happen" / "Unlikely scenario" | "Should not" is not "cannot." What conditions would cause it? |
| No evidence for probability | "Has not happened before" | "Has not happened" requires "and here is why it cannot" to justify Low |
### Countermeasures
1. **Evidence requirement**: Every probability rating must include supporting evidence. No evidence = cannot rate Low.
2. **Quantification mandate**: At least one impact dimension must be quantified (dollars, hours, affected users, SLA percentage).
3. **Distribution check**: If >60% of risks are rated the same level, force re-assessment. Real risk profiles are not uniform.
4. **Challenge-down test**: For every "Low" rating, ask: "What would need to be true for this to be Medium?" If the answer describes current conditions, it is Medium.
5. **Pre-mortem frame**: Instead of "What is the risk?", ask "Assume this failed. What went wrong?" Different framing produces different (often higher) severity ratings.
---
## Failure Mode 3: Generic Templates Without Context
### The Problem
LLMs fill templates with plausible-sounding but generic content. The process documentation, change request, or vendor evaluation looks complete but contains no information specific to the actual organization, system, or situation.
Generic tells:
- Risk mitigations that apply to any organization ("implement monitoring")
- Process steps that describe any process ("review and approve")
- Vendor evaluations that could describe any vendor ("strong market position")
- Change requests with no system-specific impact analysis
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| Substitutable content | "Implement appropriate monitoring" | Would this sentence change if the system/vendor/process were different? |
| No proper nouns | Risk register with no system names | Does the document reference specific systems, teams, people? |
| Template language | "As appropriate" / "as needed" / "relevant stakeholders" | These are placeholders, not content. |
| Missing metrics | "Measure performance" | Which metric? What target? How to collect? |
| Universal truths | "Ensure high availability" | Would any organization disagree? If nobody would disagree, it carries no information. |
| No exceptions | Process with zero edge cases | Every real process has exceptions. Zero exceptions = template, not documentation. |
### Countermeasures
1. **Proper noun density**: Every operations document should reference specific systems, teams, tools, and people by name. If it could describe any organization, it does not describe yours.
2. **Substitution test**: Replace the subject (vendor name, system name, process name) with a different one. If the document still reads correctly, it is generic.
3. **Exception requirement**: Every process document must include at least 3 exceptions/edge cases. If the LLM produces zero, push back: "What goes wrong? What is the unusual case?"
4. **Metrics specificity**: Every metric must have a named KPI, a numeric target, and a collection method. "Track performance" is not a metric.
---
## Failure Mode 4: Missing Rollback and Recovery
### The Problem
LLMs consistently produce success-path-only content. Runbooks describe what happens when everything works. Change requests describe the implementation but not the undo. Risk assessments describe mitigations but not what to do when mitigations fail.
This is optimism bias in document form. The document assumes the happy path because describing failure modes is harder and the LLM has no operational experience to draw from.
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| No rollback section | Change request with implementation plan but no undo | Is there a tested procedure to reverse this change? |
| Success-only steps | Runbook steps without failure handling | What happens when step 3 fails? Is that documented? |
| No escalation paths | Procedure with no "if this does not work" guidance | When should the operator stop and call someone? |
| Rollback as afterthought | "If something goes wrong, undo the changes" | Are rollback steps as specific as implementation steps? |
| Untested rollback | "Rollback: revert to previous version" | Has this rollback been tested? How long does it take? |
| Missing non-reversible markers | Steps that cannot be undone are not called out | Which steps are irreversible? What compensating actions exist? |
### Countermeasures
1. **Symmetry requirement**: For every implementation step, require a corresponding rollback step. If rollback is not possible, require explicit documentation of why and what compensating action exists.
2. **Rollback specificity parity**: Rollback steps must be as specific as implementation steps. "Undo the migration" fails the same vagueness test as "Run the migration."
3. **Non-reversible marking**: Any step that cannot be undone must be flagged prominently with compensating actions documented.
4. **Escalation completeness**: Every procedure must define when to stop trying and escalate, with contact information and message template.
5. **Rollback time estimation**: Every rollback plan must include estimated duration. "Roll back" without timing is not planning.
---
## Failure Mode 5: Artificial Completeness
### The Problem
LLMs produce documents that look complete — all sections filled, all tables populated, all checkboxes checked — without having the information to fill them. This is more dangerous than obvious incompleteness because it passes casual review.
The LLM invents plausible-sounding data to fill templates: made-up SLA numbers, invented vendor capabilities, fabricated process metrics, placeholder people in RACI matrices.
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| Suspiciously complete | All fields filled on first draft without user input | Were all these values provided, or were some generated? |
| Round numbers | "99.9% uptime" / "$100,000 impact" / "50 users affected" | Are these measured values or estimates? |
| Named people without verification | "Owner: John Smith" | Was this person confirmed as owner, or was the name generated? |
| Metrics without sources | "Error rate: 2.3%" | Where does this number come from? Can you verify it? |
| Consistent formatting | Every risk scored, every field filled, no gaps | Gaps in real assessments are normal. Perfect completeness is suspicious. |
### Countermeasures
1. **Mark unknowns explicitly**: It is better to write "[UNKNOWN — verify with ops team]" than to fabricate a value. Require LLM to mark any value not directly provided by the user.
2. **Source attribution**: Every metric, number, or factual claim must cite its source. No source = flagged for verification.
3. **Confidence tagging**: Tag each section as "verified" (user provided data), "estimated" (LLM inferred from context), or "placeholder" (needs verification). Never tag fabricated data as verified.
4. **Gap tolerance**: Incomplete documents with honest gaps are better than complete documents with hidden fabrications. Enforce a culture where "[TODO]" is professional and fabrication is failure.
---
## Failure Mode 6: Shallow Compliance Artifacts
### The Problem
LLMs produce compliance documentation that satisfies the structure of a framework without satisfying the intent. Controls are documented as existing without evidence. Policies reference procedures that do not exist. Risk assessments list mitigations that have never been implemented.
This is checkbox theater: the document exists, the controls are listed, the boxes are checked — and none of it reflects reality.
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| Controls without evidence | "RBAC is implemented" | Where is the evidence? Access review logs? Configuration screenshots? |
| Policies without procedures | "Data retention policy exists" | Where is the procedure that enacts the policy? Who executes it? |
| Mitigations without implementation | "Automated scanning is in place" | When was the last scan? What were the results? Who reviews them? |
| Status "Complete" without date | "Control implemented" | When? By whom? Last verified when? |
| Generic control descriptions | "Appropriate security controls are in place" | Which controls? Configured how? Verified how? |
### Countermeasures
1. **Evidence chain**: Every control must link to evidence. Evidence must be dated and attributed. "Control exists" without evidence = control does not exist for audit purposes.
2. **Procedure validation**: For every documented policy, verify the implementing procedure exists, is current, and has been executed within its review cycle.
3. **Last-verified dates**: Every control must have a last-verified date and a next-review date. Undated controls are unverified controls.
4. **Reality sampling**: Pick 3 controls at random and verify they work as documented. If any fail, the entire compliance artifact is suspect.
---
## Failure Mode 7: Copy-Paste Mitigations
### The Problem
LLMs reuse the same mitigation language across different risks. "Implement monitoring and alerting" appears as the mitigation for network outage, data loss, security breach, and performance degradation. These are not mitigations — they are noise masquerading as action.
### Detection Criteria
| Indicator | Example | Test |
|-----------|---------|------|
| Repeated mitigations | Same "monitoring and alerting" for 3+ risks | Is the monitoring different for each risk, or is it the same phrase? |
| Non-specific monitoring | "Monitor the situation" | Monitor what? What threshold triggers action? Who acts? |
| Unbounded "improve" | "Improve documentation" | Which documentation? What improvement? By when? |
| Activity without outcome | "Conduct regular reviews" | How often? Who reviews? What happens with findings? |
| Mitigation = control | "Implement access controls" | This is a control, not a mitigation plan. What specific access control? For what? By when? |
### Countermeasures
1. **Uniqueness check**: Compare mitigations across the risk register. If >2 risks share identical mitigation language, each must be differentiated for its specific risk.
2. **SMART test**: Every mitigation must be Specific, Measurable, Assignable, Relevant, and Time-bound. "Monitor the situation" fails all five.
3. **Action-outcome pairing**: Every mitigation must state both the action AND the expected outcome. "Implement monitoring" -> "Configure PagerDuty alert for error rate >5%, reviewed by on-call engineer within 15 minutes, expected to reduce MTTR from 2 hours to 30 minutes."
4. **Owner and date**: No mitigation without a named owner and a completion date. Unowned mitigations are aspirations, not plans.
---
## Cross-Mode Quality Checklist
Run this checklist against every operations output before delivery.
```
## Operations Output Quality Gate
### Specificity
- [ ] No steps use "ensure", "verify", "check" without specifying HOW
- [ ] All commands include full paths, flags, and user context
- [ ] All references to systems/people/tools use proper nouns
- [ ] Could someone unfamiliar with the context follow this document?
### Risk Honesty
- [ ] No "Low" probability rating without supporting evidence
- [ ] All impact ratings include at least one quantified dimension
- [ ] Risk distribution is not artificially uniform
- [ ] "Has not happened" is not used as sole justification for "Low"
### Completeness Integrity
- [ ] Unknown values marked as [UNKNOWN] or [TODO], not fabricated
- [ ] All metrics cite their source
- [ ] Named people/roles are verified, not generated
- [ ] Gaps are honest, not filled with generic content
### Recovery Planning
- [ ] Rollback procedure exists for every changeable state
- [ ] Rollback steps are as specific as implementation steps
- [ ] Non-reversible steps are flagged with compensating actions
- [ ] Escalation paths include contact info and trigger criteria
### Context Specificity
- [ ] Substitution test passed (change the subject — does the doc still work? If yes, too generic)
- [ ] At least 3 exceptions/edge cases documented per process
- [ ] Mitigations are unique per risk, not copy-pasted
- [ ] Templates are filled with real data, not placeholder patterns
```
references/operations/process-documentation.md
---
title: Process Documentation — Mapping, RACI, Bottleneck Analysis, Optimization
domain: operations
level: 3
skill: operations
---
# Process Documentation Reference
> **Scope**: Business process documentation from capture through optimization. Covers process mapping techniques, RACI matrix construction and enforcement, bottleneck analysis methodology, and process optimization frameworks. Use when documenting existing processes, building SOPs, identifying inefficiencies, or planning capacity.
> **Generated**: 2026-05-05 — Processes drift. Review SOPs quarterly or when incidents reveal the documented process diverges from reality.
---
## Overview
Process documentation fails for one reason: it describes how the process should work instead of how it actually works. The idealized version lives in a wiki nobody reads. The real process — with its workarounds, exceptions, and tribal knowledge — lives in people's heads. When those people leave, the process leaves with them.
Document reality first. Optimize second. A perfect SOP for a process nobody follows is worse than a messy SOP that matches what people actually do.
---
## Process Mapping
### Capture Methodology
Start by interviewing the people who do the work. Not their managers. Not the original process designer. The people who run it today.
**Questions to ask:**
| Question | What It Reveals |
|----------|----------------|
| "Walk me through the last time you did this." | The actual process, not the idealized version |
| "What do you do differently when [exception]?" | Edge cases that break the documented flow |
| "Where do you usually get stuck or wait?" | Bottlenecks and handoff failures |
| "What do you wish someone had told you when you started?" | Tribal knowledge that should be documented |
| "What step do you sometimes skip? What happens?" | Process steps that may be unnecessary or poorly designed |
| "Who do you go to when something goes wrong?" | Undocumented escalation paths |
| "What tools do you use that aren't in the official process?" | Shadow IT and workarounds |
### Process Map Components
Every process map must include:
| Component | Description | Example |
|-----------|-------------|---------|
| **Trigger** | What initiates the process | Customer submits ticket. End of month. Alert fires. |
| **Steps** | Sequential actions with decision points | Triage, assign, investigate, resolve, verify, close |
| **Decision points** | Binary branches in the flow | Severity > P2? Yes -> escalate. No -> standard queue. |
| **Handoffs** | Where work moves between people/teams | Triage team -> assigned engineer |
| **Inputs** | What each step needs to start | Ticket details, access credentials, runbook |
| **Outputs** | What each step produces | Diagnosis, fix, verification result, resolution note |
| **Exceptions** | What happens outside the normal flow | Customer escalates. Duplicate ticket. Known issue. |
| **Timing** | How long each step typically takes | Triage: 5min. Investigation: 15min-4hr. Resolution: varies. |
| **Roles** | Who performs each step | L1 support, on-call engineer, team lead |
### ASCII Process Flow Template
For text-based documentation:
```
[Trigger] -> [Step 1] -> <Decision?> --Yes--> [Step 2a] -> [Step 3]
|
No
|
v
[Step 2b] -> [Step 3]
[Step 3] -> <Verification?> --Pass--> [Step 4: Close]
|
Fail
|
v
[Step 3: Rework] -> <Verification?> ...
```
### Multi-Swim-Lane Template
For processes crossing teams:
```
| Team A | Team B | Team C |
|-----------------|-----------------|-----------------|
| [Request] | | |
| | | | |
| v | | |
| [Triage] ------>| [Assign] | |
| | | | |
| | v | |
| | [Investigate] | |
| | | | |
| | <Escalate?> --->| [Review] |
| | | | | |
| | No | v |
| | v | [Approve/Deny] |
| | [Resolve] <----| | |
| | | | |
| [Notify] <------| [Close] | |
```
---
## RACI Matrix
### RACI Definitions
| Role | Definition | Rule |
|------|-----------|------|
| **R** — Responsible | Does the work | At least one per step. Can be multiple people. |
| **A** — Accountable | Owns the outcome. Makes the final call. | Exactly ONE per step. This is non-negotiable. |
| **C** — Consulted | Provides input before the work is done | Two-way communication. Keep this list short. |
| **I** — Informed | Notified after the work is done | One-way communication. Can be broad. |
### RACI Construction Rules
1. **One A per step.** If two people are accountable, nobody is. This is the most violated rule and the most important.
2. **R without A = work without ownership.** Every R needs an A above it who cares if it gets done.
3. **A without R = accountability without capability.** The accountable person must have authority over the responsible person.
4. **Too many Cs = slow process.** Every C is a person who must be consulted before proceeding. Each one adds latency.
5. **Too many Rs = diffused responsibility.** If everyone is responsible, nobody feels ownership.
### RACI Template
```
## RACI Matrix: [Process Name]
| Step | Team Lead | Engineer | QA | PM | Support | Ops |
|------|-----------|----------|----|----|---------|-----|
| 1. Receive request | I | | | R | A | |
| 2. Triage and prioritize | C | | | A | R | |
| 3. Assign work | | R | | A | I | |
| 4. Execute | | R | C | A | | I |
| 5. Review/QA | | C | R | A | | |
| 6. Deploy | | R | | I | | A |
| 7. Verify in production | | C | R | I | | A |
| 8. Close and notify | I | | | | R | A |
```
### RACI Validation Checklist
- [ ] Every step has exactly one A
- [ ] Every step has at least one R
- [ ] No person is both R and A on the same step (unless team of one)
- [ ] C count per step is ≤3 (more than 3 = bottleneck)
- [ ] No person is A on >50% of steps (overloaded decision-maker)
- [ ] Every R has a named person or role, not "TBD"
- [ ] The A for each step can actually make decisions (has authority)
- [ ] I recipients have a defined notification method
---
## SOP Template
### Standard Operating Procedure Structure
```
## Process: [Name]
**Owner**: [Person/Team — the single accountable party for this process]
**Last Updated**: [YYYY-MM-DD]
**Review Cadence**: [Quarterly / Semi-annually / Annually]
**Version**: [X.Y]
### Purpose
[One paragraph: why this process exists, what problem it solves, what outcome it produces]
### Scope
- **Included**: [What this process covers]
- **Excluded**: [What this process does NOT cover — prevent scope creep]
- **Related processes**: [Links to upstream/downstream processes]
### Trigger
[What initiates this process: event, schedule, request, threshold]
### Prerequisites
- [Access/permissions required]
- [Tools/systems required]
- [Data/inputs required]
- [Training/certification required]
### RACI Matrix
[See RACI section above]
### Procedure
#### Step 1: [Name]
- **Who**: [Role]
- **When**: [Trigger or timing relative to previous step]
- **Input**: [What this step needs]
- **Action**: [Exact instructions — specific enough for someone unfamiliar]
- **Output**: [What this step produces]
- **Duration**: [Typical time: "5-10 minutes"]
- **Exception**: [What to do if it does not go as expected]
#### Step 2: [Name]
[Same structure]
### Exceptions and Edge Cases
| Scenario | Frequency | Handling |
|----------|-----------|---------|
| [Exception] | [Common/Rare] | [What to do] |
### Metrics
| Metric | Target | Measurement Method | Review Cadence |
|--------|--------|--------------------|----------------|
| Cycle time | [Target] | [How to measure] | [Weekly/Monthly] |
| Error rate | [Target] | [How to measure] | [Weekly/Monthly] |
| Throughput | [Target] | [How to measure] | [Weekly/Monthly] |
### History
| Version | Date | Author | Change |
|---------|------|--------|--------|
| 1.0 | [Date] | [Name] | Initial version |
```
---
## Bottleneck Analysis
### Identifying Bottlenecks
A bottleneck is the step that constrains the throughput of the entire process. Improving any other step does not increase overall throughput until the bottleneck is resolved.
**Detection methods:**
| Method | How | Signal |
|--------|-----|--------|
| Queue length | Measure work-in-progress before each step | Longest queue = bottleneck |
| Wait time | Measure time between step completion and next step start | Longest wait = handoff bottleneck |
| Utilization | Measure how busy each role/resource is | >90% utilization = capacity bottleneck |
| Rework rate | Measure how often each step's output is rejected | Highest rework = quality bottleneck |
| Dependency count | Count how many steps must complete before this step can start | Most dependencies = coordination bottleneck |
### Bottleneck Classification
| Type | Cause | Example | Fix Approach |
|------|-------|---------|-------------|
| **Capacity** | Not enough people/resources | One reviewer for all code reviews | Add capacity, parallelize, batch |
| **Handoff** | Work stalls between teams | Dev-to-QA handoff takes 2 days average | Reduce handoffs, co-locate, automate handoff |
| **Approval** | Waiting for sign-off | Every change needs VP approval | Delegate authority, batch approvals, set SLAs |
| **Information** | Missing data blocks progress | Can't start without spec that's always late | Pull information upstream, provide templates, set deadlines |
| **Skill** | Only one person can do this step | Only Alice knows the billing system | Cross-train, document, reduce specialization dependency |
| **Tool** | Tooling is slow or manual | Manual data entry between systems | Automate, integrate, replace tool |
### Bottleneck Analysis Template
```
## Bottleneck Analysis: [Process Name]
**Date**: [YYYY-MM-DD]
**Analyst**: [Name]
### Process Metrics
| Step | Avg Duration | Wait Before | Queue Depth | Rework Rate | Resource Utilization |
|------|-------------|-------------|-------------|-------------|---------------------|
| [Step 1] | [time] | [time] | [count] | [%] | [%] |
| [Step 2] | [time] | [time] | [count] | [%] | [%] |
### Identified Bottlenecks (ranked by impact)
1. **[Step/Handoff]**: [Type] bottleneck
- Evidence: [What data shows this is the constraint]
- Impact: [How much this slows the overall process]
- Root cause: [Why this step is constrained]
- Recommendation: [Specific fix]
- Estimated improvement: [Time/throughput gain]
### Constraint Chain
[Sometimes fixing bottleneck #1 reveals bottleneck #2 was always there]
Current: Step 3 (capacity) -> Step 5 (approval) -> Step 7 (skill)
After fixing Step 3: Step 5 becomes the new constraint
```
---
## Process Optimization
### Waste Identification Framework
Based on lean methodology, adapted for knowledge work.
| Waste Type | Definition | Knowledge Work Examples | Detection |
|-----------|-----------|------------------------|-----------|
| **Waiting** | Time spent idle, in queues, or blocked | Waiting for approval. Waiting for spec. Waiting for environment. | Measure queue time between steps. |
| **Rework** | Doing work that was already done (due to defects) | Bug found in production. Spec changed after implementation. | Count iterations per work item. |
| **Handoffs** | Transferring work between people/teams | Dev to QA. Support to engineering. Design to development. | Count handoffs per process. Each one = context loss. |
| **Over-processing** | Doing more work than necessary | Detailed report nobody reads. Three levels of approval for minor changes. | Ask: "What happens if we skip this step?" |
| **Motion** | Unnecessary movement/switching | Context-switching between tasks. Navigating between tools. | Count tool switches and context changes per task. |
| **Inventory** | Work-in-progress that is not moving | 50 open tickets, 30 untouched for 2+ weeks. | Count WIP. If WIP > throughput × cycle time, you have inventory waste. |
| **Overproduction** | Making more than needed or before needed | Building features nobody requested. Creating reports in advance. | Track feature usage. Track report readership. |
### Optimization Strategies
| Strategy | When to Use | Example |
|----------|------------|---------|
| **Eliminate** | Step adds no value | Remove approval step for changes under $500 |
| **Automate** | Step is deterministic and repetitive | Auto-assign tickets based on component tags |
| **Parallelize** | Independent steps run sequentially | Run security review and code review simultaneously |
| **Simplify** | Step is more complex than necessary | Replace 5-page change request with structured form |
| **Combine** | Multiple steps could be one | Merge triage and assignment into single step |
| **Batch** | Small items processed one-at-a-time inefficiently | Batch minor approvals into daily review |
| **Shift left** | Defects found late, expensive to fix | Add linting to pre-commit hooks instead of code review |
### Before/After Comparison Template
```
## Process Optimization: [Name]
### Current State
- Total steps: [N]
- Total cycle time: [X hours/days]
- Handoffs: [N]
- Approval gates: [N]
- Known pain points: [List]
### Proposed Changes
| Change | Type | Step Affected | Expected Impact |
|--------|------|--------------|-----------------|
| [Change] | Eliminate/Automate/etc. | Step [N] | Save [X] time per cycle |
### Future State
- Total steps: [N] (was [N])
- Total cycle time: [X] (was [X])
- Handoffs: [N] (was [N])
- Approval gates: [N] (was [N])
### Impact Summary
| Metric | Before | After | Improvement |
|--------|--------|-------|-------------|
| Cycle time | [X] | [X] | [X]% faster |
| Error rate | [X]% | [X]% | [X]% reduction |
| Cost per cycle | $[X] | $[X] | $[X] saved |
| Throughput | [X]/week | [X]/week | [X]% increase |
### Implementation Plan
| Phase | Action | Owner | Timeline |
|-------|--------|-------|----------|
| 1 | [Quick win] | [Name] | [Date] |
| 2 | [Medium effort] | [Name] | [Date] |
| 3 | [Longer-term] | [Name] | [Date] |
```
---
## Capacity Planning Integration
### Resource-to-Process Mapping
When processes are documented, map resource consumption for capacity planning.
| Process | Frequency | Duration | Roles Required | Peak Load |
|---------|-----------|----------|---------------|-----------|
| [Process A] | Daily | 2hr | 1 engineer, 1 PM | Month-end |
| [Process B] | Weekly | 4hr | 2 engineers | Q4 |
| [Process C] | On-demand | 1hr avg | 1 support | After releases |
### Utilization Impact Analysis
When optimizing a process, quantify the capacity freed:
```
Current: Process takes 4hr/week from senior engineer (10% utilization)
Optimized: Process takes 1hr/week (2.5% utilization)
Freed: 3hr/week = 7.5% capacity returned to project work
Annual value: 3hr × 48 weeks × $75/hr loaded cost = $10,800
```
---
## Process Documentation Failure Modes
| Failure Mode | Symptom | Fix |
|-------------|---------|-----|
| Aspirational documentation | SOP describes how it should work, not how it does | Interview practitioners. Document current state first. Optimize second. |
| RACI with multiple As | Two people "accountable" per step | Enforce one A per step. Resolve ownership disputes before documenting. |
| Exception-free processes | No edge case handling documented | Ask "What do you do when X goes wrong?" for every step. |
| Orphan SOPs | Created once, never updated, nobody reads it | Embed process docs in the workflow (linked from tickets, dashboards). Review cadence enforced. |
| Over-documented | 50-page SOP for a 5-step process | Match documentation depth to process complexity. Runbooks for ops. SOPs for business processes. |
| Documenting tools, not outcomes | "Click the blue button, then the green button" | Document what to accomplish, not how to navigate a UI (UIs change). |
| No metrics | Process documented but no way to know if it works | Every SOP needs at least one measurable metric with a target. |
| Copy-paste SOPs | Same template filled identically for different processes | Each process has unique exceptions, timing, and failure modes. Generic templates produce generic docs. |
references/operations/risk-assessment.md
---
title: Risk Assessment — Probability x Impact Matrix, Categories, Mitigation Planning, Residual Risk
domain: operations
level: 3
skill: operations
---
# Risk Assessment Reference
> **Scope**: Operational risk assessment for projects, processes, vendors, and changes. Covers the probability x impact scoring matrix, risk category taxonomy, mitigation planning with ownership, residual risk tracking, and compliance-adjacent risk management. Use when identifying, scoring, or mitigating operational risks. For strategic/CEO-level risk, see `skills/business/business-ops/references/risk-assessment.md`.
> **Generated**: 2026-05-05 — Risk assessments are time-sensitive. Reassess when conditions change, not on a fixed schedule alone.
---
## Overview
Risk assessment fails in two predictable ways. First: risks get identified but never scored, so everything is "High" and nothing is prioritized. Second: risks get scored but mitigations are vague ("monitor the situation"), so the risk register becomes documentation theater.
This reference enforces structured scoring with evidence requirements and mitigation specificity. A risk rated "Low" without rationale is as suspicious as a risk rated "High" — both need justification.
---
## Probability x Impact Matrix
### The Matrix
| | **Negligible Impact** | **Minor Impact** | **Moderate Impact** | **Major Impact** | **Severe Impact** |
|---|---|---|---|---|---|
| **Almost Certain (>90%)** | Medium | High | High | Critical | Critical |
| **Likely (60-90%)** | Low | Medium | High | High | Critical |
| **Possible (30-60%)** | Low | Medium | Medium | High | High |
| **Unlikely (10-30%)** | Low | Low | Medium | Medium | High |
| **Rare (<10%)** | Low | Low | Low | Medium | Medium |
### Probability Scoring Guide
Do not guess. Score based on evidence.
| Rating | Probability | Evidence Required |
|--------|------------|-------------------|
| Almost Certain | >90% | Has happened multiple times. Current conditions make recurrence near-guaranteed. |
| Likely | 60-90% | Has happened before in similar circumstances. Contributing factors are present. |
| Possible | 30-60% | Could happen. Some contributing factors exist. No direct precedent but plausible scenario. |
| Unlikely | 10-30% | Theoretically possible but few contributing factors present. Has not happened here. |
| Rare | <10% | Requires multiple simultaneous failures or unprecedented conditions. |
**Scoring discipline**: "It hasn't happened yet" does NOT justify "Rare." Ask: "What would need to be true for this to happen?" If the answer involves conditions that currently exist, move the probability up.
### Impact Scoring Guide
Quantify impact in at least one measurable dimension.
| Rating | Financial | Operational | Reputational | Compliance |
|--------|-----------|------------|--------------|------------|
| Severe | >$1M loss or >20% revenue impact | Complete service outage >24h. Data loss. | National media coverage. Customer exodus. | License revocation. Criminal liability. |
| Major | $100K-$1M loss or 5-20% revenue | Service degraded >4h. Significant data exposure. | Industry press coverage. Key customer loss. | Regulatory fine. Formal investigation. |
| Moderate | $10K-$100K loss or 1-5% revenue | Service degraded 1-4h. Limited data exposure. | Customer complaints. Social media attention. | Audit finding. Remediation required. |
| Minor | $1K-$10K loss or <1% revenue | Brief disruption <1h. No data loss. | Internal complaints. No external visibility. | Minor non-conformance. Self-remediated. |
| Negligible | <$1K loss | Inconvenience only. No service impact. | No visibility. | Documentation gap only. |
Adjust thresholds to your organization's scale. A $10K loss is "Negligible" for a Fortune 500 and "Major" for a 10-person startup.
---
## Risk Categories
### Category Taxonomy
| Category | Subcategories | Example Risks |
|----------|--------------|---------------|
| **Operational** | Process, People, Systems, Suppliers | Key-person dependency. Single point of failure in infrastructure. Manual process error rate. |
| **Financial** | Budget, Revenue, Cost, Liquidity | Vendor price escalation beyond budget. Project cost overrun. Revenue shortfall from delay. |
| **Compliance** | Regulatory, Contractual, Internal Policy | GDPR data handling violation. Missed SLA triggering penalty. Audit finding with remediation deadline. |
| **Strategic** | Market, Competition, Technology, Timing | Technology shift making current stack obsolete. Competitor launch. Market contraction. |
| **Reputational** | Customer, Partner, Public, Employee | Data breach notification. Service outage during peak. Public security incident. |
| **Security** | Data, Access, Infrastructure, Third-Party | Unauthorized access via compromised vendor. Unpatched vulnerability exploited. Credential leak. |
### Risk Identification by Category
Structured prompts to systematically find risks. Work through each applicable category.
**Operational Risks:**
- What happens if person X is unavailable for 4 weeks?
- Which systems have no redundancy?
- Which manual processes have the highest error rate?
- What vendor dependencies have no fallback?
- What happens during peak load?
**Financial Risks:**
- What costs are variable and could spike?
- What revenue assumptions could be wrong?
- What happens if a key customer churns?
- Where are we locked into pricing that could change?
- What is the cost of a 3-month delay?
**Compliance Risks:**
- What regulations apply and when was last audit?
- Which controls are documented but not tested?
- What data crosses jurisdictional boundaries?
- What contractual SLAs are we close to breaching?
- When do certifications expire?
**Security Risks:**
- What third parties have access to our data?
- When was the last penetration test?
- What happens if credentials leak?
- Which systems lack audit logging?
- What is the mean time to detect an intrusion?
---
## Risk Register Template
### Individual Risk Entry
```
### Risk: [ID] — [Short Title]
**Category**: [Operational | Financial | Compliance | Strategic | Reputational | Security]
**Description**: [What could happen — specific, not generic]
**Trigger**: [What event or condition would cause this risk to materialize]
**Probability**: [Almost Certain | Likely | Possible | Unlikely | Rare] — [Evidence/rationale]
**Impact**: [Severe | Major | Moderate | Minor | Negligible] — [Quantified: $X, Y hours downtime, Z users affected]
**Risk Level**: [Critical | High | Medium | Low] — derived from matrix
**Owner**: [Named person, not "the team"]
**Status**: [Open | Mitigating | Mitigated | Accepted | Closed]
**Mitigation Plan**:
1. [Specific action] — [Owner] — [Due date]
2. [Specific action] — [Owner] — [Due date]
**Residual Risk After Mitigation**: [Level] — [Why this level is acceptable]
**Review Date**: [When to reassess this risk]
```
### Summary Risk Register Table
```
| ID | Risk | Category | Prob | Impact | Level | Owner | Status | Mitigation | Residual |
|----|------|----------|------|--------|-------|-------|--------|------------|----------|
| R-001 | [Title] | Operational | Likely | Major | High | [Name] | Open | [Summary] | Medium |
| R-002 | [Title] | Financial | Possible | Moderate | Medium | [Name] | Mitigating | [Summary] | Low |
| R-003 | [Title] | Security | Unlikely | Severe | High | [Name] | Open | [Summary] | Medium |
```
---
## Mitigation Planning
### Mitigation Strategy Types
| Strategy | When to Use | Example |
|----------|------------|---------|
| **Avoid** | Eliminate the risk by not doing the activity | Cancel the vendor integration that introduces compliance exposure |
| **Reduce** | Lower probability or impact | Add redundancy to eliminate single point of failure |
| **Transfer** | Shift risk to another party | Insurance, SLA penalties, contractual indemnification |
| **Accept** | Consciously decide to carry the risk | Low-impact risk where mitigation cost exceeds potential loss |
| **Share** | Distribute risk across parties | Joint venture, consortium, shared infrastructure |
### Mitigation Quality Test
A mitigation is specific and actionable or it is not a mitigation.
| Bad Mitigation | Good Mitigation | Why |
|---------------|----------------|-----|
| "Monitor the situation" | "Set up PagerDuty alert for error rate >5%. Assign to on-call. Review weekly in ops standup." | Observable. Measurable. Owned. |
| "Improve documentation" | "Write runbook for database failover by 2026-06-01. Assign to J. Smith. Validate via tabletop exercise." | Specific deliverable. Due date. Validation method. |
| "Train the team" | "Schedule 2-hour hands-on session for all on-call engineers by Q3. Test with simulation. Track completion." | Specific format. Audience. Deadline. Verification. |
| "Reduce dependency" | "Implement read replica by 2026-07-15 to eliminate single DB dependency. Test failover monthly." | Concrete action. Timeline. Ongoing validation. |
| "Be more careful" | NOT a mitigation | Behavior-based mitigations do not survive 3am incidents |
### Mitigation Tracking
For each mitigation:
| Field | Required | Description |
|-------|----------|-------------|
| Action | Yes | Specific deliverable or change |
| Owner | Yes | Named person (not team or role) |
| Due date | Yes | Calendar date, not "next quarter" |
| Status | Yes | Not Started / In Progress / Complete / Overdue |
| Verification | Yes | How to confirm the mitigation is working |
| Effectiveness review | Yes | Date to assess whether the mitigation actually reduced risk |
---
## Residual Risk Tracking
### What Is Residual Risk?
Residual risk = risk that remains after mitigations are applied. Every mitigation reduces risk; none eliminates it entirely.
```
Initial Risk Level: High (Likely x Major)
Mitigation: Add read replica; implement automated failover
Residual Risk Level: Medium (Unlikely x Major)
Rationale: Failover eliminates single-point-of-failure. Impact unchanged
because any DB outage during failover still causes brief disruption.
Acceptance: Medium is within risk appetite. Reviewed by [Name] on [Date].
```
### Residual Risk Decision Framework
| Residual Level | Action Required |
|---------------|----------------|
| Critical | Unacceptable. Additional mitigations required. Escalate to leadership. |
| High | Requires explicit acceptance by risk owner's manager. Document rationale. |
| Medium | Acceptable with monitoring. Define review cadence (monthly or quarterly). |
| Low | Acceptable. Review annually or when conditions change. |
### Risk Appetite Statement
Define organizational risk appetite before assessing. Without it, every risk assessment is subjective.
```
## Risk Appetite
**Overall posture**: [Risk-averse | Balanced | Risk-tolerant]
**By category**:
| Category | Appetite | Rationale |
|----------|----------|-----------|
| Operational | Balanced | Accept calculated risks to move faster |
| Financial | Risk-averse | Protect cash flow; avoid commitments >$X without board approval |
| Compliance | Zero tolerance | Regulatory violations are existential |
| Security | Risk-averse | Data breaches have outsized reputational impact |
| Strategic | Risk-tolerant | Market position requires calculated bets |
```
---
## Compliance-Adjacent Risk Management
### Mapping Risks to Controls
When risks overlap with compliance requirements, map them to the applicable control framework.
| Risk | Compliance Framework | Control | Evidence Required |
|------|---------------------|---------|-------------------|
| Unauthorized data access | SOC 2 CC6.1 | Role-based access control | Access review logs, quarterly access recertification |
| Data loss | ISO 27001 A.12.3 | Backup and recovery | Backup logs, restore test results |
| Unpatched vulnerability | PCI DSS 6.2 | Patch management | Scan reports, patch deployment records |
| Insider threat | SOC 2 CC6.3 | Separation of duties | Role matrix, approval logs |
### Audit Evidence for Risk Management
Maintain evidence that risk management is active, not just documented.
| Evidence Type | What It Proves | Collection Frequency |
|--------------|----------------|---------------------|
| Risk register with dates | Risks are tracked over time | Updated per event + quarterly review |
| Mitigation status updates | Mitigations are progressing | Monthly |
| Residual risk reviews | Risk acceptance is deliberate | Quarterly |
| Incident-to-risk mapping | Incidents trigger risk reassessment | Per incident |
| Risk appetite review | Appetite is current, not stale | Annually |
---
## Risk Assessment Failure Modes
| Failure Mode | Symptom | Fix |
|-------------|---------|-----|
| Everything is High | No prioritization possible | Force-rank: at most 20% of risks can be High/Critical. The rest need differentiation. |
| Optimism bias | Probability consistently underestimated | For each "Low" probability: name the evidence. "Has not happened" requires "and here is why it cannot." |
| Impact without numbers | "High impact" with no quantification | Require at least one measurable dimension: dollars, hours, users, SLA. |
| Orphan mitigations | Mitigations with no owner or due date | No mitigation without owner + date. Ownerless mitigations do not exist. |
| Risk register as artifact | Created once, never updated | Schedule quarterly reviews. Trigger reviews on incidents. Mark register with last-reviewed date. |
| Treating issues as risks | Conflating current problems with future risks | Issues are happening now (action items). Risks might happen (mitigation plans). Separate them. |
| Copy-paste risks | Generic risks not tailored to context | Every risk must reference the specific project, system, or decision. "Data breach" is not a risk. "Unauthorized access to customer PII via compromised vendor API key" is. |
| Mitigation theater | "Monitor the situation" as a mitigation | Mitigations must be verifiable. If you cannot test whether the mitigation is working, it is not a mitigation. |
references/operations/runbook-authoring.md
---
title: Runbook Authoring — Step Structure, Verification, Rollback, Escalation
domain: operations
level: 3
skill: operations
---
# Runbook Authoring Reference
> **Scope**: Step-by-step runbook construction for operational procedures. Covers step structure with mandatory components, verification checklists, rollback procedures (whole-runbook and per-step), escalation paths, and troubleshooting tables. Use when authoring new runbooks or reviewing existing ones for completeness.
> **Generated**: 2026-05-05 — Validate commands and paths against your actual infrastructure.
---
## Overview
Runbooks fail for one reason: the author knew things the reader does not. The author writes "run the script" because they know which script, where it lives, what user runs it, and what the output looks like. The reader — a new hire at 3am during an incident — knows none of that.
Every runbook authoring decision should pass one test: **Could someone who has never touched this system follow this step right now, with no Slack messages and no phone calls?**
---
## Step Structure — The Five Required Components
Every step in a runbook MUST have all five components. A step missing any component is incomplete.
### 1. Action (What To Do)
The exact command, UI action, or manual procedure. No ambiguity.
| Bad | Good | Why |
|-----|------|-----|
| "Run the migration script" | `cd /opt/deploy && ./migrate.sh --env=prod --dry-run` as user `deploy` | Specifies path, script, flags, and user |
| "Check the logs" | `journalctl -u api-server --since "5 minutes ago" \| grep ERROR` | Specifies which logs, time window, filter |
| "Restart the service" | `sudo systemctl restart api-server.service` then wait 30s | Specifies exact command and timing |
| "Update the config" | Edit `/etc/app/config.yaml`: set `max_connections: 500` (was 200) | Specifies file, key, new value, old value |
| "Notify the team" | Post to #ops-incidents: "DB failover in progress. ETA 15min." | Specifies channel and message content |
Rules:
- Include the full path. `./script.sh` means nothing without the directory.
- Include the user context. `sudo`, `su - deploy`, or "as your own user" — state it.
- Include flags. `--dry-run` vs `--execute` is the difference between a test and production data loss.
- Include wait times. "Wait 30 seconds" is a step component, not implied.
### 2. Expected Result (What Should Happen)
What the operator should see if the step succeeded. Concrete, observable output.
| Bad | Good |
|-----|------|
| "It should work" | Output: `Migration complete. 47 rows updated. 0 errors.` |
| "Service is healthy" | `curl -s localhost:8080/health` returns `{"status":"ok","uptime":">0"}` within 5s |
| "Logs look normal" | `journalctl -u api-server -n 5` shows `INFO` lines, no `ERROR` or `WARN` |
| "Database is accessible" | `psql -h db-primary -U app -c 'SELECT 1'` returns `1` within 2s |
Rules:
- Show exact output when possible. Operators compare what they see against what the runbook says.
- Include timing. "Returns within 5s" catches hangs that "returns ok" misses.
- Include negative checks. "No ERROR lines in the last 60 seconds" catches delayed failures.
### 3. Failure Handling (What To Do When It Goes Wrong)
What the operator does if the expected result does not appear. This is NOT optional.
| Failure | Action |
|---------|--------|
| Command returns non-zero exit code | Check stderr. If "connection refused," verify the service is running: `systemctl status api-server` |
| Output shows unexpected row count | STOP. Do not proceed. This indicates data inconsistency. Jump to Escalation. |
| Timeout (no response in 30s) | Retry once. If second attempt also times out, check network: `ping db-primary` and `telnet db-primary 5432` |
| Permission denied | Verify you are running as the correct user. Check: `whoami` should return `deploy` |
| Partial success | Note which items succeeded in the runbook history. Proceed to the next step only if documented as safe to continue with partial results. |
Rules:
- Differentiate between retryable and non-retryable failures.
- State when to STOP. Not every failure should be worked around.
- State when to escalate. "If you've spent more than 10 minutes on this step, escalate."
### 4. Verification (How To Confirm Success)
A separate check — distinct from the expected result — that confirms the step achieved its purpose.
| Step Purpose | Verification |
|-------------|-------------|
| Deploy new version | `curl -s /version` returns the new build hash |
| Database migration | `psql -c "SELECT COUNT(*) FROM schema_migrations WHERE version='20240515'"` returns 1 |
| Config change | `grep max_connections /etc/app/config.yaml` shows `500` AND `systemctl show api-server --property=ActiveState` shows `active` |
| Cache flush | Request a known cached resource; response header `X-Cache: MISS` confirms fresh fetch |
Rules:
- Verification checks the outcome, not just the action. "The command ran" is not verification. "The system now behaves correctly" is.
- Include both positive and negative verification where applicable. Service is up AND old version is not running.
### 5. Per-Step Rollback (How To Undo This Step)
How to reverse this specific step if later steps fail and the whole runbook needs to be unwound.
| Step | Rollback |
|------|----------|
| Deploy new version | `kubectl rollout undo deployment/api -n production` |
| Database migration | `cd /opt/deploy && ./migrate.sh --env=prod --direction=down --version=20240515` |
| Config change | Restore from backup: `cp /etc/app/config.yaml.bak /etc/app/config.yaml && systemctl restart api-server` |
| DNS change | Revert CNAME in Route 53 to previous value. TTL is 300s; allow 5min for propagation. |
Rules:
- Not all steps are reversible. State "NOT REVERSIBLE — see Rollback section for compensating action" when applicable.
- Include timing. How long does the rollback take to propagate?
- Rollback should be tested. An untested rollback is a hope, not a plan.
---
## Verification Checklists
### Pre-Execution Checklist
Run before starting any runbook. If any item fails, STOP.
```
## Pre-Execution Checklist
- [ ] Correct environment confirmed (production/staging/dev)
- [ ] Required access verified (SSH, console, database, admin panels)
- [ ] Required tools available (kubectl, psql, aws, etc.)
- [ ] Maintenance window active (if required)
- [ ] Stakeholders notified of upcoming work
- [ ] Rollback procedure reviewed and understood
- [ ] Monitoring dashboards open (list specific URLs)
- [ ] Communication channel open (#ops-incidents or equivalent)
- [ ] Previous run's history reviewed for known issues
- [ ] Backup/snapshot taken if applicable
```
### Post-Execution Checklist
Run after completing all steps. Confirms the runbook achieved its purpose.
```
## Post-Execution Checklist
- [ ] All verification steps passed
- [ ] Monitoring shows normal behavior for 15+ minutes post-change
- [ ] No new alerts triggered
- [ ] Stakeholders notified of completion
- [ ] Runbook history updated with date, operator, and observations
- [ ] Any deviations from the runbook documented
- [ ] Maintenance window closed (if applicable)
- [ ] Temporary access/permissions revoked
```
### Review Checklist (For Runbook Authors)
Use when writing or reviewing a runbook for completeness.
```
## Runbook Review Checklist
- [ ] Every step has all 5 components (action, expected, failure, verify, rollback)
- [ ] No step requires knowledge not present in the document
- [ ] Prerequisites list all required access, tools, and prior state
- [ ] Rollback section covers full unwinding, not just last step
- [ ] Escalation paths have names, contact methods, and trigger criteria
- [ ] Troubleshooting table covers failure modes from each step
- [ ] Someone unfamiliar with the system could follow this at 3am
- [ ] Commands use absolute paths, explicit users, and complete flags
- [ ] Expected results include concrete output, not "it should work"
- [ ] Timing constraints stated (wait times, propagation delays, TTLs)
```
---
## Rollback Procedures
### Per-Step vs. Whole-Runbook Rollback
Two levels of rollback. Both are required.
| Level | When Used | Structure |
|-------|-----------|-----------|
| Per-step | Later step fails; need to undo this step as part of unwinding | Component 5 of each step |
| Whole-runbook | Catastrophic failure or post-completion problems; need to return to pre-runbook state | Dedicated section at end of runbook |
### Whole-Runbook Rollback Template
```
## Rollback Procedure
### When to Roll Back
- [Specific trigger: metric exceeds threshold, error rate above X%, user reports Y]
- [Time limit: "If not resolved within 30 minutes of completion, roll back"]
### Pre-Rollback
- [ ] Confirm rollback is the right action (not just a transient issue)
- [ ] Notify stakeholders: "Rolling back [change]. Estimated time: [X minutes]"
- [ ] Open monitoring dashboards
### Rollback Steps (in reverse order)
1. [Undo Step N]: [exact command]
- Verify: [check]
2. [Undo Step N-1]: [exact command]
- Verify: [check]
3. [Continue in reverse...]
### Post-Rollback Verification
- [ ] System returned to pre-change state
- [ ] All services healthy
- [ ] No data loss or corruption
- [ ] Monitoring confirms normal behavior for 15+ minutes
### Post-Rollback Actions
- [ ] Notify stakeholders of completed rollback
- [ ] Document what went wrong in the runbook history
- [ ] Create incident ticket if applicable
- [ ] Schedule post-mortem for the failed change
```
### Non-Reversible Steps
Some steps cannot be undone. Identify these before execution.
| Non-Reversible Action | Compensating Action |
|----------------------|---------------------|
| Email sent to customers | Send correction/update email |
| Data deleted from production | Restore from backup (document backup location and restore procedure) |
| Third-party API call (payment, notification) | Manual reversal through vendor portal or support ticket |
| Published DNS change (cached globally) | Revert record; wait for TTL expiry. Document TTL value. |
| Schema migration that drops columns | Restore from pre-migration backup. Document backup timing. |
Rules:
- Non-reversible steps must be called out prominently in the runbook
- Require explicit confirmation before executing non-reversible steps
- Compensating actions must be documented even if they are manual
---
## Escalation Paths
### Escalation Trigger Criteria
Define when to stop working independently and escalate. Time-box troubleshooting.
| Trigger | Action |
|---------|--------|
| Step fails after 2 retry attempts | Escalate to on-call engineer |
| Troubleshooting exceeds 10 minutes with no progress | Escalate to on-call engineer |
| Data inconsistency detected | STOP immediately. Escalate to data owner + on-call |
| Customer-facing impact confirmed | Escalate to incident commander + comms |
| Uncertainty about whether to proceed | Escalate. Asking is always better than guessing. |
### Escalation Contact Template
```
## Escalation Contacts
| Situation | Primary Contact | Method | Backup Contact | Method |
|-----------|----------------|--------|----------------|--------|
| System outage | [Name] | PagerDuty / [phone] | [Name] | Slack DM / [phone] |
| Data issue | [Name] | Slack #data-oncall | [Name] | [phone] |
| Security incident | [Name] | PagerDuty (P1) | Security team | security@company.com |
| Customer impact | [Name] | Slack #incidents | [Name] | [phone] |
| Unknown/general | On-call engineer | PagerDuty | Engineering manager | [phone] |
```
### Escalation Message Template
When escalating, include all of these. Missing context wastes responder time.
```
**What**: [Runbook name] failed at Step [N]
**When**: [Timestamp] [Timezone]
**Where**: [Environment: prod/staging] [Region/cluster]
**Symptom**: [What you observed]
**Expected**: [What should have happened]
**Actions taken**: [What you tried, including retry attempts]
**Current state**: [Is the system degraded? Is the change partially applied?]
**Urgency**: [Customer impact? Data at risk? Time-sensitive?]
```
---
## Troubleshooting Tables
### Structure
Every runbook should include a troubleshooting table mapping symptoms to causes and fixes.
```
## Troubleshooting
| Symptom | Likely Cause | Diagnostic Command | Fix |
|---------|-------------|-------------------|-----|
| Connection refused | Service not running | `systemctl status api-server` | `systemctl start api-server` |
| Permission denied | Wrong user context | `whoami` | `su - deploy` or `sudo -u deploy` |
| Timeout after 30s | Network/firewall issue | `telnet db-primary 5432` | Check security groups / iptables |
| Unexpected row count | Prior failed migration | `SELECT * FROM schema_migrations ORDER BY version DESC LIMIT 5` | Run missing migration or restore from backup |
| Pods in CrashLoopBackOff | Config error or resource limits | `kubectl logs pod/api-xxx -n production --tail=50` | Fix config, then `kubectl rollout restart` |
```
### Common Cross-Runbook Issues
These appear in many runbooks. Include the relevant ones.
| Issue | Diagnostic | Resolution |
|-------|-----------|------------|
| DNS not resolving | `dig +short hostname` | Check DNS records. Allow for TTL propagation. |
| SSL certificate expired | `openssl s_client -connect host:443 2>/dev/null \| openssl x509 -noout -dates` | Renew certificate. Check auto-renewal config. |
| Disk full | `df -h /path` | Identify and remove old logs/artifacts. Expand volume if recurring. |
| OOM killed | `dmesg \| grep -i oom` | Increase memory limits or fix memory leak. |
| Clock skew | `date` vs NTP source | `sudo systemctl restart ntp` or `chronyc makestep` |
---
## Runbook Metadata and Lifecycle
### Required Metadata
Every runbook header must include:
```
## Runbook: [Task Name]
**Owner:** [Team/Person]
**Frequency:** [Daily/Weekly/Monthly/As Needed/Event-Driven]
**Last Updated:** [YYYY-MM-DD]
**Last Run:** [YYYY-MM-DD]
**Last Reviewed:** [YYYY-MM-DD]
**Review Cadence:** [Quarterly]
**Environment:** [Production/Staging/Both]
```
### Run History
Append to the history table after every execution.
```
## Run History
| Date | Operator | Duration | Outcome | Notes |
|------|----------|----------|---------|-------|
| 2026-05-01 | jsmith | 25min | Success | No issues |
| 2026-04-15 | mdoe | 40min | Success | Step 3 slow; network latency |
| 2026-04-01 | jsmith | 55min | Rolled back | Step 5 failed. See INC-1234. |
```
### Review Triggers
Review the runbook when any of these occur:
| Trigger | Action |
|---------|--------|
| Infrastructure change (new cluster, DB migration, cloud region) | Validate all commands and paths |
| Incident involving this procedure | Update with lessons learned |
| 90 days since last review | Full review per the Review Checklist |
| New team member runs it for the first time | Shadow run with review |
| Dependency change (upgraded tool, API version, OS) | Validate commands and expected output |
---
## Runbook Authoring Failure Modes
| Failure Mode | Example | Fix |
|-------------|---------|-----|
| The "just" step | "Just restart the service" | Specify exactly which service, how, and verify it came back. |
| Assumed knowledge | "SSH to the usual box" | Name the host, document the SSH command, specify which key. |
| Missing failure paths | Steps only describe success | Add failure handling and when-to-escalate for every step. |
| Stale commands | Scripts that moved or flags that changed | Date-stamp the runbook. Review quarterly. Run in staging first. |
| Copy-paste from chat | Slack thread pasted as-is | Restructure into the 5-component format. Validate every command. |
| Rollback as afterthought | "If something goes wrong, undo the changes" | Specify exact undo commands in reverse order with verification. |
| Single-author syndrome | Only the author can follow it | Have someone else run it. Fix where they get stuck. |
| No timing information | "Wait for it to finish" | "Wait approximately 60-90 seconds. Timeout after 3 minutes." |
references/operations/vendor-management.md
---
title: Vendor Management — Scorecard, Due Diligence, Contract Triggers, Performance Monitoring
domain: operations
level: 3
skill: operations
---
# Vendor Management Reference
> **Scope**: Operational vendor management lifecycle — evaluation scorecards, due diligence checklists, contract review triggers, and ongoing performance monitoring. Covers new vendor evaluation, renewal decisions, and vendor comparison. For strategic build-vs-buy decisions, see `skills/business/business-ops/references/vendor-evaluation.md`.
> **Generated**: 2026-05-05 — Vendor landscapes shift. Reassess scorecards annually and contracts at each renewal.
---
## Overview
Vendor management fails in three places. At selection: comparing feature lists instead of evaluating fitness for your specific context. At contracting: not reading the termination clause, the auto-renewal window, or the price escalation formula. At operation: not measuring performance until renewal, when it is too late to negotiate from strength.
This reference structures the vendor lifecycle so each phase produces artifacts that feed the next.
---
## Vendor Evaluation Scorecard
### Scoring Framework
Score each dimension 1-10 independently before discussing with stakeholders. Consensus-first scoring anchors to the loudest voice.
| Dimension | Weight | Score (1-10) | Scoring Guide |
|-----------|--------|-------------|---------------|
| Functional fit | 5x | ___ | Solves the actual problem without significant workarounds. 10 = drop-in fit. 1 = requires custom development. |
| Total cost of ownership | 4x | ___ | All-in cost over 3 years including hidden costs. 10 = predictable, within budget. 1 = unpredictable, escalating. |
| Integration complexity | 4x | ___ | Effort to connect to existing stack. 10 = <1 day. 1 = months of custom work. |
| Support quality | 3x | ___ | Response time, escalation path, resolution quality. 10 = dedicated support, <1h response. 1 = community forums only. |
| Security / compliance | 3x | ___ | Certifications match requirements. 10 = exceeds all requirements. 1 = no certifications, no audit reports. |
| Data portability | 3x | ___ | Can you export data in a usable format? 10 = full export, standard formats, API access. 1 = no export, proprietary format. |
| Company stability | 2x | ___ | Financial health, customer retention, market position. 10 = profitable, growing, established. 1 = pre-revenue, single investor. |
| Contract flexibility | 2x | ___ | Termination, pricing, scaling terms. 10 = month-to-month, no lock-in. 1 = multi-year, auto-renew, exit penalties. |
| **Weighted Total** | **26x** | **___/260** | |
**Interpretation:**
- 80%+ (208+): Strong fit. Proceed with contract negotiation.
- 60-79% (156-207): Acceptable with known gaps. Document gaps as contract conditions.
- 40-59% (104-155): Marginal. Only proceed if no better alternative exists.
- <40% (<104): Do not proceed. Gaps are structural, not negotiable.
### Comparison Matrix
When evaluating 2+ vendors side-by-side:
| Criterion | Vendor A | Vendor B | Vendor C | Winner |
|-----------|----------|----------|----------|--------|
| Functional fit | | | | |
| TCO (Year 3) | $___/yr | $___/yr | $___/yr | |
| Integration effort | ___days | ___days | ___days | |
| Support SLA | ___hr response | ___hr response | ___hr response | |
| Data export | [format] | [format] | [format] | |
| Contract term | ___months | ___months | ___months | |
| **Scorecard total** | ___/260 | ___/260 | ___/260 | |
---
## Total Cost of Ownership
### TCO Components
License price is not TCO. Calculate all-in cost.
| Component | Year 1 | Year 2 | Year 3 | Notes |
|-----------|--------|--------|--------|-------|
| License / subscription | $ | $ | $ | Per-seat, flat, usage-based? Price escalation clause? |
| Implementation | $ | — | — | Integration, data migration, configuration |
| Training | $ | $ | $ | Initial training + ongoing for new hires |
| Support / maintenance | $ | $ | $ | Included or add-on? Tier pricing? |
| Internal staff time | $ | $ | $ | Admin, maintenance, troubleshooting hours x loaded cost |
| Customization | $ | $ | $ | Custom development, API work, workflow adaptation |
| Infrastructure | $ | $ | $ | Additional hosting, storage, bandwidth |
| Compliance | $ | $ | $ | Audit costs, additional controls required |
| Exit costs | — | — | $ | Data migration, contract termination, replacement overlap |
| **Total** | **$** | **$** | **$** | |
| **Cumulative** | **$** | **$** | **$** | |
### Hidden Cost Checklist
Costs frequently missed in vendor evaluations:
- [ ] Price escalation at renewal (typical: 5-10% annual, often buried in terms)
- [ ] Overage charges for exceeding usage tiers
- [ ] Premium support tier required for production SLA
- [ ] Additional seats for admin/service accounts
- [ ] SSO/SAML as paid add-on (common in SaaS)
- [ ] API rate limits requiring paid upgrade
- [ ] Data storage beyond included allocation
- [ ] Professional services for initial setup
- [ ] Certification/compliance add-ons
- [ ] Currency fluctuation on non-USD contracts
- [ ] Opportunity cost of internal resources spent on vendor management
- [ ] Cost of vendor downtime to your operations (SLA credits rarely cover actual losses)
---
## Due Diligence Checklist
### Financial Stability
| Check | Source | Red Flag |
|-------|--------|----------|
| Funding / revenue status | Crunchbase, press releases, SEC filings | Runway <18 months. Revenue declining. Recent layoffs >20%. |
| Customer base size and diversity | Case studies, G2/Capterra, reference calls | <50 customers. Single-industry concentration. |
| Key customer concentration | Ask directly | >30% revenue from one customer (they leave, vendor is at risk) |
| Leadership stability | LinkedIn, press | C-suite turnover in last 12 months |
| Acquisition signals | Industry news, investor activity | Actively seeking acquisition (roadmap becomes irrelevant) |
### Security Posture
| Check | Source | Red Flag |
|-------|--------|----------|
| SOC 2 Type II report | Request directly | No report. Type I only. Report >12 months old. |
| Penetration test results | Request directly (summary) | Never conducted. Findings not remediated. Won't share summary. |
| Incident history | Public disclosures, news | Breach in last 24 months with poor response. |
| Data encryption | Technical documentation | No encryption at rest. No TLS in transit. Shared keys. |
| Access controls | Technical documentation, SOC 2 | No RBAC. No MFA for admin. Shared credentials. |
| Subprocessor list | DPA / privacy documentation | Unclear data handling chain. Subprocessors in restricted jurisdictions. |
### Contract and Legal
| Check | What to Look For | Red Flag |
|-------|-----------------|----------|
| Auto-renewal clause | Notice period, renewal terms | Auto-renew with <30-day cancellation window |
| Termination rights | For cause, for convenience, notice period | No termination for convenience. Excessive notice period. |
| Price escalation | Annual increase caps, CPI adjustments | Uncapped increases. No ceiling on usage-based pricing. |
| Data ownership | Who owns data created in the platform | Vendor claims rights to aggregated/derived data |
| Data portability | Export formats, migration support, timeline | No export. Proprietary format only. Charges for export. |
| SLA and remedies | Uptime guarantee, measurement, credits | SLA credits capped at <monthly fee. Measured monthly (hides daily outages). |
| Liability caps | Limitation of liability | Liability capped at 1x annual fees (inadequate for data loss) |
| Indemnification | IP infringement, data breach | No indemnification for breaches caused by vendor |
| Insurance | Cyber liability, E&O | No cyber liability insurance |
### Reference Checks
Do not accept only vendor-provided references. They are curated.
| Reference Type | Questions to Ask |
|---------------|-----------------|
| Vendor-provided (3 minimum) | What surprised you? What took longer than expected? Would you choose them again? |
| Self-sourced (find on G2, LinkedIn) | Why did you choose them? What are the real downsides? How is support when things break? |
| Churned customers (ask vendor who left) | Why did you leave? What was the exit process like? What would have kept you? |
| Similar-scale organizations | How does it perform at our scale? What breaks as you grow? |
---
## Contract Review Triggers
Events that should trigger a contract review outside the normal renewal cycle.
| Trigger | Action |
|---------|--------|
| Vendor acquired by another company | Review continuity, roadmap, pricing commitments. Prepare exit plan. |
| Major security incident at vendor | Assess exposure. Review incident response. Consider SLA implications. |
| Price increase >5% at renewal | Benchmark against alternatives. Negotiate or initiate RFP. |
| SLA breach (3+ in rolling 12 months) | Document pattern. Negotiate improved terms or credits. Evaluate alternatives. |
| Vendor layoffs >15% of workforce | Assess impact on support, roadmap, stability. Prepare contingency. |
| Your usage model changes significantly | Review pricing fit. May need different tier or different vendor. |
| Regulatory change affecting data handling | Review vendor compliance posture. Update DPA if needed. |
| Vendor changes terms unilaterally | Review impact. Exercise termination rights if material. |
| Key contact/champion at vendor leaves | Rebuild relationship. Assess organizational support depth. |
| Your team reports declining satisfaction | Survey users. Document specific issues. Use as negotiation leverage. |
---
## Performance Monitoring
### Vendor Performance Scorecard (Ongoing)
Review quarterly. Score 1-5 for each dimension.
| Dimension | Q1 | Q2 | Q3 | Q4 | Trend | Notes |
|-----------|----|----|----|----|-------|-------|
| Uptime / availability | | | | | | Measure against SLA |
| Support responsiveness | | | | | | Time to first response, time to resolution |
| Support quality | | | | | | Resolution rate, escalation frequency |
| Feature delivery | | | | | | Roadmap items delivered on time |
| Communication quality | | | | | | Proactive notification, transparency |
| Billing accuracy | | | | | | Invoice errors, unexpected charges |
| **Average** | | | | | | |
### SLA Tracking
```
## SLA Performance: [Vendor Name]
| SLA Metric | Committed | Actual | Status | Credit Due |
|-----------|-----------|--------|--------|------------|
| Uptime | 99.9% | ___% | Pass/Fail | $___ |
| Response time (P1) | 1 hour | ___hr | Pass/Fail | $___ |
| Response time (P2) | 4 hours | ___hr | Pass/Fail | $___ |
| Resolution time (P1) | 4 hours | ___hr | Pass/Fail | $___ |
| Resolution time (P2) | 24 hours | ___hr | Pass/Fail | $___ |
```
### Performance Review Meeting Agenda
Quarterly review with vendor account team:
1. **SLA review**: Performance against commitments. Credits owed.
2. **Support review**: Ticket volume, resolution quality, escalation patterns.
3. **Roadmap update**: Features delivered, upcoming features, timeline changes.
4. **Usage review**: Current vs. projected. Tier optimization opportunities.
5. **Issues log**: Open items from previous review. New issues.
6. **Relationship health**: Communication quality, responsiveness, partnership.
7. **Contract items**: Upcoming renewal, pricing discussion, term changes.
---
## Vendor Risk Assessment
### Concentration Risk
| Question | Threshold | Action if Exceeded |
|----------|-----------|-------------------|
| What % of a critical process depends on this vendor? | >80% | Document alternative. Create exit plan. |
| Can you operate for 48 hours without this vendor? | No | Create business continuity plan for vendor outage. |
| How long to migrate to an alternative? | >6 months | Begin evaluating alternatives now. Do not wait for crisis. |
| What data is locked in this vendor's platform? | Any critical data | Verify export capability. Test export quarterly. |
### Exit Planning
Every vendor relationship should have an exit plan. Not because you plan to leave, but because you might need to.
```
## Vendor Exit Plan: [Vendor Name]
### Trigger Criteria
- [What would cause us to exit: acquisition, breach, cost, performance]
### Data Migration
- **Data to export**: [What data, what format, what volume]
- **Export method**: [API, bulk export, manual, vendor-assisted]
- **Export tested**: [Yes/No — when last tested]
- **Migration target**: [Where data goes next]
### Replacement Options
| Option | Readiness | Migration Effort | Cost Delta |
|--------|-----------|-----------------|------------|
| [Vendor B] | Evaluated / Not evaluated | [weeks] | [+/-$X/yr] |
| [Build in-house] | Feasible / Not feasible | [weeks] | [+/-$X/yr] |
| [Manual process] | Temporary only | [immediate] | [+$X/yr in labor] |
### Contract Obligations
- **Notice period**: [X days]
- **Termination fee**: [$X]
- **Data deletion timeline**: [X days post-termination]
### Timeline
- **Best case**: [X weeks from decision to full migration]
- **Worst case**: [X weeks, accounting for data complexity and testing]
```
---
## Vendor Management Failure Modes
| Failure Mode | Symptom | Fix |
|-------------|---------|-----|
| Feature-list comparison | Choosing the vendor with the longest feature list | Score on fitness for your specific use case, not feature count |
| Ignoring exit costs | No exit plan until you need to leave | Build exit plan during evaluation. Test data export before signing. |
| Renewal surprise | Auto-renewed at higher price, nobody noticed | Calendar the renewal window -90 days. Review performance before renewal. |
| Single-threaded relationship | One person manages the vendor; they leave | Document vendor relationship. At least two people attend QBRs. |
| SLA credit neglect | Vendor misses SLA, nobody claims credits | Automate SLA tracking. File claims within contractual window. |
| Demo-driven evaluation | Chose based on the sales demo, not real testing | Require POC with your data, your integrations, your scale |
| Ignoring churned customers | Only spoke to vendor-curated references | Actively seek out customers who left. Their reasons are the real risks. |
| TCO = license price | Did not account for implementation, training, maintenance, exit | Use the full TCO template. Every component, every year. |
references/product-management.md
# Product Management
Umbrella skill for PM workflows: specs, roadmaps, stakeholder comms, research synthesis, competitive analysis, metrics review, sprint planning, and product brainstorming. Each mode loads its own reference files on demand.
---
## Mode Detection
Classify into one mode before proceeding.
| Mode | Signal Phrases | Reference |
|------|---------------|-----------|
| **SPEC** | write spec, PRD, feature requirements, acceptance criteria, user stories | `references/product-management/spec-writing.md` |
| **ROADMAP** | roadmap, prioritize, Now/Next/Later, reprioritize, timeline, OKR alignment | `references/product-management/roadmap-planning.md` |
| **STAKEHOLDER** | stakeholder update, status report, executive brief, launch announcement | (inline templates) |
| **RESEARCH** | synthesize research, interview analysis, user feedback, personas, thematic analysis | `references/product-management/research-synthesis.md` |
| **COMPETITIVE** | competitive brief, competitor analysis, battle card, positioning, win/loss | (inline templates) |
| **METRICS** | metrics review, KPI, funnel analysis, retention, cohort, dashboard, OKR scoring | `references/product-management/metrics-review.md` |
| **SPRINT** | sprint planning, backlog grooming, capacity, sprint goal, carryover | (inline templates) |
| **BRAINSTORM** | brainstorm, explore problem, stress-test idea, thinking partner, assumption testing | (Socratic — see below) |
If the request spans modes, pick the primary mode and note the secondary.
---
## Workflow by Mode
### SPEC Mode
**Load**: `references/product-management/spec-writing.md`, `references/product-management/llm-pm-failure-modes.md`
1. **Understand** — Accept any input: feature name, problem statement, user request, vague idea.
2. **Gather context** — Ask conversationally (not a wall of questions):
- User problem and who experiences it
- Target users / segments
- Success metrics (how will we know it worked?)
- Constraints: technical, timeline, regulatory, dependencies
- Prior art: attempted before? Existing solutions?
3. **Generate PRD** with these sections:
| Section | Content |
|---------|---------|
| Problem Statement | 2-3 sentences. Who, how often, cost of not solving. Grounded in evidence. |
| Goals | 3-5 measurable outcomes. Outcomes not outputs. |
| Non-Goals | 3-5 explicit exclusions with rationale. |
| User Stories | "As a [specific type], I want [capability] so that [benefit]." Group by persona. Include edge cases. |
| Requirements | P0 (must-have), P1 (nice-to-have), P2 (future). Each with acceptance criteria. |
| Success Metrics | Leading (days-weeks) and lagging (weeks-months). Specific targets with measurement method. |
| Open Questions | Tagged by owner (eng, design, legal, data). Blocking vs non-blocking. |
| Timeline | Hard deadlines, dependencies, phasing. |
4. **Review** — Offer iteration, expansion, follow-up artifacts (design brief, ticket breakdown).
**Acceptance criteria format**: Given/When/Then or checklist. Cover happy path, error cases, edge cases. No ambiguous words ("fast", "intuitive") without concrete definitions.
**Scope management**: Write explicit non-goals. Any scope addition requires a scope removal or timeline extension. Separate v1 from v2. Time-box investigations.
### ROADMAP Mode
**Load**: `references/product-management/roadmap-planning.md`, `references/product-management/llm-pm-failure-modes.md`
1. **Current state** — Get existing roadmap (paste, describe, or build from scratch).
2. **Determine operation**:
| Operation | Inputs | Key Actions |
|-----------|--------|-------------|
| Add item | Name, priority, effort, timeframe, owner, dependencies | Suggest placement based on priorities and capacity |
| Update status | Item + new status (not started / in progress / at risk / blocked / completed / cut) | For at-risk/blocked: require blocker + mitigation |
| Reprioritize | What changed (strategy shift, new data, resource change) | Apply framework (RICE, ICE, MoSCoW, Value/Effort). Show before/after. |
| Move timeline | Why (scope change, dependency slip, resource constraint) | Identify downstream impacts. Flag hard-deadline conflicts. |
| Create new | Timeframe, format preference, initiative list | Use Now/Next/Later unless user specifies otherwise |
3. **Generate** — Status overview, items grouped by timeframe/theme, risks/dependencies, change summary.
4. **Follow up** — Offer audience-specific formatting, change communication drafts.
**Capacity rule**: When adding to roadmap, always ask "What comes off?" Roadmaps are zero-sum against capacity.
### STAKEHOLDER Mode
1. **Update type**: Weekly / Monthly / Launch / Ad-hoc
2. **Audience detection**:
| Audience | Frame | Length |
|----------|-------|--------|
| Executives | Outcome-focused, G/Y/R status, strategic alignment | < 300 words |
| Engineering | Technical detail, links to PRs/tickets, decisions needed with options | As needed |
| Cross-functional | Impact on their team, asks with deadlines, input opportunities | Medium |
| Customers | Benefits-focused, no jargon, honest timelines | Short |
| Board | Metrics-driven, risk-focused, strategic | Very concise |
3. **Generate** using audience-appropriate template.
4. **Risk communication** — Use ROAM framework (Resolved, Owned, Accepted, Mitigated). Every risk comes with: clear statement, quantified impact, likelihood with evidence, mitigation plan, specific ask.
**Executive update rule**: Lead with conclusion, not journey. "We shipped X and it moved Y" not "we had 14 standups." Status color reflects reality, not optimism.
**Gate**: Update draft exists. Audience-appropriate framing verified (no engineering jargon in exec updates, no hand-waving in engineering updates). Every risk has a ROAM classification.
### RESEARCH Mode
**Load**: `references/product-management/research-synthesis.md`, `references/product-management/llm-pm-failure-modes.md`
1. **Gather inputs** — Accept any combination: pasted text, uploaded files, described findings.
2. **Process** — For each source extract: observations, verbatim quotes, behaviors (vs stated preferences), pain points, positive signals, context.
3. **Thematic analysis**:
- Familiarize -> Initial coding -> Theme development -> Theme review -> Theme refinement -> Report
- Affinity mapping: one observation per note, let clusters emerge, split large clusters.
- Triangulation: methodological, source, temporal. Findings supported by multiple sources are stronger.
4. **Priority matrix**:
| | High Impact | Low Impact |
|---|---|---|
| **High Frequency** | Top priority | Quality-of-life |
| **Low Frequency** | Segment-specific | Note and deprioritize |
5. **Generate synthesis**: Research overview, 5-8 key findings (with evidence, frequency, impact, confidence), user segments/personas, opportunity areas, actionable recommendations, open questions.
**Critical rule**: Distinguish behaviors from stated preferences. Behavioral data always outweighs what users say they want. Quote attribution uses participant type ("Enterprise admin, 200-person team"), never names.
### COMPETITIVE Mode
1. **Scope** — Which competitor(s)? Full comparison or specific area? What decision does this inform?
2. **Research** — Product pages, pricing, changelogs, customer reviews (G2, Capterra), job postings (strategic signals), community discussions.
3. **Generate brief**:
- Competitor overview (company, positioning, momentum)
- Feature comparison matrix (Strong/Adequate/Weak/Absent ratings)
- Positioning analysis (For [target] who [need], [Product] is a [category] that [benefit])
- Honest strengths and weaknesses
- Opportunities and threats
- Strategic implications: build/accelerate/deprioritize, differentiate vs parity, positioning adjustments
4. **Competitive set levels**: Direct (same problem, same way), Indirect (same problem, different way), Adjacent (could expand into your space), Substitute (entirely different approach including "do nothing").
**Honesty rule**: Dismissing competitors makes analysis useless. Rate based on real product experience and customer feedback, not marketing claims. Be honest about where competitors lead.
**Gate**: Competitive brief exists with feature matrix, positioning analysis, and strategic implications. At least one honest "they lead here" finding present.
### METRICS Mode
**Load**: `references/product-management/metrics-review.md`, `references/product-management/llm-pm-failure-modes.md`
1. **Gather data** — Get metrics with comparison data (previous period, targets). Ask about known events (launches, incidents, seasonality).
2. **Organize** — Use metrics hierarchy:
| Level | Purpose | Examples |
|-------|---------|---------|
| North Star | Core value delivered | WAU completing core workflow |
| L1 (Health) | Lifecycle stages | Acquisition, Activation, Engagement, Retention, Monetization, Satisfaction |
| L2 (Diagnostic) | Drill-down | Funnel steps, feature adoption, segment breakdowns, performance |
3. **Analyze** — For each metric: current value, trend, vs target, rate of change, anomalies. Identify correlations, leading indicators, segment-driven aggregate trends.
4. **Generate review**: Summary (2-3 sentences), scorecard table, trend analysis, bright spots, areas of concern, recommended actions, caveats.
5. **Goal-setting support**: OKRs (2-3 objectives, 2-4 KRs each, outcomes not outputs, 70% completion = target for stretch). Target-setting: baseline -> benchmark -> trajectory -> effort -> confidence.
**Context rule**: Absolute numbers without comparison are useless. Always show vs previous period, vs target, vs benchmark. Small fluctuations are noise — focus on meaningful changes.
### SPRINT Mode
1. **Gather**: Team roster + availability, sprint length, prioritized backlog, carryover, dependencies.
2. **Capacity calculation**: Available days minus overhead (meetings, on-call, PTO). Rule of thumb: 60-70% of time on planned work.
3. **Allocation**: 70% planned features, 20% tech health, 10% unplanned buffer.
4. **Generate sprint plan**:
- Sprint goal (one sentence)
- Capacity table (person, available days, allocation, notes)
- Sprint backlog (P0 must-ship, P1 should-ship, P2 stretch)
- Planned capacity vs sprint load (target 70-80%)
- Risks with impact and mitigation
- Definition of done
- Key dates (start, mid-sprint check, demo, retro)
5. **Carry over honestly** — If something did not ship, understand why before re-committing.
**Gate**: Sprint plan exists with goal, capacity table, prioritized backlog, and load vs capacity check showing 70-80% target.
### BRAINSTORM Mode (Socratic)
This mode is fundamentally different. The PM does not get a deliverable. They get a thinking partner. **Be opinionated. Push back. Bring unexpected angles. Challenge assumptions.**
**Session principles**:
- Apply frameworks to the specific problem at hand
- Generate, evaluate, then discuss before handing over
- Challenge ideas actively — push back with reasons
- Hold divergent exploration open before converging
- Explore multiple directions before committing
**Sub-modes** — Detect which fits and shift as conversation evolves:
| Sub-mode | When | Approach |
|----------|------|----------|
| Problem Exploration | PM has a problem area, not a defined problem | Ask "who has this problem?" and "what are they doing today?" Map the ecosystem. Distinguish symptoms from root causes. |
| Solution Ideation | Problem is well-defined, need options | Generate 5-7 distinct approaches before evaluating. Include one "do the opposite" and one "remove something." Resist early convergence. |
| Assumption Testing | PM has a direction, needs stress-testing | List every assumption (stated + unstated). Find the riskiest one. Suggest the cheapest test. Play devil's advocate. |
| Strategy Exploration | Big bets, positioning, direction | Map possible moves. Think in bets (odds, payoff). Consider second-order effects and competitive responses. |
**Session rhythm**: Frame -> Diverge -> Provoke -> Converge -> Capture.
**Ideation techniques**:
- Constraint removal: "What if no technical/budget/political constraints?"
- Analogies: "How does [another industry] solve this?"
- Inversion: "How would we make this worse?" Then reverse.
- Decomposition: Break into subproblems, solve independently, recombine.
- User hat-switching: Power user? New user? Admin? Someone who hates the product?
**Frameworks as thinking tools** (use when they help, not as templates):
| Framework | Structure | Failure Mode |
|-----------|-----------|-------------|
| **HMW** | "How might we [outcome] for [user] without [constraint]?" | Too broad ("improve onboarding") or too narrow ("add tooltip to step 3") |
| **JTBD** | "When [situation], I want to [motivation] so I can [outcome]." | Functional jobs are easy; emotional and social jobs are often more powerful. Ask "what did they fire?" |
| **Opportunity Solution Tree** | Outcome -> Opportunities (from research) -> Solutions (multiple per opportunity) -> Experiments (cheapest test) | Opportunities must trace to evidence, not imagination. One solution per opportunity = not enough exploration. |
| **First Principles** | State assumption -> Break to fundamentals -> Question each -> Rebuild | Use when team is stuck in incrementalism |
| **OODA** | Observe -> Orient -> Decide -> Act -> loop | Most teams get stuck in Orient. OODA says: orient with what you have, act, let next cycle correct. |
| **Reverse Brainstorming** | "How to make this worse?" -> List -> Reverse each | When team is stuck; people are better at identifying wrong than imagining right. |
**Provocation prompts**:
- "What is the strongest argument against this?"
- "Who would hate this and why?"
- "What are we not seeing?"
- "What if the opposite were true?"
- "What is the 10x more ambitious version?"
**Gate**: Session produced at least one challenge the PM hadn't considered. Captured decisions/next-steps documented. Frameworks used as thinking tools, not dumped as checklists.
---
## LLM Failure Modes in PM Work
See `references/product-management/llm-pm-failure-modes.md` for the complete failure mode catalog (vague specs, fabricated research, generic competitive analysis, metrics without context, happy-path-only specs, framework regurgitation, scope creep enablement). Universal failure modes in `skills/shared-patterns/llm-domain-failure-modes-base.md`.
---
## Prioritization Frameworks (Cross-Mode Reference)
Used in SPEC, ROADMAP, and SPRINT modes.
| Framework | Formula / Method | Best For |
|-----------|-----------------|----------|
| **RICE** | (Reach x Impact x Confidence) / Effort | Large backlog, quantitative comparison |
| **ICE** | Impact x Confidence x Ease (1-10 each) | Quick prioritization, early-stage |
| **MoSCoW** | Must / Should / Could / Won't | Scoping a release, forcing prioritization conversations |
| **Value vs Effort** | 2x2 matrix: Quick Wins, Big Bets, Fill-ins, Money Pits | Visual prioritization in team sessions |
**Failure modes**: Using a framework as a rubber stamp for a decision already made. If the RICE score does not match intuition, investigate why — do not just adjust the inputs until it does.
---
## Output Conventions
- Markdown with clear headers. Scannable. Busy stakeholders read headers and bold text.
- Tables for comparisons, scorecards, feature matrices.
- Status labels: **Done**, **On Track**, **At Risk**, **Blocked**, **Not Started**.
- Executive content: < 300 words. Engineering content: as detailed as needed.
- Every recommendation is specific enough to act on. "Improve onboarding" is not actionable. "Add progress indicator to setup flow" is.
references/product-management/llm-pm-failure-modes.md
# LLM Failure Modes in Product Management
Where LLMs systematically fail at PM tasks. Loaded across all modes as a guardrail reference.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers product management-specific failures only.
---
## Why This File Exists
LLMs are fluent generators. Fluency is dangerous in PM work because PM artifacts look correct when they are not. A spec with vague acceptance criteria passes a casual read. Fabricated user quotes sound plausible. A roadmap that is a feature list with dates looks like strategic planning.
This reference catalogs the specific failure modes, their signatures, and the defenses against each.
---
## Failure Mode 1: Vague Specifications
### What Happens
The LLM produces a spec that reads well but contains requirements no engineer can implement against. Requirements use subjective language ("intuitive", "fast", "user-friendly") without measurable definitions. Acceptance criteria are either missing or restate the requirement in different words.
### Signatures
| Signal | Example |
|--------|---------|
| Subjective adjectives | "The interface should be intuitive and responsive" |
| Restated requirements as AC | Requirement: "Support file upload." AC: "Users can upload files." |
| Missing edge cases | Happy path only. No error states, empty states, or boundary conditions. |
| Implementation-free | No consideration of what happens at scale, under load, or with bad input |
| Ambiguous scope | "Support major file formats" — which ones? |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Ban ambiguous words | Maintain a banned-word list: intuitive, fast, user-friendly, responsive, scalable, secure — require concrete definitions for each |
| Given/When/Then enforcement | Every requirement gets at least one Given/When/Then acceptance criterion |
| Edge case prompting | Explicitly prompt for: error states, empty states, boundary conditions, concurrent access, permission failures, offline/degraded |
| Quantity check | Each P0 requirement should have 3-5 acceptance criteria covering happy path + edge cases |
| Engineer review test | "Could an engineer implement this without asking any clarifying questions?" If no, the spec is incomplete. |
---
## Failure Mode 2: Fabricated Research Data
### What Happens
The LLM generates plausible-sounding user quotes, statistics, persona details, or research findings that are not grounded in any data the user provided. Because the LLM is trained on real research examples, fabricated data looks authentic.
### Signatures
| Signal | Example |
|--------|---------|
| Unsourced quotes | "As one user put it, 'I just wish the export worked better'" — when no user said this |
| Precise statistics from nowhere | "67% of users reported frustration with the onboarding flow" — when no survey was conducted |
| Rich persona details without data | "Sarah, 34, marketing manager at a 200-person SaaS company, uses the product 3x daily" — entirely invented |
| Consistent findings | All findings align perfectly with a clean narrative. Real data is messy. |
| Missing uncertainty | No confidence levels, no "we don't know", no contradictions |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Source citation requirement | Every finding must cite a specific source the user provided. No exceptions. |
| Quote verification | Every verbatim quote must trace to user-provided data. If the user did not provide interview transcripts, there are no quotes. |
| Confidence level mandate | Every finding gets High/Medium/Low confidence. If data is thin, say so. |
| Contradiction check | If all findings align too neatly, flag it. Real research has contradictions. |
| Explicit unknowns | Require an "Open Questions" section listing what the data does NOT answer. |
| Never generate synthetic data | The LLM never invents statistics, quotes, or persona attributes. If the data is not there, the finding is not there. |
### The Core Rule
**Every claim in a research synthesis must trace to something the user provided.** If the user gave 5 interview notes, findings come from those 5 interviews. If the user gave survey results, statistics come from that survey. The LLM synthesizes, interprets, and structures — it does not invent.
---
## Failure Mode 3: Generic Competitive Analysis
### What Happens
The LLM produces a competitive brief that reads like a marketing brochure mashup. Feature comparisons are based on publicly available feature lists rather than real product experience. Strengths and weaknesses are generic ("strong brand", "large customer base"). Strategic implications are obvious ("we should differentiate").
### Signatures
| Signal | Example |
|--------|---------|
| Feature list comparison | Checking boxes for "has feature X" without assessing quality or depth |
| Marketing language | Using competitors' own positioning language uncritically |
| No evidence sources | Claims without citations (customer reviews, analyst reports, real usage) |
| Balanced to the point of useless | Every competitor has "strengths and weaknesses" but nothing actionable |
| Missing "so what" | Analysis without strategic implications |
| Competitor dismissal | Downplaying competitor strengths to make our position look better |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Evidence requirement | Every claim about a competitor cites a source: customer review, analyst report, pricing page, job posting |
| Quality rating | Rate capabilities as Strong/Adequate/Weak/Absent, not just present/absent |
| Honest assessment | Rate based on real product experience and customer feedback, not marketing claims. Be honest about where competitors lead. |
| Strategic implications mandate | The brief must end with specific recommendations: build/accelerate/deprioritize, differentiate vs parity |
| Customer perspective | Frame comparison through what buyers evaluate, not internal categories |
| Shelf-life awareness | Note the date. Flag areas that change fast. Competitive analysis gets stale quickly. |
---
## Failure Mode 4: Metrics Without Context
### What Happens
The LLM presents numbers without comparison baselines, statistical awareness, or acknowledgment of uncertainty. A metric is reported as "good" or "bad" without reference to industry benchmarks, previous periods, or targets. Small fluctuations are treated as meaningful trends. Correlation is implied to be causation.
### Signatures
| Signal | Example |
|--------|---------|
| Naked numbers | "DAU is 15,000" — is that good? Bad? Up? Down? |
| False precision | "Retention improved by 0.3%" on a sample of 200 users |
| Missing sample sizes | Percentages without denominators |
| Causation claims | "We shipped feature X and retention improved" without experimental evidence |
| Cherry-picked timeframes | Choosing the period that tells the best story |
| Vanity framing | "We have 100,000 total signups!" — cumulative metric that only goes up |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Comparison mandate | Every metric shows: current value, previous period, target, benchmark |
| Sample size reporting | Always report the n. Flag small samples explicitly. |
| Statistical significance | For A/B tests: require p < 0.05. For metric movements: distinguish signal from noise. |
| Causation guard | Never claim causation from observational data. Use language like "correlated with", "coincided with", "may be related to" |
| Consistent timeframes | Same comparison period across all metrics. No mixing. |
| Rate over cumulative | Use rate metrics (DAU, weekly conversion) over cumulative (total signups ever) |
| Segment check | Always ask: does the aggregate mask segment-specific trends? |
---
## Failure Mode 5: Roadmaps Disconnected from Strategy
### What Happens
The LLM produces a feature list with dates and calls it a roadmap. Items lack strategic rationale. No connection to company goals, OKRs, or user research. The "roadmap" answers "what are we building?" but not "why are we building it?" or "what are we choosing NOT to build?"
### Signatures
| Signal | Example |
|--------|---------|
| Feature list format | Items described as features, not outcomes or opportunities |
| No "why" | Items have descriptions but no strategic justification |
| Missing non-goals | No mention of what was explicitly excluded |
| No capacity awareness | More items than the team can build, with no acknowledgment |
| No prioritization rationale | Items are listed but not ranked with defensible logic |
| No dependency mapping | Cross-team dependencies unmentioned |
### Defenses
| Defense | Implementation |
|---------|---------------|
| OKR linkage | Every item answers: "Which Key Result does this move?" |
| Theme organization | Group items under strategic themes, not just timeframes |
| Non-goals section | Explicitly list what is NOT on the roadmap and why |
| Capacity check | Compare total estimated effort against available capacity. Flag overcommitment. |
| Prioritization framework | Apply RICE, ICE, or MoSCoW with visible scores/rationale |
| Dependency mapping | List all cross-team, external, and sequential dependencies |
| "What comes off?" rule | Adding an item requires naming what is deprioritized or delayed |
---
## Failure Mode 6: Happy-Path-Only Specifications
### What Happens
The LLM describes only the successful path through a feature. Error handling, edge cases, failure modes, empty states, permission boundaries, and concurrent access scenarios are omitted. The spec looks complete because the happy path is well-described, but engineers discover gaps during implementation.
### Signatures
| Signal | Example |
|--------|---------|
| No error states | What happens when the API fails? Network timeout? Invalid input? |
| No empty states | What does the user see before any data exists? |
| No permission model | What can viewers see vs editors vs admins? |
| No concurrent access | What if two users edit the same thing simultaneously? |
| No offline/degraded behavior | What happens with poor connectivity? |
| No limits | What happens at maximum items? Maximum characters? Zero items? |
### Defenses
**Mandatory edge case checklist for every user story**:
- [ ] **Error state**: What happens when the primary action fails?
- [ ] **Empty state**: What appears before any data exists?
- [ ] **Boundary conditions**: Behavior at max/min/zero values
- [ ] **Permission variations**: Different views per role
- [ ] **Concurrent access**: Multiple users acting simultaneously
- [ ] **Undo/recovery**: Can the action be reversed?
- [ ] **Offline/degraded**: Behavior under poor connectivity
- [ ] **Loading state**: What appears during async operations?
- [ ] **Partial failure**: What if part of the operation succeeds and part fails?
---
## Failure Mode 7: Framework Regurgitation
### What Happens
The LLM dumps framework definitions (RICE, MoSCoW, JTBD, HMW, OST) as if listing them is the same as applying them. The PM gets a textbook explanation of RICE scoring instead of actual RICE scores for their specific initiatives. Frameworks are presented as templates to fill in rather than thinking tools to apply.
### Signatures
| Signal | Example |
|--------|---------|
| Framework definition instead of application | "RICE stands for Reach, Impact, Confidence, Effort..." without scoring anything |
| Multiple frameworks without selection | Listing 5 frameworks without recommending which fits this situation |
| Generic examples | "For example, a collaboration tool might..." instead of using the actual product |
| Framework as deliverable | The output IS the framework template, not a populated artifact |
| Framework shopping | Trying each framework until one gives the desired answer |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Apply, do not explain | If using RICE, produce actual scores for actual initiatives. Not a tutorial on RICE. |
| Select one framework | Choose the framework that fits the situation. Justify the choice. Do not list all options. |
| Use specific data | Populate with the user's actual numbers, products, and context. |
| Framework = means, not end | The deliverable is the prioritized list, not the framework. |
| Anti-gaming | If RICE scores do not match intuition, investigate the mismatch. Do not adjust inputs to force the desired ranking. |
---
## Failure Mode 8: Scope Creep Enablement
### What Happens
The LLM accepts every feature request and expands scope without surfacing tradeoffs. When a user says "could we also add X?" the LLM integrates X into the spec without noting the capacity impact, timeline extension, or what would need to be cut. The spec grows in every iteration.
### Signatures
| Signal | Example |
|--------|---------|
| Additive-only iteration | Every revision adds scope, nothing is ever removed |
| No tradeoff surfacing | New features accepted without capacity or timeline impact |
| Missing non-goals | The spec has no "won't do" section |
| v1 = everything | No phasing, no MVP definition, everything is "must have" |
| Stakeholder pleasing | Different stakeholders' wishes all accommodated without prioritization |
### Defenses
| Defense | Implementation |
|---------|---------------|
| "What comes off?" rule | Every scope addition requires naming what is deprioritized |
| Non-goals section | Mandatory in every spec. Reviewed after each scope change. |
| Capacity check | Total effort vs available capacity. Flag overcommitment immediately. |
| P0 challenge | "If everything is P0, nothing is P0." Challenge every must-have. |
| Phase separation | Clear v1 / v2 boundary. v2 is not "later" — it is "explicitly not now." |
| Investigation time-box | "If we cannot resolve X in 2 days, we cut it." |
---
## Cross-Cutting Defense: The Verification Habit
Across all failure modes, the root cause is the same: **the LLM generates plausible output that passes a casual read.** The defense is systematic verification at the point of generation, not after.
**Before delivering any PM artifact, verify**:
1. **Specificity test**: Could someone act on this without asking clarifying questions?
2. **Source test**: Does every factual claim trace to something the user provided?
3. **Completeness test**: Are edge cases, error states, and failure modes addressed?
4. **Tradeoff test**: Are tradeoffs made explicit, not hidden?
5. **Context test**: Are numbers presented with comparison, baseline, and uncertainty?
6. **Strategic test**: Does this connect to a goal, OKR, or strategic theme?
7. **Honesty test**: Does this acknowledge what we do not know?
If any test fails, fix it before delivering. Do not rationalize past it.
references/product-management/metrics-review.md
# Metrics Review Reference
Deep reference for product metrics analysis, goal-setting, and dashboard design. Loaded by METRICS mode.
---
## Product Metrics Hierarchy
### North Star Metric
The single metric capturing core value delivered. Must be:
| Property | Test |
|----------|------|
| Value-aligned | Moves when users get more value |
| Leading | Predicts long-term business success |
| Actionable | Product team can influence it |
| Understandable | Everyone in the company gets it |
**Examples by product type**:
| Product Type | North Star | Why |
|-------------|------------|-----|
| Collaboration tool | Weekly active teams with 3+ contributing members | Measures collaborative value, not just logins |
| Marketplace | Weekly transactions completed | Measures successful exchange, not just visits |
| SaaS platform | Weekly active users completing core workflow | Measures real usage, not just opening the app |
| Content platform | Weekly engaged reading/viewing time | Measures attention, not clicks |
| Developer tool | Weekly deployments using the tool | Measures integration into real workflow |
### L1 Metrics — Health Indicators
5-7 metrics that together paint complete product health across the user lifecycle.
#### Acquisition
| Metric | What It Answers |
|--------|----------------|
| New signups / trial starts | Are new users finding us? |
| Signup conversion rate | Are visitors converting? |
| Channel mix | Where are new users coming from? |
| Cost per acquisition | What are we paying per new user (paid channels)? |
#### Activation
| Metric | What It Answers |
|--------|----------------|
| Activation rate | % of new users completing the key action predicting retention |
| Time to activate | How long from signup to activation? |
| Setup completion rate | % completing onboarding steps |
| First value moment | When users first experience core value |
**Defining activation**: Look at retained vs churned users. What actions did retained users take that churned did not? The activation event should be strongly predictive of long-term retention and achievable within the first session or few days.
#### Engagement
| Metric | What It Answers |
|--------|----------------|
| DAU / WAU / MAU | How many active users at each timeframe? |
| DAU/MAU ratio (stickiness) | What fraction of monthly users return daily? (>0.5 = daily habit, <0.2 = infrequent) |
| Core action frequency | How often users do the thing that matters most |
| Session depth | How much users do per session |
| Feature adoption | % using key features |
**Defining "active"**: A login? A page view? A core action? Different definitions tell different stories. Choose deliberately and document.
#### Retention
| Metric | What It Answers |
|--------|----------------|
| D1 / D7 / D30 / D90 retention | % returning after 1 day / 1 week / 1 month / 3 months |
| Cohort retention curves | How retention evolves per signup cohort |
| Churn rate | % of users or revenue lost per period |
| Resurrection rate | % of churned users who return |
**Reading retention curves**:
| Shape | Diagnosis |
|-------|-----------|
| Steep initial drop, then flat | Activation problem. Users who survive the first few days stick. |
| Steady decline without flattening | Engagement problem. No habit formed. |
| Flat early, then gradual decline | Initial value delivered but not sustained. |
| Improving across cohorts | Product improvements working. |
#### Monetization
| Metric | What It Answers |
|--------|----------------|
| Free to paid conversion | Are users upgrading? |
| MRR / ARR | What is recurring revenue? |
| ARPU / ARPA | Revenue per user/account? |
| Expansion revenue | Growth from existing customers? |
| Net revenue retention | Revenue retained including expansion + contraction? |
#### Satisfaction
| Metric | What It Answers |
|--------|----------------|
| NPS | Would users recommend? |
| CSAT | Are users satisfied with specific interactions? |
| Support ticket volume | Are issues increasing or decreasing? |
| App store ratings | What do public reviews say? |
### L2 Metrics — Diagnostic
Used to investigate L1 changes. Load on demand, not reviewed routinely.
- Funnel conversion at each step
- Feature-level usage and adoption
- Segment breakdowns (plan, company size, geography, role)
- Performance metrics (load time, error rate, API latency)
- Content-specific engagement (which features/pages drive engagement)
---
## KPI Frameworks
### OKRs (Objectives and Key Results)
**Objectives**: Qualitative, aspirational, time-bound, directional, memorable.
**Key Results**: Quantitative, specific, time-bound, outcome-based, 2-4 per Objective.
**Example**:
```
Objective: Make our product indispensable for daily workflows
KR1: Increase DAU/MAU from 0.35 to 0.50
KR2: Increase D30 retention for new users from 40% to 55%
KR3: 3 core workflows with >80% task completion rate
```
**Scoring**: 0.0-0.3 = missed, 0.4-0.6 = progress, 0.7-1.0 = achieved.
**Failure modes**:
| Failure Mode | Problem |
|-------------|---------|
| Too many OKRs (>3 objectives) | Focus diluted, nothing gets done well |
| Output KRs ("ship X features") | Measures activity not impact |
| Sandbagged targets | 100% confidence = not ambitious. Target 70% completion. |
| No mid-period review | Discover off-track too late to correct |
| Dishonest grading | Defeats the purpose of the system |
### Target-Setting Process
| Step | Action | Failure Mode |
|------|--------|-------------|
| 1. Baseline | Establish current reliable value | Setting targets without knowing where you start |
| 2. Benchmark | Check comparable products / industry | Ignoring context (B2B vs B2C retention norms differ 3x) |
| 3. Trajectory | What is current trend? | Setting 6% target when metric already improving 5%/month |
| 4. Effort | How much investment behind this? | Ambitious target with no allocated resources |
| 5. Confidence | Commit (high confidence) vs stretch (ambitious) | Single target with no range |
---
## Funnel Analysis
### Building a Funnel
1. Define the sequence of steps users take to reach the outcome
2. Measure conversion at each step
3. Identify the biggest drop-off points — these are highest-leverage opportunities
### Funnel Metrics
| Metric | Formula | Use |
|--------|---------|-----|
| Step conversion | (Users completing step N) / (Users completing step N-1) | Identify where users drop off |
| Overall conversion | (Users completing final step) / (Users entering funnel) | Overall efficiency |
| Time between steps | Median time from step N to step N+1 | Identify friction points |
| Drop-off rate | 1 - step conversion | Quantify the problem |
### Funnel Analysis Rules
- **Segment everything**: Different user types have wildly different funnels. Enterprise vs SMB, mobile vs desktop, new vs returning.
- **Time-bound**: Define the window. "Signs up and activates within 7 days" is different from "ever activates."
- **Beware averages**: Mean time-to-activate can be misleading if the distribution is bimodal. Use median or percentiles.
- **Attribution**: If you changed multiple things, you cannot attribute funnel changes to one change. Run A/B tests for causal claims.
### Common Funnel Failure Modes
| Failure Mode | Problem | Fix |
|-------------|---------|-----|
| Too many steps | Overwhelms analysis | Focus on 5-7 key decision points |
| Undefined "active" | Different queries give different answers | Document the definition precisely |
| No segmentation | Hides segment-specific problems | Always break down by key dimensions |
| Correlation = causation | "Users who do X retain better" does not mean X causes retention | Design experiments to test causal hypotheses |
---
## Cohort Analysis
### What Cohort Analysis Reveals
Cohort analysis groups users by when they joined (or performed an action) and tracks behavior over time.
**Key insight**: Aggregate metrics mask cohort effects. Overall retention could be declining even as each new cohort retains better — if growth is slowing, older (worse-retaining) cohorts dominate the average.
### Cohort Table Format
```
Week 0 Week 1 Week 2 Week 3 Week 4
Jan W1 100% 45% 32% 28% 25%
Jan W2 100% 48% 35% 30% --
Jan W3 100% 52% 38% -- --
Jan W4 100% 55% -- -- --
```
**Reading the table**:
- **Rows** (left to right): How a single cohort retains over time. Looking for the curve to flatten.
- **Columns** (top to bottom): How the same retention point improves across cohorts. Are newer cohorts better?
- **Diagonal**: All cohorts at the same calendar time. Reveals seasonal or product-wide effects.
### Cohort Analysis Best Practices
| Practice | Why |
|----------|-----|
| Use behavioral cohorts, not just time | "Users who activated in week 1" reveals more than "users who signed up in January" |
| Compare cohorts before/after changes | Did the onboarding redesign improve D7 retention? Compare pre and post cohorts. |
| Wait for maturity | A 2-week-old cohort cannot tell you about D30 retention. Be patient. |
| Control for externalities | Holidays, marketing campaigns, outages all affect cohorts differently |
---
## Statistical Pitfalls
### The Big Five
| Pitfall | What Happens | Defense |
|---------|-------------|---------|
| **Small sample noise** | Drawing conclusions from tiny datasets | Report sample sizes. Flag n<100 findings as directional only. |
| **Survivorship bias** | Analyzing only users who stayed, ignoring those who left | Include churned/inactive users in analysis. Study drop-offs. |
| **Simpson's Paradox** | Aggregate trend reverses when segmented | Always segment. A flat overall metric can hide one segment growing and another shrinking. |
| **Cherry-picking timeframes** | Choosing the period that tells the story you want | Use consistent comparison periods. Show multiple timeframes. |
| **Vanity metrics** | Metrics that always go up but indicate nothing | Total signups ever, total page views. Use rate metrics instead. |
### Statistical Significance
- For A/B tests: require 95% confidence (p < 0.05) before declaring a winner
- For small samples: be cautious. A 0.1 point NPS change on n=50 is noise.
- For metric movements: distinguish signal from noise. Weekly fluctuations of 5-10% are often normal variance.
- Report confidence intervals, not just point estimates when possible.
### Correlation vs Causation
"Users who use Feature X retain better" does NOT mean Feature X causes retention.
Possible explanations:
1. Feature X causes retention (what you hope)
2. Retained users discover Feature X over time (reverse causation)
3. Power users both use Feature X and retain better (confounding variable)
4. Feature X and retention both correlate with company size (spurious correlation)
**How to establish causation**: Run an A/B experiment. Randomly expose users to Feature X and measure retention difference.
---
## Review Cadences
### Weekly (15-30 min)
| What | Depth |
|------|-------|
| North Star | Current value, WoW change |
| Key L1 metrics | Notable movements only |
| Active experiments | Results, statistical significance |
| Anomalies | Unexpected spikes or drops |
| Alerts | Anything triggered |
**Action threshold**: Investigate if something looks off. Otherwise note and move on.
### Monthly (30-60 min)
| What | Depth |
|------|-------|
| Full L1 scorecard | MoM trends |
| OKR progress | On track / at risk / off track |
| Cohort analysis | Newer cohorts improving? |
| Feature adoption | Recent launches performing? |
| Segment analysis | Divergence between segments? |
**Action**: Identify 1-3 areas to investigate or invest in. Update priorities if metrics reveal new information.
### Quarterly (60-90 min)
| What | Depth |
|------|-------|
| OKR scoring | Grade the quarter honestly |
| L1 trends over quarter | Direction and rate of change |
| Year-over-year comparison | Seasonal adjustment, long-term trajectory |
| Competitive context | Market shifts, competitor movements |
| What worked / what did not | Attribution of results to actions |
**Action**: Set next-quarter OKRs. Adjust product strategy based on data.
---
## Dashboard Design
### Principles
| Principle | Implementation |
|-----------|---------------|
| Question-first | What decisions does this dashboard support? Design backwards from the decision. |
| Information hierarchy | North Star most prominent. L1 next. L2 on drill-down. |
| Context over numbers | Every number shows: current value, comparison, trend direction. |
| Fewer metrics | 5-10 that matter. Everything else in detailed reports. |
| Consistent timeframes | Same period for all metrics. No mixing daily and monthly. |
| Visual status | Green (on track), Yellow (attention), Red (off track). |
| Actionability | Every metric is something the team can influence. |
### Layout
```
┌────────────────────────────────────────┐
│ NORTH STAR: [Metric] + trend + target │
├──────────┬──────────┬──────────────────┤
│ Acquisition │ Activation │ Engagement │
│ [metrics] │ [metrics] │ [metrics] │
├──────────┬──────────┬──────────────────┤
│ Retention │ Revenue │ Satisfaction │
│ [metrics] │ [metrics]│ [metrics] │
├────────────────────────────────────────┤
│ Active Experiments / Recent Launches │
├────────────────────────────────────────┤
│ L2 drill-down (on demand) │
└────────────────────────────────────────┘
```
### Dashboard Failure Modes
| Failure Mode | Why It Fails |
|-------------|-------------|
| Vanity metrics | Total signups ever, total page views — always go up, indicate nothing |
| Too many metrics | If it requires scrolling, cut metrics |
| No comparison | Raw numbers without previous period or target |
| Stale dashboards | Not updated or reviewed in months |
| Output metrics | Tickets closed, PRs merged instead of user/business outcomes |
| One dashboard for all | Execs, PMs, and engineers need different views |
### Alerting
| Alert Type | Trigger | Example |
|-----------|---------|---------|
| Threshold | Metric crosses critical boundary | Error rate > 1%, conversion < 5% |
| Trend | Sustained decline over multiple periods | 3 consecutive weeks of retention decline |
| Anomaly | Significant deviation from expected range | Traffic 50% below predicted |
**Alert hygiene**:
- Every alert is actionable. If you cannot do anything, do not alert.
- Review and tune regularly. False positives train people to ignore all alerts.
- Every alert has an owner. Who responds when it fires?
- Not everything is P0. Set severity levels.
---
## Metric Scorecard Template
```
## Product Metrics Review: [Period]
### Summary
[2-3 sentences: overall health, most notable change, key callout]
### Scorecard
| Metric | Current | Previous | Change | Target | Status |
|--------|---------|----------|--------|--------|--------|
| [North Star] | | | | | |
| Acquisition: Signups | | | | | |
| Activation: Rate | | | | | |
| Engagement: DAU/MAU | | | | | |
| Retention: D30 | | | | | |
| Revenue: MRR | | | | | |
| Satisfaction: NPS | | | | | |
### Trend Analysis
[For each metric worth discussing: what happened, why, one-time or sustained]
### Bright Spots
- [What is going well]
### Areas of Concern
- [What needs attention]
### Recommended Actions
1. [Specific investigation / experiment / investment / alert]
### Caveats
- [Data quality issues, comparability notes, missing metrics]
```
references/product-management/roadmap-planning.md
# Roadmap Planning Reference
Deep reference for roadmap creation, updates, and prioritization. Loaded by ROADMAP mode.
---
## Roadmap Frameworks
### Now / Next / Later
The simplest and most effective format for most teams.
| Horizon | Timeframe | Confidence | Content |
|---------|-----------|-----------|---------|
| **Now** | Current sprint/month | High — committed | Active work. Scoped. Owners assigned. |
| **Next** | 1-3 months | Medium — planned | Prioritized, not started. Good confidence in what, less in when. |
| **Later** | 3-6+ months | Low — directional | Strategic bets and opportunities. Scope and timing flexible. |
**When to use**: Most teams, most of the time. Avoids false precision on dates. Good for leadership and external communication.
**Failure modes**: Treating "Later" as a dumping ground. Items in Later should still tie to strategy — they are directional bets, not a wish list.
### Quarterly Themes
2-3 themes per quarter representing strategic investment areas.
```
Q3 2026 Themes:
├── Enterprise Readiness
│ ├── SSO / SAML support
│ ├── Audit logging
│ └── Role-based access control
├── Activation Improvements
│ ├── Guided onboarding v2
│ ├── Template gallery
│ └── First-run experience redesign
└── Platform Extensibility
├── Public API v2
└── Webhook improvements
```
**When to use**: When you need strategic alignment visibility. Good for planning meetings and executive communication.
**Failure modes**: Themes that are too broad ("Make product better") or too narrow (a single feature disguised as a theme).
### OKR-Aligned Roadmap
Map items directly to Objectives and Key Results.
| Objective | Key Result | Initiatives | Expected Impact |
|-----------|-----------|------------|----------------|
| Make product indispensable for daily workflows | Increase DAU/MAU from 0.35 to 0.50 | Notification improvements, mobile app, quick actions | +0.08 DAU/MAU |
| | Increase D30 retention from 40% to 55% | Onboarding v2, activation flow, email re-engagement | +10pp retention |
**When to use**: Organizations that run on OKRs. Creates clear accountability between what you build and what you measure.
**Failure modes**: Initiatives without expected impact estimates. If you cannot estimate impact, the link to the KR is speculative.
### Timeline / Gantt
Calendar-based view showing start dates, end dates, durations, parallelism, and dependencies.
**When to use**: Execution planning with engineering. Identifying scheduling conflicts.
**When NOT to use**: External communication. Creates false precision expectations. "Shipping SSO on March 15" becomes a promise.
---
## Prioritization Frameworks
### RICE Score
**Formula**: (Reach x Impact x Confidence) / Effort
| Dimension | Definition | Scale |
|-----------|-----------|-------|
| **Reach** | Users/customers affected in a time period | Concrete numbers (e.g., 500/quarter) |
| **Impact** | Needle movement per person reached | 3=massive, 2=high, 1=medium, 0.5=low, 0.25=minimal |
| **Confidence** | How confident in reach + impact estimates | 100%=data-backed, 80%=some evidence, 50%=gut feel |
| **Effort** | Person-months (eng + design + all functions) | Concrete estimate |
**When to use**: Large backlog, need quantitative defensibility.
**Failure modes**: Gaming confidence scores to make preferred initiatives win. If RICE does not match intuition, investigate why — do not adjust inputs until it does.
### ICE Score
**Formula**: Impact x Confidence x Ease (each 1-10)
Simpler than RICE. Ease is inverse of effort (higher = easier to build).
**When to use**: Quick prioritization. Early-stage products. Insufficient data for RICE.
### MoSCoW
| Category | Definition | Decision Test |
|----------|-----------|---------------|
| **Must** | Roadmap fails without these. Non-negotiable. | "Would the quarter be a failure without this?" |
| **Should** | Important and expected. Delivery viable without. | High-priority fast follows. |
| **Could** | Desirable. Lower priority. | Include only if capacity allows. |
| **Won't** | Explicitly out of scope this period. | List for clarity. |
**When to use**: Scoping a release or quarter. Forcing prioritization conversations with stakeholders.
### Value vs Effort Matrix
| | Low Effort | High Effort |
|---|---|---|
| **High Value** | Quick Wins — do first | Big Bets — plan carefully |
| **Low Value** | Fill-ins — spare capacity | Money Pits — do not do |
**When to use**: Visual prioritization in team sessions. Building shared understanding of tradeoffs.
---
## OKR Alignment
### Writing Product OKRs
**Objectives**: Qualitative, aspirational, time-bound (quarterly/annually), directional.
**Key Results**: Quantitative, specific, time-bound, outcome-based (not output-based), 2-4 per Objective.
| Weak KR | Problem | Strong KR |
|---------|---------|-----------|
| "Launch onboarding v2" | Output, not outcome | "Increase activation rate from 30% to 50%" |
| "Ship 10 features" | Activity, not impact | "3 core workflows achieve >80% task completion" |
| "Improve NPS" | No target, no timeline | "Increase NPS from 32 to 45 by end of Q3" |
**Scoring**: 0.0-0.3 = missed, 0.4-0.6 = progress, 0.7-1.0 = achieved. 70% completion is the target for stretch OKRs.
**Failure modes**:
- Too many OKRs (2-3 objectives max)
- KRs that are sandbagged (100% confidence = not ambitious enough)
- KRs that measure effort instead of results
- Not reviewing at mid-period
- Not grading honestly at end of period
### Roadmap-to-OKR Mapping
Every roadmap item should answer: "Which Key Result does this move?"
Items that cannot answer this question are one of:
1. **Strategically orphaned** — cut them or justify with a separate rationale
2. **Infrastructure/tech debt** — legitimate but frame as enabling future KRs
3. **Reactive work** — customer escalations, compliance requirements (acceptable, but track the ratio)
**Healthy ratio**: 70%+ of roadmap items directly tied to OKRs. 20% tech health. 10% unplanned buffer.
---
## Dependency Mapping
### Dependency Types
| Type | Description | Example | Risk Level |
|------|------------|---------|-----------|
| Technical | Feature B requires infra from Feature A | "API v2 needs auth service refactor" | Medium |
| Team | Requires work from another team | "Need design review from Design team" | High |
| External | Waiting on vendor, partner, third party | "Stripe Connect certification" | Very High |
| Knowledge | Need research/investigation results first | "User testing results before building flow" | Medium |
| Sequential | Must ship A before starting B | "Billing integration before usage-based pricing" | Medium |
### Managing Dependencies
| Rule | Implementation |
|------|---------------|
| List explicitly | Every dependency visible in roadmap |
| Assign owner | Someone responsible for resolving each dependency |
| Set "need by" date | When does the dependent item need this resolved? |
| Build buffer | Dependencies are highest-risk items. Pad them. |
| Flag cross-team early | Cross-team coordination requires lead time |
| Contingency plan | What if the dependency slips? |
### Reducing Dependencies
Before accepting a dependency, ask:
- Can we build a simpler version that avoids it?
- Can we parallelize with an interface contract or mock?
- Can we sequence differently to move the dependency earlier?
- Can we absorb the work into our team?
---
## Capacity Planning
### Estimating Capacity
```
Raw capacity = engineers x days in sprint
Overhead = meetings + on-call + interviews + holidays + PTO
Available capacity = Raw capacity - Overhead
Planned capacity = Available capacity x 0.65 (60-70% rule)
```
### Allocation Model
| Category | % | Purpose |
|----------|---|---------|
| Planned features | 70% | Roadmap items advancing strategic goals |
| Technical health | 20% | Tech debt, reliability, performance, DX |
| Unplanned | 10% | Buffer for urgent issues, quick wins, requests |
**Adjustments by context**:
| Situation | Shift |
|-----------|-------|
| New product | More features, less tech debt |
| Mature product | More tech debt and reliability |
| Post-incident | More reliability, fewer features |
| Rapid growth | More scalability and performance |
### Capacity vs Ambition
- If commitments exceed capacity, something must give
- Do not solve capacity problems by pretending people can do more — cut scope
- When adding to roadmap, always ask "What comes off?"
- Commit to fewer things and deliver reliably > overcommit and disappoint
---
## Communicating Roadmap Changes
### Change Triggers
- New strategic priority from leadership
- Customer feedback / research that changes priorities
- Technical discovery that changes estimates
- Dependency slip from another team
- Resource change (team grows, shrinks, key person leaves)
- Competitive move requiring response
### Communication Framework
| Step | Action |
|------|--------|
| 1. Acknowledge | Be direct about what is changing and why |
| 2. Explain | What new information drove this decision? |
| 3. Show tradeoff | What was deprioritized? What slips? |
| 4. Present new plan | Updated roadmap with changes reflected |
| 5. Acknowledge impact | Who is affected? Stakeholders expecting deprioritized items hear it directly. |
### Avoiding Roadmap Whiplash
- Do not change for every piece of new information. Have a threshold.
- Batch updates at natural cadences (monthly, quarterly) unless truly urgent.
- Distinguish "roadmap change" (strategic reprioritization) from "scope adjustment" (normal execution refinement).
- Track change frequency. Frequent changes may signal unclear strategy, not responsiveness.
---
## Theme-Based Planning
### Building Themes
Themes represent strategic investment areas, not individual features.
**Good theme characteristics**:
- Tied to a business outcome or user need cluster
- Broad enough to contain multiple initiatives
- Narrow enough to be a coherent narrative
- Communicates WHY you are investing, not just WHAT you are building
| Bad Theme | Why | Good Theme |
|-----------|-----|-----------|
| "Build stuff" | No strategic signal | "Enterprise readiness" |
| "SSO" | Too narrow (single feature) | "Security and compliance foundations" |
| "Make users happy" | Too vague, no direction | "Reduce time-to-first-value for new users" |
### Theme-to-Initiative Mapping
```
Theme: Reduce time-to-first-value for new users
├── Guided onboarding flow (Now)
├── Template gallery (Now)
├── Sample data environment (Next)
├── Onboarding email sequence (Next)
└── AI-assisted setup (Later)
```
Each initiative under a theme inherits the strategic rationale. When stakeholders ask "why are we building X?" the answer is the theme.
### Balancing Themes
- 2-3 themes per quarter (max)
- Themes should be roughly balanced in investment unless explicitly stated otherwise
- At least one theme should address existing user retention/satisfaction (not all new growth)
- Review theme balance against business priorities each quarter
---
## Roadmap Review Cadences
| Cadence | Purpose | Depth | Attendees |
|---------|---------|-------|-----------|
| Weekly | Execution check, surface blockers | Status updates on Now items | PM + eng lead |
| Monthly | Trend review, priority adjustments | Now/Next reprioritization | Product team + stakeholders |
| Quarterly | Strategic review, OKR scoring, next-quarter planning | Full roadmap reassessment | Product + eng + design + leadership |
### Quarterly Planning Process
1. **Score** — Grade previous quarter OKRs honestly
2. **Review** — What worked, what did not, what changed
3. **Input** — Gather strategic context: company priorities, customer feedback, competitive intel, tech debt backlog
4. **Draft** — Propose themes, initiatives, and OKRs for next quarter
5. **Negotiate** — Stakeholder review, capacity check, priority tradeoffs
6. **Commit** — Finalize roadmap and OKRs, communicate broadly
references/product-management/spec-writing.md
# Spec Writing Reference
Deep reference for writing feature specifications and PRDs. Loaded by SPEC mode.
---
## PRD Section Guide
### Problem Statement
The foundation. Everything else flows from this.
**Structure**:
- What is the user problem? (2-3 sentences)
- Who experiences it and how often?
- What is the cost of not solving it? (user pain, business impact, competitive risk)
- What evidence grounds this? (research, support data, metrics, customer feedback)
**Failure modes**:
| Bad | Why | Better |
|-----|-----|--------|
| "Users need better onboarding" | No problem defined, no evidence | "42% of new users abandon setup before connecting their first integration (Mixpanel, Q1). Support tickets about 'getting started' are our #2 category." |
| "Enterprise customers want SSO" | States a solution, not a problem | "Enterprise IT teams cannot enforce their security policies without centralized auth. Three $100K+ prospects cited this as a blocker in Q4 pipeline reviews." |
| "We should improve performance" | No specificity, no who, no impact | "Page load time for the dashboard averages 4.2s on mobile (target: <2s). Users with >50 items see 8s+ loads. Mobile DAU is 30% below desktop DAU per capita." |
### Goals
**Rules**:
- 3-5 specific, measurable outcomes
- Each answers: "How will we know this succeeded?"
- Outcomes, not outputs: "reduce time to first value by 50%" not "build onboarding wizard"
- Distinguish user goals (what users get) from business goals (what the company gets)
**Examples**:
| Weak | Strong |
|------|--------|
| "Improve user experience" | "Reduce median time-to-first-success from 12 minutes to 5 minutes within 30 days of launch" |
| "Increase engagement" | "Increase D7 retention for new users from 35% to 50%" |
| "Drive revenue" | "Convert 15% of free users to paid within 60 days of activation" |
### Non-Goals
As important as goals. Prevent scope creep during implementation.
**Rules**:
- 3-5 explicit exclusions
- Adjacent capabilities out of scope for this version
- Brief rationale for each: not enough impact, too complex, separate initiative, premature
**Example**:
```
Non-Goals:
- Mobile app support (separate initiative, Q3 roadmap)
- Admin bulk operations (low request volume, <5 tickets/month)
- Real-time collaboration (requires infrastructure investment beyond this scope)
- Custom branding per workspace (enterprise-only need, will address in enterprise tier work)
```
---
## User Story Patterns
### Format
"As a [specific user type], I want [capability] so that [benefit]."
### INVEST Criteria
| Criterion | Test |
|-----------|------|
| **I**ndependent | Can be developed and delivered alone |
| **N**egotiable | Details can be discussed, not a contract |
| **V**aluable | Delivers value to the user (not just the team) |
| **E**stimable | Team can roughly estimate effort |
| **S**mall | Completable in one sprint |
| **T**estable | Clear verification path |
### Common Failure Modes
| Pattern | Problem | Fix |
|---------|---------|-----|
| Too vague | "As a user, I want the product to be faster" | What specifically? Which workflow? What target? |
| Solution-prescriptive | "As a user, I want a dropdown menu" | Describe the need: "I want to select from my saved filters" |
| No benefit | "As a user, I want to click a button" | Why? What does clicking accomplish? |
| Too large | "As a user, I want to manage my team" | Break into: invite members, set roles, remove members, view activity |
| Internal focus | "As engineering, we want to refactor the database" | This is a task, not a user story. What user outcome does the refactor enable? |
| Missing edge cases | Only happy path described | Add: error states, empty states, boundary conditions, permission failures |
### Edge Case Prompts
For every user story, explicitly address:
- **Empty state**: What does the user see before any data exists?
- **Error state**: What happens when the action fails? Network error? Validation error? Permission denied?
- **Boundary conditions**: What happens at limits? Max items? Max characters? Zero items? Concurrent access?
- **Permission variations**: What does a viewer see vs an editor vs an admin?
- **Undo/recovery**: Can the user reverse this action? What happens if they do?
- **Offline/degraded**: What happens with poor connectivity? Partial data?
---
## Acceptance Criteria Methodology
### Given/When/Then Format
```
Given [precondition or context]
When [action the user takes]
Then [expected outcome]
```
**Rules**:
- Cover happy path, error cases, and edge cases
- Be specific about expected behavior, not implementation
- Include negative test cases (what should NOT happen)
- Each criterion is independently testable
- Ban ambiguous words without definition
### Ambiguity Elimination
| Ambiguous | Specific |
|-----------|----------|
| "fast" | "< 200ms p95 response time" |
| "user-friendly" | "completes task in < 3 clicks, 0 help-text references needed" |
| "intuitive" | "80% of test users complete without guidance on first attempt" |
| "responsive" | "renders correctly at 320px-2560px viewport width" |
| "scalable" | "handles 10K concurrent users with < 500ms p99 latency" |
| "secure" | "encrypted at rest (AES-256), in transit (TLS 1.3), audit logged" |
### Checklist Format (Alternative)
```
- [ ] Admin can enter SSO provider URL in organization settings
- [ ] Team members see "Log in with SSO" on login page
- [ ] SSO login creates account if none exists (email match)
- [ ] SSO login links to existing account on email match
- [ ] Failed SSO shows error message with retry option and support link
- [ ] Admin can disable SSO (members fall back to email/password)
- [ ] SSO removal does NOT delete existing accounts
```
---
## Requirements Prioritization
### P0 / P1 / P2 Framework
| Priority | Definition | Test |
|----------|-----------|------|
| **P0 (Must-Have)** | Feature cannot ship without these. Minimum viable. | "Would we not ship without this?" If no, P0. |
| **P1 (Nice-to-Have)** | Significantly improves experience. Core use case works without them. | Fast follow-ups after launch. |
| **P2 (Future)** | Out of scope for v1. Design should support them later. | Architectural insurance — guide decisions now. |
**Discipline rules**:
- Be ruthless about P0s. Tighter must-have list = faster ship + faster learning.
- If everything is P0, nothing is P0. Challenge every must-have.
- P1s are things you are confident you will build soon, not a wish list.
- P2s prevent accidental architecture decisions that make future work hard.
### MoSCoW Cross-Reference
| MoSCoW | Maps to | Usage |
|--------|---------|-------|
| Must have | P0 | Non-negotiable commitments |
| Should have | P1 | Important, expected, but delivery is viable without |
| Could have | Below P1 | Include only if capacity allows |
| Won't have | Explicit exclusion | Out of scope this version |
---
## Scope Management
### Recognizing Scope Creep
- Requirements added after spec approval
- "Small" additions accumulating into a larger project
- Building features no user asked for ("while we are at it...")
- Launch date moving without explicit re-scoping
- Stakeholders adding requirements without removing anything
### Prevention Tactics
| Tactic | Implementation |
|--------|---------------|
| Explicit non-goals | Every spec has them |
| Scope addition = scope removal | Any add comes with a remove or timeline extension |
| v1 / v2 separation | Clear boundary in spec |
| Problem statement check | Review spec against original problem. Does everything serve it? |
| Investigation time-box | "If we cannot figure out X in 2 days, we cut it" |
| Parking lot | Capture good ideas that are out of scope |
---
## Success Metrics Definition
### Leading Indicators (Change in Days-Weeks)
| Metric | What It Measures |
|--------|-----------------|
| Adoption rate | % of eligible users who try the feature |
| Activation rate | % who complete the core action |
| Task completion rate | % who accomplish their goal |
| Time to complete | Duration of core workflow |
| Error rate | How often users hit errors or dead ends |
| Feature usage frequency | How often users return to the feature |
### Lagging Indicators (Change in Weeks-Months)
| Metric | What It Measures |
|--------|-----------------|
| Retention impact | Does this feature improve retention? |
| Revenue impact | Does this drive upgrades, expansion, or new revenue? |
| NPS / satisfaction | Does this improve user sentiment? |
| Support ticket reduction | Does this reduce support load? |
| Competitive win rate | Does this help win more deals? |
### Target-Setting Rules
- Specific: "50% adoption within 30 days" not "high adoption"
- Based on comparables: similar features, industry benchmarks, explicit hypotheses
- Two thresholds: "success" and "stretch"
- Measurement method defined: what tool, what query, what time window
- Evaluation cadence defined: 1 week, 1 month, 1 quarter post-launch
---
## Output Structure
```markdown
# [Feature Name] — Product Requirements Document
## Problem Statement
[2-3 sentences. Who, evidence, cost of not solving.]
## Goals
1. [Measurable outcome tied to user/business metric]
2. ...
## Non-Goals
1. [Exclusion] — [rationale]
2. ...
## User Stories
### [Persona 1]
- As a [type], I want [capability] so that [benefit]
- AC: Given... When... Then...
- AC: Given... When... Then...
### [Persona 2]
- ...
## Requirements
### P0 — Must-Have
| Requirement | Acceptance Criteria | Dependencies |
|------------|-------------------|--------------|
| ... | ... | ... |
### P1 — Nice-to-Have
| Requirement | Acceptance Criteria | Dependencies |
|------------|-------------------|--------------|
| ... | ... | ... |
### P2 — Future Considerations
| Requirement | Notes |
|------------|-------|
| ... | ... |
## Success Metrics
### Leading Indicators
| Metric | Target | Measurement | Evaluation |
|--------|--------|-------------|------------|
| ... | ... | ... | ... |
### Lagging Indicators
| Metric | Target | Measurement | Evaluation |
|--------|--------|-------------|------------|
| ... | ... | ... | ... |
## Open Questions
| Question | Owner | Blocking? |
|----------|-------|-----------|
| ... | ... | ... |
## Timeline
| Milestone | Date | Dependencies |
|-----------|------|-------------|
| ... | ... | ... |
```
references/productivity.md
# Productivity
Umbrella skill for personal and team productivity: task decomposition, daily/weekly planning, meeting optimization, status updates, goal setting, and focus management. Each mode loads its own reference files on demand.
---
## Mode Detection
Classify into one mode before proceeding.
| Mode | Signal Phrases | Reference |
|------|---------------|-----------|
| **TASK** | add task, prioritize tasks, task list, what's on my plate, decompose work, batch tasks | `references/productivity/task-management.md` |
| **PLAN** | daily plan, plan my day, time blocks, plan my week, energy mapping | `references/productivity/daily-weekly-planning.md` |
| **MEETING** | meeting agenda, optimize meeting, meeting audit, cancel this meeting, async alternative | `references/productivity/meeting-optimization.md` |
| **STATUS** | status update, standup, weekly update, stakeholder update, progress report | `references/productivity/status-updates.md` |
| **REVIEW** | weekly review, retro, retrospective, reflect on week, monthly review | `references/productivity/daily-weekly-planning.md` |
| **GOAL** | set goals, OKRs, quarterly goals, goal progress, key results | `references/productivity/daily-weekly-planning.md` |
If the request spans modes, pick the primary mode and note the secondary.
---
## Workflow by Mode
### TASK Mode
**Load**: `references/productivity/task-management.md`, `references/productivity/llm-productivity-failure-modes.md`
1. **Capture** — Accept tasks in any format: freeform text, bullet lists, pasted meeting notes, vague intentions. Extract actionable items.
2. **Decompose** — Apply vertical slicing (because horizontal slices create work that cannot ship independently):
| Slice Quality | Example |
|---------------|---------|
| Good (vertical) | "User can upload a CSV and see a preview" — shippable alone |
| Weak (horizontal) | "Build the upload API" — requires the UI to deliver value |
3. **Estimate** — Assign time estimates using the 1/2/4-hour bucketing system (because finer granularity creates false precision, coarser loses planning value). Tasks over 4 hours need decomposition.
4. **Prioritize** — Apply the appropriate framework based on context:
| Context | Framework | Why |
|---------|-----------|-----|
| Personal daily work | Eisenhower (urgent/important) | Fast, intuitive, separates reactive from proactive |
| Backlog with many items | ICE (Impact/Confidence/Ease) | Quantitative ranking without heavy data requirements |
| Team sprint planning | Weighted scoring against goals | Defensible, transparent to stakeholders |
5. **Organize** — Group by context (because context-switching between unrelated tasks costs 15-25 minutes per switch). Batch similar work: all emails together, all code reviews together, all writing together.
**Gate**: Every task has an action verb, a completion condition, and a time estimate. Vague items like "think about marketing" get reframed as "Draft 3 marketing channel options with pros/cons (2h)."
### PLAN Mode
**Load**: `references/productivity/daily-weekly-planning.md`, `references/productivity/llm-productivity-failure-modes.md`
1. **Gather constraints** — Ask for (conversationally, not as a wall of questions):
- Calendar commitments for the day/week
- Hard deadlines
- Energy level and known energy patterns (because matching task difficulty to energy state increases completion rates)
- Carryover from yesterday
2. **Select top priorities** — Identify the Top 3 outcomes for the day (because more than 3 priorities means zero priorities). Apply the "if only these 3 things got done, would today feel successful?" test.
3. **Build time blocks** — Map tasks to calendar slots:
| Block Type | When to Schedule | Duration |
|------------|-----------------|----------|
| Deep work (creation, analysis) | Peak energy hours (usually morning) | 90-120 min |
| Reactive work (email, Slack, reviews) | Low energy hours (usually post-lunch) | 30-60 min batches |
| Admin/maintenance | End of day | 30 min |
| Buffer | Between blocks | 15 min minimum |
4. **Identify conflicts** — Flag when calendar meetings fragment deep work blocks. Surface the cost: "You have 3 meetings between 9-12, leaving zero uninterrupted blocks during your peak hours."
5. **Generate plan** — Output a concrete, time-blocked plan with the Top 3 outcomes highlighted.
**Gate**: Plan accounts for actual calendar (not aspirational free time). Deep work blocks are at least 90 minutes. Buffers exist between blocks. Total planned work does not exceed available hours minus 20% (because unplanned work always appears).
### MEETING Mode
**Load**: `references/productivity/meeting-optimization.md`, `references/productivity/llm-productivity-failure-modes.md`
1. **Determine operation**:
| Operation | What to Do |
|-----------|-----------|
| Audit existing meeting | Apply the 5P framework: Purpose, Participants, Preparation, Process, Payoff |
| Design new agenda | Build outcome-driven agenda with time allocations and decision types |
| Convert to async | Draft async alternative with decision framework and deadline |
| Optimize recurring meeting | Analyze frequency, attendance, decision output vs time spent |
2. **For audits** — Calculate meeting cost (participants x hourly rate x duration x frequency). Surface the number because most people underestimate it. A weekly 1-hour meeting with 8 people at $75/hr costs $31,200/year.
3. **For agendas** — Every agenda item gets:
- **Type**: Decision, Discussion, Information, or Brainstorm (because different types need different facilitation)
- **Owner**: Who presents/facilitates this item
- **Time**: Allocated minutes
- **Pre-read**: What participants should review before the meeting
- **Outcome**: What "done" looks like for this item
4. **For async conversion** — Apply the async-first decision tree:
- Can this be a document with comments? Do that instead.
- Does this need real-time debate? Keep the meeting, shorten it.
- Does this need a decision from one person? Send them a 1-page memo with a deadline.
**Gate**: Every meeting has a stated purpose that could not be achieved async. Every agenda item has a type, owner, time allocation, and defined outcome. Information-only meetings are flagged for conversion to async.
### STATUS Mode
**Load**: `references/productivity/status-updates.md`, `references/productivity/llm-productivity-failure-modes.md`
1. **Detect audience** — Different audiences need different framing:
| Audience | Frame | Length | Lead With |
|----------|-------|--------|-----------|
| Manager (1:1) | Progress + blockers + asks | 3-5 bullets | What you need from them |
| Team (standup) | Yesterday/Today/Blockers | 60 seconds spoken | Blockers first |
| Stakeholders | Outcomes + timeline + risks | 1 page | Business impact |
| Executives | Red/Yellow/Green + decisions needed | < 200 words | Decisions needed |
2. **Gather inputs** — Ask for:
- What shipped or progressed since last update
- What is blocked and by whom
- What decisions are needed (and from whom)
- Timeline changes (and why)
3. **Generate update** using the Progress/Plans/Problems format:
- **Progress**: Completed outcomes (not activities). "Shipped search indexing, 40% faster queries" beats "worked on search."
- **Plans**: Next period's commitments with confidence levels.
- **Problems**: Blockers with specific asks. "Need API access from Platform team by Friday to unblock integration testing" beats "waiting on dependencies."
4. **For standups** — Optimize for brevity:
- Lead with blockers (because that is the only part the team can act on in real time)
- State completed items as outcomes, not activities
- State today's focus as the single most important deliverable
**Gate**: Status update exists. Outcomes framed as results (not activities). Every problem has a specific ask with a named owner and deadline. Executive updates are under 200 words.
### REVIEW Mode
**Load**: `references/productivity/daily-weekly-planning.md`, `references/productivity/llm-productivity-failure-modes.md`
1. **Collect** — Gather data from the period:
- What was planned vs what actually happened
- Tasks completed, deferred, or abandoned
- Unplanned work that appeared
- Calendar analysis: time in meetings vs deep work vs reactive work
2. **Process** — For each incomplete item:
| Outcome | Action |
|---------|--------|
| Deferred (still relevant) | Reschedule with honest time estimate |
| Deferred (no longer relevant) | Remove — carrying dead tasks creates cognitive overhead |
| Blocked | Identify the specific unblock action and owner |
| Abandoned (scope changed) | Archive with reason |
3. **Reflect** — Surface patterns (because reviews that skip reflection are just task lists):
- What type of work consistently gets deferred? (This reveals priority misalignment or estimation failures)
- Where did unplanned work come from? (This reveals process gaps or boundary issues)
- Which commitments to others were met vs missed? (This reveals reliability patterns)
- What was the ratio of deep work to reactive work? (Target: at least 40% deep work)
4. **Decide** — Identify 1-3 concrete adjustments for the next period. Specific and testable: "Block 9-11am as no-meeting time" not "do more deep work."
5. **For retrospectives** — Facilitate with:
- What went well (keep doing)
- What could improve (change one thing)
- Action items (assigned, with deadlines)
- Separate observations from emotions from actions (because conflating them derails retros)
**Gate**: Review compares planned vs actual. At least one pattern is surfaced from the data. Adjustments are specific and testable (not aspirational). Dead tasks are removed, not carried forward indefinitely.
### GOAL Mode
**Load**: `references/productivity/daily-weekly-planning.md`, `references/productivity/llm-productivity-failure-modes.md`
1. **Determine scope**: Quarterly OKRs, annual goals, project milestones, or personal development goals.
2. **Structure goals** using the outcome hierarchy:
| Level | Timeframe | Format | Example |
|-------|-----------|--------|---------|
| Vision | 1-3 years | Narrative | "Become the team's go-to person for data infrastructure" |
| Objective | Quarter | Qualitative outcome | "Make the data pipeline reliable enough that on-call is boring" |
| Key Result | Quarter | Measurable milestone | "Reduce pipeline failures from 12/month to 2/month" |
| Initiative | Weeks | Concrete project | "Add circuit breakers to the 5 highest-failure-rate jobs" |
3. **Validate each goal** against:
- **Measurability**: How will you know it is done? (Binary completion or metric target)
- **Influence**: Do you control the outcome, or does it depend on others? (Goals you do not control are hopes, not goals — reframe as the actions within your control)
- **Tension**: Does this goal conflict with another goal? (Surface tradeoffs explicitly)
- **Stretch calibration**: 70% confidence of achievement = good stretch. 100% = sandbagging. 30% = aspirational wish.
4. **Connect to daily work** — Map goals down to weekly themes and daily tasks. Goals that do not connect to this week's work are not goals yet — they are intentions.
**Gate**: Every goal has a measurable completion condition. Goals connect to at least one concrete next action. Conflicting goals have explicit tradeoff decisions. Quarterly goals have monthly check-in milestones.
---
## LLM Failure Modes in Productivity Work
See `references/productivity/llm-productivity-failure-modes.md` for the complete failure mode catalog (aspirational planning, unestimated tasks, generic advice, agendaless meetings, shallow reviews, activity-based status). Universal failure modes in `skills/shared-patterns/llm-domain-failure-modes-base.md`.
---
## Prioritization Frameworks (Cross-Mode Reference)
Used in TASK, PLAN, and GOAL modes.
| Framework | Method | Best For |
|-----------|--------|----------|
| **Eisenhower** | 2x2: Urgent/Important. Do (U+I), Schedule (I), Delegate (U), Drop (neither). | Personal daily prioritization |
| **ICE** | Impact x Confidence x Ease (1-10 each) | Quick ranking of a medium-sized backlog |
| **Weighted Scoring** | Score items against 3-5 criteria with explicit weights | Team decisions requiring transparency and defensibility |
| **Time-to-Value** | Prioritize by shortest path to delivering user value | When facing analysis paralysis on a long backlog |
Apply frameworks to the specific situation. Producing a framework explanation instead of an applied prioritization is a failure mode (see `references/productivity/llm-productivity-failure-modes.md`).
---
## Output Conventions
- Markdown with clear headers. Scannable by someone with 30 seconds.
- Tables for comparisons, schedules, and priority matrices.
- Time blocks in `HH:MM - HH:MM` format with task and estimated duration.
- Status labels: **Done**, **In Progress**, **Blocked**, **Deferred**, **Dropped**.
- Executive-facing content: < 200 words. Team-facing: as detailed as needed.
- Every recommendation is specific enough to act on today. "Improve focus" is not actionable. "Block 9-11am as no-meeting deep work time" is.
---
## Reference Loading Table
| Mode | Primary Reference | Secondary Reference |
|------|------------------|-------------------|
| TASK | `references/productivity/task-management.md` | `references/productivity/llm-productivity-failure-modes.md` |
| PLAN | `references/productivity/daily-weekly-planning.md` | `references/productivity/llm-productivity-failure-modes.md` |
| MEETING | `references/productivity/meeting-optimization.md` | `references/productivity/llm-productivity-failure-modes.md` |
| STATUS | `references/productivity/status-updates.md` | `references/productivity/llm-productivity-failure-modes.md` |
| REVIEW | `references/productivity/daily-weekly-planning.md` | `references/productivity/llm-productivity-failure-modes.md` |
| GOAL | `references/productivity/daily-weekly-planning.md` | `references/productivity/llm-productivity-failure-modes.md` |
references/productivity/daily-weekly-planning.md
# Daily & Weekly Planning Reference
Deep reference for daily planning, weekly reviews, monthly reconciliation, and goal setting. Loaded by PLAN, REVIEW, and GOAL modes.
---
## Daily Planning
### The Top 3 Method
Each day has three outcomes that define success. Everything else is bonus.
**Selection criteria**:
- "If only these 3 things got done, would today feel successful?"
- At least one Top 3 item should advance a weekly or quarterly goal (because otherwise the urgent crowds out the important every single day)
- Maximum one reactive/maintenance item in the Top 3
**Format**:
```
Today's Top 3:
1. [Most important outcome] — [time estimate] — [scheduled block]
2. [Second outcome] — [time estimate] — [scheduled block]
3. [Third outcome] — [time estimate] — [scheduled block]
```
### Energy Mapping
Match task difficulty to energy state. Most knowledge workers have predictable energy patterns.
| Energy Level | Typical Window | Best Task Types | Protect This Time |
|-------------|---------------|-----------------|-------------------|
| **Peak** | First 2-3 hours after starting | Creative work, complex analysis, writing, architecture decisions | Yes — this is your highest-value production window |
| **Sustained** | Mid-morning to early afternoon | Collaborative work, reviews, structured tasks | Partially — meetings here are acceptable |
| **Recovery** | Post-lunch, late afternoon | Email, admin, routine tasks, planning tomorrow | No — use this for low-stakes batch processing |
| **Second wind** | Some people get one; 60-90 min late afternoon | Varies by person | If present, use for a focused sprint on one task |
**Personalization**: Ask about the person's actual patterns. Default to the table above as a starting point, then adjust. Someone who works 6am-2pm has a different map than someone who works 10am-6pm.
### Time-Blocking Template
```
[Pre-work ritual] 05 min Review Top 3, check calendar
[Deep work block 1] 90 min Top 3 item #1
[Buffer] 15 min
[Deep work block 2] 90 min Top 3 item #2
[Lunch] 30-60 min
[Meeting block] As scheduled
[Reactive batch] 30 min Email, Slack, reviews
[Deep work block 3] 60 min Top 3 item #3
[Admin batch] 30 min Expense reports, scheduling, misc
[End-of-day wrap] 10 min Review progress, plan tomorrow's Top 3
```
**Adaptation rules**:
- If calendar has meetings before 11am, move deep work block 1 to the earliest available slot. Flag the fragmentation cost.
- If more than 3 hours of meetings, only 2 Top 3 items are realistic. Acknowledge this explicitly.
- Buffer time is not optional — unplanned work always appears. Plans without buffers fail by noon.
### Planning Failure Modes with Corrections
| Failure Mode | What Goes Wrong | Do Instead |
|-------------|----------------|------------|
| Planning 8 productive hours | Unplanned work, energy dips, and transition costs consume 20-40% | Plan for 5-6 hours of focused output. Leave the rest as buffer. |
| No deep work blocks | Day becomes entirely reactive — lots of activity, few outcomes | Schedule at least one 90-min uninterrupted block before checking email or Slack |
| Top 5 (or 7, or 10) priorities | With more than 3, everything feels equally important and nothing gets the focus it needs | Limit to 3. If a 4th is truly critical, one of the original 3 was not actually Top 3. |
| Planning without checking calendar | Plan calls for 4 hours of deep work but calendar has 5 hours of meetings | Build the plan around the calendar, not despite it |
| Same plan every day regardless of energy | Monday morning energy is different from Friday afternoon energy | Adjust ambition to energy. Monday: tackle the hardest thing. Friday: review, plan, tie up loose ends. |
| Skipping the end-of-day review | Tomorrow starts cold with no momentum | 10-min wrap-up: what finished, what carries over, what is tomorrow's #1 |
---
## Weekly Review
### The 5-Step Process
Run weekly (Friday afternoon or Monday morning). Takes 30-60 minutes. This is the most important productivity habit — it is the feedback loop that makes everything else work.
#### Step 1: Collect
Gather all open loops into one place:
- Inbox (email, chat, notifications)
- Meeting notes from the week
- Sticky notes, scraps, mental to-dos
- Commitments made to others
- Ideas that surfaced during the week
**Goal**: Empty your head. Every "I should..." becomes a written item.
#### Step 2: Process
For each collected item, make one decision:
| Decision | Criteria | Action |
|----------|----------|--------|
| **Do** | < 2 minutes | Do it now during the review |
| **Defer** | Takes longer, you are the right person | Add to task list with time estimate |
| **Delegate** | Someone else is better positioned | Send to them with clear ask and deadline |
| **Drop** | Not important enough to act on | Delete. If it matters, it will come back. |
#### Step 3: Organize
- Update task list: remove completed items, add new ones, update priorities
- Review calendar for next week: identify preparation needed for meetings
- Check goals: is this week's planned work moving the quarterly needle?
#### Step 4: Review
Compare planned vs actual for the past week:
| Metric | What It Shows |
|--------|--------------|
| **Completion rate** | % of planned tasks completed. Target: 70-80%. Below 60% = over-planning. Above 90% = under-challenging. |
| **Carryover count** | Tasks that rolled from last week. More than 3 chronic carryovers = task is either too big, not important, or blocked. |
| **Unplanned ratio** | % of work that was not on the plan. Above 40% = either poor planning or a reactive role that needs different planning strategies. |
| **Deep work hours** | Hours spent in uninterrupted focused work. Target: 15+ hours/week for individual contributors. |
| **Meeting load** | Hours in meetings. Above 50% of work hours = meeting problem, not productivity problem. |
#### Step 5: Decide
Based on the review data, make 1-3 adjustments:
- Adjustments must be specific and testable ("Block Tuesday and Thursday 9-11am as no-meeting time" not "do more deep work")
- Try one adjustment for 2 weeks before adding another (because stacking changes obscures which one helped)
- If the same issue appears 3 weeks in a row with no improvement, the adjustment is not working — try a different approach
### Weekly Review Checklist
```
[ ] Inboxes at zero (email, chat, notifications processed)
[ ] All commitments to others tracked with deadlines
[ ] Task list reflects reality (no stale items, no missing items)
[ ] Next week's calendar reviewed, prep tasks identified
[ ] Top 3 outcomes identified for next week
[ ] One adjustment from this week's review data
```
---
## Monthly/Quarterly Goal Reconciliation
### Monthly Check-In (30 minutes)
1. **Progress scan**: For each quarterly goal, assess Red/Yellow/Green:
| Status | Meaning | Action |
|--------|---------|--------|
| **Green** | On track for quarterly target | Continue current approach |
| **Yellow** | Behind pace, but recoverable | Identify the specific bottleneck. Adjust weekly plans to allocate more time. |
| **Red** | Significantly behind, at risk | Decide: double down (and what gets cut?), reduce scope, or abandon with reason |
2. **Goal relevance check**: Has anything changed that makes a goal irrelevant? Strategy shifts, role changes, new information. Dropping a goal because circumstances changed is good judgment. Dropping a goal because it is hard is avoidance — distinguish the two.
3. **Next month's themes**: What 2-3 themes should dominate next month's weekly plans to move goals forward?
### Quarterly Review (60-90 minutes)
1. **Score each goal**: Met / Partially Met / Missed
2. **Analyze misses**: Why? Categories:
| Miss Reason | Pattern | Correction |
|-------------|---------|------------|
| Too ambitious | Goal was a 30% confidence bet, not 70% | Calibrate goal difficulty. 70% confidence = good stretch. |
| Crowded out | Urgent work consumed the time | Build goal work into the weekly plan as non-negotiable blocks |
| Wrong goal | Goal did not align with what actually mattered | Improve goal-setting process. Involve stakeholders earlier. |
| Dependencies failed | Blocked by others | Identify dependencies at goal-setting time. Build relationships to unblock. |
3. **Set next quarter**: 2-3 objectives, each with 2-4 measurable key results. The objectives should feel uncomfortable but achievable (70% confidence). If everything feels easy, you are sandbagging. If everything feels impossible, you are dreaming.
---
## Goal-Setting Methodology
### The Outcome Hierarchy
| Level | Timeframe | Format | Test |
|-------|-----------|--------|------|
| **Vision** | 1-3 years | Narrative description of desired state | "Would I recognize this if I saw it?" |
| **Objective** | Quarter | Qualitative outcome statement | "Is this an outcome, not an output?" |
| **Key Result** | Quarter | Measurable milestone | "Can I unambiguously determine if this is met?" |
| **Initiative** | Weeks | Concrete project or workstream | "Does this directly move a Key Result?" |
| **Task** | Hours-days | Single action with completion condition | "Can this be done in one sitting?" |
### Writing Good Key Results
| Quality | Example | Problem |
|---------|---------|---------|
| **Good** | "Reduce average API response time from 450ms to 200ms" | Clear baseline, clear target, measurable |
| **Good** | "Ship 3 customer-requested integrations selected by quarterly survey results" | Countable, traceable to customer input |
| **Weak** | "Improve API performance" | No baseline, no target, not measurable |
| **Weak** | "Launch new integrations" | How many? Which ones? How do we know we are done? |
| **Bad** | "Make the API fast" | Subjective, unmeasurable, no completion condition |
### Goal Conflict Resolution
When goals pull in different directions:
1. Name the conflict explicitly: "Shipping faster conflicts with improving test coverage"
2. Decide which wins this quarter (because trying to optimize for both means optimizing for neither)
3. Set a constraint on the losing side: "Ship faster, but test coverage does not drop below 70%"
4. Revisit next quarter — the losing priority becomes next quarter's focus if the constraint is holding
---
## Day-of-Week Patterns
Different days have different energy and context. Use this as a starting template, then adjust to the person.
| Day | Energy Pattern | Good For | Caution |
|-----|---------------|----------|---------|
| **Monday** | High starting energy, fresh from weekend | Hardest creative work, biggest decisions, week planning | Avoid filling with catch-up meetings — protect the energy |
| **Tuesday-Wednesday** | Peak productivity days for most | Deep work, key deliverables, important collaborations | These are your most valuable days. Guard them. |
| **Thursday** | Steady but starting to wind down | Reviews, feedback, collaborative refinement, decisions | Good day for closing loops started earlier in the week |
| **Friday** | Lower energy, end-of-week fatigue | Weekly review, loose ends, planning next week, learning | Do not start new complex work on Friday — finish things |
references/productivity/llm-productivity-failure-modes.md
# LLM Failure Modes in Productivity Work
Where LLMs systematically fail at productivity tasks. Loaded across all modes as a guardrail reference.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers productivity-specific failures only.
---
## Why This File Exists
LLMs are optimistic planners and fluent generators. Both traits are dangerous in productivity work. An aspirational daily plan looks motivating until it collides with a real calendar. A task list without time estimates feels productive until nothing gets prioritized. Generic advice sounds wise until it fails to adapt to a specific person's constraints.
This reference catalogs the specific failure modes, their signatures, and the defenses against each.
---
## Failure Mode 1: Aspirational Planning
### What Happens
The LLM generates a daily or weekly plan that ignores real constraints: existing calendar commitments, energy fluctuations, commute time, meal breaks, transition costs between tasks. The plan assumes 8 hours of peak-energy productive time. The user follows it for 2 hours, falls behind, and abandons the plan entirely.
### Signatures
| Signal | Example |
|--------|---------|
| No calendar awareness | Plan schedules 4 hours of deep work when 3 hours of meetings are already booked |
| No buffer time | Back-to-back blocks with zero transition time |
| Peak energy all day | Creative work scheduled for 4pm as if it is the same as 9am |
| 100% utilization | Every minute planned, leaving zero room for unplanned work |
| Aspirational task count | 12 items on a daily plan when 3-5 is realistic |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Calendar-first planning | Start with the calendar. Build the plan around committed time, not on top of it. |
| Buffer mandate | Minimum 15-minute buffer between blocks. Minimum 20% of day unplanned. |
| Energy-aware scheduling | Ask about energy patterns. Schedule hard work during peak energy, routine work during recovery. |
| Task count ceiling | Daily plans have a maximum of 3 Top priority items. Plans with more than 3 priorities have zero priorities. |
| Utilization cap | Plan for 5-6 productive hours, not 8. The remaining time absorbs meetings, transitions, and surprises. |
---
## Failure Mode 2: Unestimated Task Lists
### What Happens
The LLM produces a beautifully organized task list with categories, priorities, and descriptions — but no time estimates. Without estimates, the user cannot plan a day around the list. They pick tasks by gut feel, run out of time, and carry tasks forward indefinitely. The list grows. Morale shrinks.
### Signatures
| Signal | Example |
|--------|---------|
| No time estimates | Tasks listed with titles and descriptions only |
| No priority markers | All tasks appear equally important |
| No completion conditions | "Research competitors" — when is this done? |
| Infinite list | 40+ items with no triage or grouping |
| No distinction between quick and deep | A 5-minute email reply next to a 4-hour architecture doc |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Mandatory estimates | Every task gets a time bucket: 1h, 2h, or 4h. Tasks under 1h batch together. Tasks over 4h decompose. |
| Completion conditions | Every task has a "done when" clause. "Research competitors" becomes "Write a 1-page comparison of the top 3 competitors' pricing models." |
| Priority assignment | Use Eisenhower (personal) or ICE (backlog). Unprioritized lists are incomplete. |
| List ceiling | Active task list stays under 15 items. Anything beyond goes to a backlog or "someday" list. If you cannot maintain 15 active items, you cannot maintain 40. |
| Size visibility | When presenting a task list, include total estimated hours. "This list totals 28 hours — that is roughly 5 days of focused work." |
---
## Failure Mode 3: Generic Productivity Advice
### What Happens
The LLM dispenses widely-known productivity tips without adapting them to the user's specific situation. "Try the Pomodoro technique." "Use the Eisenhower matrix." "Eat the frog — do the hardest thing first." This advice is correct in general and useless in particular.
### Signatures
| Signal | Example |
|--------|---------|
| Named techniques without adaptation | "Have you tried the Pomodoro technique?" without knowing the user's work involves 3-hour deep coding sessions that Pomodoro would fragment |
| One-size-fits-all | Same advice for a manager with 6 hours of meetings and an IC with 6 hours of deep work |
| Technique-dropping | Mentioning 5 techniques without applying any to the user's specific situation |
| Context-free tips | "Block time for deep work" without knowing the user's calendar or team culture |
| Book recommendations instead of solutions | "Check out Getting Things Done!" when the user asked for help planning their Tuesday |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Context-first | Gather the user's actual constraints (calendar, role, energy patterns, team size, meeting load) before recommending anything. |
| Apply, do not recommend | Instead of "use the Eisenhower matrix," take their actual tasks and sort them into the matrix. Show the result, not the technique name. |
| Adapt to role | A manager's productivity system (meeting optimization, delegation, decision velocity) is fundamentally different from an IC's (focus protection, batch processing, deep work blocks). |
| One recommendation at a time | Pick the single highest-leverage change for this person and implement it. Five tips is an article. One implemented change is progress. |
| Test against specifics | Before giving advice, check: "Does this recommendation account for [specific constraint the user mentioned]?" If not, revise it. |
---
## Failure Mode 4: Meeting Agendas Without Decision Outcomes
### What Happens
The LLM creates a meeting agenda that lists topics to discuss but does not specify what decisions the meeting should produce. The result is a well-organized meeting that covers three topics, generates discussion, and concludes without anyone knowing what was decided or what happens next.
### Signatures
| Signal | Example |
|--------|---------|
| Topic-only agenda | "1. Discuss Q2 roadmap. 2. Review hiring plan. 3. Budget update." — what decisions? |
| No time allocations | Topics listed without time budgets, so item 1 consumes 50 minutes of a 60-minute meeting |
| No item types | Discussion items, decision items, and information items all treated identically |
| No pre-read | Attendees arrive cold. The first 15 minutes are context-setting that should have happened async. |
| No defined outcome | "Agenda" is really a topic list. After the meeting, there is no artifact. |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Item typing | Every agenda item is labeled: Decision, Discussion, Information, or Brainstorm. Different types need different facilitation. |
| Outcome definition | Each item has a "done when" statement. "Q2 roadmap — done when top 3 priorities are rank-ordered and each has an owner." |
| Time allocation | Each item gets allocated minutes that sum to the meeting duration minus 5 minutes (for wrap-up). |
| Pre-read requirement | Information-heavy items have pre-read materials distributed 24h before. Meeting time is for questions and decisions, not presentations. |
| Action item capture | Designate a note-taker. Decisions and action items captured during the meeting with owners and deadlines. |
---
## Failure Mode 5: Task-List Reviews
### What Happens
The LLM produces a "weekly review" that is just a task list with checkmarks. What was done. What was not done. Move undone items to next week. This is task management, not review. There is no reflection, no pattern identification, no adjustment to the system. The same failures repeat week after week.
### Signatures
| Signal | Example |
|--------|---------|
| Done/Not-done only | "Completed 7 of 12 tasks. Carrying 5 forward." No analysis of why. |
| No planned vs actual comparison | Tasks listed without comparing to what was originally planned |
| No pattern identification | The same type of task gets deferred every week and nobody notices |
| No systemic questions | Review does not ask "why do I consistently overcommit?" or "why does admin work always expand?" |
| Carryover without examination | Tasks moved forward every week for a month with no examination of whether they matter |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Planned vs actual comparison | Show what was planned for the week alongside what actually happened. The gap is the data. |
| Pattern questions | Ask: What type of work consistently gets deferred? Where did unplanned work come from? Which commitments to others were met vs missed? |
| Carryover limit | A task that carries forward 3+ times gets one of: decompose (too big), escalate (blocked), drop (not actually important). Indefinite carryover is not allowed — it creates a guilt backlog that drains energy. |
| Adjustment mandate | Every review produces 1-2 specific, testable adjustments. "Block Tuesday morning for deep work" not "do more deep work." |
| Metrics tracking | Track completion rate, carryover count, unplanned work ratio, deep work hours. Trends over 4+ weeks reveal systemic issues that single-week reviews miss. |
---
## Failure Mode 6: Activity-Based Status Updates
### What Happens
The LLM writes status updates framed as activities ("worked on", "spent time", "had meetings about") rather than outcomes ("shipped", "decided", "unblocked"). Activity framing signals effort. Outcome framing signals impact. Stakeholders receiving activity-based updates cannot tell if the project is progressing or just consuming time.
### Signatures
| Signal | Example |
|--------|---------|
| Activity verbs | "Worked on search feature" instead of "Shipped search indexing — queries 40% faster" |
| No measurable progress | "Continued integration work" — are we 30% done or 90% done? |
| Meeting attendance as progress | "Met with design team about the new flow" — what was decided? |
| No blockers surfaced | Everything sounds on track even when it is not |
| No decisions needed | Status update does not ask the reader for anything — purely informational |
### Defenses
| Defense | Implementation |
|---------|---------------|
| Outcome verb requirement | Use "shipped", "decided", "unblocked", "measured", "validated" instead of "worked on", "discussed", "continued" |
| Measurable progress | Include percentage, count, or milestone: "3 of 5 API endpoints migrated. On track for completion Wednesday." |
| Decision surfacing | Every update includes decisions made (for the record) and decisions needed (for the reader to act on) |
| Honest risk reporting | Status colors reflect reality. If it is yellow, say yellow. Early warnings preserve trust. Late surprises destroy it. |
| Ask inclusion | Status updates that do not ask the reader for anything are FYIs, not updates. If no action is needed, say so explicitly — that is itself useful information. |
---
## Cross-Cutting Defense: The Specificity Test
Across all failure modes, the root cause is the same: the LLM generates plausible, well-structured output that lacks specificity. The defense is to check every recommendation, plan, and task against the person's actual situation.
**Before delivering any productivity artifact, verify**:
1. **Calendar test**: Does this plan account for the real calendar, not an imaginary free day?
2. **Energy test**: Does this schedule match the person's stated or likely energy patterns?
3. **Estimate test**: Does every task have a time estimate and a completion condition?
4. **Adaptation test**: Would this recommendation change if the person's role, schedule, or team size were different? If not, it is probably generic.
5. **Actionability test**: Can the person start on this within the next 24 hours? If not, it is aspirational.
6. **Carryover test**: Are deferred tasks examined for patterns, or just moved forward?
7. **Outcome test**: Are results framed as outcomes (impact) or activities (effort)?
If any test fails, revise the artifact before delivering. Productive-looking output that does not survive contact with a real calendar is worse than no output — it consumes the user's time to produce and their trust when it fails.
references/productivity/meeting-optimization.md
# Meeting Optimization Reference
Deep reference for meeting audits, agenda design, async conversion, and recurring meeting optimization. Loaded by MEETING mode.
---
## The 5P Audit Framework
Every meeting should pass this framework. If it fails on Purpose, the meeting should not exist.
| P | Question | Pass Criteria | Fail Signal |
|---|----------|--------------|-------------|
| **Purpose** | What decision or outcome does this meeting produce? | Specific, articulable outcome that requires real-time interaction | "Touch base," "sync up," "stay aligned" with no concrete deliverable |
| **Participants** | Does every attendee have a role (decide, input, inform)? | Each person either makes a decision, provides input that changes the outcome, or needs the information to act | More than 2 people who are "just listening" — they should get notes instead |
| **Preparation** | What should attendees review beforehand? | Pre-read exists and is distributed 24h+ before | Attendees arrive cold. The first 15 minutes are context-setting. |
| **Process** | How will the meeting be run? | Agenda with time allocations, facilitation plan, decision method | Freeform discussion that ends when time runs out |
| **Payoff** | What leaves the room? | Decisions documented, action items assigned with owners and deadlines | "Good discussion" with no recorded outcomes |
### Applying the Audit
For each meeting under review:
1. State the Purpose in one sentence. If you need two sentences, the meeting has two purposes — split it or pick one.
2. List every Participant with their role. Anyone without a clear role is optional — make them optional.
3. Identify the Preparation gap. If none exists, create a pre-read document.
4. Design the Process. Allocate time per agenda item. Identify the decision method for each decision item.
5. Define the Payoff. What artifacts does this meeting produce? Who captures them?
---
## Meeting Types and Templates
### Decision Meeting
**Purpose**: Arrive at a specific decision with the people who have authority and context.
| Element | Specification |
|---------|--------------|
| **Duration** | 30 min (extend to 45 only if multiple decisions) |
| **Attendees** | Decision maker + 2-4 people with critical context. Informed parties get notes. |
| **Pre-read** | 1-page decision memo: context, options (3 max), recommendation, tradeoffs |
| **Agenda** | 5 min: Confirm everyone read the memo. 15 min: Debate the options. 5 min: Make the decision. 5 min: Document and assign next steps. |
| **Output** | Decision recorded. Rationale documented. Action items assigned. |
**Key principle**: The meeting is for debate and decision, not for presenting context. Context goes in the pre-read. If attendees have not read it, spend 5 minutes on a verbal summary, then proceed — do not re-present the full document.
### Discussion Meeting
**Purpose**: Explore a topic, gather perspectives, develop shared understanding.
| Element | Specification |
|---------|--------------|
| **Duration** | 45-60 min |
| **Attendees** | People with diverse perspectives on the topic. 4-8 optimal. |
| **Pre-read** | Framing document: what we know, what we are exploring, specific questions to answer |
| **Agenda** | 5 min: Frame the discussion with specific questions. 35-45 min: Structured discussion (round-robin or breakout). 5-10 min: Synthesize key takeaways, identify follow-ups. |
| **Output** | Key insights documented. Open questions listed. Next steps and owners assigned. |
### Information Meeting
**Purpose**: Communicate information that requires Q&A or immediate reaction.
| Element | Specification |
|---------|--------------|
| **Duration** | 15-25 min maximum |
| **Attendees** | Only people who need to act on the information |
| **Pre-read** | The information itself (document, dashboard, announcement) |
| **Agenda** | 5 min: Highlight what changed and why it matters. 10-15 min: Q&A. 5 min: Action items. |
| **Output** | Questions answered. Action items captured. |
**Async-first check**: If no Q&A is expected and no immediate action is needed, this should be an email/document with a deadline for questions. Convert to async.
### Brainstorm Meeting
**Purpose**: Generate ideas on a specific problem.
| Element | Specification |
|---------|--------------|
| **Duration** | 45-60 min |
| **Attendees** | 3-6 people with diverse perspectives. More than 6 fragments the conversation. |
| **Pre-read** | Problem statement with constraints, prior attempts, and what "good" looks like |
| **Agenda** | 10 min: Restate problem and constraints. 20 min: Diverge — generate ideas (no evaluation). 15 min: Converge — cluster, evaluate, select top 3. 10 min: Define next steps for top ideas. |
| **Output** | Ranked list of ideas. Top 3 with assigned owners for next steps. |
**Facilitation rule**: Separate divergent and convergent thinking. Evaluating ideas during generation kills creativity. Generate first, evaluate second.
---
## Async-First Decision Tree
Before scheduling any meeting, apply this filter:
```
Does this require real-time interaction?
├── NO: Can it be a document with comments?
│ ├── YES → Write the document. Set a comment deadline. Skip the meeting.
│ └── NO: Can it be a recorded video + async feedback?
│ ├── YES → Record a Loom/video. Set a feedback deadline.
│ └── NO → Schedule the meeting (rare path).
└── YES: Does it require back-and-forth debate?
├── YES → Schedule a Decision Meeting (30 min)
└── NO: Does it require brainstorming?
├── YES → Schedule a Brainstorm Meeting (45-60 min)
└── NO → Schedule an Information Meeting (15-25 min)
```
### Situations That Genuinely Need Meetings
| Situation | Why Async Fails |
|-----------|----------------|
| Sensitive feedback (performance, conflict) | Tone and nuance get lost in text. Real-time allows reading reactions. |
| Multi-party negotiation | Async threads create misalignment. Real-time converges faster. |
| Creative brainstorming | Idea building requires rapid back-and-forth energy. |
| Crisis response | Speed of coordination matters more than documentation quality. |
| Relationship building (1:1s) | Trust is built through presence, not documents. |
### Situations That Should Almost Always Be Async
| Situation | Better Async Format |
|-----------|-------------------|
| Status updates | Written update in a shared channel with a template |
| FYI announcements | Email or document with a "questions by [date]" deadline |
| Document reviews | Comments on the document, consolidated by the author |
| Progress check-ins | Dashboard or weekly written summary |
| Information sharing | Recorded walkthrough video (< 10 min) |
---
## Meeting Cost Calculator
### Formula
```
Annual cost = (# participants) x (avg hourly cost) x (duration in hours) x (frequency per year)
```
### Common Examples
| Meeting | Participants | Duration | Frequency | Annual Cost (@$75/hr) |
|---------|-------------|----------|-----------|----------------------|
| Weekly team standup | 8 | 30 min | 50/year | $15,000 |
| Weekly team sync | 8 | 60 min | 50/year | $30,000 |
| Biweekly sprint planning | 6 | 90 min | 26/year | $17,550 |
| Monthly all-hands | 30 | 60 min | 12/year | $27,000 |
| Daily standup | 5 | 15 min | 250/year | $15,625 |
| Weekly 1:1 | 2 | 30 min | 50/year | $3,750 |
**How to use the number**: The cost itself does not determine whether a meeting is worthwhile. A $30,000/year meeting that produces weekly alignment saving $100,000 in rework is a bargain. A $3,750/year 1:1 that produces no decisions or growth conversations is waste. The number makes the cost visible so the value can be assessed honestly.
### Hidden Costs Not in the Formula
| Cost | Impact |
|------|--------|
| **Context-switching** | 15-25 min before and after each meeting to context-switch. A 30-min meeting actually costs ~60-75 min. |
| **Fragmentation** | A meeting at 10:30am breaks a morning into two sub-90-min blocks, neither long enough for deep work. |
| **Preparation** | Time spent preparing for the meeting (reviewing docs, creating slides). |
| **Recovery** | Emotionally draining meetings (conflict, bad news) reduce productivity for hours afterward. |
---
## Recurring Meeting Optimization
### Audit Protocol for Recurring Meetings
For each recurring meeting, ask:
| Question | If the Answer Is No |
|----------|--------------------|
| Did this meeting produce a decision or action item in the last 3 occurrences? | Reduce frequency or cancel |
| Is every regular attendee actively participating? | Make passive attendees optional; send them notes |
| Could the meeting be 50% shorter with better preparation? | Create a pre-read requirement and shorten |
| Is the frequency right? | Weekly meetings where nothing changes week-to-week should be biweekly. Daily meetings where updates take 2 minutes should be async. |
### Frequency Decision Guide
| Signal | Current Frequency | Recommendation |
|--------|------------------|---------------|
| "Nothing new since last time" is common | Weekly | Move to biweekly |
| Updates take < 5 min, no discussion | Daily | Move to async check-in |
| Decisions pile up between meetings | Biweekly | Move to weekly, or allow ad-hoc decision meetings |
| Meeting consistently runs over time | Any | Either extend (with tighter agenda) or split into two focused meetings |
| Attendance is declining | Any | Purpose has eroded — re-audit with 5P framework |
---
## Action Item Capture
### During the Meeting
Designate one person (not the facilitator) to capture:
- **Decision**: What was decided, by whom, with what rationale
- **Action item**: What needs to happen, who owns it, by when
- **Open question**: What was not resolved, who will follow up
### Action Item Format
```
[ ] [Action verb] [specific deliverable] — Owner: [name] — Due: [date]
```
**Good**: "Draft the API migration plan with timeline and resource requirements — Owner: Sarah — Due: March 8"
**Weak**: "Sarah to look into the API thing" — No deliverable, no deadline, ambiguous scope.
### Post-Meeting Protocol
Within 30 minutes of meeting end:
1. Send notes to all attendees and relevant stakeholders
2. Action items added to the appropriate task tracking system
3. Decisions recorded in the relevant decision log or document
If this does not happen, the meeting's value dissipates within 24 hours as people's memories diverge.
---
## Meeting-Free Time Protection
### Why It Matters
Knowledge workers need 3-4 hours of uninterrupted time daily for deep work. A single meeting in the middle of a morning block reduces that block's productive output by 40-60% (because of context-switching overhead on both sides).
### Protection Strategies
| Strategy | Implementation | Tradeoff |
|----------|---------------|----------|
| **No-meeting mornings** | Block 9am-12pm on the team calendar. Meetings only after noon. | Works well for individual deep work. May conflict with cross-timezone collaboration. |
| **No-meeting days** | Designate one full day (often Wednesday or Friday) with zero meetings. | Highest deep work yield. Requires team-wide commitment. |
| **Focus blocks** | Each person blocks 2-3 hour chunks as "busy" and declines meetings during them. | Most flexible. Lowest compliance without cultural support. |
| **Meeting windows** | All meetings happen in designated windows (e.g., 1-5pm). | Clear structure. Limits scheduling flexibility. |
### Defending Focus Time
When someone requests a meeting during protected focus time:
1. Offer an alternative time within your meeting windows
2. Offer an async alternative ("Can I review the document and send comments by EOD instead?")
3. If neither works, accept the meeting but acknowledge the tradeoff: "I can attend, but this means [deliverable] moves to tomorrow."
Making the cost visible builds organizational awareness. Over time, people stop booking over focus time because they see the downstream impact.
references/productivity/status-updates.md
# Status Updates Reference
Deep reference for status communication across audiences, standup optimization, retrospective facilitation, and change communication. Loaded by STATUS mode.
---
## Status Update Templates by Audience
### Manager (1:1 Update)
**Frame**: What you need from them + what they need to know about your work.
```
## This Week
**Shipped**
- [Outcome 1] — [impact or next step]
- [Outcome 2] — [impact or next step]
**In Progress**
- [Work item] — [% complete, expected done date]
**Blocked / Need Your Help**
- [Blocker] — I need [specific ask] by [date] to unblock [deliverable]
**Heads Up**
- [Risk or upcoming change] — no action needed yet, but [context]
```
**Key principle**: Lead with what you need from your manager. That is the actionable part. They can skim the rest.
### Team (Standup / Daily Update)
**Frame**: What the team can act on right now.
```
**Blockers** (if any — lead with these)
- [Blocker] — need [specific help] from [person]
**Done yesterday**
- [Outcome] (not "worked on X" — what finished?)
**Today's focus**
- [Single most important deliverable]
```
**Duration**: 60 seconds spoken, 3-5 lines written. Anything longer belongs in a separate conversation.
**Standup optimization rules**:
- Lead with blockers because that is the only part the team can act on in real time
- Say what finished (outcomes), not what you touched (activities)
- State one focus, not a task list — the team needs to know your priority, not your full schedule
- If you have nothing blocked and nothing interesting to report, say "No blockers, continuing on [project]" — do not invent content
### Stakeholder Update
**Frame**: Business impact + timeline + risks + decisions needed.
```
## [Project Name] — Status Update [Date]
### Summary
[2-3 sentences: what happened, what it means, what comes next]
### Progress
| Milestone | Status | Notes |
|-----------|--------|-------|
| [Milestone 1] | Done | [Key outcome or metric] |
| [Milestone 2] | On Track | [Expected completion date] |
| [Milestone 3] | At Risk | [Why and what we're doing about it] |
### Risks
| Risk | Impact | Likelihood | Mitigation |
|------|--------|-----------|------------|
| [Risk 1] | [What happens if it materializes] | High/Med/Low | [What we're doing] |
### Decisions Needed
- [Decision] — from [person/group] — by [date] — to unblock [what]
### Next Period
- [Key deliverable 1] — [expected date]
- [Key deliverable 2] — [expected date]
```
### Executive Update
**Frame**: Red/Yellow/Green + decisions needed. Executives scan, they do not read.
```
## [Project Name] — [GREEN/YELLOW/RED]
**One-line summary**: [What happened and why it matters to the business]
**Key metric**: [Number] vs [target] ([trend direction])
**Decision needed**: [Specific decision] by [date]
— Option A: [one sentence + tradeoff]
— Option B: [one sentence + tradeoff]
— Recommendation: [which and why]
**Next milestone**: [What] by [when]
```
**Rules for executive updates**:
- Under 200 words. If it is longer, it will not be read.
- Lead with conclusion, not journey. "We shipped X and it moved Y" — not "we had 14 standups and 3 sprint reviews."
- Status color reflects reality, not optimism. Yellow means "at risk with a mitigation plan." Red means "we need help." Calling a red project yellow to avoid a difficult conversation makes the eventual conversation worse.
- Every risk has a mitigation. Surfacing a risk without a plan is an alarm, not an update.
---
## The Progress/Plans/Problems Format
The universal format that works for any audience with adjustment to detail level.
### Progress (What Shipped)
**Frame as outcomes, not activities.**
| Activity Framing (avoid) | Outcome Framing (use) |
|--------------------------|----------------------|
| "Worked on the search feature" | "Shipped search indexing — queries are 40% faster" |
| "Had meetings about the launch plan" | "Aligned on March 15 launch date with marketing and engineering" |
| "Reviewed the Q2 roadmap" | "Finalized Q2 roadmap: 3 themes, 12 initiatives, capacity-matched" |
| "Investigated the performance issue" | "Identified the root cause: N+1 queries on the dashboard. Fix shipping tomorrow." |
**Why this matters**: Activity framing signals effort. Outcome framing signals impact. Stakeholders care about impact. Team members care about impact. Your future self reviewing these updates cares about impact.
### Plans (What Is Coming)
- State commitments with confidence levels:
- **Committed**: "Will ship by Friday" (90%+ confidence)
- **Planned**: "Targeting next week" (70% confidence)
- **Stretch**: "If capacity allows" (< 50% confidence)
- Include dependencies: "Shipping the integration depends on API access from Platform team, expected Wednesday."
- Flag timeline changes with reasons: "Originally targeting March 1, now March 8 due to [specific reason]."
### Problems (What Is Blocked or At Risk)
Every problem gets three parts:
1. **What**: Clear statement of the blocker or risk
2. **Impact**: What happens if it is not resolved (deadline miss, quality degradation, downstream delay)
3. **Ask**: Specific request with a named owner and deadline
**Good problem statement**: "The staging environment has been down since Monday, blocking integration testing. If not resolved by Wednesday, the March 15 launch date is at risk. Need DevOps to prioritize the fix — I've escalated to [name]."
**Weak problem statement**: "Staging is down, which is causing issues."
---
## Retrospective Facilitation
### Format: Start/Stop/Continue
The simplest effective retro format.
| Category | Prompt | Example |
|----------|--------|---------|
| **Start** | What should we begin doing? | "Start writing decision memos before scheduling decision meetings" |
| **Stop** | What should we stop doing? | "Stop extending sprint scope after planning. Additions go to next sprint." |
| **Continue** | What is working well? | "Continue the Thursday code review pairing — it's catching bugs faster" |
### Facilitation Protocol
1. **Set the stage** (5 min): State the retro's scope (last sprint, last month, specific project). Remind the team that the goal is improvement, not blame.
2. **Gather data** (10 min): Each person writes their Start/Stop/Continue items silently. One item per note. No discussion during writing.
3. **Cluster** (10 min): Group similar items. Name each cluster. Identify the top 3 clusters by dot voting.
4. **Discuss** (20 min): For each top cluster, discuss:
- What is happening? (observations)
- Why is it happening? (root cause, not symptoms)
- What can we change? (specific action, not aspiration)
5. **Decide** (5 min): Pick 1-2 action items. Each has an owner and a deadline. More than 2 actions from a single retro dilutes focus.
### Retro Failure Modes with Corrections
| Failure Mode | What Goes Wrong | Do Instead |
|-------------|----------------|------------|
| Blame-oriented discussion | People defend instead of improving | Frame as system problems: "What about our process allowed this?" not "Who caused this?" |
| Action items without owners | Nothing happens between retros | Every action item has a named owner and a specific deadline |
| Too many action items | Nothing gets priority attention | Maximum 2 action items per retro. Less is more. |
| Skipping the retro because "we're busy" | The team loses its only improvement feedback loop | Keep the retro. Shorten it to 15 minutes if needed. A short retro beats no retro. |
| Same issues every retro | Actions from previous retros are not tracked | Start each retro by reviewing last retro's action items. Did they happen? Did they help? |
| Only discussing recent events | Recency bias ignores systemic patterns | Explicitly ask about patterns: "Is this the first time, or does this keep happening?" |
---
## Change Communication
When plans, timelines, or scope change, communicate proactively.
### Change Communication Template
```
## Change: [What Changed]
**What**: [Specific change in one sentence]
**Why**: [Root cause — honest, not spin]
**Impact**: [What this means for stakeholders — timeline, scope, deliverables]
**Mitigation**: [What we're doing to minimize impact]
**New plan**: [Updated timeline or scope with confidence level]
**Ask**: [What you need from the audience, if anything]
```
### Communication Timing
| Change Type | When to Communicate | To Whom |
|-------------|-------------------|---------|
| Timeline slip (< 1 week) | Within 24 hours | Direct team, manager |
| Timeline slip (> 1 week) | Same day you know | Manager, stakeholders, dependent teams |
| Scope reduction | Before it is finalized | Stakeholders who requested the cut items |
| Scope increase | Before accepting | Manager (for capacity impact), team (for workload) |
| Risk materialized | Immediately | Anyone affected + anyone who can help |
| Strategy change | As soon as decided | Full team + stakeholders |
**Key principle**: Communicate bad news early. The earlier you surface a problem, the more options exist to address it. Late surprises destroy trust faster than early warnings.
---
## Update Cadence Guide
| Cadence | When to Use | Format |
|---------|------------|--------|
| **Daily** (standup) | Active sprint, high-velocity project, crisis response | Blockers/Done/Focus — 60 seconds |
| **Weekly** | Standard project cadence, manager 1:1s | Progress/Plans/Problems — 1 page |
| **Biweekly** | Steady-state projects, low-change-rate work | Summary with key metrics and risks |
| **Monthly** | Executive reporting, long-running programs | Scorecard + narrative + decisions needed |
| **Ad hoc** | Significant changes, milestones, escalations | Change communication template |
**Rule of thumb**: Update at the cadence of decision-making. If decisions happen weekly, update weekly. If the team makes daily decisions, update daily. Updating more frequently than decisions are made creates noise. Updating less frequently creates information gaps.
references/productivity/task-management.md
# Task Management Reference
Deep reference for task decomposition, prioritization, state management, and batch processing. Loaded by TASK mode.
---
## Task Decomposition
### Vertical Slicing
Every task should represent a complete slice of value — something that delivers a result on its own, even if small.
| Slice Type | Test | Example |
|------------|------|---------|
| **Vertical (good)** | "If this is the only thing done today, does it deliver a result?" | "Write the intro section of the proposal and send for early feedback" |
| **Horizontal (split further)** | "Does this require other tasks to be useful?" | "Research competitors" — useful only when synthesized into something |
**Splitting technique**: Take any large task and ask "What is the smallest version that delivers a visible result?" Then do that version first.
| Before | After |
|--------|-------|
| "Redesign the dashboard" | "Replace the dashboard header with the new layout and ship it" |
| "Write the Q2 report" | "Draft the executive summary with key metrics and circulate for input" |
| "Migrate the database" | "Migrate the users table, validate row counts, update the read path" |
### Timeboxing Rules
| Bucket | Use For | If Over Budget |
|--------|---------|---------------|
| **1 hour** | Single-action tasks: reply, review, quick fix | Already right-sized |
| **2 hours** | Focused work: drafting, analysis, implementation | Check if it can split into two 1h tasks |
| **4 hours** | Deep work: design sessions, complex writing, architecture | Maximum single-task size. Split if larger. |
| **> 4 hours** | Decompose further | This is a project, not a task. Break into 1-4h subtasks with individual completion conditions. |
**Why these buckets**: Finer granularity (15-min, 30-min) creates false precision — estimation error on knowledge work is 50-200%. Coarser (half-day, full-day) loses planning value. The 1/2/4 system is precise enough to plan a day, loose enough to absorb variance.
---
## Priority Frameworks Applied
### Eisenhower Matrix (Personal Daily Use)
| | Urgent | Not Urgent |
|---|---|---|
| **Important** | **Do now.** Client deadline, production incident, blocking someone. | **Schedule.** Strategic work, skill building, relationship investment. This quadrant is where career growth lives. |
| **Not Important** | **Delegate or timebox.** Most email, many meetings, routine requests. Set a 30-min window, then stop. | **Drop.** Busywork, low-value notifications, meetings without agendas. Saying no here funds the Important/Not-Urgent quadrant. |
**The key insight**: Most people spend their day in Urgent (both rows). The difference between productive and busy is how much time goes to Important/Not-Urgent. Target: 30% of your day in that quadrant.
### ICE Scoring (Backlog Prioritization)
Score each item 1-10 on three dimensions:
| Dimension | 1-3 (Low) | 4-6 (Medium) | 7-10 (High) |
|-----------|-----------|--------------|-------------|
| **Impact** | Marginal improvement, few people affected | Noticeable improvement, moderate reach | Significant outcome, many people affected |
| **Confidence** | Gut feel, no supporting data | Some evidence, analogous experience | Data-backed, validated approach |
| **Ease** | Multi-week, many dependencies, new territory | Days to a week, known approach | Hours to a day, straightforward path |
**ICE Score** = Impact x Confidence x Ease. Sort descending. Do the high-score items first.
**Failure mode defense**: If you find yourself adjusting scores to justify a preferred item, stop. The score is showing you something. Investigate the mismatch between your intuition and the numbers — that gap contains information.
### Weighted Scoring (Team Decisions)
When the team needs transparent, defensible prioritization:
1. Choose 3-5 criteria relevant to your goals (e.g., revenue impact, user satisfaction, strategic alignment, effort, risk)
2. Assign weights that sum to 100% (force tradeoff: if everything is weighted equally, the weights carry no information)
3. Score each item 1-5 on each criterion
4. Weighted score = sum of (score x weight) across criteria
Present the table. Let the team discuss where scores diverge — that disagreement is the valuable part.
---
## Task State Machine
Every task lives in exactly one state:
```
CAPTURED → DEFINED → SCHEDULED → IN PROGRESS → DONE
↓ ↓ ↓
DEFERRED BLOCKED DROPPED
↓ ↓
(re-enter DEFINED when ready)
```
| State | Entry Condition | Exit Condition |
|-------|----------------|----------------|
| **CAPTURED** | Task exists in any form | Has action verb + completion condition + time estimate |
| **DEFINED** | Passes the clarity test (see below) | Assigned to a specific day or time block |
| **SCHEDULED** | On a specific day's plan | Work begins |
| **IN PROGRESS** | Active work happening | Completion condition met, or state change |
| **DONE** | Completion condition verified | Archive after 1 week |
| **DEFERRED** | Consciously postponed with a review date | Review date arrives → re-enter DEFINED |
| **BLOCKED** | Cannot proceed. Blocker identified with owner. | Blocker resolved → re-enter SCHEDULED |
| **DROPPED** | Deliberately abandoned with reason | Terminal state |
**Clarity test** (DEFINED entry gate): "Could someone else pick this up and know exactly what 'done' looks like?" If not, the task is still CAPTURED.
---
## Dependency Tracking
### Dependency Types
| Type | Example | Tracking Method |
|------|---------|----------------|
| **Blocking** | "Cannot start X until Y is done" | List blocker explicitly. Assign owner to unblock. |
| **Informing** | "X would benefit from Y's output, but can proceed without it" | Note the dependency. Start X; incorporate Y when available. |
| **External** | "Waiting on vendor response / approval / access" | Set a follow-up date. Escalate if no response by date. |
### Waiting-On Protocol
For every external dependency:
1. Document what you are waiting for, from whom, since when
2. Set a follow-up date (default: 3 business days)
3. On follow-up date: ping once. If no response in 24h, escalate or find an alternative path.
4. Carry the waiting item visibly — do not let it disappear into a backlog
---
## Batch Processing Patterns
### Context-Switching Cost
Switching between unrelated tasks costs 15-25 minutes of reorientation. Three switches in a morning can consume an hour of productive time.
**Mitigation**: Group similar tasks into batches and process them in dedicated windows.
| Batch Type | Tasks | Optimal Window |
|------------|-------|---------------|
| **Communication** | Email replies, Slack threads, review requests | 30-min blocks, 2-3x/day |
| **Review** | Code reviews, document reviews, PR feedback | Single 60-min block |
| **Creation** | Writing, design, coding, analysis | 90-120 min uninterrupted blocks |
| **Admin** | Expense reports, scheduling, tool setup | 30 min end-of-day |
| **Planning** | Task triage, calendar review, priority updates | 15 min start-of-day |
### Processing Rules
- Process batches in order of energy requirement: creation first (high energy), then review (medium), then communication and admin (low).
- Set a hard stop for each batch. Communication expands to fill available time — timebox it.
- When an item in a batch requires deep thought, pull it out and schedule it as its own deep work block. Do not let one complex item derail the batch.
---
## Task Extraction from Conversations
When processing meeting notes, chat threads, or email:
| Signal | Task Type | Example |
|--------|-----------|---------|
| "I'll..." / "I can..." | Commitment you made | "I'll send the updated numbers by Friday" → Task: Send updated numbers (due Friday) |
| "Can you..." / "Please..." | Request received | "Can you review the proposal?" → Task: Review proposal (owner: you) |
| "We should..." / "We need to..." | Team action item | "We should update the docs" → Task: Update docs (owner: TBD — clarify) |
| "Let's follow up on..." | Follow-up | "Let's follow up next week" → Task: Follow up on [topic] (due: next week) |
| Deadline mentioned | Deadline task | "The board meeting is March 15" → Task: Prepare board materials (due: March 13, 2-day buffer) |
**Extraction rule**: Surface extracted tasks for confirmation. Present them, ask which to add. Automatically adding tasks from ambiguous signals ("we should...") creates noise.
references/risk-assessment.md
---
title: Risk Assessment — Pre-Mortem Analysis, Scenario Planning, Probability-Weighted Outcomes, Black Swan Identification
domain: strategic-decision
level: 3
skill: strategic-decision
---
# Risk Assessment Reference
> **Scope**: Structured risk assessment for CEO-level strategic decisions — pre-mortem templates, scenario planning matrices, probability-weighted expected value, and black swan identification. Use AFTER the decision matrix has selected an option and BEFORE committing resources. Does NOT cover project execution risk (see skills/project-evaluation/references/feasibility-scoring.md).
> **Version range**: Framework-agnostic — applies to market entry, partnership, acquisition, and major investment decisions.
> **Generated**: 2026-04-09 — validate probability estimates against current market data; gut-feel priors need external calibration.
---
## Overview
Strategic decisions are made with incomplete information. Risk assessment does not eliminate uncertainty — it makes uncertainty explicit so it can be priced into the decision. The most common failure is a decision that was scored correctly under a best-case scenario but was never tested against realistic downside cases. Pre-mortem analysis, scenario planning, and probability weighting are tools for stress-testing the winning option before commitment. Black swan identification catches the tail risks that structured frameworks miss because they are, by definition, outside the range of normal planning.
---
## Pre-Mortem Analysis Template
Run this AFTER selecting the winning option from the decision matrix, BEFORE committing budget or headcount.
**Setup**: Tell your team "Assume 18 months have passed. This decision failed. Not partially — badly. Revenue is down, a key relationship broke, or we reversed course publicly. What happened?"
```
## Pre-Mortem: [Decision Name]
## Date: [YYYY-MM-DD]
## Decision being stress-tested: [1 sentence description of the chosen option]
### Round 1: Individual brainstorm (silent, 10 minutes)
Each participant writes their top 3 failure modes independently.
(No discussion until all have written — anchoring bias prevention)
---
### Failure Mode Catalog (consolidated)
#### Failure Mode 1
- Scenario: [What happened in the world? Internal failure, external shock, or assumption error?]
- Which assumption was wrong: [The specific belief that was embedded in the original decision]
- Early warning signal: [What we could observe in months 1-3 that signals this is occurring]
- Detection trigger: [Specific metric, event, or threshold that would confirm this path]
- Mitigation: [What we add to the plan NOW to reduce probability or impact]
- Probability (1-10%): [Low/Medium/High or rough % if you have data]
#### Failure Mode 2
- Scenario: ___
- Which assumption was wrong: ___
- Early warning signal: ___
- Detection trigger: ___
- Mitigation: ___
- Probability: ___
#### Failure Mode 3
- Scenario: ___
- Which assumption was wrong: ___
- Early warning signal: ___
- Detection trigger: ___
- Mitigation: ___
- Probability: ___
### High-impact, low-probability failures (see black swan section)
- [Scenario that team said "this will never happen"]
- [Scenario that was immediately dismissed as "too extreme"]
### Pre-mortem summary
- Does any failure mode change the recommended option? [Yes / No]
- If Yes: which option becomes preferred, or what condition must be true before proceeding?
- Mitigations added to execution plan: [numbered list]
- New monitoring metrics added: [numbered list]
```
**Worked example** (SaaS company deciding to expand into EU market):
Failure Mode 1: Regulatory delay.
- Scenario: GDPR compliance requirements are stricter than legal review estimated. Full compliance requires 9 months, not 3. Launch delayed, sales team hired but generating no pipeline.
- Assumption wrong: Legal team used a competitor's compliance timeline without accounting for our data architecture.
- Early warning signal: Month 2 — data residency audit reveals gaps not in original scope.
- Mitigation added: Hire EU-based DPO before hiring sales team, not after.
Failure Mode 2: Currency risk compounds with sales cycle length.
- Scenario: EUR/USD shifts unfavorably during a 6-month enterprise sales cycle. Deals signed in EUR produce 12% less USD revenue than projected. Finance model breaks.
- Assumption wrong: EUR pricing was set at a fixed USD equivalent without currency hedging.
- Mitigation added: EUR contracts include currency adjustment clause at >8% movement.
---
## Scenario Planning Matrix
Use when the decision outcome depends heavily on external conditions that you cannot control.
**Step 1**: Identify the 2-3 most uncertain external variables (not the decision variables — the world variables).
```
## Scenario Matrix: [Decision Name]
### Key Uncertainties (choose the 2 most impactful and uncertain)
- Uncertainty A: [e.g., "Regulatory environment — favorable vs. unfavorable"]
- Uncertainty B: [e.g., "Market adoption speed — fast vs. slow"]
### Four Scenarios
| | Uncertainty A: Favorable | Uncertainty A: Unfavorable |
|-----|--------------------------|---------------------------|
| **Uncertainty B: Fast** | SCENARIO 1: Best case | SCENARIO 2: Regulatory headwind |
| **Uncertainty B: Slow** | SCENARIO 3: Patient build | SCENARIO 4: Worst case |
```
### Scenario Detail Sheets
For each scenario, fill in:
```
## Scenario [N]: [Name]
## Probability estimate: ___%
### What the world looks like
- [Uncertainty A description for this scenario]
- [Uncertainty B description for this scenario]
- [Any secondary effects]
### Our outcome in this scenario
- Revenue impact vs. plan: [+X% / -X% / +$Xm / -$Xm]
- Timeline impact: [ahead by X months / delayed by X months]
- Strategic position: [stronger / weaker / neutral] — why?
### Required adaptations
- [What we do differently if this scenario emerges]
- [Decision triggers: if signal Y appears, execute adaptation Z]
### Acceptable threshold
- Is this scenario acceptable given our constraints? [Yes / No / Conditional]
- If No: what would need to be true to make it acceptable?
```
**Probability check**: All four scenario probabilities must sum to 100%. If you cannot assign probabilities, the uncertainty is not well-defined enough for scenario planning — break it down further.
### Scenario Summary Table
| Scenario | Probability | Year 1 Revenue Impact | Strategic Outcome | Verdict |
|----------|-------------|----------------------|------------------|---------|
| 1: Best case | ___% | $___ | ___ | Accept |
| 2: Regulatory headwind | ___% | $___ | ___ | Accept / Conditional |
| 3: Patient build | ___% | $___ | ___ | Accept / Conditional |
| 4: Worst case | ___% | $___ | ___ | Accept / Reject |
**Decision rule**: If the weighted average outcome (probability × impact) is positive AND the worst-case scenario is survivable (company does not fail, strategic position does not collapse permanently), proceed with the decision. If worst-case is fatal, add structural protections before committing.
---
## Probability-Weighted Outcome Framework
Use when comparing a decision under uncertainty against a safer alternative.
```
## Probability-Weighted Expected Value: [Decision Name]
### Option A: [The strategic decision]
| Outcome | Probability | Value | Weighted Value |
|---------|-------------|-------|----------------|
| Upside (best 25%) | ___% | $___ | $___ |
| Base case | ___% | $___ | $___ |
| Downside (worst 25%) | ___% | $___ | $___ |
| Catastrophic | ___% | $___ | $___ |
| **Expected Value** | 100% | — | **$___** |
### Option B: [The safer alternative or "do nothing"]
| Outcome | Probability | Value | Weighted Value |
|---------|-------------|-------|----------------|
| Upside | ___% | $___ | $___ |
| Base case | ___% | $___ | $___ |
| Downside | ___% | $___ | $___ |
| **Expected Value** | 100% | — | **$___** |
### Comparison
Expected value difference (A - B): $___
Variance difference (A - B): $___ [Option A has higher/lower variance]
Recommendation: [choose based on EV if variance is acceptable; choose B if downside of A is survivability-threatening]
```
**Calibration pitfall**: Teams systematically overestimate "Upside" probability. Apply the outside view: across similar decisions in your industry, what fraction actually hit the upside case? Use that as your prior, then adjust for specifics.
---
## Black Swan Identification
Standard risk frameworks miss low-probability, high-impact events because they focus on the expected range of outcomes. Black swans are definitionally outside that range. This section forces deliberate attention to the tail.
### Black Swan Taxonomy for Strategic Decisions
| Category | Examples | How to Test for It |
|----------|----------|-------------------|
| **Technology discontinuity** | A new technology makes your core product irrelevant or dramatically reduces cost of entry | "What AI/automation capability, if it existed at 10× current performance, would make our strategy moot?" |
| **Regulatory shock** | A new law or enforcement action eliminates or transforms the market | "In which jurisdiction could a single regulator shut down our core model?" |
| **Key dependency failure** | A supplier, platform, or partner stops existing or changes terms dramatically | "What happens if [critical vendor/platform] doubles prices or exits the market?" |
| **Talent concentration** | Loss of 1-2 key people makes the strategy unexecutable | "If [critical person] left tomorrow, could we still execute?" |
| **Market definition shift** | Customers solve the problem in a completely different way | "What non-obvious substitute could eliminate the problem we're solving?" |
| **Macro shock** | Recession, currency crisis, pandemic, or geopolitical event changes the operating environment | "At what macroeconomic scenario does our plan require renegotiation?" |
### Black Swan Assessment Template
```
## Black Swan Log: [Decision Name]
For each category above, complete:
### Category: [Technology Discontinuity / Regulatory / etc.]
- Candidate event: [Specific scenario that would qualify]
- Probability: [Extremely low / Unmeasurable / "Will not happen in our planning horizon"]
- Impact if it occurs: [Catastrophic / Major / Manageable]
- Lead time to respond: [Days / Months / Years]
- Pre-commitment hedge: [Is there a low-cost action we can take now that reduces exposure?]
- Yes: ___
- No: [Accept the risk explicitly — document the conscious choice]
### Black Swan Investment Decision
- Total budget for black swan hedges: $___
(Rule of thumb: 2-5% of decision budget for hedges; more than 5% = the "black swan" is not actually a tail risk)
- Hedges selected: ___
- Hedges explicitly declined: ___ [and why — this is as important as the ones accepted]
```
---
## Risk Register Integration
After pre-mortem, scenario planning, and black swan identification, consolidate into a risk register.
| Risk ID | Description | Probability (1-5) | Impact (1-5) | Risk Score (P×I) | Owner | Mitigation | Status |
|---------|-------------|------------------|--------------|-----------------|-------|------------|--------|
| R01 | ___ | ___ | ___ | ___ | ___ | ___ | Open |
| R02 | ___ | ___ | ___ | ___ | ___ | ___ | Mitigated |
| R03 | ___ | ___ | ___ | ___ | ___ | ___ | Accepted |
**Risk scoring interpretation**:
- Score 1-4: Accept (monitor only)
- Score 5-9: Mitigate (specific action required)
- Score 10-16: Address before committing (blocker if unresolved)
- Score 20-25: Escalate to board level or reverse decision
**Status definitions**:
- Open: Risk identified, no mitigation taken
- Mitigated: Specific action taken that reduces probability or impact
- Accepted: Explicitly decided not to mitigate — risk is within tolerance
- Closed: Risk condition no longer exists
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Pre-mortem as validation exercise
**What it looks like**: Team runs pre-mortem but all proposed failure modes are minor and easily dismissed. The session feels like box-checking.
**Why wrong**: If your pre-mortem does not generate at least one uncomfortable insight — something that changes the plan or makes the team nervous — it was not conducted with intellectual honesty. Good pre-mortems hurt a little.
**Do instead**: Assign the most skeptical person in the room to lead Round 1. Give them explicit permission to name the failure mode everyone is privately thinking but not saying. The facilitator must reject any failure mode that does not name which assumption was wrong.
**Fix**: Require the most skeptical person in the room to lead Round 1. If no one is playing devil's advocate, assign the role explicitly. The facilitator should reject any failure mode that does not include "which assumption was wrong."
### Scenario planning without probability assignment
**What it looks like**: Four beautifully written scenarios with no probabilities. All scenarios treated as equally likely in discussion.
**Why wrong**: If all scenarios are equally likely, the expected value calculation cannot be done. The team defaults to discussing the most vivid scenario (usually worst-case) and either overcorrects or dismisses it.
**Do instead**: Assign probabilities before discussing any scenario's implications. Rough estimates ("~60%, ~20%, ~15%, ~5%") are sufficient. The disagreement that surfaces during assignment is more valuable than the scenarios themselves.
**Fix**: Require probability assignments before discussing scenario implications. Probabilities do not need to be precise — "~60%, ~20%, ~15%, ~5%" is enough. The act of assigning forces explicit disagreement about likelihood, which is more valuable than the scenarios themselves.
### Black swan conflated with low-probability planned risk
**What it looks like**: "The black swan scenario is that we only close 50% of our projected pipeline."
**Why wrong**: 50% pipeline miss is a normal risk, not a black swan. It belongs in the scenario planning matrix. Black swans are outside the planning range, not just the pessimistic end of the normal range.
**Do instead**: Apply the three-part black swan test before labeling any risk as one: it would not appear in a standard risk register, it falls outside the range of outcomes the team has experienced or planned for, and it would require a fundamentally different response rather than an adjusted plan. Everything else goes in the scenario matrix.
**Fix**: For a risk to qualify as a black swan, it must satisfy: (1) would not appear in a standard risk register, (2) is outside the range of outcomes the team has experienced or planned for, and (3) would require a fundamentally different response, not just an adjusted plan.
### Risk register without owners
**What it looks like**: 12-item risk register, all marked "Team" as owner.
**Why wrong**: Team ownership is no ownership. No one monitors the risk; no one detects early warning signals; no one escalates.
**Do instead**: Assign every risk to a single named individual before the register is considered complete. If no one will accept ownership of a high-probability risk, escalate it to the decision-maker immediately — that refusal is itself a signal worth surfacing.
**Fix**: Every risk in the register must have a single named individual as owner. If no one will own it, escalate the risk to decision-maker level immediately — unowned high-probability risks are the most dangerous items in any plan.
---
## Detection Commands Reference
```bash
# Check if pre-mortem was run for a decision (documentation audit)
grep -r "pre-mortem\|premortem\|failure mode" docs/decisions/ 2>/dev/null | head -20
# Check if risk register has unowned risks
# (For structured risk files in CSV format)
python3 -c "
import csv, sys
with open('risk_register.csv') as f:
reader = csv.DictReader(f)
unowned = [r for r in reader if not r.get('Owner') or r['Owner'] in ('Team', 'TBD', '')]
print(f'Unowned risks: {len(unowned)}')
for r in unowned:
print(f' {r[\"Risk ID\"]}: {r[\"Description\"]}')
" 2>/dev/null || echo "No risk_register.csv found — check decision docs"
# Calculate expected value from scenario probabilities
python3 -c "
scenarios = [
('Upside', 0.25, 500000),
('Base', 0.50, 200000),
('Downside', 0.20, -50000),
('Catastrophic', 0.05, -300000),
]
ev = sum(p * v for _, p, v in scenarios)
prob_check = sum(p for _, p, _ in scenarios)
print(f'Expected Value: \${ev:,.0f}')
print(f'Probability sum: {prob_check:.2f} (must be 1.0)')
"
```
---
## See Also
- `decision-matrices.md` — weighted scoring and pre-mortem integration with option selection
- `strategic-frameworks.md` — Porter's Five Forces and SWOT for market-level risk context
- `skills/project-evaluation/references/feasibility-scoring.md` — risk-adjusted feasibility for project execution risk
references/roi-frameworks.md
---
title: ROI Frameworks — Calculation Templates, Effort Estimation, Risk-Adjusted NPV
domain: project-evaluation
level: 3
skill: project-evaluation
---
# ROI Frameworks Reference
> **Scope**: ROI calculation templates, effort estimation methods (T-shirt sizing, three-point estimation, story points), and risk-adjusted NPV analysis for project evaluation. Use when determining whether a project is worth the investment, and how to compare multiple projects.
> **Version range**: Framework-agnostic — applies to software, content, and business projects equally.
> **Generated**: 2026-04-09 — validate hourly rate assumptions against current market rates before using.
---
## Overview
ROI calculations fail most often from underestimating cost (planning fallacy) and overestimating value (optimism bias). The frameworks here combat both: three-point estimation forces pessimistic scenarios into view; risk-adjusted NPV applies a confidence haircut to value claims. The goal is not to predict the future precisely — it is to identify which projects are clearly worth doing, which are clearly not, and which genuinely require more information before deciding.
---
## ROI Calculation Template
Fill this out for every project before the go/no-go verdict.
```
## ROI Analysis: [Project Name]
### Value Delivered
#### Direct Value
- Revenue generated: $___ / quarter (source: ___)
- Cost savings: $___ / quarter (source: ___)
- Time saved × fully-loaded rate: ___ hours/quarter × $___/hour = $___ / quarter
- Annual direct value: $___
#### Indirect Value
- Strategic capability enabled: [describe, estimate order of magnitude if possible]
- Team learning / skill building: [hours × future productivity multiplier]
- Positioning / brand value: [low / medium / high — qualitative]
- Options value (enables future projects): [list what this unlocks]
### Total Cost
#### Build / Implementation Cost
- Engineering effort (from estimation below): ___ hours × $___/hour = $___
- Design effort: ___ hours × $___/hour = $___
- QA / testing: ___ hours × $___/hour = $___
- Infrastructure / tooling: $___
- One-time total: $___
#### Ongoing Cost (Annual)
- Maintenance engineering: ___ hours/month × 12 × $___/hour = $___
- Infrastructure recurring: $___/year
- Support burden: ___ hours/month × 12 × $___/hour = $___
- Annual ongoing total: $___
#### Opportunity Cost
- What the team CANNOT do while working on this: [list 1-2 projects]
- Estimated value of foregone projects: $___
### ROI Calculation
| | Year 1 | Year 2 | Year 3 |
|-|--------|--------|--------|
| Value delivered | $___ | $___ | $___ |
| Implementation cost | -$___ | — | — |
| Ongoing cost | -$___ | -$___ | -$___ |
| Opportunity cost | -$___ | — | — |
| **Net value** | **$___** | **$___** | **$___** |
| **Cumulative** | **$___** | **$___** | **$___** |
Payback period: ___ months (when cumulative net value first turns positive)
3-Year ROI: (Total value - Total cost) / Total cost × 100 = ____%
```
---
## Effort Estimation Methods
### Method 1: T-Shirt Sizing
Fast and useful for early-stage project evaluation or when comparing many projects.
| Size | Hours Range | Definition | Examples |
|------|------------|------------|---------|
| XS | 1-8 hours | Single task, no unknowns, solo work | Fix a bug, write a post, add a field to a form |
| S | 8-40 hours | Small feature or initiative, well-understood | Simple API endpoint, 1-week content sprint |
| M | 40-160 hours | Medium feature, some unknowns | New user flow, data migration, 1-month content series |
| L | 160-400 hours | Large feature, significant complexity | New product vertical, major architecture change |
| XL | 400-1000 hours | Program-level work, multiple unknowns | New product line, major platform migration |
| XXL | 1000+ hours | Full initiative requiring multiple teams | Company pivots, multi-year roadmap items |
**T-shirt sizing rules**:
1. Size against the MVP scope, not the final vision
2. If you cannot place it in a size, it is not scoped well enough yet
3. Anything L or above requires three-point estimation before committing
### Method 2: Three-Point Estimation (PERT)
Use for any project M or larger. Forces consideration of realistic pessimistic scenario.
```
## Three-Point Estimate: [Work Package Name]
Optimistic (O): ___ hours
(Best realistic case — no surprises, team is available, dependencies ready)
Most Likely (M): ___ hours
(What you would bet money on if you had to — moderate friction expected)
Pessimistic (P): ___ hours
(What happens when 1-2 things go wrong — unexpected complexity, key person absent)
PERT estimate: (O + 4M + P) / 6 = ___ hours
Standard deviation: (P - O) / 6 = ___ hours (uncertainty range)
95% confidence range: PERT ± 2σ = ___ to ___ hours
```
**Worked example** (new onboarding flow, mid-level team):
```
O: 40 hours (team knows this codebase, no major surprises)
M: 80 hours (need to refactor old flow, some unknown edge cases)
P: 160 hours (auth library has undocumented behaviors; mobile edge cases)
PERT: (40 + 4×80 + 160) / 6 = 520 / 6 ≈ 87 hours
σ: (160 - 40) / 6 = 20 hours
95% range: 47 to 127 hours
```
Report as: "87 hours most likely, 95% confident it lands in 47-127 hours."
### Method 3: Reference Class Forecasting
Most accurate for software projects. Uses historical data instead of bottom-up estimation.
```
## Reference Class: [Project Type]
Find 3 similar past projects:
1. [Project name]: planned ___ hours, actual ___ hours, ratio: ___
2. [Project name]: planned ___ hours, actual ___ hours, ratio: ___
3. [Project name]: planned ___ hours, actual ___ hours, ratio: ___
Average overrun ratio: ___
(If no history: use 1.5× for known tech/team, 2.0× for novel tech/team)
Current estimate (bottom-up or T-shirt): ___ hours
Reference-class adjusted: ___ hours × overrun ratio = ___ hours
```
---
## Risk-Adjusted NPV Template
Use when comparing projects with different risk profiles. Accounts for probability of success.
```
## Risk-Adjusted NPV: [Project Name]
### Success probability estimate
- Technical risk (failure probability): ___%
- Market risk (lower adoption than expected): ___%
- Execution risk (team/timeline failure): ___%
- Combined success probability (multiply complements):
P(success) = (1 - tech_risk) × (1 - market_risk) × (1 - execution_risk)
P(success) = ___
### Expected value calculation
Base case NPV (from ROI template): $___
P(success) × Base case NPV = $___
Failure cost (sunk cost if project abandoned):
- Resources consumed: $___
- Opportunity cost: $___
P(failure) × Failure cost = $___
Risk-adjusted NPV = Expected value - Failure cost = $___
```
**Comparison table** (when evaluating multiple projects):
| Project | Base NPV | P(success) | Risk-Adj NPV | Effort | RICE Score |
|---------|---------|-----------|--------------|--------|-----------|
| Project A | $50K | 85% | $42.5K | 3 mo | 14.2 |
| Project B | $80K | 50% | $40K | 4 mo | 10.0 |
| Project C | $20K | 95% | $19K | 1 mo | 19.0 |
Project C wins on RICE despite lowest base NPV. Lower risk + shorter effort = highest return per month invested.
---
## Planning Fallacy Mitigation Techniques
| Technique | When to Use | Application |
|-----------|-------------|-------------|
| **Reference class forecasting** | Any project M or larger | Compare against 3+ similar past projects; use their actual-to-planned ratio |
| **Add 50% buffer to novel work** | First time team does this type of project | If team has never done X before, multiply estimate by 1.5 automatically |
| **Separate estimation from planning** | Before timeline commitment | Have someone who did no planning estimate independently; compare |
| **Pre-mortem** | Before final commitment | "We're at launch day and we're 2 months late. What happened?" |
| **Explicit unknowns list** | During estimation | List every assumption. Each assumption is a risk that should inflate the pessimistic estimate. |
---
## Story Points to Hours Conversion (For Agile Teams)
When estimates are in story points and you need hours for ROI calculation.
```
## Velocity Calibration
Team's average velocity: ___ points/sprint
Sprint length: ___ weeks
Team size: ___ engineers
Hours per sprint (total): team_size × sprint_length × 40 hours/week = ___ hours
Hours per story point: total_hours_per_sprint / velocity = ___ hours/point
Story point estimate: ___ points
Converted to hours: ___ points × ___ hours/point = ___ hours
```
**Warning**: Story point conversion is only meaningful for teams with 6+ sprints of stable velocity. New teams or teams with changing composition should not use this conversion.
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Planning fallacy (systematic underestimation)
**What it looks like**: "This will take 2 weeks." Team has never shipped a project of this complexity in under 6 weeks.
**Detection**: Compare the current estimate against historical actuals for similar projects. If this team has a history of 2× overruns, this estimate needs to be doubled before being used in ROI calculations.
**Do instead**: Apply reference class forecasting before finalizing any estimate. Name the 3 most similar past projects and their actual completion times. Use that historical baseline, not the optimistic inside view, as the input to your ROI cost calculation.
**Fix**: Use reference class forecasting. Require the estimator to name the 3 most similar past projects and their actual completion times before finalizing any estimate.
### ROI calculated at fantasy scope
**What it looks like**: ROI analysis uses the full product vision (V3 feature set) but effort estimate uses MVP scope (V1 feature set). Value is dramatically overstated.
**Why wrong**: The ROI analysis and the effort estimate must use the same scope definition.
**Do instead**: Before finalizing, verify that the "value delivered" section and the "build cost" section reference the same project definition. Write the scope label explicitly in both sections so the mismatch is visible if it exists.
**Fix**: Explicitly verify that the "value delivered" section and the "build cost" section reference the same project definition before finalizing.
### Opportunity cost excluded from total cost
**What it looks like**: ROI shows positive return. But the team has 3 other projects that must be delayed to work on this. Those delayed projects have their own value.
**Why wrong**: Every yes is a no to something else. Excluding opportunity cost inflates apparent ROI.
**Do instead**: Add a "what won't get done" row to the cost section of every ROI model. Name the specific projects or work that gets displaced. If you cannot name them, the capacity assumption in the model is wrong.
**Fix**: Always include a "what won't get done" row in the cost section. If you cannot name what gets delayed, the capacity assumption is wrong.
### Ignoring ongoing cost in 3-year model
**What it looks like**: Year 1 implementation cost is $50K. Year 2 and Year 3 show $0 cost because "it's built."
**Why wrong**: Every system requires maintenance. Every piece of content requires updating. Even "done" projects consume support time.
**Do instead**: Apply the 15-20% rule as a floor: Year 2 and Year 3 each carry at minimum 15% of Year 1 implementation cost as ongoing maintenance. For content, use 20-30% of creation cost. Any lower figure requires explicit justification.
**Rule**: Year 2-3 minimum ongoing cost = 15-20% of Year 1 implementation cost for software projects. For content: 20-30% of creation cost for curation and updating.
---
## See Also
- `feasibility-scoring.md` — feasibility scoring models and go/no-go decision framework
- `skills/build-vs-buy/references/tco-framework.md` — TCO analysis for technology procurement decisions
- `skills/strategic-decision/references/decision-matrices.md` — decision matrices for comparing projects
references/sales.md
# Sales
Umbrella skill for sales execution: call preparation, pipeline health analysis, outreach drafting, competitive intelligence, and forecasting. Each mode loads its own references on demand. Detects the mode from the request, loads the right context, and executes the appropriate workflow.
**Scope**: Revenue-facing workflows where the user is preparing for, executing, or following up on sales activities. Use csuite for strategic business decisions. Use research-pipeline for deep multi-source research. Use voice-writer for content generation.
---
## Mode Detection
Classify the request into exactly one mode. If the request spans multiple, choose the primary and note the secondary.
| Mode | Signal Phrases | Primary Activity |
|------|---------------|-----------------|
| **CALL-PREP** | "prep me for", "meeting with", "call with", "get ready for", "before my call" | Research + agenda + questions for an upcoming meeting |
| **PIPELINE** | "pipeline review", "deal health", "stale deals", "pipeline hygiene", "which deals" | Analyze pipeline health, prioritize deals, flag risks |
| **OUTREACH** | "draft outreach", "cold email", "reach out to", "write email to", "LinkedIn message" | Research prospect then draft personalized message |
| **COMPETITIVE** | "competitive intel", "battlecard", "how do we compare", "vs competitor", "competitor research" | Analyze competitors, build positioning, talk tracks |
| **FORECAST** | "forecast", "gap to quota", "commit vs upside", "pipeline coverage", "will I hit my number" | Weighted forecast with scenarios and gap analysis |
| **CALL-SUMMARY** | "call notes", "summarize call", "follow up email", "action items from call", "what happened on the call" | Extract action items, draft follow-up, log summary |
| **RESEARCH** | "research company", "look up", "intel on", "who is", "tell me about" | Company/person research for sales context |
---
## Reference Loading Table
Load only the references required by the detected mode. Always load `references/sales/llm-sales-failure-modes.md` for any mode.
| Mode | Load These References |
|------|----------------------|
| CALL-PREP | `references/sales/call-prep.md`, `references/sales/llm-sales-failure-modes.md` |
| PIPELINE | `references/sales/pipeline-analysis.md`, `references/sales/llm-sales-failure-modes.md` |
| OUTREACH | `references/sales/outreach-patterns.md`, `references/sales/llm-sales-failure-modes.md` |
| COMPETITIVE | `references/sales/competitive-intelligence.md`, `references/sales/llm-sales-failure-modes.md` |
| FORECAST | `references/sales/pipeline-analysis.md`, `references/sales/llm-sales-failure-modes.md` |
| CALL-SUMMARY | `references/sales/call-prep.md`, `references/sales/llm-sales-failure-modes.md` |
| RESEARCH | `references/sales/call-prep.md`, `references/sales/competitive-intelligence.md`, `references/sales/llm-sales-failure-modes.md` |
---
## Workflow: CALL-PREP
**Framework**: GATHER -> RESEARCH -> SYNTHESIZE -> DELIVER
**Phase 1: GATHER** -- Collect meeting context from the user.
Ask for: company name, meeting type (discovery/demo/negotiation/check-in/QBR), attendees (names + titles), user's goal for the call, any context they want to share (paste notes, emails, prior interactions).
Accept whatever they provide. Missing fields become research targets, not blockers.
**Gate**: Company name known. Meeting type classified.
**Phase 2: RESEARCH** -- Web research to fill gaps.
Search for: company news (last 30 days), funding/leadership changes, attendee LinkedIn profiles, company product/service description, industry context. Extract: what the company does, recent trigger events, attendee backgrounds, hiring signals.
Source every company fact from search results. State 'not found' for missing data. Do not invent revenue figures, employee counts, or funding rounds that were not found in search results. See `references/sales/llm-sales-failure-modes.md`.
**Gate**: Company profile assembled from verified sources. Gaps explicitly noted.
**Phase 3: SYNTHESIZE** -- Build the prep brief.
Produce: account snapshot table, attendee profiles with talking points, context/history summary, suggested agenda tailored to meeting type, discovery questions targeting understanding gaps, potential objections with responses.
Meeting type shapes the output:
- **Discovery**: questions > talking. Focus on qualification signals.
- **Demo**: tailored examples for their use case. Focus on technical requirements.
- **Negotiation**: objection handling, value justification. Focus on path to agreement.
- **Check-in/QBR**: value delivered, expansion opportunities. Focus on renewal signals.
**Gate**: All sections populated. No placeholder text. No fabricated details.
**Phase 4: DELIVER** -- Present the formatted brief.
Output a structured markdown brief: Account Snapshot (table), Attendee Profiles, Context & History, Suggested Agenda (numbered), Discovery Questions (5-7), Potential Objections (table: objection | response), Internal Notes.
---
## Workflow: PIPELINE
**Framework**: INGEST -> SCORE -> PRIORITIZE -> DELIVER
**Phase 1: INGEST** -- Get pipeline data.
Accept: CSV upload (preferred), pasted deal descriptions, or verbal pipeline summary. Required fields per deal: name, amount, stage, close date. Helpful: last activity date, owner, primary contact, created date.
If the user describes deals verbally, structure into a deal table before analysis.
**Gate**: Deals structured in tabular format. At minimum: name, amount, stage, close date.
**Phase 2: SCORE** -- Health assessment on four dimensions.
| Dimension | Weight | Red Flag |
|-----------|--------|----------|
| Stage Progression | 25 | Same stage 30+ days |
| Activity Recency | 25 | No activity 14+ days |
| Close Date Accuracy | 25 | Close date in the past |
| Contact Coverage | 25 | Single-threaded (one contact) |
Score each deal on each dimension. Compute pipeline health score (0-100). Identify risk flags: stale deals, stuck deals, past close dates, single-threaded deals, missing data.
Do not fabricate activity dates or deal history the user did not provide. If a field is missing, flag it as a hygiene issue rather than assuming a value.
**Gate**: Health score computed. Risk flags enumerated. Hygiene issues listed.
**Phase 3: PRIORITIZE** -- Rank deals and generate action plan.
Default weighting: Close Date (30%), Deal Size (25%), Stage (20%), Activity (15%), Risk (10%). User can override: "focus on big deals" or "I need quick wins."
Classify deals into: Close This Week, Close This Month, Nurture, Consider Removing.
Generate top 3 priority actions with: deal name, reason it is priority, specific next step, dollar impact.
**Gate**: Deals ranked. Top 3 actions identified. Dead-weight deals flagged.
**Phase 4: DELIVER** -- Output the pipeline review.
Sections: Pipeline Health Score (table), Priority Actions This Week (top 3), Deal Prioritization Matrix (by time horizon), Risk Flags (stale/stuck/past-date/single-threaded tables), Hygiene Issues (table), Pipeline Shape (by stage, by month, by size), Recommendations, Deals to Consider Removing.
---
## Workflow: OUTREACH
**Framework**: RESEARCH -> HOOK -> DRAFT -> DELIVER
**Phase 1: RESEARCH** -- Always research before drafting. Never send generic outreach.
Parse the request: extract person name, company, title/role, email/LinkedIn if provided.
Search for: person background (LinkedIn, bio), company overview, recent news/trigger events, shared connections or interests, company hiring signals.
Must find before drafting: who they are (title, background), what the company does, a personalization hook (trigger event, their content, mutual connection, company initiative).
**Gate**: Target identified. Company understood. At least one genuine personalization hook found.
**Phase 2: HOOK** -- Select the personalization angle.
Priority order:
1. Trigger event (funding, hiring, product launch, news) -- most timely
2. Mutual connection -- social proof
3. Their content (post, article, talk) -- shows research
4. Company initiative -- relevant to their priorities
5. Role-based pain point -- least personal, last resort
Ground every opening in specific research. Reference a finding only true of this prospect. See `references/sales/outreach-patterns.md` for anti-template patterns.
**Gate**: Hook selected with source. Opening line drafted from real research.
**Phase 3: DRAFT** -- Write the message.
Email structure (AIDA): Personalized opening (hook), Interest (their problem in 1-2 sentences), Desire (brief proof point from a similar company), Action (clear, low-friction CTA).
Rules: under 150 words. No markdown formatting (no bold, no headers). Short paragraphs (2-3 sentences). One value prop. One CTA. Plain text that looks natural in any email client.
Also produce: 2-3 subject line alternatives (under 50 chars, no spam words), LinkedIn connection request (under 300 chars, no pitch), follow-up sequence (Day 3, Day 7, Day 14 breakup).
**Gate**: Email under 150 words. No markdown formatting. Opening references specific research. CTA is one clear ask.
**Phase 4: DELIVER** -- Output the outreach package.
Sections: Research Summary (target, hook, goal), Email Draft (subject, body), Subject Alternatives, LinkedIn Message, Why This Approach (table: element | based on), Follow-up Sequence.
---
## Workflow: COMPETITIVE
**Framework**: SCOPE -> RESEARCH -> ANALYZE -> DELIVER
**Phase 1: SCOPE** -- Define the competitive landscape.
Gather: user's company and product, competitors to analyze (1-5), any specific deals where they are competing, known strengths/weaknesses.
If first interaction, ask for seller context. On subsequent invocations, confirm stored context.
**Gate**: Seller context established. Competitor list defined.
**Phase 2: RESEARCH** -- Systematic research per competitor.
For each competitor, search for: product features, pricing model, recent announcements (90 days), product updates/changelog, G2/Capterra reviews, customer base, careers/hiring signals.
Also search: "[Your company] vs [Competitor]" for existing comparisons.
Never fabricate pricing, features, or customer claims. If pricing is not publicly available, state "pricing not publicly listed" rather than guessing. See `references/sales/llm-sales-failure-modes.md`.
**Gate**: Each competitor researched from public sources. No fabricated claims.
**Phase 3: ANALYZE** -- Build competitive positioning.
For each competitor, produce: where they win (with your counter), where you win (with proof points), pricing intelligence, talk tracks (early mention / displacement / late addition), objection handling, landmine questions (questions that expose their weaknesses without badmouthing).
Build a comparison matrix: Feature | You | Competitor 1 | Competitor 2 | ...
Acknowledge where competitors are strong. Credibility comes from honesty about relative strengths.
**Gate**: Comparison matrix complete. Talk tracks per competitor. Landmine questions identified.
**Phase 4: DELIVER** -- Output the competitive analysis.
Sections: Comparison Matrix (feature grid), Per-Competitor Battlecards (profile, differentiators, talk tracks, objections, landmines), Refresh Recommendations.
---
## Workflow: FORECAST
**Framework**: INGEST -> WEIGHT -> SCENARIO -> DELIVER
**Phase 1: INGEST** -- Gather pipeline and targets.
Required: pipeline deals (CSV, pasted, or described), quota number, period end date, already-closed amount.
Required per deal: name, amount, stage, close date. Helpful: last activity date, owner, account name.
**Gate**: Pipeline data structured. Quota and timeline known.
**Phase 2: WEIGHT** -- Apply stage probabilities and risk adjustments.
Default stage probabilities (user can override):
| Stage | Probability |
|-------|------------|
| Closed Won | 100% |
| Negotiation / Contract | 80% |
| Proposal / Quote | 60% |
| Evaluation / Demo | 40% |
| Discovery / Qualification | 20% |
| Prospecting / Lead | 10% |
Risk adjustments: no activity 14+ days (-10%), close date in past (-20%), single-threaded (-10%), stage 30+ days stuck (-15%). These compound.
Classify each deal: Commit (high confidence, would stake forecast on it) or Upside (could close but has risk).
Do not inflate forecast numbers. Better to under-promise than to include deals the user knows are unlikely. See `references/sales/llm-sales-failure-modes.md`.
**Gate**: Weighted forecast calculated. Commit vs upside classified. Risk adjustments applied.
**Phase 3: SCENARIO** -- Build three scenarios and gap analysis.
| Scenario | Method |
|----------|--------|
| Best Case | All deals close as expected |
| Likely Case | Stage-weighted probabilities with risk adjustments |
| Worst Case | Only commit deals close |
Gap analysis: quota minus (closed + likely forecast) = gap. For each gap dollar: identify acceleration candidates, revival candidates, new pipeline needed.
Coverage ratio: open pipeline / remaining quota. Below 2x is risky. 3x is healthy.
**Gate**: Three scenarios calculated. Gap quantified. Coverage ratio stated.
**Phase 4: DELIVER** -- Output the forecast.
Sections: Summary Table (quota, closed, pipeline, weighted, gap, coverage), Forecast Scenarios (table), Pipeline by Stage (table with probability and weighted value), Commit vs Upside (deal-level tables with reasons), Risk Flags (table), Gap Analysis (options to close), Recommendations.
---
## Workflow: CALL-SUMMARY
**Framework**: EXTRACT -> STRUCTURE -> DRAFT -> DELIVER
**Phase 1: EXTRACT** -- Parse call notes or transcript.
Accept: pasted notes (bullet points, rough notes, stream of consciousness), full transcript, or verbal description of what happened.
Extract: attendees (names, titles), call type (discovery/demo/negotiation/check-in), key discussion points, customer priorities stated, objections/concerns raised, competitor mentions, action items (with owners), agreed next steps, deal impact signals.
**Gate**: Key discussion points identified. Action items extracted with owners.
**Phase 2: STRUCTURE** -- Organize into internal summary.
Produce: Call Summary header (company, date, attendees, type), Key Discussion Points, Customer Priorities, Objections/Concerns (with status: addressed/open), Competitive Intel, Action Items (table: owner | action | due date), Next Steps, Deal Impact assessment.
**Gate**: Summary structured. All action items have owners and dates.
**Phase 3: DRAFT** -- Write customer follow-up email.
Rules: plain text only (no markdown, no bold, no headers). Concise. Reference key discussion points. List commitments made. State clear next step with timeline. Professional but not stiff.
Only include commitments the user explicitly stated were made. Do not infer or invent commitments from the summary. If uncertain, ask.
**Gate**: Email is plain text. Under 200 words. Next step stated with date.
**Phase 4: DELIVER** -- Output both artifacts.
Sections: Internal Summary (full structured summary for your team), Customer Follow-Up Email (ready to send), CRM Update Suggestions (stage change, next step field, activity log).
---
## Workflow: RESEARCH
**Framework**: PARSE -> SEARCH -> SYNTHESIZE -> DELIVER
**Phase 1: PARSE** -- Identify the research target.
Classify: company research, person research, competitor research, or pre-meeting research. Extract: company name/domain, person name and title, research purpose.
**Gate**: Target identified. Research type classified.
**Phase 2: SEARCH** -- Systematic web research.
Company searches: homepage, news (last 90 days), funding, careers, product, customers. Person searches: LinkedIn profile, background, recent activity. Domain-based: extract company from domain, then run company searches.
**Gate**: Sources gathered. No fabricated details.
**Phase 3: SYNTHESIZE** -- Build the research profile.
Produce: Quick Take (2-3 sentences: who they are, why they might need you, best angle), Company Profile (table), Recent News (with relevance), Hiring Signals, Key People (with talking points), Qualification Signals (positive/concerns/unknown), Recommended Approach (entry point, opening hook, discovery questions).
Never state employee count, revenue, or funding as fact unless found in search results. Mark uncertain data points explicitly.
**Gate**: Profile complete. All facts sourced. Gaps noted.
**Phase 4: DELIVER** -- Output the research brief.
Sections: Quick Take, Company Profile (table), Recent News, Hiring Signals, Key People (with talking points), Qualification Signals, Recommended Approach, Sources.
---
## LLM Failure Modes
See `references/sales/llm-sales-failure-modes.md` for the complete failure mode catalog (fabricated company details, invented history, generic outreach, hallucinated financials, optimistic forecasting, template personalization, fake competitor claims, over-promising, ignoring disqualification signals). Universal failure modes in `skills/shared-patterns/llm-domain-failure-modes-base.md`.
---
## Output Format Specifications
### All Modes
- Structured markdown with clear headers
- Tables for structured data (company profiles, deal lists, scoring)
- Bullet lists for action items and recommendations
- No motivational framing ("Great news!" / "Exciting opportunity!")
- No atmosphere text. Every sentence carries information.
### Customer-Facing Output (Outreach, Follow-Up Emails)
- Plain text only. No markdown formatting.
- Under 200 words for follow-ups, under 150 for cold outreach.
- Short paragraphs (2-3 sentences max).
- One clear CTA per message.
### Internal Output (Summaries, Analysis, Forecasts)
- Use tables for multi-dimensional data.
- Include "Sources" or "Based On" attribution for factual claims.
- Flag uncertainty explicitly: "Not confirmed" / "Estimate" / "Based on [source]".
references/sales/call-prep.md
# Call Preparation Reference
Deep methodology for sales call preparation. Covers research checklists, meeting-type frameworks, question design, and objection mapping.
---
## Research Checklist
Run this checklist before synthesizing any call prep brief. Each item must be either completed (with source) or explicitly marked "not found."
### Company Research (Required)
| Research Item | Search Query | What to Extract | Red Flag If Missing |
|--------------|-------------|-----------------|---------------------|
| Company overview | "[Company] about" | What they do, who they serve, how they make money | Cannot proceed |
| Industry | "[Company] industry" | Sector, sub-sector, market position | Low-priority gap |
| Company size | "[Company] employees" OR site:linkedin.com | Employee count range, growth trajectory | Note as unknown |
| Recent news | "[Company] news" (last 30 days) | Announcements, press releases, trigger events | Proceed without |
| Funding | "[Company] funding crunchbase" | Last round, amount, investors, date | Proceed without |
| Leadership changes | "[Company] CEO/CTO/CRO appointed" | New executives, departures | Proceed without |
| Product/service | "[Company] product" | Core offering, pricing model, target market | Cannot proceed |
| Hiring signals | "[Company] careers" | Open roles, growth areas, tech stack clues | Proceed without |
| Competitors | "[Company] vs" OR "[Company] alternatives" | Named competitors, positioning | Proceed without |
### Attendee Research (Required per person)
| Research Item | Search Query | What to Extract |
|--------------|-------------|-----------------|
| Current role | "[Name] [Company] LinkedIn" | Title, tenure, responsibilities |
| Background | "[Name] LinkedIn" | Prior companies, education, career trajectory |
| Content | "[Name] [Company] blog/podcast/talk" | Published content, speaking topics, public opinions |
| Mutual connections | Cross-reference with user's network | Shared contacts, shared schools, shared companies |
| Role in deal | Infer from title + meeting context | Decision maker / Champion / Evaluator / Influencer / Blocker |
### Deal Context (From user input only -- never fabricate)
| Context Item | Source | How It Shapes Prep |
|-------------|--------|-------------------|
| Prior interactions | User provides | Reference in agenda, track commitments |
| Open commitments | User provides | List as follow-up items |
| Known objections | User provides | Prepare responses, structure agenda to address |
| Competitor involvement | User provides | Load competitive intelligence reference |
| Budget signals | User provides | Calibrate proposal scope |
| Timeline signals | User provides | Adjust urgency of next steps |
---
## Meeting Type Frameworks
Each meeting type has a different optimal structure. The meeting type determines agenda shape, question ratio, and output emphasis.
### Discovery Call
**Objective**: Understand their world. Qualify the opportunity.
**Agenda structure** (talk/listen ratio: 20/80):
| Phase | Duration | Activity |
|-------|----------|----------|
| Open | 5 min | Build rapport. Reference trigger event or research finding. |
| Context | 5 min | "Tell me about [their situation]. What prompted this conversation?" |
| Pain exploration | 15 min | Deep dive on stated challenges. Ask "why" twice. |
| Impact quantification | 5 min | "What does this cost you? What happens if nothing changes?" |
| Process | 5 min | "Who else is involved? What does your evaluation process look like?" |
| Next steps | 5 min | Propose specific next step with date. |
**Qualification signals to listen for**:
| Signal Type | Positive | Negative |
|------------|----------|----------|
| Budget | Discusses budget ranges, asks about pricing | "We don't have budget for this" |
| Authority | Decision maker present, or clear path to them | "I'd need to check with..." (vague) |
| Need | Describes specific pain with impact | Generic interest, no urgency |
| Timeline | Active project, defined deadline | "Maybe next year" |
| Champion | Advocates for you, shares internal context | Passive information gathering |
**Discovery question design**: Questions should be open-ended, follow a logical sequence (situation -> problem -> impact -> need), and uncover information the prospect hasn't volunteered. Avoid leading questions that telegraph your solution.
| Question Category | Example Pattern | Purpose |
|------------------|-----------------|---------|
| Situation | "Walk me through how you currently handle X" | Understand baseline |
| Problem | "What breaks down when Y happens?" | Identify pain |
| Impact | "What does that cost you in terms of Z?" | Quantify pain |
| Need-payoff | "If you could solve X, what would that unlock?" | Link to value |
| Process | "Who else would need to be involved in evaluating this?" | Map decision process |
| Timeline | "What's driving the timeline on this?" | Assess urgency |
### Demo / Presentation
**Objective**: Show relevant capability. Get feedback. Advance to next stage.
**Agenda structure** (talk/listen ratio: 60/40):
| Phase | Duration | Activity |
|-------|----------|----------|
| Recap | 5 min | Summarize what you learned in discovery. Confirm priorities. |
| Demo: Priority 1 | 10 min | Show the capability that addresses their top pain. Get reaction. |
| Demo: Priority 2 | 10 min | Second use case. Tie to their stated needs. |
| Demo: Priority 3 | 5 min | Third use case or technical depth if requested. |
| Q&A | 10 min | Address questions. Surface objections. |
| Next steps | 5 min | Propose evaluation plan, POC, or proposal. |
**Demo failure modes** (things that kill demos):
| Failure Mode | Why It Fails | Instead |
|-------------|-------------|---------|
| Feature tour | Prospect doesn't care about features they didn't ask for | Show only what maps to their stated needs |
| Happy path only | Feels rehearsed, doesn't build confidence | Show edge cases they will encounter |
| No interaction | Prospect zones out after 10 minutes of watching | Pause every 5 minutes: "Is this relevant to your situation?" |
| Jargon dump | Alienates non-technical attendees | Use their language, not your product vocabulary |
| Skipping setup context | Prospect doesn't understand why you're showing this | Start each section: "You mentioned X. Here's how we handle that." |
### Negotiation / Proposal Review
**Objective**: Address concerns. Justify value. Close gaps to agreement.
**Agenda structure** (talk/listen ratio: 40/60):
| Phase | Duration | Activity |
|-------|----------|----------|
| Confirm proposal | 5 min | "Have you had a chance to review? What questions came up?" |
| Address concerns | 15 min | Work through each objection systematically |
| Value reinforcement | 10 min | Tie back to their ROI, their pain, their timeline |
| Terms discussion | 10 min | Pricing, contract, implementation |
| Path to close | 5 min | "What needs to happen for you to move forward?" |
### Check-in / QBR
**Objective**: Demonstrate value delivered. Surface expansion opportunities.
**Agenda structure** (talk/listen ratio: 50/50):
| Phase | Duration | Activity |
|-------|----------|----------|
| Results review | 10 min | Metrics, wins, value delivered since last QBR |
| Their feedback | 10 min | "What's working? What's not?" |
| Roadmap alignment | 10 min | Upcoming features relevant to their needs |
| Expansion | 5 min | New use cases, additional teams, upsell |
| Next quarter | 5 min | Priorities, success metrics, meeting cadence |
---
## Objection Mapping Framework
Structure objections by category. Prepare responses that acknowledge the concern, reframe, and provide evidence.
### Common Objection Categories
| Category | Example Objections | Response Framework |
|----------|-------------------|-------------------|
| **Price** | "Too expensive" / "Over budget" / "Competitor is cheaper" | Acknowledge -> Reframe as total cost (include switching cost, ramp time, hidden fees) -> Quantify value against their specific pain |
| **Timing** | "Not now" / "Next quarter" / "Too busy" | Acknowledge -> Ask what changes next quarter -> Quantify cost of delay -> Offer phased approach |
| **Status quo** | "Current solution works fine" / "We built our own" | Acknowledge -> Ask about specific pain points you uncovered -> Quantify manual effort / opportunity cost |
| **Authority** | "Need to check with my boss" / "Committee decision" | Acknowledge -> Offer to present to committee -> Ask what their recommendation will be -> Equip champion |
| **Trust** | "Too new" / "Unproven" / "Security concerns" | Acknowledge -> Reference similar customers -> Offer POC -> Provide security documentation |
| **Competition** | "We're also looking at X" | Acknowledge -> Ask what they like about X -> Position on differentiated strengths -> Offer comparison |
### Objection Response Pattern
```
1. ACKNOWLEDGE: "I understand that [their concern]. That's important."
2. CLARIFY: "Can you help me understand specifically what [aspect] concerns you?"
3. REFRAME: "What we've seen with similar companies is [evidence-based reframe]."
4. EVIDENCE: "For example, [Customer X] had the same concern and [outcome]."
5. CHECK: "Does that address your concern, or is there more to it?"
```
Never dismiss an objection. Never argue. The goal is to understand the real concern behind the stated objection.
---
## Post-Call Workflow
After every call, capture:
| Item | Example | Why |
|------|---------|-----|
| Key decisions made | "Agreed to 30-day POC starting Feb 1" | Tracks deal progression |
| Action items (yours) | "Send security questionnaire by Friday" | Commitments to honor |
| Action items (theirs) | "Share technical requirements doc" | Follow-up triggers |
| Objections surfaced | "Concerned about migration timeline" | Shapes next interaction |
| Competitive intel | "Also evaluating Competitor X for reporting" | Informs positioning |
| Next meeting | "Technical deep-dive with engineering team, Feb 10" | Calendar action |
| Deal stage change | "Move from Discovery to Evaluation" | CRM update |
| Risk signals | "VP seemed disengaged after pricing slide" | Early warning |
---
## Attendee Role Classification
Classify each attendee to shape how you engage them.
| Role | Characteristics | How to Engage |
|------|----------------|---------------|
| **Decision Maker** | Signs the contract. Controls budget. | Focus on ROI, risk, strategic fit. Respect their time. |
| **Champion** | Internal advocate. Wants you to win. | Equip with internal talking points. Share ammunition. |
| **Evaluator** | Technical validation. Tests the product. | Deep technical answers. Don't oversell. Honesty builds trust. |
| **Influencer** | Shapes opinion but doesn't decide. | Understand their agenda. Align your pitch to what makes them look good. |
| **Blocker** | Opposes the deal. Has competing priorities. | Acknowledge their concerns directly. Find out what would change their mind. |
| **User** | Will use the product daily. | Show ease of use. Address workflow concerns. Listen to pain points. |
Multiple roles can attend the same meeting. Adjust your agenda to give each role the information they need. If the audience is mixed, lead with the decision maker's priorities and address technical depth in Q&A or a follow-up session.
---
## Information Density Rules for Call Prep Output
| Section | Target Length | Must Include | Must Not Include |
|---------|-------------|-------------|-----------------|
| Account Snapshot | 6-8 row table | Company, industry, size, status, last touch | Unverified financials |
| Attendee Profile | 4-6 lines per person | Name, title, background, talking point | Assumptions about personality |
| Suggested Agenda | 5-7 items | Opening reference, core topics, next steps | Generic "introductions" unless new participants |
| Discovery Questions | 5-7 questions | Gap-targeted questions | Leading questions that telegraph your solution |
| Objection Table | 2-4 rows | Likely objections based on context | Every possible objection |
references/sales/competitive-intelligence.md
# Competitive Intelligence Reference
Competitive analysis framework, battlecard structure, positioning strategy, and landmine question design. The goal: help sellers win deals against specific competitors with evidence-based positioning.
---
## Competitive Research Protocol
### Research Checklist Per Competitor
Run these searches for each competitor. Mark each as completed or "not publicly available."
| Research Area | Search Queries | What to Extract |
|--------------|---------------|-----------------|
| Product | "[Competitor] product features" | Core capabilities, recent additions, gaps |
| Pricing | "[Competitor] pricing" | Model (per seat, usage, flat), tiers, enterprise pricing |
| Positioning | "[Competitor] about" site:competitor.com | How they describe themselves, target market, value props |
| Recent releases | "[Competitor] changelog OR product updates OR releases" (90 days) | What they shipped, direction signals |
| Reviews | "[Competitor] reviews G2 OR Capterra OR TrustRadius" | Customer sentiment, common complaints, praised features |
| Customers | "[Competitor] customers" OR "[Competitor] case study" | Logos, industries, use cases |
| Hiring | "[Competitor] careers" | Growth areas (hiring for X = investing in X) |
| Funding | "[Competitor] funding crunchbase" | Stage, amount, investors, runway signals |
| Comparisons | "[Competitor] vs" | Third-party comparisons, feature matrices |
| Weaknesses | "[Competitor] problems OR issues OR limitations" | Known issues, migration stories |
### Source Quality Hierarchy
| Source | Trust Level | Notes |
|--------|------------|-------|
| Competitor's own website/docs | High for features, low for weaknesses | They don't advertise limitations |
| G2/Capterra verified reviews | High | Real users, verified purchase |
| Independent analyst reports | High | Gartner, Forrester have methodology |
| Customer case studies | Medium | Self-selected success stories |
| Reddit/HN discussions | Medium | Authentic but anecdotal |
| Blog comparisons | Low-Medium | Often affiliate-driven or outdated |
| Your sales team's field intel | High for deal context | But may have confirmation bias |
**Never cite a source you did not find in search results.** If a claim cannot be sourced, mark it as "unverified field intel" or drop it.
---
## Battlecard Structure
One battlecard per competitor. Each battlecard follows this structure.
### 1. Competitor Profile
| Field | Content |
|-------|---------|
| Company | Name, website |
| Founded | Year |
| Funding | Stage + total raised |
| Employees | Count or range |
| Target Market | Who they sell to (size, industry, role) |
| Pricing Model | How they charge |
| Market Position | Leader / Challenger / Niche / Emerging |
### 2. What They Sell
2-3 sentences. What their product does, who it serves, how they position it. Use their own language (from their website) to show you understand their pitch.
### 3. Recent Releases (Last 90 Days)
| Date | Release | Strategic Signal |
|------|---------|-----------------|
| [Date] | [Feature/Product] | [What this tells you about their direction] |
Each release tells you where they're investing. Multiple releases in one area = strategic priority. No releases in an area = potential weakness or deprioritization.
### 4. Where They Win
Be honest. Credibility with prospects comes from acknowledging competitor strengths.
| Area | Their Advantage | Your Counter |
|------|----------------|-------------|
| [Area] | [Specific strength with evidence] | [How to handle when prospect raises it] |
**Counter strategies**:
- **Acknowledge and redirect**: "You're right, they're strong in X. The question is whether X is the critical factor for your use case."
- **Reframe the criteria**: "X matters, but [related capability] is where most teams spend their time."
- **Concede and differentiate**: "They do X well. Where we differ is Y, which [evidence] shows matters more for [their situation]."
### 5. Where You Win
| Area | Your Advantage | Proof Point |
|------|---------------|-------------|
| [Area] | [Specific strength] | [Customer quote, metric, case study] |
Proof points must be real. Verifiable customer results. Named companies (with permission). Specific metrics. If no proof point exists for a claimed advantage, mark it as "advantage without public proof point."
### 6. Pricing Intelligence
| Dimension | Their Model | Your Model | Positioning |
|-----------|-----------|-----------|------------|
| Base pricing | [Their price point or model] | [Your price point] | [How to discuss] |
| Hidden costs | [Implementation, training, add-ons] | [Your all-in cost] | [Surface their hidden costs] |
| Contract terms | [Length, exit clauses] | [Your terms] | [Flexibility advantage if applicable] |
| Discounting | [Known discount patterns] | [Your approach] | [Don't race to bottom] |
If competitor pricing is not publicly available, state "not publicly listed" rather than guessing. Guessed pricing destroys credibility if the prospect knows the real number.
### 7. Talk Tracks
Scenario-based positioning for different deal situations.
**When they come up early in evaluation**:
```
Acknowledge them. Position the evaluation criteria. "Good company to evaluate.
Here's what I'd suggest looking at when you compare us: [your strong dimensions].
Most teams in your situation find [dimension] is what separates the options."
```
**When prospect currently uses them (displacement)**:
```
Don't badmouth. Ask about their experience. "What's working well? Where do you
wish it did more?" Let them articulate the gaps. Then connect those gaps to your
strengths. Never say "they can't do X" -- say "teams that need X usually find..."
```
**When added late to evaluation (you're the incumbent threat)**:
```
Emphasize switching cost, relationship depth, and roadmap alignment. "You've
invested in [your product]. Here's what we're building toward [roadmap]. The
question is whether the delta justifies the migration cost and timeline."
```
### 8. Objection Handling
| Objection | Response |
|-----------|---------|
| "Competitor is cheaper" | [Reframe to TCO: implementation, training, ongoing cost. Or: "What's included at that price?"] |
| "Competitor has [feature]" | [If true: acknowledge and redirect. If partial: clarify scope. If false: correct gently with source.] |
| "Competitor is bigger/more established" | [Reframe: size != fit. Your advantage in [responsiveness/focus/speed].] |
| "We already use Competitor" | [Ask about gaps. Quantify friction. Propose parallel evaluation or phased migration.] |
### 9. Landmine Questions
Questions that expose competitor weaknesses without badmouthing. Ask these during discovery or evaluation to set criteria in your favor.
**Design principle**: A landmine question asks the prospect to evaluate a dimension where you're strong and the competitor is weak. The prospect discovers the gap themselves.
| Your Strength | Their Weakness | Landmine Question |
|--------------|---------------|------------------|
| [Capability] | [Their gap] | "How important is [capability] to your evaluation?" |
| [Performance] | [Their limitation] | "What performance requirements do you have for [dimension]?" |
| [Integration] | [Their ecosystem gap] | "Which tools in your stack need to integrate with this?" |
| [Support] | [Their support model] | "What level of support does your team need during implementation?" |
| [Security] | [Their compliance gap] | "What compliance certifications are required for your organization?" |
**Rules for landmine questions**:
- Frame as discovery questions, not gotchas
- Never mention the competitor in the question
- Let the prospect connect the dots
- Have follow-up questions ready when they answer
- If they don't know, it's still useful: "Worth checking with [Competitor] how they handle this"
---
## Comparison Matrix Design
Build a feature-level comparison grid for the user's reference.
### Matrix Structure
| Category | Capability | You | Competitor A | Competitor B |
|----------|-----------|-----|-------------|-------------|
| [Core] | [Feature 1] | [Status] | [Status] | [Status] |
| [Core] | [Feature 2] | [Status] | [Status] | [Status] |
| [Advanced] | [Feature 3] | [Status] | [Status] | [Status] |
| [Integration] | [Feature 4] | [Status] | [Status] | [Status] |
### Status Values
| Status | Meaning | Display |
|--------|---------|---------|
| Full | Feature exists, mature, no caveats | Green check |
| Partial | Feature exists with limitations | Yellow circle |
| Beta | Feature exists but not production-ready | Orange dot |
| Roadmap | Planned but not shipped | Gray clock |
| None | Not available | Red X |
| Unknown | Cannot determine from public sources | Question mark |
**Honesty rules**:
- Mark your own features honestly (Partial is better than lying about Full)
- Mark competitor features at their strongest interpretation from public sources
- Mark "Unknown" rather than guessing "None"
- Include categories where competitors are stronger
---
## Win/Loss Analysis Framework
When the user has win/loss data (from CRM or experience), analyze for patterns.
### Win Pattern Analysis
| Dimension | What to Analyze |
|-----------|----------------|
| Deal size | Do you win more often in certain deal sizes? |
| Industry | Are certain verticals stronger for you? |
| Buyer persona | Which roles champion you? |
| Entry point | How did the deal start? (Inbound, outbound, referral) |
| Evaluation criteria | Which criteria predict your wins? |
| Competitor | Which competitors do you beat consistently? |
### Loss Pattern Analysis
| Dimension | What to Analyze |
|-----------|----------------|
| Loss reason | Price, features, relationship, timing, status quo? |
| Stage of loss | Early disqualification vs late-stage loss? |
| Competitor | Which competitor wins the deals you lose? |
| Missing capability | What feature/capability gap caused the loss? |
| Process | Did you have access to decision maker? |
### Pattern-to-Action Mapping
| Pattern | Action |
|---------|--------|
| Consistently lose on price to Competitor X | Reframe to TCO. Develop pricing counter-talk track. |
| Win when Champion is [Role] | Target that role in prospecting. |
| Lose when evaluation starts at [feature] | Set evaluation criteria early. Plant landmines before formal eval. |
| Win in [Industry], lose in [Industry] | Focus GTM on strong verticals. Develop weak-vertical playbook. |
---
## Competitive Monitoring Cadence
| Activity | Frequency | What to Check |
|----------|-----------|--------------|
| Product page scan | Monthly | New features, messaging changes |
| Pricing page check | Monthly | Pricing model changes, new tiers |
| G2/Capterra reviews | Monthly | New reviews, trend changes, competitor response |
| Job postings | Monthly | New roles signal investment areas |
| News/press | Weekly (automated) | Announcements, funding, partnerships |
| Changelog/blog | Bi-weekly | Product releases, strategic direction |
| Analyst reports | Quarterly | Market positioning shifts |
| Win/loss review | Quarterly | Pattern updates from recent deals |
---
## Positioning Strategy Frameworks
### Two-Dimensional Positioning Map
Pick two dimensions where you can differentiate. Plot yourself and competitors.
Common dimension pairs:
- Ease of use vs. Power/Depth
- Price vs. Feature completeness
- Speed to deploy vs. Customizability
- Point solution vs. Platform
- SMB-focused vs. Enterprise-focused
The goal: find dimensions where you occupy a unique position. If you're in the same quadrant as a competitor, you're competing on execution, not positioning.
### Category Creation
When you can't win in an existing category, define a new one.
| Existing Category | Category Creation | Positioning Advantage |
|------------------|-------------------|---------------------|
| "CRM" | "Revenue Intelligence Platform" | Shifts from feature comparison to vision |
| "Project Management" | "Work OS" | Broadens scope beyond direct competitors |
| "Monitoring" | "Observability Platform" | Reframes the problem |
Category creation works when: you have a genuine capability expansion, the prospect has needs beyond the existing category, the new category has enough market validation to not seem made up.
Category creation fails when: it's just relabeling the same product, the prospect thinks in the existing category, you can't deliver on the broader promise.
---
## Competitive Intelligence Ethics
| Do | Don't |
|----|-------|
| Use public information | Use stolen/leaked confidential data |
| Acknowledge competitor strengths | Fabricate competitor weaknesses |
| Ask prospects about their experience | Pressure prospects to share competitor pricing |
| Position on your merits | Badmouth competitors by name |
| Cite verified reviews | Cherry-pick unrepresentative reviews |
| Update battlecards with real field data | Invent customer quotes or results |
references/sales/llm-sales-failure-modes.md
# LLM Sales Failure Modes
Where LLMs fail in sales work. This reference catalogs specific failure patterns, their causes, and the guardrails that prevent them. Load this reference for every sales mode.
> **Shared base**: Universal LLM failure modes (hallucination, overconfidence, generic output, arithmetic errors, stale knowledge) are documented in `skills/shared-patterns/llm-domain-failure-modes-base.md`. This file covers sales-specific failures only.
---
## Why Sales Is Especially Dangerous for LLMs
Sales output goes directly to prospects and customers. A fabricated detail in internal documentation is a bug. A fabricated detail in a customer-facing email is a credibility-destroying event that can kill a deal and damage a company's reputation.
Sales work requires:
- **Factual accuracy** about companies, people, and products (LLMs hallucinate these)
- **Specific personalization** that cannot be generic (LLMs default to generic)
- **Numeric precision** in forecasts and financials (LLMs guess and round)
- **Honest uncertainty** about unknowns (LLMs fill gaps with plausible fiction)
- **Tone calibration** per relationship stage (LLMs default to cheerful formality)
Every failure mode below has been observed in production sales AI output.
---
## Failure Mode Catalog
### 1. Company Detail Fabrication
**What happens**: LLM generates plausible-sounding but fictional company details: revenue figures, employee counts, founding dates, office locations, product features, customer lists.
**Why it happens**: The LLM has training data about many companies. It pattern-matches to produce a "typical" company profile that sounds right but may contain details from other companies, outdated information, or pure fabrication.
**Example**:
```
BAD: "Acme Corp, founded in 2015, has 450 employees and $40M ARR."
(None of these numbers were in search results -- LLM fabricated plausible values)
GOOD: "Acme Corp: [founded year not found]. Employee count: approximately 200-500
based on LinkedIn estimate. Revenue: not publicly disclosed."
```
**Guardrail**: Every company fact must have a source (search result, user input, or public filing). State "not found" or "not publicly available" for missing data. Never fill gaps with plausible estimates unless explicitly labeled as estimates.
---
### 2. Invented Relationship History
**What happens**: LLM fabricates prior interactions, meeting history, or relationship context that never occurred.
**Why it happens**: The call prep or pipeline context suggests an existing relationship. The LLM generates "what would have happened" based on typical sales interactions, presenting fiction as history.
**Example**:
```
BAD: "In your last call with Sarah, she mentioned concerns about implementation
timeline and asked about SOC 2 compliance."
(User never provided this information -- LLM invented a plausible prior call)
GOOD: "No prior interaction history provided. First contact assumptions apply."
```
**Guardrail**: Only reference history the user explicitly provided (pasted notes, described interactions, uploaded transcripts). Never assume or generate prior interactions. If the user says "prep me for a follow-up call," ask what happened in the previous interaction.
---
### 3. Generic Outreach Disguised as Personalization
**What happens**: LLM produces email that has the structure of personalized outreach but uses interchangeable content. The "personalization" could apply to any prospect.
**Why it happens**: The LLM learned the pattern of personalized email (specific opening + pain point + proof + CTA) but fills slots with generic content rather than actual research findings.
**Detection test**: Replace the company name and person name with any other prospect. If the email still reads naturally, it's generic.
**Example**:
```
BAD: "Hi Sarah, I noticed you're doing great work at Acme Corp. Companies in
your industry often struggle with scaling their operations efficiently.
We've helped similar companies achieve better results."
GOOD: "Hi Sarah, saw Acme's Q3 announcement about expanding into APAC.
Teams opening new regions usually hit data residency complexity first --
that's where [specific capability] helped [Named Customer] cut their
compliance timeline from 6 months to 6 weeks."
```
**Guardrail**: The opening sentence must reference a fact discovered in research that is only true of this specific person or company. The proof point must name a real customer with a real result. Apply the substitution test before output.
---
### 4. Financial Data Hallucination
**What happens**: LLM generates specific financial figures (revenue, ARR, growth rate, valuation, market cap) that were not found in search results.
**Why it happens**: Financial data appears frequently in training data. The LLM generates numbers that "feel right" for a company of that size/stage. These numbers are frequently wrong by 2-10x.
**Example**:
```
BAD: "With estimated ARR of $25M and 40% year-over-year growth..."
(These numbers were not in any search result -- LLM extrapolated)
GOOD: "Revenue not publicly disclosed. Last known funding: Series B,
$30M in 2023 (source: TechCrunch). Revenue likely in $10-50M range
based on stage, but this is speculative."
```
**Guardrail**: Financial figures only from verified sources: SEC filings, press releases with specific numbers, Crunchbase (funding data), earnings reports. All estimates must be explicitly labeled as estimates with reasoning. Never present an estimate as a fact.
---
### 5. Forecast Optimism Bias
**What happens**: LLM inflates forecast probabilities, includes unlikely deals to make the forecast look healthier, and avoids recommending deal removal.
**Why it happens**: LLMs are trained on helpful, positive-toned text. "This deal is likely to close" pattern-matches better with the training distribution than "This deal should be removed from your pipeline."
**Example**:
```
BAD: "Based on the strong engagement signals, this deal has a good chance
of closing this quarter."
(Deal has no activity in 30 days, close date pushed twice, single-threaded)
GOOD: "This deal has three critical risk factors: 30 days inactive,
close date pushed twice, single contact. Recommend qualifying out
or pushing to next quarter."
```
**Guardrail**: Use standard stage probabilities with risk adjustments (see pipeline-analysis.md). Risk adjustments only reduce probability, never increase it. Always flag deals for removal when evidence warrants. Include "Deals to Consider Removing" section in every pipeline review.
---
### 6. Template-Shaped Email Output
**What happens**: LLM uses markdown formatting in customer-facing emails: **bold text**, # headers, * bullet points. These render as literal characters in email clients.
**Why it happens**: Markdown is the LLM's native formatting language. It defaults to markdown for all structured output, including emails.
**Example**:
```
BAD:
**What we discussed:**
- **Budget**: $50K allocated
- **Timeline**: Q2 implementation
- **Next steps**: Schedule technical review
GOOD:
Here's what we discussed:
- Budget: $50K allocated
- Timeline: Q2 implementation
- Next steps: Schedule technical review
```
**Guardrail**: All customer-facing email output must use plain text only. No asterisks, no hash symbols, no square brackets, no formatted bullet points. Use plain dashes for lists. Test: copy the output into a plain text editor. Does it look like a normal email? If you see * or # characters, reformat.
---
### 7. Fabricated Competitive Claims
**What happens**: LLM invents competitor pricing, features, limitations, or customer claims that are not backed by search results.
**Why it happens**: The LLM has some training data about competitors but not enough to be accurate. It fills gaps with plausible-sounding claims that may be outdated, incorrect, or fabricated.
**Example**:
```
BAD: "Competitor X charges $50/seat/month and doesn't support SSO on
their standard plan."
(Neither claim was verified in search results)
GOOD: "Competitor X: Pricing not publicly listed (enterprise quote required).
SSO support: Available on Enterprise tier per their pricing page
(source: competitor.com/pricing, accessed 2026-05-01)."
```
**Guardrail**: Every competitive claim must link to a source. If the source is a search result, note the URL and access date. If the source is user-provided field intel, label it as such. If a claim cannot be sourced, either drop it or mark it as "unverified -- confirm before using in customer conversations."
---
### 8. Over-Promising in Follow-Up Emails
**What happens**: LLM includes commitments in follow-up emails that the user did not authorize: deliverables, timelines, introductions, discounts.
**Why it happens**: The LLM generates what a "good follow-up" looks like based on training data. Good follow-ups include specific commitments. The LLM generates plausible commitments that the seller may not have made.
**Example**:
```
BAD: "As discussed, I'll send over a custom pricing proposal by Friday
and schedule a call with our CTO for a technical deep-dive."
(User only committed to sending a case study)
GOOD: "As discussed, I'll send over the [Industry] case study this week.
Let me know if a technical deep-dive with our engineering team would
be useful as a next step."
```
**Guardrail**: Follow-up emails only include commitments the user explicitly stated were made. If the user described the call but didn't specify commitments, ask: "What did you commit to sending or doing?" Do not infer commitments from the conversation summary.
---
### 9. Ignoring Disqualification Signals
**What happens**: LLM always finds reasons to pursue a deal, never recommends walking away or disqualifying.
**Why it happens**: The LLM's training data rewards helpfulness and optimism. "We should pursue this" matches the helpful pattern better than "Walk away from this deal." Sales training materials emphasize persistence.
**Example**:
```
BAD: "While there are some challenges, this deal still has potential
and I recommend continued engagement."
(Prospect explicitly said they have no budget, no timeline, and are
under contract with a competitor for 2 more years)
GOOD: "Disqualification recommended. Three disqualifying factors:
1. No budget allocated
2. No timeline for evaluation
3. Locked into competitor contract through 2028
Recommend: Add to nurture list. Re-engage 6 months before contract
renewal (early 2028)."
```
**Guardrail**: Define disqualification criteria and apply them without hesitation. When 2+ disqualifying factors are present, recommend disqualification first, with a path to re-engage later. Disqualifying a bad deal is more valuable than wasting time on it.
**Disqualification criteria**:
| Factor | Threshold |
|--------|----------|
| Budget | No budget AND no path to budget |
| Authority | Cannot access decision maker after 2 attempts |
| Need | No identified pain point after discovery |
| Timeline | No evaluation window in next 6 months |
| Fit | Product does not solve their stated problem |
| Competition | Locked into competitor contract with no exit |
| Engagement | No response after full follow-up sequence |
---
### 10. Tone Miscalibration
**What happens**: LLM uses the same tone regardless of deal stage, relationship depth, or prospect seniority.
**Why it happens**: The LLM defaults to a "professional but friendly" tone that works for initial outreach but feels wrong at other stages.
| Stage | Correct Tone | LLM Default | Problem |
|-------|-------------|-------------|---------|
| Cold outreach | Direct, value-focused | Over-formal | Feels stiff, unnatural |
| Discovery | Curious, consultative | Solution-pushing | Premature pitching |
| Negotiation | Confident, flexible | Eager to please | Undermines pricing power |
| Follow-up | Brief, specific | Cheerful, wordy | Wastes prospect's time |
| Re-engagement | Low-pressure, new value | Guilt-inducing | "Just checking in" pattern |
| Executive comms | Concise, strategic | Feature-detailed | Misses the altitude |
**Guardrail**: Calibrate tone to deal stage and audience seniority. Executives get shorter, more strategic communication. Technical evaluators get deeper, more precise language. Early-stage prospects get value-first, question-driven outreach. Late-stage prospects get direct, action-oriented communication.
---
## Meta-Guardrails
These rules apply across all sales modes.
### The Verification Rule
Before outputting any factual claim about a company, person, product, or financial metric, answer: "Where did I learn this?" If the answer is "from my training data" or "it seems likely," do not present it as fact. Either:
1. Search for verification
2. Label it as an assumption
3. Omit it
### The Substitution Test
For personalized content (outreach, call prep): Replace the prospect's name and company with any other. If the content still reads naturally, the personalization is fake. Rewrite.
### The Commitment Audit
For follow-up emails: List every commitment in the email. Cross-reference against what the user explicitly stated. Remove any commitment not confirmed by the user.
### The Negativity Test
For pipeline reviews and forecasts: Count the number of positive vs. negative assessments. If every deal gets a positive assessment, the analysis is biased. Force at least one "consider removing" recommendation per review.
### The Source Test
For competitive intelligence: Every claim about a competitor must have a source notation. "Source: competitor.com/pricing" or "Source: G2 review, Oct 2025" or "Source: user-provided field intel." Claims without sources get marked "unverified" or removed.
---
## Failure Mode Quick Reference
| # | Failure Mode | One-Line Prevention |
|---|-------------|-------------------|
| 1 | Company detail fabrication | Every fact needs a source. "Not found" is better than wrong. |
| 2 | Invented relationship history | Only reference what the user explicitly provided. |
| 3 | Generic outreach | Substitution test: if another prospect fits, rewrite. |
| 4 | Financial hallucination | Only from SEC filings, press releases, or Crunchbase. Label estimates. |
| 5 | Forecast optimism | Standard probabilities. Risk adjustments only decrease. |
| 6 | Markdown in emails | Plain text only. No asterisks, no headers. |
| 7 | Fake competitive claims | Every claim needs a source URL or "unverified" label. |
| 8 | Over-promising in follow-ups | Only include commitments the user confirmed. |
| 9 | Ignoring disqualification | 2+ disqualifying factors = recommend walking away. |
| 10 | Tone miscalibration | Match tone to deal stage and audience seniority. |
references/sales/outreach-patterns.md
# Outreach Patterns Reference
Personalized outreach frameworks, anti-template patterns, channel selection, and follow-up sequences. The core principle: research first, draft second. Generic outreach is noise.
---
## The Anti-Template Principle
Templates are the enemy of effective outreach. A template-shaped email triggers the prospect's spam filter (mental and literal) because it reads like every other sales email.
**The test**: Could this email opening apply to any prospect at any company? If yes, it is a template. Rewrite it.
| Template-Shaped (Bad) | Research-Driven (Good) |
|----------------------|----------------------|
| "I hope this email finds you well" | [Delete this line entirely] |
| "I noticed you work at Company" | "Saw Company's Series B announcement last week" |
| "Congrats on your new role" | "Congrats on the VP Eng role -- saw you came from Stripe's platform team" |
| "I'm reaching out because we help companies like yours" | "Your CTO mentioned scaling challenges in that Podcast X interview" |
| "We're an AI-powered platform that helps..." | [Lead with their problem, not your product] |
| "I wanted to introduce myself" | [Nobody cares who you are yet. Lead with value.] |
---
## Personalization Hierarchy
Hooks ranked by effectiveness. Use the highest-tier hook available.
### Tier 1: Trigger Events (Most Effective)
A trigger event is a change that creates urgency or relevance.
| Trigger | Why It Works | Example Opening |
|---------|-------------|----------------|
| Funding round | New capital = new priorities, new hires, new projects | "With the Series C closing last month, curious how you're thinking about [topic]" |
| Leadership hire | New leader = new direction, eager to make impact | "Saw you joined as VP Engineering -- teams I work with in their first 90 days often..." |
| Product launch | Signals priorities and investment areas | "Noticed the launch of [Product Feature]. Teams scaling [related capability] often..." |
| Acquisition | Integration creates new problems | "With the [Company] acquisition closing, integration planning usually surfaces..." |
| Earnings mention | Public signal of strategic priorities | "In last quarter's earnings, [CEO] mentioned [priority]. That's where..." |
| Job posting | Hiring for X = investing in X | "Saw you're hiring [X]. Teams building that capability often need..." |
### Tier 2: Mutual Connections
| Connection Type | Opening Pattern |
|----------------|----------------|
| Shared contact | "[Name] mentioned you'd be a good person to talk to about [topic]" |
| Shared company | "We overlapped at [Company] -- I was in [department] from [years]" |
| Shared school | Brief mention only if natural. Don't force it. |
| Shared community | "Enjoyed your post in [community/forum] about [topic]" |
### Tier 3: Their Content
| Content Type | Opening Pattern |
|-------------|----------------|
| Blog post | "Your piece on [topic] resonated -- especially [specific point]" |
| Podcast appearance | "Caught your episode on [podcast]. Your point about [X] stuck with me" |
| Conference talk | "Saw your talk at [event] on [topic]. The point about [X] was spot on" |
| LinkedIn post | "Your post about [topic] sparked a thought about [related angle]" |
### Tier 4: Company Initiative
| Signal | Opening Pattern |
|--------|----------------|
| Strategic initiative | "Noticed [Company] is investing in [initiative]" |
| Industry trend | "Teams in [industry] are dealing with [trend]. Curious how it's affecting [Company]" |
| Regulatory change | "[Regulation] is changing the game for [industry]. Wondering how you're approaching..." |
### Tier 5: Role-Based Pain (Last Resort)
Use only if Tiers 1-4 yield nothing. Least personal but still better than no personalization.
| Role | Common Pain | Opening Pattern |
|------|-------------|----------------|
| VP Engineering | Team productivity, technical debt, hiring | "Engineering leaders I talk to are spending [X%] of time on [pain]" |
| CTO | Architecture decisions, scaling, security | "CTOs scaling from [stage] to [stage] often hit [specific challenge]" |
| VP Sales | Pipeline quality, forecast accuracy, rep productivity | "Sales leaders in [industry] are seeing [specific trend]" |
| CFO | Spend visibility, vendor consolidation, ROI justification | "Finance teams consolidating vendors often discover [insight]" |
---
## Email Construction Rules
### Structure: AIDA (Attention, Interest, Desire, Action)
| Section | Purpose | Max Length | Rule |
|---------|---------|-----------|------|
| Subject line | Get the email opened | Under 50 chars | No spam words (free, guaranteed, act now). Personalized. |
| Opening | Prove you researched them | 1-2 sentences | Must reference specific research finding |
| Interest | Connect to their problem | 1-2 sentences | Their challenge, not your product |
| Desire | Brief proof point | 1 sentence | Similar company, specific result |
| CTA | One clear ask | 1 sentence | Low-friction. Question format. |
Total email: under 150 words. 5 sentences is ideal. 7 sentences is the hard max.
### Subject Line Patterns
| Pattern | Example | When to Use |
|---------|---------|-------------|
| [Their initiative] + [angle] | "Notion's AI scaling + a thought" | Trigger event available |
| Question format | "How [Company] handles [X]?" | Research reveals gap |
| Mutual connection | "[Name] suggested I reach out" | Warm intro available |
| Specific + short | "Re: your [conference] talk" | Content-based hook |
**Subject line failure modes** (cause low open rates):
| Bad Subject | Why It Fails |
|------------|-------------|
| "Quick question" | Overused. Feels manipulative. |
| "Touching base" | No value signal. Screams sales. |
| "Partnership opportunity" | Vague. Every spam email says this. |
| "I'd love to connect" | Nobody opens this. |
| "[Company] + [Your Company]" | Only works if your brand is known to them |
| ALL CAPS or excessive punctuation | Spam filter trigger |
### Formatting Rules
| Rule | Reason |
|------|--------|
| No markdown (no **bold**, no *italic*) | Renders as literal asterisks in many email clients |
| No HTML headers or bullet formatting | Plain text looks natural, formatted text looks templated |
| Short paragraphs (2-3 sentences) | Mobile readability. Most email is read on phones. |
| Plain dashes for lists | Renders correctly everywhere |
| No images or logos in cold outreach | Triggers spam filters. Adds load time. |
| No tracking pixels in first touch | Builds trust. Some prospects notice tracking. |
---
## Channel Selection Logic
```
1. Verified email available?
YES -> Email preferred (higher response rate for cold outreach)
Also prepare LinkedIn backup for follow-up
NO -> Go to step 2
2. LinkedIn profile found?
YES -> LinkedIn connection request (no pitch)
Follow-up message template for after connection
NO -> Go to step 3
3. Warm intro possible?
YES -> Suggest mutual connection outreach first
Draft message for the connector
NO -> Company website contact form (lowest priority)
```
### Channel-Specific Rules
| Channel | Length Limit | Tone | Structure |
|---------|-------------|------|-----------|
| Cold email | 150 words | Professional, direct | AIDA |
| LinkedIn connection request | 300 chars | Casual, no pitch | Personal hook + "would love to connect" |
| LinkedIn follow-up message | 200 words | Conversational | Value-first, then soft transition to why |
| Warm intro request | 100 words | Brief for the connector | Context for connector + what you'd like to discuss |
---
## Follow-Up Sequence Design
### Timing
| Touch | Timing | Purpose |
|-------|--------|---------|
| Initial email | Day 0 | Full AIDA outreach |
| Follow-up 1 | Day 3 | New angle. Not "just checking in." |
| Follow-up 2 | Day 7 | Different value prop or proof point |
| Follow-up 3 | Day 14 | Break-up email (creates urgency through scarcity) |
| LinkedIn touch | Day 5-10 | Parallel channel. Connection request or engage with their content. |
### Follow-Up Failure Modes
| Bad Follow-Up | Why It Fails | Instead |
|--------------|-------------|---------|
| "Just following up" | No new value. Feels like nagging. | Add new information, angle, or resource. |
| "Bumping this to the top of your inbox" | Patronizing. They saw it. | Offer a different reason to engage. |
| "Did you get my last email?" | Passive-aggressive. | New angle, acknowledge they're busy. |
| Same message resent | Lazy. Obviously automated. | Each follow-up must add new value. |
| "I know you're busy but..." | Undermines your importance. | Lead with value, not apology. |
### Follow-Up Templates by Touch
**Day 3 -- New Angle**:
```
Hi [Name],
[New piece of information: recent article, case study, data point].
Thought this might be relevant given [their situation].
[Same CTA from original email, rephrased].
[Signature]
```
**Day 7 -- Different Proof**:
```
Hi [Name],
[Different customer story or result]. They were dealing with
[similar challenge to prospect].
Happy to share how they approached it if useful.
[Signature]
```
**Day 14 -- Break-Up**:
```
Hi [Name],
Haven't heard back, so I'll assume the timing isn't right.
If [topic] becomes a priority, happy to reconnect.
[Signature]
```
The break-up email consistently has the highest response rate in the sequence because it removes pressure.
---
## Scenario-Specific Templates
### Cold Outreach (No Prior Relationship)
```
Subject: [Their initiative] + [your angle]
Hi [Name],
[Personal hook from Tier 1-3 research].
[1 sentence on their likely challenge based on role/company].
[Brief proof: "We helped [Similar Company] achieve [Result]".]
Worth a 15-min call to see if relevant?
[Signature]
```
### Warm Outreach (Mutual Connection / Prior Meeting)
```
Subject: Following up from [context]
Hi [Name],
[Reference how you know them or who connected you].
[Why reaching out now -- their trigger event].
[Specific value you can offer].
[Low-friction CTA]
[Signature]
```
### Re-Engagement (Went Dark)
```
Subject: [Short, curiosity-driven]
Hi [Name],
[Acknowledge time passed. No guilt.]
[New reason to reconnect -- their news or your news].
[Simple question to reopen dialogue]
[Signature]
```
### Post-Event Follow-Up
```
Subject: Great meeting you at [Event]
Hi [Name],
[Specific detail from your conversation -- proves it's not mass-sent].
[Value-add: article, intro, or resource related to what you discussed].
[Soft CTA for next conversation]
[Signature]
```
### Referral Request
```
Subject: Quick ask
Hi [Name],
[Mention shared success or positive experience].
Working with a few other [industry] companies on [topic].
Anyone in your network dealing with [specific challenge]?
Happy to return the favor.
[Signature]
```
---
## Outreach Quality Checklist
Run this checklist on every drafted email before presenting to the user.
| Check | Pass Criteria |
|-------|-------------|
| Opening is specific | References a fact only true of this prospect |
| Under word limit | Cold: under 150. Follow-up: under 100. |
| One CTA | Exactly one ask. Not two. Not zero. |
| No markdown | No asterisks, no headers, no bullet formatting |
| No spam words | No "free," "guaranteed," "limited time," "act now" |
| No feature dump | Product mentioned in one sentence max |
| Subject under 50 chars | Short, specific, no spam triggers |
| Plain text formatting | Would look natural pasted into Gmail |
| No fake personalization | Nothing that could apply to any prospect |
| Proof point is real | Customer story or result is verifiable |
---
## Response Rate Benchmarks
For calibrating expectations with the user.
| Metric | Cold Email | LinkedIn | Warm Intro |
|--------|-----------|----------|-----------|
| Open rate (email) | 30-50% | N/A | 60-80% |
| Reply rate | 5-15% | 10-25% | 30-50% |
| Meeting booked rate | 2-5% | 5-10% | 15-25% |
| Positive reply rate | 1-3% | 3-8% | 10-20% |
These are baselines for well-researched, personalized outreach. Mass-blast templated email performs 3-5x worse.
references/sales/pipeline-analysis.md
# Pipeline Analysis Reference
Pipeline health scoring, deal prioritization, forecast methodology, and risk detection. Used by both PIPELINE and FORECAST modes.
---
## Pipeline Health Scoring Model
Score pipeline on four dimensions. Each dimension is 0-25 points. Total health score is 0-100.
### Dimension 1: Stage Progression (0-25)
Measures whether deals are moving through stages at a healthy velocity.
| Scoring Rule | Points |
|-------------|--------|
| No deals stuck in same stage 30+ days | 25 |
| 1-2 deals stuck 30+ days | 20 |
| 3-5 deals stuck 30+ days | 15 |
| 6-10 deals stuck 30+ days | 10 |
| 10+ deals stuck OR more than 50% of pipeline stuck | 5 |
**Expected stage velocity** (industry baseline, B2B SaaS):
| Stage Transition | Healthy Duration | Warning | Critical |
|-----------------|-----------------|---------|----------|
| Prospecting -> Discovery | 7-14 days | 21 days | 30+ days |
| Discovery -> Evaluation | 14-21 days | 30 days | 45+ days |
| Evaluation -> Proposal | 7-14 days | 21 days | 30+ days |
| Proposal -> Negotiation | 7-14 days | 21 days | 30+ days |
| Negotiation -> Closed | 7-21 days | 30 days | 45+ days |
These are defaults. Enterprise deals move slower. SMB faster. Ask the user about their typical cycle length and adjust.
### Dimension 2: Activity Recency (0-25)
Measures engagement momentum across the pipeline.
| Scoring Rule | Points |
|-------------|--------|
| All deals have activity within 7 days | 25 |
| 1-2 deals with no activity 14+ days | 20 |
| 3-5 deals silent 14+ days | 15 |
| 6-10 deals silent 14+ days | 10 |
| 10+ deals silent OR more than 40% silent | 5 |
**Activity types that count** (in descending signal strength):
| Activity | Signal Strength | Notes |
|----------|----------------|-------|
| Customer-initiated email/call | Very strong | They are engaged |
| Scheduled meeting occurred | Strong | Active dialogue |
| Seller-initiated email with reply | Moderate | Two-way communication |
| Seller-initiated email, no reply | Weak | One-way, may be ghosting |
| Internal note only | None | Does not count as customer activity |
### Dimension 3: Close Date Accuracy (0-25)
Measures whether close dates reflect reality or wishful thinking.
| Scoring Rule | Points |
|-------------|--------|
| No deals with close date in the past | 25 |
| 1-2 deals past close date | 20 |
| 3-5 deals past close date | 15 |
| 6-10 deals past close date | 10 |
| 10+ past OR more than 30% past | 5 |
**Close date integrity signals**:
| Signal | Interpretation |
|--------|---------------|
| Close date pushed 3+ times | Deal may be zombie. Qualify hard. |
| Close date matches end-of-quarter | Often placeholder, not real commitment |
| Close date within 2 weeks but no meeting scheduled | Unlikely to close on time |
| Close date aligns with stated event (contract renewal, budget cycle) | Legitimate anchor |
### Dimension 4: Contact Coverage (0-25)
Measures multi-threading -- the number of contacts engaged per deal.
| Scoring Rule | Points |
|-------------|--------|
| All deals have 2+ active contacts | 25 |
| 1-3 deals single-threaded | 20 |
| 4-6 deals single-threaded | 15 |
| 7+ deals single-threaded | 10 |
| More than 50% single-threaded | 5 |
**Why single-threading kills deals**:
- Champion leaves company: deal dies immediately
- Champion goes on leave: deal stalls with no alternate path
- Champion loses internal influence: no backup advocate
- Champion's priorities shift: nobody else carries your case
**Multi-threading targets by deal size**:
| Deal Size | Minimum Contacts | Ideal |
|-----------|-----------------|-------|
| Under $25K | 1 (acceptable) | 2 |
| $25K-100K | 2 | 3-4 |
| $100K-500K | 3 | 5+ |
| $500K+ | 4 | 7+ across departments |
---
## Deal Prioritization Framework
### Default Weighting
| Factor | Weight | Scoring Method |
|--------|--------|---------------|
| Close Date | 30% | Inverse days to close: deals closing soonest score highest |
| Deal Size | 25% | Normalized to largest deal in pipeline |
| Stage | 20% | Later stage = higher score |
| Activity | 15% | Days since last activity (inverse) |
| Risk | 10% | Composite of risk flags (fewer = higher) |
### Alternative Weightings (User-Triggered)
| User Says | Adjustment |
|-----------|-----------|
| "Focus on big deals" | Deal Size -> 40%, Close Date -> 20% |
| "I need quick wins" | Close Date -> 40%, Stage -> 25%, Deal Size -> 15% |
| "Fix my pipeline hygiene" | Risk -> 30%, Activity -> 25%, Close Date -> 20% |
| "Help me hit my number" | Close Date -> 35%, Deal Size -> 30%, Stage -> 20% |
### Time Horizon Classification
| Category | Criteria | Action Required |
|----------|---------|-----------------|
| **Close This Week** | Close date within 7 days AND stage >= Proposal | Focus time. Daily action. Remove blockers. |
| **Close This Month** | Close date within 30 days AND stage >= Evaluation | Keep warm. Weekly touchpoint. Track progress. |
| **Nurture** | Close date 30-90 days OR stage <= Discovery | Periodic check-in. Build relationship. Share value. |
| **Consider Removing** | Past close date 30+ days, no activity 30+ days, pushed 3+ times | Qualify out or mark closed-lost. Pipeline inflation hurts forecasting. |
---
## Forecast Methodology
### Stage-Weighted Forecast
Base calculation per deal: `Amount * Stage Probability * Risk Adjustment = Weighted Value`
Default stage probabilities:
| Stage | Base Probability | Notes |
|-------|-----------------|-------|
| Closed Won | 100% | Already booked |
| Verbal Commit | 90% | Verbal yes, awaiting paperwork |
| Negotiation / Contract | 80% | Terms being discussed |
| Proposal / Quote | 60% | Proposal delivered, awaiting response |
| Evaluation / Demo | 40% | Active evaluation |
| Discovery / Qualification | 20% | Exploring fit |
| Prospecting / Lead | 10% | Early stage, low confidence |
### Risk Adjustments
Apply after base probability. These compound multiplicatively.
| Risk Factor | Adjustment | Detection |
|------------|-----------|-----------|
| No activity 14+ days | -10% (multiply by 0.90) | Last activity date vs today |
| No activity 30+ days | -25% (multiply by 0.75) | Last activity date vs today |
| Close date in past | -20% (multiply by 0.80) | Close date < today |
| Close date pushed 3+ times | -30% (multiply by 0.70) | Requires user confirmation |
| Single-threaded | -10% (multiply by 0.90) | Only one contact |
| No next step defined | -15% (multiply by 0.85) | Missing from deal record |
| Champion left company | -50% (multiply by 0.50) | Requires user confirmation |
| Competitor displacement | -15% (multiply by 0.85) | Incumbent competitor present |
**Example**: $100K deal in Proposal stage, no activity 14 days, single-threaded.
Base: 60% -> Risk: 0.60 * 0.90 * 0.90 = 0.486 -> Weighted: $48,600.
### Commit vs. Upside Classification
| Category | Criteria | Forecast Treatment |
|----------|---------|-------------------|
| **Commit** | Stage >= Negotiation AND no critical risk flags AND rep would bet on it | Include in worst-case scenario |
| **Best Case** | Stage >= Evaluation AND active engagement | Include in best-case only |
| **Upside** | Everything else in pipeline | Exclude from forecast, note for pipeline generation |
Rules for commit classification:
- Never commit a deal with close date in the past
- Never commit a deal with no activity 30+ days
- Never commit a deal the user says has low confidence
- Ask the user: "Would you bet your forecast on this deal?" If hesitation, it's upside.
### Three-Scenario Model
| Scenario | Formula | Use |
|----------|---------|-----|
| **Best Case** | Sum of all deal amounts at stage probability (no risk adjustment) | Upper bound. Optimistic. |
| **Likely Case** | Sum of weighted values (stage probability * risk adjustments) | Planning target. Most realistic. |
| **Worst Case** | Sum of Commit deals only | Floor. What you can count on. |
### Gap Analysis
```
Gap = Quota - Closed to Date - Likely Case Forecast
Coverage Ratio = Open Pipeline / (Quota - Closed to Date)
```
| Coverage Ratio | Assessment |
|---------------|-----------|
| 4x+ | Strong. Focus on execution, not pipeline generation. |
| 3x | Healthy. Standard coverage for B2B. |
| 2x-3x | Tight. Need to accelerate deals and maintain pipeline generation. |
| 1x-2x | At risk. Pipeline generation is urgent. Every deal matters. |
| Below 1x | Critical. Cannot hit target from existing pipeline alone. |
### Gap-Closing Strategies
| Strategy | When to Use | Expected Impact |
|----------|------------|----------------|
| **Accelerate** | Deals in late stage that can close faster | Pull in revenue from future periods |
| **Upsize** | Existing deals with expansion potential | Increase deal value without new pipeline |
| **Revive** | Stalled deals with prior engagement | Lower cost than new pipeline |
| **Generate** | Coverage below 2x | New pipeline at early stages |
| **Negotiate** | Quota discussion with management | Adjust target if pipeline is structurally thin |
---
## Pipeline Shape Analysis
Healthy pipelines have a funnel shape -- more deals in early stages, fewer in late stages. Inverted funnels signal pipeline generation problems.
### Ideal Distribution (B2B SaaS)
| Stage Group | % of Total Deals | % of Total Value |
|------------|-----------------|------------------|
| Early (Prospecting, Discovery) | 40-50% | 20-30% |
| Mid (Evaluation, Proposal) | 30-35% | 35-45% |
| Late (Negotiation, Contract) | 15-25% | 30-40% |
### Shape Diagnostics
| Shape | What It Means | Recommended Action |
|-------|--------------|-------------------|
| **Top-heavy** (many early, few late) | Deals are entering but not progressing | Review qualification criteria. Are early-stage deals real? |
| **Bottom-heavy** (few early, many late) | Pipeline generation has stalled | Immediate prospecting push. Pipeline will dry up in 60-90 days. |
| **Uniform** (even distribution) | Generally healthy but inspect individual deals | Focus on accelerating mid-stage deals |
| **Bimodal** (many early + many late, hollow middle) | Conversion problem in mid-stages | Diagnose why deals stall in evaluation/proposal |
| **Single-deal dependent** | One deal is >40% of pipeline value | Extreme risk concentration. Generate pipeline urgently. |
---
## Hygiene Audit Checklist
Flag these issues automatically when reviewing pipeline data.
| Issue | Detection Rule | Impact | Fix |
|-------|---------------|--------|-----|
| Missing close date | Close date field empty | Cannot forecast | Add realistic close date |
| Missing amount | Amount field empty or $0 | Cannot forecast | Estimate or qualify |
| Missing next step | No next action recorded | Stall risk | Define next action |
| Missing contact | No primary contact | Cannot engage | Assign contact |
| Duplicate deals | Same account + similar amount + overlapping dates | Inflated pipeline | Merge or close duplicate |
| Stale stage | Same stage 45+ days | Over-represented in forecast | Advance, regress, or close |
| Zombie deal | Past close date + no activity 30+ days + pushed 2+ times | Dead weight | Close lost |
| Orphaned deal | Owner left company or changed role | No one working it | Reassign |
---
## Reporting Cadence
| Review Type | Frequency | Focus |
|------------|-----------|-------|
| Deal inspection | Daily | Top 3-5 deals. What changed? What's the next action? |
| Pipeline review | Weekly | Full pipeline health. Prioritization. Risk flags. |
| Forecast call | Bi-weekly or weekly | Commit/upside. Gap analysis. Scenario update. |
| Pipeline shape | Monthly | Stage distribution. Generation vs. close rate. |
| Deep clean | Quarterly | Full hygiene audit. Remove dead weight. Reset close dates. |
references/strategic-frameworks.md
---
title: Strategic Frameworks — Porter's Five Forces, SWOT, OKR Alignment
domain: strategic-decision
level: 3
skill: strategic-decision
---
# Strategic Frameworks Reference
> **Scope**: Porter's Five Forces checklist, SWOT scoring templates, and OKR alignment matrices for CEO-level strategic decisions. Use when analyzing market position, competitive dynamics, or organizational direction-setting.
> **Version range**: Framework-agnostic — these frameworks are stable; the templates here adapt them for small-team use.
> **Generated**: 2026-04-09 — validate competitive data freshness before using Five Forces output.
---
## Overview
Porter's Five Forces is misused more often than it is useful. Most applications list generic observations per force without weighting them or drawing a conclusion. SWOT analyses similarly produce four lists that nobody acts on. This reference contains opinionated, scored versions of each framework that produce a decision-ready output, not a slide deck. OKR alignment matrices answer the question practitioners actually need: "Does this strategic option move our OKRs, or does it just sound strategic?"
---
## Porter's Five Forces — Scored Checklist
**Purpose**: Assess industry attractiveness and your competitive position within it. Run before major market entry or expansion decisions.
Rate each force 1-5: 1 = very low pressure on you (good), 5 = very high pressure (bad).
### Force 1: Threat of New Entrants
| Factor | Your Assessment (1-5) | Evidence |
|--------|-----------------------|----------|
| Capital requirements to enter | ___ | ___ |
| Economies of scale advantage for incumbents | ___ | ___ |
| Network effects protecting incumbents | ___ | ___ |
| Regulatory/compliance barriers | ___ | ___ |
| Brand loyalty / switching costs for customers | ___ | ___ |
| **Force Score (average)** | **___** | |
Score 1-2: High barriers, you are protected. Score 4-5: Low barriers, expect new competitors.
### Force 2: Bargaining Power of Suppliers
| Factor | Your Assessment (1-5) | Evidence |
|--------|-----------------------|----------|
| Number of alternative suppliers | ___ | ___ |
| Switching cost to change supplier | ___ | ___ |
| Supplier concentration vs. industry concentration | ___ | ___ |
| Supplier's ability to forward-integrate | ___ | ___ |
| **Force Score (average)** | **___** | |
### Force 3: Bargaining Power of Buyers
| Factor | Your Assessment (1-5) | Evidence |
|--------|-----------------------|----------|
| Buyer concentration (few large buyers vs many small) | ___ | ___ |
| Switching cost for buyers | ___ | ___ |
| Price sensitivity of buyers | ___ | ___ |
| Buyer's ability to backward-integrate | ___ | ___ |
| Availability of substitutes for buyers | ___ | ___ |
| **Force Score (average)** | **___** | |
### Force 4: Threat of Substitutes
| Factor | Your Assessment (1-5) | Evidence |
|--------|-----------------------|----------|
| Number of substitute products/services | ___ | ___ |
| Price-performance of substitutes vs. yours | ___ | ___ |
| Buyer propensity to switch | ___ | ___ |
| **Force Score (average)** | **___** | |
### Force 5: Competitive Rivalry
| Factor | Your Assessment (1-5) | Evidence |
|--------|-----------------------|----------|
| Number and size of competitors | ___ | ___ |
| Industry growth rate | ___ | ___ |
| Product differentiation | ___ | ___ |
| Exit barriers (sunk costs, specialization) | ___ | ___ |
| **Force Score (average)** | **___** | |
### Five Forces Summary
| Force | Score (1-5) | Weight | Weighted |
|-------|-------------|--------|----------|
| Threat of New Entrants | ___ | 20% | ___ |
| Supplier Power | ___ | 15% | ___ |
| Buyer Power | ___ | 25% | ___ |
| Threat of Substitutes | ___ | 20% | ___ |
| Competitive Rivalry | ___ | 20% | ___ |
| **Overall Pressure Score** | | 100% | **___** |
**Interpretation**: Score 1.0-2.0 = attractive industry, favorable position. 2.1-3.5 = moderate pressure, monitor. 3.6-5.0 = high pressure, requires active differentiation or exit plan.
**Worked example** (indie SaaS tool in crowded productivity space):
- New Entrants: 4.2 (low capital, no network effects)
- Supplier Power: 1.5 (AWS/GCP compete for business)
- Buyer Power: 3.8 (price-sensitive, many alternatives)
- Substitutes: 4.0 (spreadsheets, free tools)
- Rivalry: 4.5 (hundreds of competitors)
- **Overall: 3.6** — high pressure; conclusion: must differentiate on vertical specificity, not features.
---
## SWOT Scoring Matrix
Replace the traditional 4-box with scored, prioritized output. Each item gets a score and an action assignment.
### Scoring Method
| Category | Items (max 5 each) | Score (1-10) | Actionable? | Owner |
|----------|--------------------|--------------|-------------|-------|
| **Strength** | Strong brand trust in niche | 8 | Amplify in sales collateral | ___ |
| **Strength** | Proprietary data set | 9 | Build moat, protect via ToS | ___ |
| **Weakness** | No enterprise sales motion | 7 | Hire AE or partner channel | ___ |
| **Weakness** | Single-founder dependency | 6 | Document processes, hire #2 | ___ |
| **Opportunity** | Regulatory change opens market | 8 | Launch compliance feature Q2 | ___ |
| **Opportunity** | Competitor product discontinued | 7 | Outreach to their user base | ___ |
| **Threat** | Large competitor entering space | 9 | Accelerate differentiation | ___ |
| **Threat** | API dependency on third party | 6 | Build abstraction layer | ___ |
**Prioritization**: Address Weaknesses with score 7+ that block top Opportunities first. Ignore Threats with score < 5 until Strengths and Opportunities are capitalized.
### SWOT → Strategic Actions Matrix
| | **Strengths (S)** | **Weaknesses (W)** |
|--|-------------------|--------------------|
| **Opportunities (O)** | **SO Strategies**: Use strengths to capture opportunities | **WO Strategies**: Overcome weaknesses to capture opportunities |
| **Threats (T)** | **ST Strategies**: Use strengths to mitigate threats | **WT Strategies**: Minimize weaknesses + avoid threats (defensive) |
Fill each quadrant with 1-2 concrete actions, not generic statements.
---
## OKR Alignment Matrix
Use before committing to a strategic initiative. Confirms the initiative actually moves the objectives, or reveals it is a distraction.
### Template
| Strategic Initiative | O1: [Objective 1] | O2: [Objective 2] | O3: [Objective 3] | Alignment Score |
|---------------------|-------------------|-------------------|-------------------|-----------------|
| Initiative A | High (3) | Medium (2) | Low (1) | 6 |
| Initiative B | Low (1) | High (3) | High (3) | 7 |
| Initiative C | None (0) | Low (1) | Medium (2) | 3 |
**Scoring**: None = 0, Low = 1, Medium = 2, High = 3. Max score = 3 × number of objectives.
**Decision rule**: Initiatives with alignment score < 30% of maximum should be questioned. They may be operationally necessary but are not "strategic" — be honest about what you are doing.
**Worked example** (content startup with 3 OKRs):
- O1: Reach 10,000 organic monthly readers by Q4
- O2: Launch paid newsletter tier by Q3
- O3: Establish 3 brand partnerships by Q4
| Initiative | O1 | O2 | O3 | Score |
|------------|----|----|-----|-------|
| Weekly long-form SEO articles | High (3) | Medium (2) | Low (1) | 6 |
| Twitter presence | Medium (2) | None (0) | Low (1) | 3 |
| Podcast series | Low (1) | Medium (2) | High (3) | 6 |
| Email list building | Medium (2) | High (3) | Medium (2) | 7 |
Conclusion: Twitter presence scores 3/9 — keep it under 2 hours/week. Email list building is the highest-leverage initiative at 7/9.
---
## Strategy Horizon Framework
For decisions that span multiple time horizons, use McKinsey's Three Horizons adapted for small teams.
| Horizon | Time Frame | Focus | Investment Level | Resource Split |
|---------|-----------|-------|-----------------|----------------|
| H1: Core | 0-12 months | Defend and extend current business | 70% of resources | Non-negotiable floor |
| H2: Adjacent | 12-36 months | Emerging opportunities with proven demand | 20% of resources | Adjustable quarterly |
| H3: Transformational | 36+ months | Options on future business models | 10% of resources | Can pause if H1 at risk |
**Trap to avoid**: Teams under pressure collapse all resources into H1 and wonder why they have no future. H3 investment survives pressure because it is small — protect it deliberately.
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Five Forces as a one-time exercise
**What it looks like**: Five Forces analysis done at company founding, cited 3 years later.
**Why wrong**: Industry forces shift. A 3-year-old analysis of supplier power in cloud infrastructure predates major pricing moves. Forces should be re-run at every major strategic inflection.
**Detection**: Check the date on any Five Forces document. Over 18 months old = must refresh before citing.
**Do instead**: Date-stamp every Five Forces analysis and schedule a refresh trigger at 18 months or at any major industry inflection (funding round, acquisition, regulatory change). Cite only analyses within that window in strategy discussions.
### SWOT without prioritization
**What it looks like**: 20-item SWOT with equal visual weight given to "strong team culture" and "no enterprise sales motion."
**Why wrong**: Without scores, SWOT becomes a list that feels like analysis but produces no action.
**Do instead**: Score every SWOT item and assign it an owner before the session ends. Archive items scoring below 5. The remaining scored, owned items are the actual action surface from the SWOT.
**Fix**: Every SWOT item must have a score and an owner. Items below score 5 are archived, not acted on.
### OKRs set after strategy, not before
**What it looks like**: Leadership decides the strategy in Q4, then writes OKRs in January to match it.
**Why wrong**: OKRs should constrain strategy choice, not validate it retroactively. If OKRs are written to match decisions already made, the alignment matrix is theater.
**Do instead**: Set OKRs for the period before running any OKR alignment matrix. The OKRs act as the constraint that filters strategic options. If OKRs do not exist when the strategy decision needs to be made, use a decision matrix instead.
**Fix**: Set OKRs for the period before running the OKR alignment matrix. If OKRs don't exist, the matrix is not useful — run the decision matrix instead.
### "We have no competitors" (Porter's Forces Denial)
**What it looks like**: Founder refuses to name competitors; "we're unique."
**Why wrong**: No competitors = no market, or no market research. Porter's Threat of Substitutes force always has content even in novel markets (the substitute is "doing nothing" or "using the old way").
**Do instead**: Reframe the question as "what alternatives does the buyer consider?" Every buyer has at least one alternative, even if it is doing nothing or continuing with the old approach. Name those alternatives and analyze them as substitutes in the Forces model.
**Fix**: Replace "competitors" with "alternatives the buyer considers." Every buyer has alternatives, even if not commercial products.
---
## See Also
- `decision-matrices.md` — weighted scoring matrices and pre-mortem templates
- `skills/competitive-intel/references/competitive-mapping.md` — competitor landscape mapping
- `skills/competitive-intel/references/market-positioning.md` — positioning maps and differentiation scoring
references/tco-framework.md
---
title: TCO Framework — Total Cost of Ownership for Build vs Buy Decisions
domain: build-vs-buy
level: 3
skill: build-vs-buy
---
# TCO Framework
> **Scope**: Total cost of ownership calculation templates, build vs buy decision scorecards, and migration cost models. Covers the first 3 years of a technology decision — the window where build vs buy trade-offs are most sensitive to get right.
> **Version range**: Framework-agnostic — applies to any software/SaaS/OSS adoption decision.
> **Generated**: 2026-04-09 — validate vendor pricing against current public pricing pages.
---
## Overview
The most common mistake in build vs buy is comparing purchase price to build cost. A SaaS that costs $500/month looks expensive next to "free to build." The analysis should compare 3-year total cost of ownership, including all hidden costs: engineering time to build and maintain, support burden, training, and opportunity cost of capacity consumed. Build almost always wins year 1 and loses years 2-3. The crossover point determines the decision.
---
## TCO Calculation Template
Fill this out for each option before any scoring. Use fully-loaded engineering cost (salary + benefits + overhead), typically 1.5-2× base salary.
### Option: [Name]
```
## Year 1 Costs
### Acquisition / Build
- Vendor license or SaaS contract: $___/month × 12 = $___
OR
- Engineering hours to build: ___ hours × $___/hour = $___
- External integrations (APIs, connectors): $___
- Infrastructure setup (servers, CI/CD, config): $___
### Onboarding
- Training / learning curve: ___ hours × $___/hour = $___
- Data migration from existing system: ___ hours × $___/hour = $___
- Documentation and runbooks: ___ hours × $___/hour = $___
Year 1 Total: $___
## Year 2-3 Costs (Annual)
### Vendor (if SaaS)
- License at expected usage tier: $___/year
- Usage overages (estimate 20% buffer): $___/year
- Upgrade / tier increase: $___/year (if usage grows)
### Build (if internal)
- Maintenance and bug fixes: ___ hours/month × 12 × $___/hour = $___
- Infrastructure (hosting, monitoring, backups): $___/year
- Security patches and dependency updates: ___ hours/quarter × 4 × $___/hour = $___
- Feature additions (roadmap items): ___ hours/year × $___/hour = $___
Annual Year 2-3 Total: $___
## 3-Year TCO Summary
| | Year 1 | Year 2 | Year 3 | 3-Year Total |
|-|--------|--------|--------|--------------|
| Vendor | $__ | $__ | $__ | $__ |
| Build | $__ | $__ | $__ | $__ |
| OSS + Customize | $__ | $__ | $__ | $__ |
```
---
## Worked Example: Authentication System
Real-but-anonymized: 8-person engineering team evaluating whether to build auth or adopt Auth0.
```
## Auth0 (Buy SaaS)
Year 1:
- Auth0 B2C plan (25K MAU): $23/month × 12 = $276
- Integration engineering: 40 hours × $150/hour = $6,000
- Documentation: 8 hours × $150/hour = $1,200
Year 1 Total: $7,476
Year 2-3 (annual):
- Auth0 at 75K MAU: $240/month × 12 = $2,880
- Maintenance: 2 hours/month × 12 × $150 = $3,600
Annual: $6,480
3-Year Total: $7,476 + $6,480 + $6,480 = $20,436
## Build from scratch
Year 1:
- JWT auth library integration: 80 hours × $150 = $12,000
- OAuth2 implementation: 120 hours × $150 = $18,000
- MFA, password reset flows: 60 hours × $150 = $9,000
- Security review: 20 hours × $150 = $3,000
Year 1 Total: $42,000
Year 2-3 (annual):
- Security patches: 8 hours/quarter × 4 × $150 = $4,800
- Feature maintenance: 4 hours/month × 12 × $150 = $7,200
- Incident response (1-2 per year estimated): 16 hours × $150 = $2,400
Annual: $14,400
3-Year Total: $42,000 + $14,400 + $14,400 = $70,800
```
**Decision**: Auth0 saves $50,364 over 3 years. Engineering team's 260 hours freed up for product differentiators. Unless data residency requirements or offline-first constraints prevent it, buy wins clearly.
---
## Build vs Buy Decision Scorecard
Score each dimension 1-10, then calculate weighted total. Customize weights for your context.
| Dimension | Weight | Build | OSS + Customize | Buy SaaS | Notes |
|-----------|--------|-------|-----------------|----------|-------|
| Fit to requirements | 5× | ___ | ___ | ___ | Does it solve the actual problem? |
| 3-Year TCO | 4× | ___ | ___ | ___ | From TCO template above |
| Operational burden | 4× | ___ | ___ | ___ | Who runs it at 3 AM? |
| Team capability | 3× | ___ | ___ | ___ | Can we build/operate it? |
| Lock-in risk | 3× | ___ | ___ | ___ | How hard to switch? |
| Time to value | 3× | ___ | ___ | ___ | Weeks until operational |
| Flexibility | 2× | ___ | ___ | ___ | Future extensibility |
| **Weighted Total** | **24×** | **___** | **___** | **___** | |
**Fit scoring guide**:
- 9-10: Solves the problem exactly, no meaningful gaps
- 7-8: Solves the problem with minor workarounds (< 5% of use cases affected)
- 5-6: Solves 70-80% of the problem, significant gaps need custom work
- 3-4: Covers the core but misses important edge cases
- 1-2: Wrong tool for the job
---
## Migration Cost Model
Run this when switching from existing solution to evaluate true total cost. Often reveals that migration cost exceeds the savings from the new option.
```
## Migration Cost Estimate
### Data Migration
- Schema mapping and transformation: ___ hours × $___/hour = $___
- Data validation and reconciliation: ___ hours × $___/hour = $___
- Historical data migration (one-time): ___ hours × $___/hour = $___
### Integration Rewiring
- API client updates (per integration): ___ integrations × ___ hours × $___/hour = $___
- Webhook reconfiguration: ___ hours × $___/hour = $___
- Auth/credential rotation: ___ hours × $___/hour = $___
### Risk and Buffer
- Rollback capability (maintain old system in parallel): ___ months × $___/month = $___
- Incident response during cutover: ___ hours × $___/hour = $___
- Buffer for unexpected complexity (add 30%): $___
Migration Total: $___
Months until break-even (savings from new system per month): ___
Break-even: ___ months
```
**Rule**: If break-even is over 24 months, the migration is marginal. If break-even is over 36 months, the migration likely destroys value — reconsider.
---
## Hidden Cost Checklist
Before finalizing TCO, confirm you have accounted for:
```
[ ] Vendor price escalation risk (SaaS prices increase 10-15% annually on average)
[ ] Usage overages (most SaaS tools have surprise usage-based tiers)
[ ] Security and compliance costs (SOC 2, GDPR compliance on vendor side — who bears audit costs?)
[ ] Contract lock-in (annual vs monthly pricing — is the discount worth the commitment?)
[ ] Support tier costs (enterprise support is often 20-25% of license cost)
[ ] Training time per new hire (ongoing, not just initial onboarding)
[ ] Integration maintenance (third-party integrations break when vendors change APIs)
[ ] Sunset risk (vendor acquisition, product discontinuation, pricing regime change)
[ ] Data export capability (can you get your data back in a usable format if you leave?)
[ ] Opportunity cost: what won't get built while this is under development/integration?
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Comparing development cost to license cost
**What it looks like**: "Auth0 costs $2,880/year. We can build it for free."
**Why wrong**: "Free to build" ignores 120+ engineering hours, security maintenance, incident response, and opportunity cost. The license is year 2+ maintenance savings, not just year 1 build savings.
**Do instead**: Run the 3-year TCO template before any build-vs-buy comparison. Put license cost and fully-loaded build cost (engineering hours, security maintenance, incident response, opportunity cost) side by side across all three years. The comparison only becomes meaningful at that scope.
**Fix**: Always run the 3-year TCO template. If the comparison is still in favor of building after year 2-3 costs are included, building may be right.
### Underestimating maintenance burden
**What it looks like**: Year 2 estimate shows zero hours for maintenance on custom-built system.
**Why wrong**: Every system has dependencies that release updates. Security vulnerabilities require patches. Load characteristics change. Features accumulate bugs over time. 10-20% of initial build effort per year is the historical baseline.
**Detection**: If Year 2 internal maintenance estimate is less than 10% of Year 1 build estimate, it is too optimistic.
**Do instead**: Use 15% of Year 1 build cost as the floor for Year 2-3 annual maintenance in every custom build TCO. Flag any estimate below that threshold and require a written justification before accepting it.
**Fix**: Use 15% of Year 1 build cost as minimum Year 2-3 annual maintenance baseline. Adjust up if the system is complex.
### TCO without usage growth modeling
**What it looks like**: SaaS cost calculated at current usage (100 users), not at expected usage in year 3 (1,000 users).
**Why wrong**: Many SaaS platforms have non-linear pricing. What looks affordable at 100 users can be prohibitive at 1,000. The opposite is also true — build costs are largely fixed.
**Do instead**: Model three usage scenarios in every TCO: current usage, 3x growth, and 10x growth. The right choice is the one that remains affordable across all three plausible futures, not just the one that looks best at today's usage.
**Fix**: Model three usage scenarios: current, 3× growth, 10× growth. Pick the option that remains affordable across all three plausible scenarios.
---
## See Also
- `vendor-evaluation.md` — vendor scorecards, RFP criteria, integration complexity scoring
- `skills/project-evaluation/references/roi-frameworks.md` — ROI calculation and risk-adjusted analysis
references/trend-analysis.md
---
title: Trend Analysis — Market Trend Detection, Technology Adoption Curves, Disruption Pattern Recognition
domain: competitive-intel
level: 3
skill: competitive-intel
---
# Trend Analysis Reference
> **Scope**: Market trend detection signals, technology adoption curve positioning, and disruption pattern recognition for competitive intelligence. Use when determining whether a market signal is a durable trend, a fad, or a disruption precursor. Complements competitive-mapping.md (which covers landscape analysis) and market-positioning.md (which covers positioning response).
> **Version range**: Framework-agnostic — applies to SaaS, consumer, content, and infrastructure markets. Technology-specific adoption benchmarks are from empirical research; validate against current analogues.
> **Generated**: 2026-04-09 — trend signals require continuous refresh; any analysis older than 6 months should be treated as hypothesis, not conclusion.
---
## Overview
Trend analysis fails most often by confusing noise for signal, or signal for sustained trend. A competitor shipping one feature is noise. A competitor changing their pricing model, hiring a new VP of Sales, and launching in a new market simultaneously is a signal cluster. Technology adoption curves explain why early signals look small and why acting on them feels premature — the correct reading of "this is still early" depends entirely on which phase of adoption you are in. Disruption pattern recognition adds the most valuable layer: identifying when a market is not just growing or shrinking, but about to be restructured.
---
## Market Trend Detection Signal Framework
### Signal Tier Classification
| Tier | Type | Reliability | Lag | Examples |
|------|------|-------------|-----|---------|
| Tier 1 — Leading | Behavior signals | Low (noisy) | 12-24 months ahead | Search volume spikes, VC investment thesis, early startup counts |
| Tier 2 — Confirming | Adoption signals | Medium | 6-12 months ahead | Enterprise pilot programs, conference session count, GitHub stars growth rate |
| Tier 3 — Lagging | Market signals | High (reliable) | 0-6 months | Gartner/Forrester coverage, mainstream press articles, incumbent product response |
**Rule**: Act on Tier 1 signals only when supported by at least one Tier 2 signal. Act on Tier 2 signals when supported by at least one Tier 3 signal. Tier 3 alone means the trend is confirmed but you may be late.
### Signal Collection Template
```
## Trend Signal Log: [Topic/Technology]
## Collection date: [YYYY-MM-DD]
## Analyst: ___
### Tier 1 Signals (Leading)
| Signal | Source | Value | Trend (↑/↓/→) | Date | Notes |
|--------|--------|-------|---------------|------|-------|
| Google Trends score (keyword) | Google Trends | ___ | ___ | ___ | Normalize to 100 = peak interest |
| GitHub repo star growth rate | GitHub | +___/month | ___ | ___ | Compare to 3-month prior |
| VC deals in space (new, 90 days) | Crunchbase/PitchBook | ___ | ___ | ___ | Count, not dollar volume |
| New GitHub repos on topic | GitHub search | ___ | ___ | ___ | repos created in last 90 days |
| ArXiv papers (if technical) | ArXiv | ___/quarter | ___ | ___ | Research = 18-36 months ahead |
### Tier 2 Signals (Confirming)
| Signal | Source | Value | Trend | Date | Notes |
|--------|--------|-------|-------|------|-------|
| Enterprise pilot count | LinkedIn Jobs, press releases | ___ | ___ | ___ | "Seeking X experience" in JDs |
| Major conference tracks/sessions | Conf agendas | ___ | ___ | ___ | New track = mainstream adjacent |
| Stack Overflow questions/week | Stack Overflow | ___ | ___ | ___ | New question velocity = active use |
| npm/PyPI download growth | Package registries | ___% MoM | ___ | ___ | Library adoption proxy |
| SaaS startups in space count | ProductHunt, AngelList | ___ | ___ | ___ | Launched last 12 months |
### Tier 3 Signals (Lagging)
| Signal | Source | Value | Trend | Date | Notes |
|--------|--------|-------|-------|------|-------|
| Gartner/Forrester mention | Analyst reports | In/Out of Hype Cycle | ___ | ___ | Position in Hype Cycle |
| Enterprise vendor product launch | Press release | Yes/No | ___ | ___ | Incumbent response signal |
| Major media coverage | TechCrunch, WSJ, NYT | ___ articles | ___ | ___ | Count in last 90 days |
### Signal Cluster Summary
Tier 1 supporting signals: ___ of ___
Tier 2 supporting signals: ___ of ___
Tier 3 supporting signals: ___ of ___
Overall trend assessment: [Emerging / Building / Confirmed / Mainstream / Declining]
```
---
## Technology Adoption Curve Positioning
Based on the Rogers Diffusion of Innovations model. Position your market accurately before setting strategy.
### Adoption Phase Reference Table
| Phase | % of Market | Characteristics | Duration (typical) | Strategic Implication |
|-------|-------------|-----------------|-------------------|----------------------|
| **Innovators** | 2.5% | Hobbyists, researchers, experimenters; willing to accept broken products | 1-3 years | Build the core; don't optimize for UX yet |
| **Early Adopters** | 13.5% | Visionaries seeking competitive advantage; pay premium; provide feedback | 2-4 years | Reference customers, case studies, category definition |
| **Early Majority** | 34% | Pragmatists; need proof, peers, integrations; risk-averse | 3-5 years | Crossing the chasm — reduce friction, add integrations |
| **Late Majority** | 34% | Conservatives; adopt because they must; price-sensitive | 3-6 years | Volume play; commoditization pressure begins |
| **Laggards** | 16% | Last to move; often forced by regulation or ecosystem pressure | Indefinite | Sustain or exit; do not invest in growth here |
### Current Phase Diagnosis
```
## Adoption Phase Assessment: [Technology/Market]
### Phase Indicators (check all that apply)
INNOVATORS phase signals:
[ ] Primary users are developers or researchers (not business buyers)
[ ] Products require technical setup; no one-click install
[ ] No vendor offers enterprise support
[ ] Community is primarily on GitHub, HN, or academic forums
[ ] Use cases are experimental, not production-critical
EARLY ADOPTER phase signals:
[ ] First conference talks appearing at major industry events
[ ] 2-5 vendors have achieved product-market fit with specific use cases
[ ] Early reference customers willing to be named publicly
[ ] Enterprise pilots (unpaid or nearly unpaid) are happening
[ ] Content about this technology is growing on technical blogs
EARLY MAJORITY phase signals:
[ ] Multiple vendors offering commercial support and SLAs
[ ] Integration with major platforms (Salesforce, AWS, etc.)
[ ] HR departments posting job descriptions requiring this skill
[ ] Industry analysts (Gartner, Forrester) have named a category
[ ] First acquisition of a player in the space
LATE MAJORITY phase signals:
[ ] Incumbent vendors (not just startups) have launched products here
[ ] Price competition is the primary marketing lever
[ ] Certification programs exist (vendor + third-party)
[ ] Community growth rate has slowed; questions on Stack Overflow more rote
[ ] Enterprise procurement (RFPs, competitive bidding) is standard
### Diagnosis
Current phase: ___
Evidence for this phase: ___
Time to next phase (estimate): ___
```
### "Crossing the Chasm" Detection
The chasm exists between Early Adopters and Early Majority. Signals that a market is AT the chasm:
```
[ ] Early adopter growth has stalled (the enthusiasts who will try anything have all tried it)
[ ] Mainstream buyers ask for references from companies like them (not just any reference)
[ ] Integration requests outpace feature requests
[ ] Support burden is growing faster than new customer acquisition
[ ] Pricing pressure emerging even though market is not yet saturated
[ ] Multiple vendors competing for the same Early Adopter segment with near-identical products
```
**Chasm crossing strategy**: Pick one vertical or use case. Own it completely with deep integrations, specific references, and specialized support. Do not try to cross the chasm with a horizontal product — it is too undifferentiated. Once the vertical is captured, use it as a beachhead.
---
## Disruption Pattern Recognition
Based on Clayton Christensen's disruption theory, extended for software and content markets. Disruption follows observable patterns before incumbents respond.
### Disruption Signal Matrix
| Pattern | Early Signal | Midpoint Signal | Late Signal (Too Late for Incumbents) |
|---------|-------------|-----------------|---------------------------------------|
| **Low-end disruption** | New entrant prices 60-80% cheaper; incumbent ignores them (wrong customers) | New entrant's product improves; incumbent's low-end customers switch | New entrant moves upmarket; incumbent loses profitability |
| **New-market disruption** | Entirely new user type that incumbent never served; product is "worse" by incumbent metrics | New-market grows; some incumbent customers defect | Incumbent's core market erodes from below and adjacent |
| **Platform disruption** | Aggregator platform pulls customers between you and your buyer | Platform starts competing with you directly for top customers | Platform controls discovery; you have no direct customer relationship |
| **Business model disruption** | Competitor offers your product's function as a feature of a larger bundle | Customers start buying the bundle instead of your standalone | Your market becomes a feature line item, not a product category |
### Disruption Diagnostic Template
```
## Disruption Risk Assessment: [Your Market]
## Date: [YYYY-MM-DD]
### Low-End Disruption Check
- Is there a competitor entering at 50%+ lower price? [Yes / No]
- Are they targeting a customer segment you consider unprofitable? [Yes / No]
- Is their product improving faster than your low-end customer's needs are increasing? [Yes / No]
If 2+ Yes: Low-end disruption risk is REAL. Assess within 6 months.
### New-Market Disruption Check
- Is there a new category of user who cannot use incumbent products (including yours) due to cost/complexity? [Yes / No]
- Is anyone building a "worse but simpler" version targeting those non-consumers? [Yes / No]
- Is that simpler product improving 20%+ annually in the dimension incumbents optimize for? [Yes / No]
If 2+ Yes: New-market disruption risk is REAL.
### Platform Disruption Check
- Does a platform (marketplace, OS, cloud provider, social network) now intermediate between you and your customer? [Yes / No]
- Is that platform showing any signal of building features that overlap with your product? [Yes / No]
- Do your customers use your product primarily through or within that platform? [Yes / No]
If 2+ Yes: Platform disruption risk is HIGH. Examine differentiation urgently.
### Business Model Disruption Check
- Is a competitor offering your core function as part of a larger suite at the same or lower price? [Yes / No]
- Are customers already buying that suite for other components, making your standalone look redundant? [Yes / No]
If 2 Yes: Business model disruption is underway. Evaluate bundling, partnership, or exit.
### Disruption Risk Summary
| Type | Risk Level | Time Horizon | Primary Action |
|------|-----------|--------------|----------------|
| Low-end | Low / Medium / High | ___ months | ___ |
| New-market | Low / Medium / High | ___ months | ___ |
| Platform | Low / Medium / High | ___ months | ___ |
| Business model | Low / Medium / High | ___ months | ___ |
```
---
## Trend Decay Recognition
Not every trend sustains. Identify decay early to avoid investing in a declining market.
| Decay Signal | Measurement | Threshold for Concern |
|-------------|-------------|----------------------|
| Search volume decline | Google Trends | >20% decline from peak over 6 months |
| GitHub star growth rate falling | Stars per month | Growth rate declining 3+ consecutive months |
| Conference tracks being removed | Agenda comparison YoY | Removal from 2+ major conferences |
| VC investments tapering | Deal count by quarter | 40%+ fewer deals than peak quarter |
| Acquisitions exceeding new entrants | M&A vs. new company formation | More acquisitions than new companies for 2+ quarters |
| Incumbent products discontinued | Product announcements | 2+ major vendors exiting the space |
**Note**: Decline in early signals (search volume) while late signals remain high (enterprise adoption) is NOT decay — it is maturation. Decay occurs when both lead and lag signals decline together.
```
## Trend Decay Audit: [Technology/Market]
## Date: ___
| Decay Signal | Current Value | Peak Value | % Change | Trend |
|-------------|--------------|------------|----------|-------|
| Google Trends score | ___ | ___ | ___% | ↑/↓ |
| GitHub stars/month | ___ | ___ | ___% | ↑/↓ |
| VC deals/quarter | ___ | ___ | ___% | ↑/↓ |
| Conference sessions count | ___ | ___ | ___% | ↑/↓ |
| New startups launched | ___ | ___ | ___% | ↑/↓ |
Decay signals active: ___ of 5
Assessment: [Growing / Maturing / Plateau / Early Decay / Late Decay]
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Treating Hype Cycle peak as confirmation of sustained growth
**What it looks like**: "Gartner put us at the Peak of Inflated Expectations — that means we're mainstream."
**Why wrong**: Gartner's Peak of Inflated Expectations precedes the Trough of Disillusionment. Being at the peak means hype has outpaced delivery and a correction is coming. Companies that raise at peak valuations and expand at peak rates are maximally exposed to the trough.
**Do instead**: Treat Hype Cycle peak placement as a signal to reduce cost structure and extend runway. The question to ask is: "How do we survive the trough?" not "How fast can we grow into this moment?"
**Fix**: Treat Hype Cycle peak as a warning to reduce cost structure and extend runway, not as a signal to accelerate.
### Confusing single-vendor activity for market trend
**What it looks like**: "Competitor X launched three features this month. This space is growing fast."
**Why wrong**: A single vendor's product velocity is company signal, not market signal. A competitor may be launching features because they are desperate, because they have a new head of product, or because a large customer demanded it. None of those explain the market.
**Do instead**: Require signal from at least 3 independent sources across different tiers before classifying something as a market trend. One vendor's activity is a Tier 2 signal at most; it needs Tier 1 corroboration before it becomes actionable.
**Fix**: Market trends require signal from at least 3 independent sources across different tiers before classification as a trend. One vendor's activity is a Tier 2 signal at most, requiring Tier 1 support to be actionable.
### Ignoring disruption until it is in your customer segment
**What it looks like**: Disruption is identified in the low-end segment. Team says "that's not our market" and takes no action.
**Why wrong**: Disruption moves upmarket. The low-end is where it starts, not where it stays. Christensen documented this across 20+ industries: the disruption that starts at the bottom reliably migrates into the incumbent's core market within 3-8 years.
**Do instead**: When low-end disruption is identified, immediately run a response window calculation: assume the disruptor improves 20% per year on the dimension that matters to your customers. Determine when it reaches your customers' threshold. That date is your deadline for a strategic response.
**Fix**: When low-end disruption is identified, run a "future competitive landscape" exercise: assume the disruptor's product improves 20% per year in the dimension you care about. At that rate, when does it reach your customer's requirements? That is your response window.
### Trend analysis without refresh cadence
**What it looks like**: Trend analysis done in Q1, cited in Q4 as the basis for strategy decisions.
**Why wrong**: Tech markets move faster than annual strategy cycles. A signal logged as "Emerging" in January can be "Confirmed" by April.
**Do instead**: Establish a refresh cadence when you publish any trend analysis: Tier 1 signals monthly, Tier 2 quarterly, full report every 6 months. Attach a freshness warning to any report before it is cited in a decision if it exceeds those thresholds.
**Fix**: Tier 1 signals: check monthly. Tier 2 signals: check quarterly. Full trend report: refresh every 6 months. Any trend report over 6 months old should include a freshness warning before it is cited in a decision.
---
## Detection Commands Reference
```bash
# Google Trends export (manual: trends.google.com > download CSV)
# Then parse:
python3 -c "
import csv
with open('trends_data.csv') as f:
reader = csv.reader(f)
rows = list(reader)
# Find peak and current
values = [int(r[1]) for r in rows[3:] if r[1].isdigit()]
peak = max(values)
current = values[-1]
print(f'Peak: {peak}, Current: {current}, % of peak: {current/peak*100:.0f}%')
" 2>/dev/null
# GitHub star growth rate for a repo
curl -s "https://api.github.com/repos/{owner}/{repo}" | python3 -c "
import json, sys
d = json.load(sys.stdin)
print(f'Stars: {d[\"stargazers_count\"]}')
print(f'Watchers: {d[\"subscribers_count\"]}')
print(f'Open issues: {d[\"open_issues_count\"]}')
print(f'Last push: {d[\"pushed_at\"]}')
print(f'Created: {d[\"created_at\"]}')
"
# Count new GitHub repos on a topic (last 90 days)
curl -s "https://api.github.com/search/repositories?q={topic}+created:>$(date -d '90 days ago' +%Y-%m-%d)&sort=stars&per_page=1" | \
python3 -c "import json,sys; d=json.load(sys.stdin); print(f'New repos (90d): {d[\"total_count\"]}')"
# Hacker News mention count for a technology
curl -s "https://hn.algolia.com/api/v1/search?query={technology}&tags=story&numericFilters=created_at_i>$(date -d '90 days ago' +%s)" | \
python3 -c "import json,sys; d=json.load(sys.stdin); print(f'HN stories (90d): {d[\"nbHits\"]}')"
# Stack Overflow question velocity
curl -s "https://api.stackexchange.com/2.3/questions?tagged={technology}&site=stackoverflow&fromdate=$(date -d '30 days ago' +%s)" | \
python3 -c "import json,sys; d=json.load(sys.stdin); print(f'SO questions (30d): {d[\"total\"]}')" 2>/dev/null
```
---
## See Also
- `competitive-mapping.md` — landscape mapping and competitor activity tracking
- `market-positioning.md` — positioning response to confirmed market trends
- `skills/strategic-decision/references/risk-assessment.md` — scenario planning for trend uncertainty
references/vendor-evaluation.md
---
title: Vendor Evaluation — Scorecards, RFP Criteria, Integration Complexity
domain: build-vs-buy
level: 3
skill: build-vs-buy
---
# Vendor Evaluation Reference
> **Scope**: Vendor evaluation scorecards, RFP criteria matrices, integration complexity scoring, and red flag detection for technology vendor decisions. Use when comparing 2+ vendors or when a buy decision has been made and vendor selection is the next step.
> **Version range**: Framework-agnostic — validate vendor-specific claims against current documentation.
> **Generated**: 2026-04-09 — SaaS market conditions change; re-run vendor evaluation annually for active contracts.
---
## Overview
Vendor evaluation fails most often by comparing features instead of evaluating fitness for a specific context. A vendor with 400 features that does not integrate with your existing auth stack is worse than one with 50 features that drops in cleanly. This reference structures evaluation around the dimensions that determine whether a vendor relationship succeeds over 2-3 years, not just at time of purchase.
---
## Vendor Evaluation Scorecard
Score each vendor 1-10 per dimension. Fill independently before comparing scores with the team.
| Dimension | Weight | Vendor A | Vendor B | Vendor C | Scoring Guide |
|-----------|--------|----------|----------|----------|---------------|
| Functional fit | 5× | ___ | ___ | ___ | Does it solve the actual problem without significant workarounds? |
| Integration complexity | 4× | ___ | ___ | ___ | Hours to connect to existing stack (10 = <8 hrs, 1 = weeks) |
| Support quality | 4× | ___ | ___ | ___ | Response time SLA, escalation path, community vs. dedicated |
| Pricing model fit | 3× | ___ | ___ | ___ | Predictability as usage grows; alignment with your cost model |
| Data ownership / portability | 3× | ___ | ___ | ___ | Can you export your data in a usable format? |
| Security / compliance | 3× | ___ | ___ | ___ | SOC 2, GDPR, HIPAA — do their certifications match your requirements? |
| Company stability | 3× | ___ | ___ | ___ | Funding, revenue growth, customer retention signals |
| API quality | 2× | ___ | ___ | ___ | REST/GraphQL maturity, rate limits, webhook reliability |
| Roadmap alignment | 2× | ___ | ___ | ___ | Are upcoming features useful to you? Direction of travel? |
| Onboarding quality | 1× | ___ | ___ | ___ | Docs, sandbox, getting-started experience |
| **Weighted Total** | **30×** | **___** | **___** | **___** | |
**Interpretation**: 80%+ of max (240+): strong fit. 60-79% (144-191): acceptable with known gaps. Below 60%: reconsider.
---
## RFP Criteria Matrix
Use this when issuing a formal RFP to 3+ vendors. Structure requirements as MUST/SHOULD/NICE for objective comparison.
### Requirements Classification
| Requirement | Priority | Vendor A | Vendor B | Vendor C | Notes |
|-------------|----------|----------|----------|----------|-------|
| **MUST HAVE (eliminators)** | | | | | |
| SSO/SAML support | MUST | Pass/Fail | Pass/Fail | Pass/Fail | |
| EU data residency | MUST | Pass/Fail | Pass/Fail | Pass/Fail | |
| 99.9% uptime SLA | MUST | Pass/Fail | Pass/Fail | Pass/Fail | |
| [Your hard requirement] | MUST | Pass/Fail | Pass/Fail | Pass/Fail | |
| **SHOULD HAVE (differentiators)** | | | | | |
| Webhook retry on failure | SHOULD | 1-5 | 1-5 | 1-5 | |
| Role-based access control | SHOULD | 1-5 | 1-5 | 1-5 | |
| Audit logging | SHOULD | 1-5 | 1-5 | 1-5 | |
| Custom domain support | SHOULD | 1-5 | 1-5 | 1-5 | |
| **NICE TO HAVE (tie-breakers)** | | | | | |
| Slack integration | NICE | Yes/No | Yes/No | Yes/No | |
| Mobile SDK | NICE | Yes/No | Yes/No | Yes/No | |
**Process**: Eliminate any vendor failing a MUST. Score SHOULD items 1-5. NICE items break ties only.
---
## Integration Complexity Scoring
Run this before scorecard scoring. Integration complexity is chronically underestimated and frequently causes "buy" decisions to cost as much as "build."
| Integration Point | Effort Level | Score (10=easy) | Notes |
|-------------------|-------------|-----------------|-------|
| Auth (SSO, OAuth2, SAML) | Low / Medium / High | ___ | Does vendor support your identity provider? |
| Data ingestion (API, ETL, webhook) | Low / Medium / High | ___ | Format compatibility, rate limits, batch support |
| Data export / sync | Low / Medium / High | ___ | Webhook vs. polling, data freshness requirements |
| User provisioning (SCIM, manual) | Low / Medium / High | ___ | Automated vs. manual user management |
| Billing / payment integration | Low / Medium / High | ___ | If applicable — does it plug into your billing system? |
| Reporting / BI integration | Low / Medium / High | ___ | Data warehouse export, native analytics gaps |
| Existing tools / workflow | Low / Medium / High | ___ | Slack, Jira, PagerDuty, internal tools |
**Effort level definitions**:
- **Low**: Works out of the box or < 8 engineering hours. Standard OAuth2, documented webhooks, stable API.
- **Medium**: 1-3 weeks of engineering time. Custom adapter needed, unstable API, limited documentation.
- **High**: 1+ months, may require professional services. Non-standard auth, batch-only data access, custom ETL.
**Integration total effort estimate**:
```
Low (×8 hrs each): ___ items × 8 = ___ hrs
Medium (×40 hrs each): ___ items × 40 = ___ hrs
High (×160 hrs each): ___ items × 160 = ___ hrs
Total: ___ hrs × $___/hr = $___
```
---
## Red Flag Detection
Run this against every finalist vendor. One critical flag = disqualify. Multiple moderate flags = negotiate contract protections.
### Critical Red Flags (disqualify immediately)
```
[ ] No data export capability (you cannot get your data back in a usable format)
[ ] Arbitration-only contract with no class action (legal risk)
[ ] No uptime SLA or SLA with credits that don't cover your actual downtime cost
[ ] "We'll add that feature" promises not in contract (vaporware negotiation)
[ ] Pricing not disclosed publicly — only available after sales call (predatory pricing)
[ ] No SOC 2 Type II when you have compliance requirements
```
### Moderate Red Flags (negotiate mitigations)
```
[ ] Single data center with no failover option
[ ] Support is email-only with 48-hour SLA (no phone/chat for critical issues)
[ ] Annual contract required at outset with no monthly trial
[ ] API rate limits that would be hit at your scale (check their public docs, not sales promises)
[ ] Changelog is sparse or months out of date (signals low development velocity)
[ ] LinkedIn shows large support/sales team but small engineering team (feature growth will stall)
[ ] G2/Capterra reviews mention recurring billing problems or difficult cancellation
[ ] GitHub issues (if OSS component) show months-old bugs unaddressed
```
### Vendor Stability Signals
| Signal | Positive | Negative |
|--------|----------|----------|
| Funding | Series B+ or profitable/bootstrapped | Seed-only with no clear path |
| Revenue transparency | Published ARR or customer count growth | No public metrics |
| Customer retention | Publicly cited NRR > 100% | Churn not discussed |
| Engineering velocity | Regular releases, active changelog | Releases monthly or less |
| Enterprise customers | Named enterprise customers in same industry | Only SMB references |
---
## Contract Negotiation Checklist
Before signing any vendor contract over $10K/year:
```
[ ] SLA with financial penalties (not just service credits)
[ ] Data portability clause: export in standard format within 30 days of termination
[ ] Price cap on renewals (e.g., "increases capped at CPI or 5%, whichever is lower")
[ ] Termination for convenience clause (can you leave with 30-day notice?)
[ ] Service level agreement for support response (not just "best effort")
[ ] Data deletion timeline after contract ends (GDPR-critical)
[ ] Audit rights if handling sensitive data
[ ] Subprocessor notification (you must be notified if they change data processors)
```
---
## Patterns to Detect and Fix
<!-- no-pair-required: section header, not an individual failure mode -->
### Buying based on demo, not hands-on trial
**What it looks like**: Sales demo shows everything working perfectly, contract signed, integration begins, and a different reality emerges.
**Why wrong**: Demos are curated. Integration complexity, actual API reliability, and support quality only reveal themselves in practice.
**Do instead**: Require a POC period before any annual contract. Define a specific integration milestone that represents real-world complexity. Sign only after that milestone is complete. Any serious vendor will accept this condition.
**Fix**: Require a proof-of-concept (POC) period with a representative integration task before committing to annual contract. "We need to complete integration milestone X before we commit" is a reasonable ask for any serious vendor.
### Evaluating current feature set, not roadmap direction
**What it looks like**: Vendor has 9/10 of features needed today. Missing 1 feature is "on the roadmap."
**Why wrong**: "On the roadmap" has no SLA. Features promised in sales negotiations are not features in contracts.
**Do instead**: Request the public changelog and last 4 release notes. Evaluate velocity and direction from those artifacts, not from verbal promises. For features that are decision-critical, require a contract addendum with a delivery date before signing.
**Fix**: Ask to see the public changelog and last 4 release notes. If the velocity is high and roadmap direction matches your needs, that is more reliable than verbal promises. Require a contract addendum for features that are decision-critical.
### Evaluating one vendor at a time (serial evaluation)
**What it looks like**: Full evaluation of Vendor A, then if not satisfied, evaluate Vendor B.
**Why wrong**: Serial evaluation introduces time pressure at Vendor B. You may sign with B because you are tired of evaluating, not because B is the best option.
**Do instead**: Run 2-3 vendor evaluations in parallel using identical RFP criteria. The coordination overhead is worth it: you eliminate time-pressure bias and can make a genuine comparison rather than a fallback decision.
**Fix**: Run parallel evaluations with 2-3 vendors simultaneously using the same RFP criteria. Takes more coordination but produces better decisions.
---
## Detection Commands Reference
```bash
# Check vendor API rate limits (if REST API)
curl -I https://api.vendor.com/v1/some-endpoint 2>&1 | grep -i "x-ratelimit"
# Check vendor uptime history
# (Manual: visit statuspage.io/<vendor> or status.<vendor>.com)
# Check GitHub activity for OSS components
curl -s "https://api.github.com/repos/{owner}/{repo}" | python3 -c "
import json,sys
d=json.load(sys.stdin)
print('Stars:', d['stargazers_count'])
print('Last push:', d['pushed_at'])
print('Open issues:', d['open_issues_count'])
"
```
---
## See Also
- `tco-framework.md` — TCO calculation templates and build vs buy decision scorecard
- `skills/build-vs-buy/SKILL.md` — full evaluation workflow with dimension weights
SKILL.md
---
name: business-ops
description: "Business operations: strategy, technology, growth, competitive intelligence, support, finance, HR, legal, operations, sales, productivity, product management."
user-invocable: false
allowed-tools:
- Read
- Write
- Bash
- Grep
- Glob
- Edit
routing:
triggers:
# Strategy/CEO
- "should we"
- "evaluate opportunity"
- "trade-off"
- "worth it"
- "invest in"
- "strategy"
# Technology/CTO
- "build vs buy"
- "vendor evaluation"
- "adopt"
- "technology choice"
- "tech stack"
# Growth/CMO
- "grow audience"
- "growth"
- "brand"
- "positioning"
- "community building"
# Competitive
- "competitive analysis"
- "market landscape"
- "differentiation"
# Evaluation
- "feasibility"
- "effort estimate"
- "ROI"
- "priority"
- "project evaluation"
- "go no go"
- "viability"
# Customer Support
- "customer support"
- "ticket triage"
- "support response"
- "knowledge base"
- "KB article"
- "escalation"
- "customer research"
# Finance
- "finance"
- "journal entry"
- "reconciliation"
- "variance analysis"
- "financial statements"
- "financial audit"
- "month-end close"
- "SOX"
# HR
- "HR"
- "human resources"
- "recruiting"
- "performance review"
- "compensation"
- "hiring"
- "onboarding"
- "org planning"
# Legal
- "legal"
- "contract review"
- "compliance check"
- "NDA"
- "legal risk"
- "legal brief"
- "vendor check"
- "german compliance"
- "DSGVO"
- "GoBD"
- "TDDDG"
- "AI Act compliance"
- "eIDAS"
# Operations
- "operations"
- "vendor review"
- "runbook"
- "process documentation"
- "risk assessment"
- "capacity plan"
- "change management"
- "compliance tracking"
# Sales
- "sales"
- "call prep"
- "pipeline review"
- "forecast"
- "draft outreach"
- "prospect research"
- "competitive intelligence"
# Productivity
- "productivity"
- "task management"
- "daily plan"
- "weekly review"
- "meeting agenda"
- "focus time"
- "goal setting"
- "status update"
- "time management"
- "prioritize tasks"
- "standup"
- "retrospective"
# Product Management
- "product management"
- "feature spec"
- "PRD"
- "roadmap"
- "stakeholder update"
- "user research"
- "sprint planning"
- "product metrics"
not_for: "micro library choices (use decision-helper), writing content, SEO of specific posts, or tactical marketing competitive analysis (use marketing) — this is executive strategy, not campaign execution. Code security audits, vulnerability scanning, or auth-flow reviews (use security-review) — only financial/accounting audit and SOX compliance. Code performance review (use reviewer-code) — this covers people performance reviews and HR operations. Software task specs, requirements, or plan-lifecycle management (use planning) — this skill prioritizes and tracks work, not specs. UX design methodology, wireframes, or accessibility audits (use design) — this handles product strategy, roadmaps, user research for feature prioritization."
complexity: Medium
category: decision-support
pairs_with:
- marketing
- data-analysis
---
# Business Operations
Umbrella skill for all business functions: executive strategy (CEO/CTO/CMO), competitive intelligence, project evaluation, customer support, finance, HR, legal, operations, sales, productivity, and product management. Each domain loads its own reference files on demand — this skill detects the mode, loads the right references, and executes the appropriate framework.
**Scope**: Business decisions and operational workflows. Use decision-helper for technical architecture micro-choices, domain agents for code, voice-writer for content, and systematic-debugging for debugging.
---
## Mode Detection
Classify the user's request into exactly one mode before proceeding. If the request spans multiple modes, choose the primary one and note the secondary.
| Mode | Signal Phrases | Reference |
|------|---------------|-----------|
| **STRATEGY** | Market entry, partnerships, resource allocation, opportunity, "should I/we", strategic pivots, investment | `references/csuite.md` |
| **TECHNOLOGY** | Build vs buy, vendor, SaaS, tech stack, architecture, adopt, technology choice | `references/csuite.md` |
| **GROWTH** | Content strategy, audience, SEO, marketing, brand, community, positioning, channel | `references/csuite.md` |
| **COMPETITIVE** | Competitor, competition, market landscape, differentiation, positioning against, market share | `references/csuite.md` |
| **EVALUATION** | Feasibility, effort estimate, ROI, priority, go/no-go, viability, "is it worth it" | `references/csuite.md` |
| **SUPPORT** | Customer support, ticket triage, support response, knowledge base, KB article, escalation | `references/customer-support.md` |
| **FINANCE** | Finance, journal entry, reconciliation, variance analysis, financial statements, audit, SOX, month-end close | `references/finance.md` |
| **HR** | HR, recruiting, performance review, compensation, hiring, onboarding, org planning | `references/hr.md` |
| **LEGAL** | Legal, contract review, compliance check, NDA, legal risk, legal brief, vendor check, DSGVO, GoBD | `references/legal.md` |
| **OPERATIONS** | Operations, vendor review, runbook, process documentation, risk assessment, capacity plan, change management | `references/operations.md` |
| **SALES** | Sales, call prep, pipeline review, forecast, draft outreach, prospect research, competitive intelligence | `references/sales.md` |
| **PRODUCTIVITY** | Productivity, task management, daily plan, weekly review, meeting agenda, focus time, goal setting, standup | `references/productivity.md` |
| **PRODUCT** | Product management, feature spec, PRD, roadmap, stakeholder update, user research, sprint planning, metrics | `references/product-management.md` |
---
## Reference Loading Table
Load references based on the detected mode. Load only the references required by the mode.
| Signal | Mode | Reference |
|--------|------|-----------|
| Market entry, partnerships, resource allocation, opportunity | STRATEGY | `references/strategic-frameworks.md`, `references/decision-matrices.md` |
| Build vs buy, vendor, SaaS, tech stack, architecture | TECHNOLOGY | `references/tco-framework.md`, `references/vendor-evaluation.md` |
| Content, audience, SEO, marketing, brand, community | GROWTH | `references/audience-segmentation.md`, `references/channel-evaluation.md` |
| Competitor, market landscape, positioning, differentiation | COMPETITIVE | `references/competitive-mapping.md`, `references/market-positioning.md` |
| Feasibility, effort, ROI, priority, go/no-go | EVALUATION | `references/feasibility-scoring.md`, `references/roi-frameworks.md` |
| Ticket triage, support response, KB article, escalation, customer research | SUPPORT | `references/customer-support.md` |
| Journal entry, reconciliation, variance, financial statements, audit, SOX | FINANCE | `references/finance.md` |
| Recruiting, performance review, compensation, hiring, onboarding, org planning | HR | `references/hr.md` |
| Contract review, compliance check, NDA, legal risk, legal brief, DSGVO, GoBD | LEGAL | `references/legal.md` |
| Vendor review, runbook, process documentation, risk assessment, capacity plan, change management | OPERATIONS | `references/operations.md` |
| Call prep, pipeline review, forecast, draft outreach, prospect research | SALES | `references/sales.md` |
| Task management, daily plan, weekly review, meeting agenda, goal setting, standup | PRODUCTIVITY | `references/productivity.md` |
| Feature spec, PRD, roadmap, stakeholder update, user research, sprint planning, metrics | PRODUCT | `references/product-management.md` |
---
## Instructions
For each mode, load the corresponding reference file for the full framework and instructions:
- **STRATEGY, TECHNOLOGY, GROWTH, COMPETITIVE, EVALUATION**: Load `references/csuite.md`
- **SUPPORT**: Load `references/customer-support.md`
- **FINANCE**: Load `references/finance.md`
- **HR**: Load `references/hr.md`
- **LEGAL**: Load `references/legal.md`
- **OPERATIONS**: Load `references/operations.md`
- **SALES**: Load `references/sales.md`
- **PRODUCTIVITY**: Load `references/productivity.md`
- **PRODUCT**: Load `references/product-management.md`
---
## Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| Too many options | 5+ options creating paralysis | Eliminate obviously inferior options first. Get to 2-4 before running full framework. |
| Not enough information | User cannot answer framing questions | Identify 2-3 critical unknowns. Recommend time-boxed research sprint before deciding. |
| Analysis paralysis | Keeps adding criteria or second-guessing | Apply reversibility test. If reversible, recommend best current option with checkpoint. |
| Emotional attachment | User has already decided, wants validation | Name the pattern directly. Ask: stress-test the choice, or genuinely evaluate all options? |
---
## References
| Reference | When to Load | Content |
|-----------|-------------|---------|
| `references/csuite.md` | Any executive strategy mode | Full C-suite decision support frameworks: STRATEGY, TECHNOLOGY, GROWTH, COMPETITIVE, EVALUATION |
| `references/strategic-frameworks.md` | STRATEGY mode | Porter's Five Forces, SWOT scoring, OKR alignment matrices |
| `references/decision-matrices.md` | STRATEGY mode | Weighted decision matrices, ICE/RICE scoring, pre-mortem templates |
| `references/tco-framework.md` | TECHNOLOGY mode | TCO templates, hidden cost checklists, migration cost models |
| `references/vendor-evaluation.md` | TECHNOLOGY mode | Vendor scorecards, RFP criteria, red flag detection, contract checklist |
| `references/audience-segmentation.md` | GROWTH mode | ICP scoring matrix, persona templates, segmentation frameworks |
| `references/channel-evaluation.md` | GROWTH mode | Channel scoring matrices, CAC/LTV models, funnel stage mapping |
| `references/competitive-mapping.md` | COMPETITIVE mode | Landscape map templates, feature matrices, activity tracker |
| `references/market-positioning.md` | COMPETITIVE mode | Positioning maps, differentiation scoring, win/loss frameworks |
| `references/feasibility-scoring.md` | EVALUATION mode | Three-dimension feasibility model, confidence calibration, decision tree |
| `references/roi-frameworks.md` | EVALUATION mode | T-shirt sizing, three-point estimation, risk-adjusted NPV |
| `references/customer-support.md` | SUPPORT mode | Triage, response drafting, KB articles, escalation, customer research |
| `references/finance.md` | FINANCE mode | Journal entries, reconciliation, variance analysis, financial statements, audit/SOX |
| `references/hr.md` | HR mode | Recruiting, performance management, compensation, org planning, people analytics |
| `references/legal.md` | LEGAL mode | Contract review, compliance, NDA triage, risk assessment, legal writing |
| `references/operations.md` | OPERATIONS mode | Runbooks, risk assessment, vendor management, process docs, change management, compliance |
| `references/sales.md` | SALES mode | Call prep, pipeline analysis, outreach, competitive intelligence, forecasting |
| `references/productivity.md` | PRODUCTIVITY mode | Task management, daily/weekly planning, meeting optimization, status updates, goals |
| `references/product-management.md` | PRODUCT mode | Feature specs, roadmaps, stakeholder updates, research synthesis, metrics, sprint planning |