Skip to main content
Core10–16 hours

AI Feature Discovery Dossier

Prepare a decision-ready AI feature dossier covering the problem, capability fit, UX, metrics, risks, economics, and experiment plan.

product discoveryAI UXmetricsunit economicsrisk management

Scenario

Task

A team proposes an AI feature without a clear value metric. You need to decide whether it should be built, how quality will be measured, and under which conditions rollout must stop.

Step-by-step execution

1. Establish problem evidence

Outcome: The problem is validated rather than invented to justify AI.

Tasks

  • Collect user evidence
  • Describe the current workflow
  • Define the non-AI baseline

Checks

  • A viable non-AI alternative is documented

2. Evaluate capability fit

Outcome: The model’s strengths and limitations are known.

Tasks

  • Choose representative tasks
  • Describe failure impact
  • Define approval needs

Checks

  • No capability promise exists without an eval plan

3. Build metrics and economics

Outcome: Value is evaluated together with quality, latency, and cost.

Tasks

  • Build the metric tree
  • Calculate cost per task
  • Define thresholds

Checks

  • Vanity metrics are not used as release gates

4. Design the rollout

Outcome: The team has a bounded experiment and rollback path.

Tasks

  • Run offline eval
  • Define a pilot cohort
  • Add monitoring
  • Set stop conditions

Checks

  • The rollout is not a big-bang release

Acceptance criteria

  • A non-AI baseline exists
  • Metrics have thresholds
  • Risks have controls
  • Cost per task is estimated
  • Rollback criteria are documented

Assessment rubric

How the result is assessed

Passing score: 72/100 · Distinction: 90/100

Problem evidence

The problem is supported by user evidence and a non-AI baseline.

25 points

Insufficient

AI is searching for a problem.

Competent

Evidence and a baseline exist.

Strong

Quantified pain, alternatives, and segment differences are documented.

Evidence required

  • ✓ Link to an artifact or code
  • ✓ README with decisions and trade-offs
  • ✓ Evidence of completed checks
  • ✓ Problem brief

Capability fit and risks

Model limits, failure impact, and controls are documented.

25 points

Insufficient

Capabilities are claimed without tests.

Competent

Representative tasks and controls exist.

Strong

An eval dataset, approval policy, and residual risk are documented.

Evidence required

  • ✓ Link to an artifact or code
  • ✓ README with decisions and trade-offs
  • ✓ Evidence of completed checks
  • ✓ Capability-risk matrix

Metrics and economics

Value, quality, safety, latency, and cost have thresholds.

30 points

Insufficient

Only vanity metrics are defined.

Competent

Release metrics have thresholds.

Strong

Sensitivity analysis, routing options, and a margin model are documented.

Evidence required

  • ✓ Link to an artifact or code
  • ✓ README with decisions and trade-offs
  • ✓ Evidence of completed checks
  • ✓ Metric tree and unit economics

Experiment and rollout

Offline eval, pilot, monitoring, and rollback are bounded.

20 points

Insufficient

A big-bang launch is planned.

Competent

Phased rollout and stop conditions exist.

Strong

A holdout, counterfactual baseline, and sunset criteria are defined.

Evidence required

  • ✓ Link to an artifact or code
  • ✓ README with decisions and trade-offs
  • ✓ Evidence of completed checks
  • ✓ Experiment plan