1. Classify data by purpose, not convenience
Outcome: Only fields required for the specific analytical job enter the pipeline.
Tasks
- →Inventory columns and derived fields
- →Mark direct and indirect identifiers and sensitive categories
- →Bind every field to an explicit purpose
- →Remove or aggregate fields without task necessity
Checks
- ✓“Might be useful” is not a purpose
- ✓Derived fields inherit privacy classification from source inputs
- ✓The minimized view is reproducible from a version-controlled transformation