Load a run → every puzzle as a tappable tile → tap to cut, tag why. Cutting a puzzle removes its solution and manifest entry with it. Export the kept run + culling log; the log carries each puzzle's spec and metrics, so every cut becomes calibration data. Amber flags are informational only — they never cut anything.
Source run
No run loaded.
Verdicts
0 kept · 0 cut · 0 total
Tiles appear below once a run is loaded.
Cumulative dataset
Every exported session's verdicts accumulate on this device (metrics only; full recipes live in each session's downloaded log). Drop the export into a Claude chat to see which specs get cut — that data is what earns a flag promotion from informational to steering.