Problem & intent
Baseline-r1 adjudication surfaced genuine unseeded defects and two imperfect baits in the X1 fixtures (tests/review-bench/). Before the next scored run-id (C2's re-run), fold the adjudicated reality back into the fixtures so recall/FP stay clean: either fix each unseeded defect in base/patch or promote it into truth.json; tighten the two baits that drew defensible flags. Full list: results/2026-07-26-baseline-r1/scores.json genuine_unseeded_excluded + notes.md.
Constraints
- Do not change any seeded bug or its truth entry (comparability across run-ids).
- The case3 fixture needs an explicit single-threaded statement (README/base comment) so the recurring TOCTOU finding is unambiguously refutable — both engines flagged it.
- case5's bounded-retention bait: state the retention bound where a reviewer will see it.
Success criteria
- A fresh engine run on the updated fixtures produces no finding that is neither truth-matched nor a deliberate bait.
- scores.json for the next run-id needs no "genuine-unseeded" escape hatch entries for known items.
Open questions
- none — mechanical fixture maintenance guided by the adjudication record.
Decision log
Problem & intent
Baseline-r1 adjudication surfaced genuine unseeded defects and two imperfect baits in the X1 fixtures (
tests/review-bench/). Before the next scored run-id (C2's re-run), fold the adjudicated reality back into the fixtures so recall/FP stay clean: either fix each unseeded defect in base/patch or promote it into truth.json; tighten the two baits that drew defensible flags. Full list:results/2026-07-26-baseline-r1/scores.jsongenuine_unseeded_excluded+ notes.md.Constraints
Success criteria
Open questions
Decision log