Skip to content

C1 spike: review methodology comparison + X1 benchmark + headless probe #28

Description

@SSFSKIM

Problem & intent

Before argus-review becomes the default review methodology, we must (a) learn what Claude's native review methods know that argus doesn't, and (b) build the measuring stick every later "no regression" claim stands on. This is child C1 of docs/doperpowers/specs/2026-07-26-claude-review-stack-roadmap.md — a SPIKE: deliverable is findings + a benchmark, never a merge to the stack.

Constraints

  • Compare argus-review v0.3.0 against: the official code-review plugin (~/.claude/plugins/marketplaces/claude-plugins-official/plugins/code-review/) and the built-in /review + /code-review sources (~/Developer/GitHub/codex_somersault/Claude Code Src/src/commands/review.ts, review/ incl. ultrareview, security-review.ts). Mechanism AND content (rubric wording, lens texts, false-positive guidance, confidence rubrics).
  • The codex-app detached path is settled by argus references/provenance.md — do NOT re-examine (human ruling; guardian review is a tool-use classifier, cloud-tasks is not the app review path).
  • The bench's argus runs go through the headless invocation path (G3's harness) — the context C4 will deploy.

Success criteria

  • G1: findings document — per methodology: mechanism map, content-level comparison, adoption candidates for argus each marked adopt/reject with reasons.
  • G2: the X1 benchmark exists (seeded-bug diffs + a few real previously-reviewed PRs; scoring = seeded-bug recall + false-positive count; codex baseline stable across two runs) and has produced baseline numbers for BOTH the codex engine and argus v0.3.0.
  • G3: headless-invocation probe — a daemon-style non-interactive argus invocation (pinned level) produces the Findings/Verdict contract captured to a file within a bounded wait. Failure flows back to the roadmap before C2 dispatches.

Open questions

  • Bench set composition/size — C1 fixes it (grow until the codex baseline is stable across two runs).
  • What "meaningful false-positive growth" means — C1 fixes it in X1; later children re-litigate none of it.

Decision log

  • Bench defined inside C1, not at C4 time: the codex baseline must be measured while codex is still the default (roadmap Decision Log).
  • Headless probe front-loaded here from C4 (external review F2): it depends on nothing downstream and de-risks the whole chain.

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:P1issue-tracker board priorityspikeissue-tracker board category: exploration spike — deliverable is findings, never a merge

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions