Merge open branches into main: fix stack (#54+#55+#56+#58), docs refresh (#57), nightly audit doc - #59
Conversation
3 vitest decision_gate.*.live.test.ts failures (missing /tmp snapshot gating), 1 Playwright upload-estimate-phase1.spec.ts timeout, and one mixDoctor.estimatePlr fallback flagged for human boundary review. No Phase 1 ground-truth mutations found. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Closes four orientation gaps in CLAUDE.md: a top-of-Architecture map that flags advisory/, experiments/, docs/, and tests/ground_truth/ as off-path; a Scripts-at-a-glance section covering scripts/ and apps/backend/scripts/; a pointer from the Phase 3 paragraph to SamplePlayback.tsx + sampleGenerationClient.ts; and a curl example plus allowlisted field paths on the csv_export.py module entry. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Oz <oz-agent@warp.dev>
Removes four classes of broken/placeholder content that made the Phase 2 recommendations read as engine output instead of producer-facing advice (audit finding #1). 1. Mix Chain card role text — `buildRoleSentence` used to fabricate "{stage phrase} by {verb}" by lowercasing the first letter of Gemini's `reason`. Because `reason` is a present-tense clause, every card produced ungrammatical splices like "Controls bass energy by ensures the extreme low-end mono…". Now renders the reason verbatim with a capitalized first letter and trailing period; HIGH-END cue suffix preserved as "(for …)". 2. Patch Framework `patchRole` — was a 7-key category-keyed fallback (`mapPatchRole`) that stamped duplicated placeholders ("Primary tone generator" on every SYNTHESIS card). Removed the field, the fallback, the JSX paragraph, and the contribution to the `inferProcessingGroup` text-concat. The category chip + per-card `whyThisWorks` already carry the bucket and explanation. 3. Patches/Mix Chain overlap — `buildPatchCards` now case-insensitively filters out devices that also appear in `mixAndMasterChain` so the Patches section stops re-listing chain devices. Synthetic fallbacks (Stereo Width, MIDI Clip Guide) are built from Phase 1 only and bypass the filter intentionally. 4. Interpretation Caution raw-JSON dump — the panel rendered `originalValue` verbatim, which for dropped recommendations is a JSON-dumped `AbletonRecommendation` (see `_stringify_warning_value` in `server_phase2.py`). Added `formatDroppedValue` helper that parses JSON-shaped values and renders a compact "device: X · parameter: Y · value: Z" summary. Non-JSON strings pass through; invalid JSON falls back to a truncated raw string. Also resolves the `$Saturator` symptom — Gemini's hallucinated `$`-prefixed device name now surfaces as a readable summary line instead of leaking through a raw JSON dump. 5. System Diagnostics dev-only leak — `validateNewFieldCoverage` emits "Phase 1 field 'X' is present… this warning is benign" coverage signals meant for the engine team, not producers. Added optional `audience?: 'dev' | 'user'` to `ValidationViolation`, marked `NEW_FIELD_UNCITED` as dev, and updated `Phase2ConsistencyReport` to filter dev-audience violations from the rendered table AND from the header counts (so the header doesn't read "5 warnings shown" above an empty table). The underlying `ValidationReport` still carries every violation for tests and offline analysis. Tripwire note: the backend grammar fix `_apply_phase2_grammar_fixes` in `server_phase2.py` only runs on the legacy `/api/phase2` endpoint, not the analysis-runs path the UI uses. After this change it becomes a redundant no-op against the new render shape; left untouched. Tests: 16 new unit cases across viewModel, UI rendering, validator, and consistency report. `npm run verify` passes (lint + 565 unit tests + build + 44 smoke tests). Live UI verified against saved run 855f1376… — all nine text-leak probes (`by ensures`, `by ducks`, `by generates`, `Primary tone generator`, `Texture and movement stage`, `Stereo placement stage`, raw `\$Saturator`, `this warning is benign`, "is present in the measurement payload but no Phase 2…") return false. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Audit Finding #3 — "the chain-of-custody citation, the product's whole differentiator, renders as a footnote." Adds a one-line `CitationHeadline` primitive inside the collapsed header of every Mix Chain / Patch / Sonic Element card so producers see the measurement evidence at scan-time without expanding. The expanded CitationBlock stays in the body unchanged — this is the collapsed-state companion, not a replacement. Each headline reads `{label} {value} →` with the arrow pointing into the device h4/h3 that follows in the title row, making the implicit "measurement justified device" statement visible — exactly the audit's literal example "Crest factor 8.2 dB → Glue Compressor." Confidence sibling pills (Solid / Workable / Rough / Unreliable, same ladder as ConfidenceBandBadge) ride alongside the value when the primary field has a paired *Confidence sibling. This preserves the chain-of-custody invariant: low-confidence measurements visibly hedge in the collapsed view too, not just after expansion. Implementation: 1. `CitationHeadline` added as a sibling export in `apps/ui/src/components/CitationBlock.tsx`. Shares the file's imports, the `CONFIDENCE_PILL_CLASSES` map, and the audit lineage. Composes existing pure helpers (`pickPhase1Value`, `formatCitedValue`, `humanizeFieldPath`, `pickPhase1Confidence`, `getConfidenceBand`, `formatBandPillLabel`) — no new resolver logic. 2. Mount points in `AnalysisResults.tsx`: - Mix Chain card header (~line 2056): between the title row and the role paragraph. - Patch card header (~line 2215): between the title row and the MetaBadgeList. - Sonic Element card header (~line 1912): between the title row and the summary paragraph. 3. Each mount is guarded by `card.phase1Fields.length > 0`. Cards without cited fields fall back to today's exact layout. Live verification against saved run 855f1376… (the audit reproducer): all 10 Mix Chain cards, 2 Patch cards, and 7 Sonic Element cards render the headline correctly with appropriate confidence pills (PUMPING STRENGTH 47% [ROUGH SKETCH 30%], SUPERSAW DETECTED [WORKABLE DRAFT 78%], ACID BASS DETECTED [SOLID SCAFFOLD 99%], NOTE TRANSCRIPTION 61% [WORKABLE DRAFT 61%], KEY C Minor [SOLID SCAFFOLD 93%], TEMPO 145 BPM [SOLID SCAFFOLD 100%]). Mobile (375px) verified — headlines fit, pills stay aligned, no overflow. Tests: 15 new cases — 10 in `tests/services/citationHeadline.test.ts` (mirrors `citationBlock.test.ts`) covering label/value/arrow render, null fallback, confidence-pill bands, value formatter parity, typography hooks. 5 in `analysisResultsUi.test.ts` covering Mix Chain / Patch / Sonic integration, empty-fields fallback, and the low-confidence Unreliable-band path. `npm run verify` passes: lint + 575 unit tests + build + 44 smoke. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ladder Audit Finding #4 — the producer used to see five competing vocabularies for "how much should I trust this number?" — HIGH/MED/LOW pills on Detected Characteristics cards, High/Moderate/Low chips on Confidence Notes, bare CONF X% text on Key/Character metric cards, SCORE X.XX badges on the BPM card (two render paths), and the canonical four-band ladder used only by the Session Musician panel. The chain-of-custody invariant in PURPOSE.md requires low-confidence measurements to produce visibly hedged advice — five vocabularies fragment that signal so the producer can't build a single mental model of trust. This PR promotes the existing four-band ladder (apps/ui/src/services/sessionMusician/confidenceBand.ts) to the canonical primitive across every confidence surface. Frontend-only; Gemini still emits HIGH/MED/LOW and the UI converts locally via a new `toConfidenceBand` normalizer. What changed: 1. New normalizer `toConfidenceBand(value)` in `confidenceBand.ts` accepts numeric 0-1 floats, 0-100 integers, Gemini string enums (HIGH/MED/LOW + High/Moderate/Low variants), and percent strings ("62%"). Returns the matching `ConfidenceBand` or null for unparseable input. The string enums map to band-mid values (HIGH→0.9, MED→0.6, LOW→0.3) so `formatBandPillLabel` reads as honest hedges. 2. `ConfidenceBandBadge` gains: - `variant: 'full' | 'compact'` — `compact` omits the copy paragraph so the pill can sit inline in card corners and metric-card footers. `full` (default) preserves the Session Musician panel behavior unchanged. - `band?: ConfidenceBand` override prop — when supplied, skips the `getConfidenceBand` round-trip. Useful when the caller pre-converts a string enum. 3. V1: Detected Characteristics cards (AnalysisResults.tsx ~1650) — bespoke HIGH/MED/LOW ternary pill replaced with `<ConfidenceBandBadge variant="compact" band={toConfidenceBand(item.confidence)} />`. `characteristicPillClass` helper kept alive only for the unrelated characteristic-name chips on the Character metric card (line 1000); flagged as a follow-up. 4. V2: Confidence Notes chips (AnalysisResults.tsx ~1187) — `toConfidenceBadges` viewModel return shape changed from `{ label, level }` (3-level legacy enum) to `{ label, band }` (canonical four-band ladder). `confidenceClass` helper deleted. 5. V3: Plain `CONF X%` text (3 sites) — replaced with `<ConfidenceBandBadge variant="compact" confidence={...} />` in: - AnalysisResults.tsx Key card footer - AnalysisResults.tsx Character card footer - MeasurementDashboard.tsx alternate Key card 6. V4: `SCORE X.XX` badges (2 sites) — replaced with band pills in AnalysisResults.tsx Tempo card and MeasurementDashboard.tsx Tempo card. `formatBpmScore` deleted from both files. Cross-Check ✓/✗ StatusBadge in MeasurementDashboard preserved — that's an agreement signal (do multiple BPM detectors agree?), not confidence. 7. Dead code retired: `normalizeConfidenceLevel`, `parseConfidenceScalar`, `confidenceClass`, two copies of `formatBpmScore`. `ConfidenceLevel` type and `MelodyInsightsViewModel.confidenceLabel` kept alive — they drive non-pill key/value text rows in Sonic Element melody insights (flagged for follow-up migration). 3-level → 4-band mismatch documented as an intentional refinement: the old `normalizeConfidenceLevel` mapped scalar 0.5-0.79 AND "medium" both to "Moderate"; the new normalizer maps "medium" to 0.6 → workable (consistent), and scalar 0.25-0.49 to "Rough sketch" (more granular than the old "Low" bucket). Tests: 20 new cases across 4 files — 1. `confidenceBand.test.ts` — 9 cases for `toConfidenceBand` (null/undefined fallback, 0-1 floats, 0-100 integers, all string enum variants, percent strings, NaN/Infinity, negative clamping). 2. `confidenceBandBadge.test.ts` — 4 cases for the compact variant and the `band` override prop (label-only when no confidence, tone override, Gemini-string routing). 3. `analysisResultsViewModel.test.ts` — rewrote the existing `toConfidenceBadges` shape assertion; added a `band: null` fallback case. 4. `analysisResultsUi.test.ts` — 5 integration cases proving each V1-V4 site renders band pill text and not the old vocabularies. Live UI verified against saved run 855f1376… — 36 band pills visible across the page (21 Solid scaffold, 7 Workable draft, 5 Rough sketch, 3 Unreliable). Zero `CONF X%` leakage. Detected Characteristics cards show the chain-of-custody hedge working as designed (Detuned Supersaw renders WORKABLE DRAFT in orange while the other 4 detections render SOLID SCAFFOLD in green). Cross-Check ✓/✗ structurally preserved (the saved run doesn't populate bpmAgreement, so it doesn't render, but the code path is untouched). `npm run verify` passes: lint + ~600 unit + build + 44 smoke. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
CHANGELOG.md "Unreleased" was missing nine shipped features (Phase 3 audition samples, URL ingest, admin DELETE bypass, source-audio route, CSV export, publicStatus, reassigned spectrogram, chain-of-custody audit overhaul, BatchedBandpass centralization, hosted runtime foundation, backend monolith split). Added them. CLAUDE.md's analyze.py CLI example only listed four flags; expanded to match the full surface (--standard, --pitch-note-only, --stem-dir, --stem-output-dir, --pitch-note-backend). Expanded the Frontend key service files list from 9 to 16 entries to cover the services that landed during the audit overhaul (httpClient, sampleGenerationClient, appliedRecommendations + userLabels, phase1Picker + phaseLabels, audioFile, fieldAnalytics + diagnosticLogs, sessionMusician helpers). apps/ui/AGENTS.md File Map likewise expanded from 4 to 14 service entries against the same list. docs/SAMPLE_GENERATION.md's "Snapshot integration" section claimed an `AnalysisRunSnapshot.stages.sampleGeneration` shape that does not exist in apps/ui/src/types/backend.ts — the snapshot tracks only measurement, pitchNoteTranslation, and interpretation. Replaced the stale TS block with a description that matches the actual on-demand artifact path (consistent with the doc's own line 47). apps/backend/AGENTS.md and apps/backend/ARCHITECTURE.md now flag symbolic_extract.py as orphaned and broken — it imports a removed BasicPitchBackend symbol from analyze.py, would raise ImportError at load, and is not referenced from any other module. Slated for removal. docs/history/README.md had `library-review-torchfx-2026-05-13.md` listed twice; removed the duplicate. docs/ARCHITECTURE_STRATEGY.md "Last updated" stamp bumped to May 2026 with a note that no shifts to the three-layer thesis or library decisions accompanied the recent feature work. Date stamps in apps/backend/AGENTS.md and apps/ui/AGENTS.md refreshed to 2026-05-17. https://claude.ai/code/session_01QEmFAUJSoRgTqPcFTfcnYW
…eader)
Closes six items from the audit's "Quick hits (low ROI, real)" list.
Each is independent; bundled into one PR because they're all small
surgical fixes.
1. QH1 — StickyNav clipped "MEASUREMENTS" at 1440px. Was
overflow-x-auto + min-w-max, which produced an invisible horizontal
scroll (macOS hides scrollbars by default; last pill looked clipped
with no scroll affordance). Switched to flex-wrap so pills flow onto
multiple rows at narrow widths; whitespace-nowrap prevents intra-
label wrap.
2. QH2 — Mix Chain card-number order badges removed. Cards are grouped
by processing stage AFTER ordering, so the order numbers appeared
out-of-sequence within each group ("1, 6, 8, 9 / 2, 4 / 5, 7 / 3 /
10"), which read as a presentation bug. The visual sequence within
each group is already meaningful; the badge added confusion without
information. Dropped.
3. QH4 — Mix Doctor Band Diagnostics table now shows the target
range alongside the optimal. Previously the column rendered only
the optimal dB (e.g. "-22.0"), so two bands with similar Delta dB
could land in different Issue buckets without the producer being
able to see why. The Issue is determined by absolute thresholds
(target.minDb / maxDb), not by the diff from optimal. Showing the
range makes the verdict legible at a glance — Sub Bass -6.4 norm,
range -16 to -8, exceeds the upper bound by 1.6 dB → TOO-LOUD;
Low Mids -14.5 norm, range -26 to -14, sits at the boundary →
OPTIMAL. Threshold logic itself was correct; this is a display fix.
4. QH5 — "Dense DAW Lab" link removed from the header. A prior audit
pass had de-emphasized it to muted text, but the new audit re-
flagged it for sitting in the primary flow's header without context
for what it is. Route stays accessible via direct URL (the
getAppViewHref helper is preserved for future re-introduction
behind a settings menu, and main.tsx still routes to the
DenseDawConcept component when activeView === 'daw-concept').
5. QH6 — header model selector visually demoted. Dropped the
"Interpretation Model" label and shrank the styling from a bordered
dropdown to a discreet text-secondary inline select. The audit
flagged the prior styling as foregrounding an AI-model choice an
intermediate producer has no basis to make. Selector stays
accessible (smoke spec requires it visible at desktop viewport;
`responsive-layout.spec.ts` and `upload-estimate-phase1.spec.ts`
both guard the `phase2-model-desktop` testid) but now reads as
background metadata. Aria-label + title attribute carry the long-
form context for assistive tech and hover discoverability.
6. QH7 — Mix Doctor section title renamed from "MixDoctor" to
"Mix Doctor". The audit flagged that the section header rendered
as "Mixdoctor" with a lowercase 'd', inconsistent with every other
section name. Root cause: `formatWord` in `utils/displayText.ts`
case-normalizes single-token CamelCase to sentence case, so
"MixDoctor" was rendering as "Mixdoctor". Two words ("Mix Doctor")
matches every other section name's pattern (Style Profile, Project
Setup, Track Layout, Mix & Master Chain, Patch Framework, etc.)
and renders cleanly through the existing pipeline.
`npm run verify` passes: lint + ~600 unit + build + 44 smoke. Live UI
verified against saved run 855f1376… with probes confirming all six
fixes:
- StickyNav.overflow-x-auto wrapper gone (pills wrap to 2 rows)
- Zero Mix Chain order badges in #section-mix-chain
- "Target (range)" header present, "Target dB" header gone
- "Dense DAW Lab" text absent from page
- "Interpretation Model" label absent from page
- "Mix Doctor" section title present, "Mixdoctor" absent
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…dline, unified confidence, quick-hits Includes the four-PR audit stack: - #54 fix(ui): stop Phase 2 results surface from leaking engine output - #55 feat(ui): surface primary citation in collapsed card headers - #56 feat(ui): unify five confidence vocabularies onto the canonical band ladder - #58 fix(ui): audit quick-hits bundle (StickyNav, Mix Chain, Mix Doctor, header)
… architecture Refresh stale docs against current code: - CHANGELOG.md Unreleased: add ~10 shipped features since PR #32 - CLAUDE.md: complete CLI flag list, add ~7 missing service files - apps/ui/AGENTS.md: complete File Map - docs/SAMPLE_GENERATION.md: drop nonexistent AnalysisRunSnapshot.stages.sampleGeneration - apps/backend/{AGENTS,ARCHITECTURE}.md: flag symbolic_extract.py as orphaned/broken - docs/history/README.md: dedupe torchfx review entry - Date stamps: bump to 2026-05-17
Adds audits/nightly-2026-05-16.md tracking 5 issues (4 test failures + 1 review item): - 3 vitest decision_gate.*.live.test.ts failures (missing /tmp snapshot gating) - 1 Playwright upload-estimate-phase1.spec.ts timeout - 1 mixDoctor.estimatePlr fallback flagged for human boundary review No Phase 1 ground-truth mutations.
slittycode
left a comment
There was a problem hiding this comment.
Verdict: APPROVE
Summary
Five frontend-only branches (UI fixes, docs refresh, nightly audit doc) landing in one merge commit. No backend code touched. All 603 unit tests pass; TypeScript is clean. The three substantive changes — unifying confidence vocabulary onto the four-band ladder, surfacing primary citations in collapsed card headers, and suppressing dev-only validator noise from the System Diagnostics panel — each have targeted test coverage.
Findings
Blocking: none.
Should fix: none.
Worth considering:
-
report.summary.errorCount/warningCountinValidationReportstill include dev-audience violations, whilePhase2ConsistencyReport.tsxcorrectly rendersuserErrorCount/userWarningCountfrom a filtereduserVisiblearray. The design is intentional and documented. If a second render site ever consumesreport.summary, updatesummaryto carry user-visible counts separately rather than patching the new consumer. Fine at one call site. -
formatDroppedValueis unexported and lives insideAnalysisResults.tsx. Covered byanalysisResultsUi.test.ts(JSON-object case + plain-string passthrough). If a second call site emerges, move it to a utility module.
Test results
603 / 603 pass (48 test files). tsc --noEmit clean.
Phase boundary check
Clean. Phase 1 measurements are never mutated or re-derived. The buildPatchCards deduplication filters abletonRecommendations against mixAndMasterChain — both Phase 2 sources. The buildRoleSentence rewrite is actually a boundary improvement: the old code spliced Gemini's reason clause into a hardcoded {prefixMap[group]} by {lowercased reason} template; the new version renders reason verbatim and only appends the (for …) suffix when Phase 1 high-end cues are present.
Generated by Claude Code
Summary
Lands six branches into main in a single PR. Three merge commits, no conflicts during merge, full UI verify gate green (lint + 603 unit tests + production build).
Merged
fix/*audit series. Since each stack PR is a single commit on top of the previous, this one merge brings:fix(ui): stop Phase 2 results surface from leaking engine output(a1218f0) — kills four classes of broken/placeholder content from Phase 2 cards.feat(ui): surface primary citation in collapsed card headers(2f2fd86) — addsCitationHeadlineto Mix Chain / Patch / Sonic Element cards.feat(ui): unify five confidence vocabularies onto the canonical band ladder(2bba337) — singletoConfidenceBandnormalizer across every confidence surface.fix(ui): audit quick-hits bundle (StickyNav, Mix Chain, Mix Doctor, header)(3d91a3f) — six independent surgical fixes from the audit's quick-hits list.docs: refresh stale items against current architecture(9ad1c4a) — CHANGELOG / CLAUDE.md / AGENTS.md / docs/* sync.audits: nightly 2026-05-16 — 5 issues(26aca23) — adds the audit doc only.When this PR lands, GitHub will auto-close the four stacked PRs (#54, #55, #56, #58) and #57.
Skipped — orphan / dead branches
Five remote branches were evaluated and skipped. All but one have no common ancestor with main (separate root commits), and merging them would delete 65k-74k lines of main while adding 10k-17k lines of older snapshots that include binary artifacts (sqlite DBs, .pptx, .docx, .mp3). Their feature work has already shipped in main via earlier PRs.
codex/ground-truth-readmeci-fix-main-workflows.github/workflows/ci.yml. Not a fix to land.claude/awesome-brahmaguptaBACKLOG.md.claude/elegant-swansonPLAN-phase1-hardening.mdplanning doc + snapshot. Plan was superseded by the work in main.codex/fix-spectral-detail-and-chordPlan stereo detail expansions,Draft plan for analyze.py updates); also includes a "fix: update subproject commit reference for awesome-brahmagupta" gitlink commit. Work shipped via other PRs.These branches are good candidates for deletion in a follow-up — not done here to preserve the option of cherry-picking anything still useful.
Test plan
cd apps/ui && npm run lint— TypeScript clean.npm run test:unit— 603/603 passing across 48 test files (includes the 30+ new tests from the fix stack:citationHeadline.test.ts,confidenceBand.test.ts,confidenceBandBadge.test.ts,phase2ConsistencyReport.test.ts, and the expandedanalysisResultsUi.test.ts/analysisResultsViewModel.test.ts/phase2Validator.test.ts).npm run build— Vite production build succeeds, 2146 modules, bundle sizes nominal.apps/backend/AGENTS.md/ARCHITECTURE.mddocs only).855f1376-ec1e-4e3e-b145-bbf6eb479296.Reviewer notes
--no-ffwith a descriptive message so the PR series is preserved ingit log --graph. Each underlying commit (a1218f0,2f2fd86,2bba337,3d91a3f,9ad1c4a,26aca23) is still individually reviewable.CLAUDE.mdandapps/ui/src/App.tsxhad touching changes from multiple branches that auto-merged cleanly via theortstrategy.docs/LAYER2_EVALUATION.mdis added from the fix stack (originally fromfix/phase2-content-leaks); not a separate change.Generated by Claude Code