From 011e7646e8938bbd420196d61ce347966b502a36 Mon Sep 17 00:00:00 2001 From: BigSimmo <87357024+BigSimmo@users.noreply.github.com> Date: Fri, 24 Jul 2026 08:09:11 +0800 Subject: [PATCH] docs: make outstanding issues the universal task ledger --- AGENTS.md | 9 +++-- docs/README.md | 1 + docs/branch-review-ledger.md | 2 +- docs/codebase-index.md | 1 + docs/outstanding-issues.md | 76 ++++++++++++++++++++++++++++++++++-- 5 files changed, 81 insertions(+), 8 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 32d9bc6c0..94d7841e0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -450,9 +450,12 @@ Run the matching planner command in `docs/productivity-workflows.md` without sid ## Outstanding-work memory (`/issues`) -`docs/outstanding-issues.md` is the durable, cross-session memory of every outstanding **task**, -**recommendation**, and **issue** for this repo. Chat context resets between sessions; that file does -not, so anything worth remembering after a session ends belongs there. +`docs/outstanding-issues.md` is the single universal, durable, cross-session ledger for every +outstanding **task**, **recommendation**, and **issue** in this repo. It owns the recommended +execution order, acuity, required capability, timing, effort, approvals, evidence, status, success +criteria, stop rules, and resolution history. Chat context resets between sessions; that file does +not, so anything worth remembering after a session ends belongs there. Do not create or maintain a +second task ledger. - When the user types `/issues`, invoke the `issues` skill (`.claude/skills/issues/SKILL.md`): read `docs/outstanding-issues.md` and state the open items back, grouped by priority. A plain `/issues` diff --git a/docs/README.md b/docs/README.md index 0b743e80c..0d1c904af 100644 --- a/docs/README.md +++ b/docs/README.md @@ -71,6 +71,7 @@ npm run docs:check-links ## Plans and workstreams (living) +- [outstanding-issues.md](outstanding-issues.md) — universal task ledger, recommended execution order, evidence, status, and resolution history - [maturity-backlog-workorders.md](maturity-backlog-workorders.md) — actionable work orders tracking the repository-maturity audit backlog - [framework-dependency-modernization-checklist.md](framework-dependency-modernization-checklist.md) — ordered Next.js 16, runtime, dependency, Turbopack, and verification migration program - [search-rag-master-plan.md](search-rag-master-plan.md) / [search-rag-master-context.md](search-rag-master-context.md) — search/RAG roadmap and shared context diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index b9b2a8b96..f8f2edb72 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -20,7 +20,7 @@ Use this ledger to prevent repeated branch and PR reviews when the reviewed HEAD | Date | Branch or ref | Reviewed HEAD | Scope | Outcome | Checks | | ---------- | -------------------------------------------------------- | ---------------------------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| 2026-07-24 | `codex/supabase-document-change-trigger` | `9c7d9edf509a51478f5bebbabcca64e3926dc877` + reviewed working diff | Document-change ingestion trigger migration, schema mirror, grants, privacy and fail-safe delivery | APPROVE. No P0-P2 finding. The trigger is update-only, acts solely on a strict JSON boolean false/absent-to-true transition, sends only the receiver's allowlisted owner-scoped fields, fails open for document writes when Vault/GUC/pg_net is unavailable, and revokes execution from public/anon/authenticated. No production URL fallback exists. Highest residual risk is deliberate pg_net at-most-once delivery; the clear-then-flip recovery and data-preserving rollback are documented, and the trigger remains inert until both the Vault secret and environment base-URL GUC are configured. | Disposable `supabase/postgres:17.6.1.127` schema replay and drift-manifest regeneration passed (16s; scratch container removed); focused schema/drift/receiver Vitest 89/89; migration-role, function-grant (30 SECURITY DEFINER functions) and owner-scope guards; production-readiness CI mode READY with expected secretless-worktree warnings; offline RAG 21 suites/307 tests; `verify:cheap` 365 files, 3,241 passed/1 skipped; static trace of receiver payload, authoritative owner-scoped reload and idempotent enqueue path. No live provider mutation or migration apply. | +| 2026-07-24 | `codex/supabase-document-change-trigger` | `9c7d9edf509a51478f5bebbabcca64e3926dc877` + reviewed working diff | Document-change ingestion trigger migration, schema mirror, grants, privacy and fail-safe delivery | APPROVE. No P0-P2 finding. The trigger is update-only, acts solely on a strict JSON boolean false/absent-to-true transition, sends only the receiver's allowlisted owner-scoped fields, fails open for document writes when Vault/GUC/pg_net is unavailable, and revokes execution from public/anon/authenticated. No production URL fallback exists. Highest residual risk is deliberate pg_net at-most-once delivery; the clear-then-flip recovery and data-preserving rollback are documented, and the trigger remains inert until both the Vault secret and environment base-URL GUC are configured. | Disposable Supabase Postgres 17.6.1.127 image replay and drift-manifest regeneration passed (16s; scratch container removed); focused schema/drift/receiver Vitest 89/89; migration-role, function-grant (30 SECURITY DEFINER functions) and owner-scope guards; production-readiness CI mode READY with expected secretless-worktree warnings; offline RAG 21 suites/307 tests; `verify:cheap` 365 files, 3,241 passed/1 skipped; static trace of receiver payload, authoritative owner-scoped reload and idempotent enqueue path. No live provider mutation or migration apply. | | 2026-07-23 | PR #1090 / `cursor/fix-phone-dock-edge-1b1d` | `761de7e9ad623b6bd8d634d849a9eb465d622e48` (merged as `09028ef217209fceb53f1122ac7738b509bce323`) | Phone safe-area and edge-to-edge search-dock UI review | MERGED. No P0-P2 finding. The branch was three commits behind, so current `origin/main` was merged before landing; the actual merge tree matched the reviewed synthetic tree. The dock remains flush to the viewport with safe-area padding inside the form, and the phone shell no longer retains the `dvh` clamp that created the Safari toolbar band. Zero actionable review threads. | `npm run ensure`; focused `ui-tools.spec.ts` phone-home and edge-to-edge scenarios: Chromium 2/2 and WebKit 2/2; refreshed hosted policy, security, unit, build, advisory UI, Production UI and required aggregate checks green; exact-head ancestry and local-main tree equality proved after merge. | | 2026-07-22 | PR #1087 / `codex/reconcile-product-truth` | `edbc2260fef59ca2fa7c6973dffb85e32354bce1` (merged as `05dc52fd8408a65117e22a6236e43252203bea92`) | Product-truth copy, account persistence and unavailable-SSO presentation | MERGED. Cross-device claims now match favourites/preferences persistence; recent searches are identified as browser-session data; the contradictory “never shared” statement is removed. All unavailable setup providers and Apple elsewhere use the connected accessible “coming soon” placeholder pattern. The single review finding was fixed, replied to and resolved. | Red DOM proof; focused 19/19; `verify:cheap` 3,220 passed / 1 skipped; `verify:ui` 265/265; PR-local build/secret scan/offline RAG; final hosted required, Production UI, policy and security checks green. No provider calls or RAG spend. | | 2026-07-22 | PR #1086 / `codex/reconcile-xlsx-budgets` | `5376880a40749b6526fd7e4603a7be9d04bc9624` (merged as `2963fba46eacd644618a588fa283f7597faa2644`) | XLSX resource-boundary review | MERGED. Enforces worksheet, non-empty-row, rendered-cell and UTF-8 output ceilings before result fragments are appended; sparse-column output is preserved. No actionable review threads. | Red 257-sheet reproducer; focused 4/4; `verify:cheap` 3,218 passed / 1 skipped; PR-local build/scan/offline RAG; hosted required/security/policy green. | diff --git a/docs/codebase-index.md b/docs/codebase-index.md index c7d597ed5..aa92ad7ad 100644 --- a/docs/codebase-index.md +++ b/docs/codebase-index.md @@ -334,6 +334,7 @@ One shared composer (`master-search-header.tsx`) serves every mode. Placement: | Full documentation index | `docs/README.md` | | Routes and modes | `docs/site-map.md` | | Search/RAG roadmap | `docs/search-rag-master-plan.md` | +| Universal task ledger | `docs/outstanding-issues.md` | | Reindex operations | `docs/reindex-runbook.md` | | Production readiness | `docs/production-readiness-checklist.md` | | Capacity / scale-up | `docs/capacity-review.md`, `docs/auth-connection-cap-runbook.md` | diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index c4b899e31..9782a446f 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -7,6 +7,10 @@ Durable, cross-session memory of everything still outstanding for this repo: ope **Rule of thumb:** if it is worth remembering after this session ends, it belongs here. +This is the repository's **single universal task ledger**. It owns recommended execution order, +acuity, required capability, timing, effort, dependencies, approvals, evidence, status, success +criteria, stop rules, and resolution history. Do not create or maintain a second task ledger. + ## How this is used - Say `/issues` in Claude Code → the skill reads this file and states the open items back, @@ -27,7 +31,70 @@ Durable, cross-session memory of everything still outstanding for this repo: ope - Resolving an item moves its row to **Resolved / archive** with the date and a one-line outcome — rows are archived, not deleted, so the history stays auditable. - +## Execution scales + +- **A1 urgent:** active safety/privacy/data-loss/release blocker; none currently. +- **A2 important:** confirmed correctness, clinical, privacy, reliability, or evaluation-integrity + work that leads its available lane. +- **A3 planned:** worthwhile work deferred to its stated trigger. +- **Optional:** do only when measured need, ownership, and cost justify it. +- **Standard:** experienced generalist. **High:** senior cross-module reasoning. **Specialist:** + database/RAG/clinical/privacy/security/evaluation expertise. **Operator:** authorised human owner. +- Effort is active work, excluding approvals, hosted/provider waits, soak, and review time. + +The order below is planning guidance, not authority to call providers, spend money, change +production, commit, push, merge, or deploy. A waiting dependency does not block an independent +executable item below it. + +## Recommended execution queue + +| Order | Source | Recommended outcome / next action | Acuity / capability | When | Active effort | Dependencies / approval | Done, verification, and stop rule | +| ----: | ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------- | ----------------------------------------------------------- | --------------------------------------- | ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 1 | `#052` | Prevent full/retry reindex overlapping a fresh agent-enrichment lease; add red single/bulk tests, then reuse `hasActiveAgentEnrichmentJob` before mutation. | A2 / High concurrency | Ready; first | 0.5–1 day | Local/offline | Fresh processing leases conflict; stale leases and enrichment mode retain behavior. Run focused route/safety tests, `verify:cheap`, production-readiness. Stop if current `main` does not reproduce it. | +| 2 | `#030` | Require distinct source identities for distinct expected comparison slots; begin with a red shared-title/two-expectation test. | A2 / High evaluation semantics | Ready; after 1 | 2–4 hours | Local/offline; protected-eval review | One source cannot satisfy both slots; existing cases stay green. Run focused matching, typecheck, `verify:cheap`. Stop before alias/retrieval/ranking changes. | +| 3 | `#019` | Prove admission evidence is lost at the comparison fallback boundary; scope the smallest correction without automatically shipping it. | A2 / Specialist RAG | Reproducer now; behavior after 9 | 0.5–1 day; fix separate | Protected review; approved baseline/post canary for behavior | Reproducer fails while packing retains both sources. Stop without behavior change if the boundary defect is not independently reproducible. | +| 4 | `#026` | Prove aged queued-without-job recovery, then activate the merged document trigger only through the approved Supabase path and configure Vault/GUC. | A2 / Specialist recovery + Operator | Recovery now; live steps in approved window | 0.5–1 day local; 1–2 hours operator | PR #1100 merged; live Supabase approval | Recovery is idempotent under crash/concurrency/repeat; migration state, controlled event, rollback, and production-readiness are recorded. Never apply raw SQL live. | +| 5 | Privacy/legal package | Execute OpenAI/Railway DPAs; decide ZDR/residency; obtain cache/subprocessor answers and APP 8 plus APP 5/1 counsel sign-off. | A2 / Operator legal/privacy | Start now; before real-patient use/privacy-approved release | 4–8 hours internal; 1–6 weeks elapsed | Signers, counsel, providers | Record countersigned evidence, IDs, dates, decisions, and approved wording in PIA/cross-border docs. Stop before public-copy changes without counsel. | +| 6 | Safety identifier config | Set a unique 32+ character `OPENAI_SAFETY_IDENTIFIER_SECRET` per environment with `OPENAI_API_KEY`, never in Git/output. | A2 / Standard local + Operator hosted | Local now; hosted approved window | 15–30 min local; 30–60 min/env | Secret manager; hosted approval | Presence-only readiness passes, HMAC pseudonymity remains, secret scans clean. Stop if governance rejects stable identifiers. | +| 7 | Operator config backlog | Reconcile query-hash, deep-probe, Supabase/OpenAI, project identity, and schedules using read-only presence checks before setting anything. | A2 / Operator platform | Next approved readiness window | 1–2 hours | Provider approval, env owners | Record status, never values; set only confirmed gaps; identity/readiness pass. Stop on target ambiguity or cross-env key reuse. | +| 8 | `#022` | Decide auditable BMJ reference policy, then review the ten highest-impact local WA documents. | A2 / Operator governance + Specialist | Decision-ready | 1–2 hours policy; 0.5–1 day reviews | Clinical authority; live writes/canary approval | Preserve third-party/unverified provenance; record reviewer/evidence/time/rationale. Stop after ten and remeasure debt. | +| 9 | `#051` + `#023` | Compare run `30018289898` with the scheduled 2026-07-26 canary/browser/label artifacts; disposition residuals without rerun. | A2 / Specialist evaluation | About 2026-07-27 02:00 AWST plus runtime | 2–4 hours | Approval to read hosted artifacts; no dispatch | Record tree/run, content/provider/latency, browser, and labeling outcomes. Stop without spend or archived lithium changes. | +| 10 | `#018` | Diagnose lithium, ADHD, metabolic residuals independently: table fast path, extractive budget, schedule-free selection. | A2 / Specialist clinical RAG | After 9; one mechanism at a time | 1–2 days diagnosis; fixes separate | Protected review; approved behavior canary | Each has a current-main fixture and scoped candidate; stop without deterministic repro or on individual canary regression. | +| 11 | `#029` | Re-enumerate fallback stubs and improve one demonstrated causal cluster at a time without weakening fail-closed behavior. | A2 / Specialist answer safety | After 9–10 | 0.5–1 day inventory; 1–3 days/fix | Approved answer validation | Replace historical 12/30 assumption; retain grounding/citation gates. Stop after each validated cluster or metric-gaming change. | +| 12 | `#001` | Reassess semantic reranking without changing default; enable only after accepted ambiguity comparison. | A2 / Specialist ranking | After 9 and rollout approval | 0.5–1 day plus canary review | Provider approval/budget | Keep off unless 36/36, recalls 1.0, zero case regressions, measured gain. Otherwise record keep-off decision. | +| 13 | `#033` | Decide whether governance metadata enters the LLM source block, distinguishing unknown from adverse status. | A3 / Specialist prompt governance | After 8–9 | 1–2 days plus eval | Better metadata, stable canary, provider approval | Prompt/serialization tests; no grounded-supported drop; zero citation failures. Stop on over-caveating/degradation. | +| 14 | `#024` | Separate Playwright `_rsc` interception failure from a real Safari defect. | A3 / High Next.js/WebKit | After 9’s matrix data or real Safari repro | 0.5–1 day | Device/provider check only if local ambiguity | Test-only fix needs discriminating repro and meaningful access-control assertions; keep Chromium green. | +| 15 | `#007` | Choose canonical Tools experience; align nav, redirects, sitemap, reachability. | A3 / Operator product + Standard frontend | Product decision window | 15–30 min decision; 0.5 day code | Product owner | One canonical path and intentional redirect with green route/UI tests. Stop on unresolved standalone need. | +| 16 | `#025` + SLO alerts | Select owned deploy/CI/ingestion/SLO channels; configure only those, ingestion after 4. | A2 / Operator integrations | Approved observability window | 1–3 hours/channel | Provider approval, secrets, destination owner | Mocked tests then one controlled non-PHI event/channel. Stop on missing owner or unsafe delivery. | +| 17 | Full release gate | Run release/clinical gates once against an exact candidate, not a moving branch. | A2 / High release + Operator | Before release/handoff needing full confidence | 0.5 day plus waits | Exact SHA, provider approval, heavy-command lock | All required gates recorded against one SHA; stop/classify first failure and do not repeat unchanged pass. | +| 18 | Staging provision | Provision dedicated Clinical KB staging on Supabase/Railway with isolated keys and synthetic/non-clinical data. | A2 / Operator + Specialist DB | After cost/ownership approval | 0.5–1 day | Billable provider approval, owner | Identity, tenancy, secrets, migrations, app/worker health, data boundary pass. Stop on target ambiguity or production data/key reuse. | +| 19 | Staging soak/rollback | Run documented soak and rollback rehearsal in dedicated staging. | A2 / Operator reliability | After 18 | 0.5 day plus soak | Staging and provider approval | SLO, rollback, data integrity, recovery evidence pass. Stop before production if unproven. | +| 20 | Production seed verification | Verify registry/differentials/medications non-empty before writes; seed only confirmed governed gaps. | A2 / Operator data + Specialist | Approved production window | 1–3 hours plus seed time | Correct owner/project, provider approval | Read-only proof precedes idempotent writes; stop if populated or target ambiguous. | +| 21 | `#011` | Switch Auth DB cap to percentage immediately before first compute scale-up. | A3 / Operator capacity | Scale-up trigger only | 30–60 min plus soak | Dashboard, planned scale-up, approval | Target only `sjrfecxgysukkwxsowpy`; record before/after and advisor/health recheck. Do not create/use staging for this task. | +| 22 | `#017` | Capture reproducible mobile/desktop production LCP, INP, CLS before performance work. | A3 / High web performance | Before 26; approved live window | 1–2 hours | Live-site approval; record route/build/throttle | Record metrics and accept/reject decision. Stop if acceptable or too noisy. | +| 23 | `#037` | Decide whether routine supported claims cap at medium trust. | A3 / Operator clinical-product + Standard frontend | Next trust-policy review | 30–60 min decision; up to 0.5 day code | Clinical/product authority | Record policy; if accepted, change flag/render expectations only and run focused tests. Stop without owner acceptance. | +| 24 | `#027` | Add uptime monitor outside GitHub/Railway only if service, budget, responder justified. | Optional / Operator SRE | Owned external-alert trigger | 1–2 hours | Vendor/privacy/owner/provider approval | Controlled non-PHI failure/recovery alerts. Stop without responder or if current monitoring accepted. | +| 25 | `#028` | Define vendor/region/retention/redaction/sampling/maps/owner before runtime error tracking. | Optional / Specialist privacy + Operator | After privacy/ownership/cost approval | 1–3 days | Privacy and provider approval | Redaction tests exclude clinical data/IDs/secrets; one non-sensitive event arrives. Stop if envelope unacceptable. | +| 26 | `#012` + `#013` + `#016` | Measure before one bundle/runtime/motion target: differentials index, catalogue chunk, prod mockup exclusion, static shell, sidebar, Therapy Compass, dialogs. | A3 / High Next.js performance | After 22 or equivalent evidence; one target | 0.5–3 days/target | Measured target, Next.js guide, UI verification | Define before/after metric, preserve behavior, run analysis/focused/`verify:cheap`/browser. Stop if gain small. | +| 27 | `#040` | Add small stable visual-regression baselines with owner/update workflow. | Optional / High visual QA | Stable surfaces and owner | 1–2 days | Stable browser, owner | Low-flake repeat runs and documented updates. Stop before blocking if churn high. | +| 28 | `#038` | Define shared comparison interaction contract before another comparison surface; keep clinical content mode-specific. | Optional / High design-system | Approved new comparison surface | 0.5–1 day | Concrete surface, product/design owner | Inventory patterns and make shared behavior testable. Stop if no new surface. | +| 29 | `#009` | Decide whether `/api/jobs` is intentional ops-only; document/test or remove. | A3 / Standard API ownership | Next API maintenance | 1–2 hours | Product/ops owner | Record owner/contract or remove unused route/tests; focused tests + `verify:cheap`. Stop if external consumers unclear. | +| 30 | `#035` | Expand conflict detection only for concrete clinically reviewed class with positive/negative fixtures. | A3 / Specialist evidence rules | Demonstrated missed conflict | 0.5–1 day design; code separate | Clinical review; provider approval only for live validation | Fixtures discriminate class without unrelated warnings. Stop if no bounded class. | +| 31 | `#036` | Decide explicit public-visibility flag vs `owner_id IS NULL` plus warning before any visibility migration. | A3 / Specialist tenancy/RLS | Before visibility/public-corpus migration | 2–4 hours decision | Security/clinical review; migration approval if adopted | Compare threat model/migration/rollback. Stop without schema if current control preferred. | +| 32 | `#039` | Converge repeated catalogue toolbar behavior only during a concrete toolbar project. | Optional / High frontend architecture | Concrete project | 0.5–1 day inventory; 1–3 days code | Product/design owner, UI verification | Prove shared behavior without flattening search semantics; stop after bounded contract. | +| 33 | `#010` | Wire a disabled Forms/Favourites placeholder only when its feature is selected. | Optional / Standard frontend-product | Per approved feature; no work now | 0.5–2 days/feature | Product scope/data contract | Action works with UI/a11y coverage. Do not batch-build unrequested placeholders. | +| 34 | `#041` | Extend Easy Read/Standard Factsheets for future patient content; no second patient mode. | Optional / High content/a11y/governance | Concrete need + source plan | 0.5–2 days/increment | Product/content/governance | Extend/test existing contracts. Stop if need/source undefined. | +| 35 | `#021` | Revisit conservative generation discards only for measured latency/waste or cheaper mechanism. | Optional / Specialist quality/cost | Evidence trigger only | 2–4 hours diagnosis; 1–2 days/candidate | Evidence and approved provider budget | Reduce cost/latency without weaker gates. Stop if cost/benefit remains poor. | + +### Queue maintenance + +1. Revalidate against current `main` before starting and remove/rewrite contradicted work. +2. Do not combine protected RAG residuals or change scores, comparators, aliases, clamps, or semantic + reranking without separate reproducers and required validation. +3. Treat provider status as a claim: verify only after approval and never store secret values. +4. Close/reclassify when a success or stop condition is met; do not preserve work for its own sake. + + ## Open items @@ -39,6 +106,7 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | ---- | --- | ----- | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | #051 | P2 | task | Stabilise the live answer-quality canary before more RAG tuning | Diagnostics landed in PR #1095: structured JSON/Markdown artifacts now record the actual checked-out SHA, run identity and latency context, and the offline trend tool separates content, provider-route and latency outcomes. First validating run `30018289898` recorded the expected tree and cost, with 36/36 retrieval green, but one report cannot establish variability; PR #1097 prevents a single failure being mislabeled as repeated. Next: compare the scheduled 2026-07-26 structured report with this run. Do not spend on an immediate retry or reapply the archived lithium guard before that comparison. | PR #1095; run `30018289898`; PR #1097; archive ref `refs/archive/rejected-rag/20260723/monitoring-subject-gate` | 2026-07-23 | | #001 | P2 | task | Semantic reranking still gated off | `RAG_SEMANTIC_RERANK_ENABLED=false` from PR #901. Do not enable until the provider-backed 36/36 retrieval-quality gate **and** an ambiguity-focused canary are explicitly approved and recorded. | `docs/process-hardening.md` (Semantic reranking rollout debt); PR #901 | 2026-07-21 | +| #052 | P2 | issue | Reindex can overlap a fresh agent-enrichment pass | Full/retry reindex preflights check `ingestion_jobs` but do not call the existing `hasActiveAgentEnrichmentJob`. A fresh `indexing_v3_agent_jobs.status='processing'` lease can therefore overlap destructive artifact work. Extend single and bulk full/retry preflights before mutation; keep stale leases non-blocking and preserve enrichment-mode behavior. | `src/lib/ingestion-mutation-safety.ts:116-159`; single/bulk reindex routes; ingestion audit 2026-07-24 | 2026-07-24 | | #005 | P3 | rec | `finalScore` saturates at clamp ceiling | Base + ~40 stacked boosts routinely exceed 1.0, so strong matches tie at 1.0 and order by an arbitrary `document_id` tiebreak. If ranking is ever revisited, break ties by the **pre-clamp** score rather than raising the `[0,1]` ceiling (downstream gates assume `[0,1]`). Ordering already sorts by the unbounded pre-clamp `rankScore` (`clinical-search.ts:1735,1927,1950-1955`), so the clamp confines only the reported confidence value, not result order. Not a defect on the current golden set; any change here is a protected RAG surface (canary required). | `docs/rag-hybrid-findings-and-todo.md` P1 item 4; `src/lib/clinical-search.ts:1735` | 2026-07-21 | | #007 | P3 | rec | `/tools` vs `/?mode=tools` parallel Tools entry points | `/tools` (standalone `ApplicationsLauncherPage`) has no inbound in-app link; the sidebar Tools item uses `/?mode=tools`. Decide the canonical entry point and wire nav consistently, or drop the standalone `/tools` page + `/applications` redirect. Currently allowlisted in `tests/route-reachability.test.ts`. | `src/app/tools/page.tsx`; `src/app/applications/route.ts` | 2026-07-21 | | #009 | P3 | rec | Confirm `/api/jobs` is intentionally server/ops-only | No client `fetch()` reaches `/api/jobs` (only tests import it). Confirm it is a deliberate ops/manual surface; if abandoned, remove it. | `src/app/api/jobs/route.ts` | 2026-07-21 | @@ -46,7 +114,6 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | #011 | P3 | task | Auth DB-connection allocation is operator-only | Supabase Auth (GoTrue) is capped at ~10 absolute DB connections (Supabase perf advisor). Switch to **percentage-based** allocation in the Supabase **dashboard** before the first compute scale-up — **not settable via SQL/MCP** (operator-owned). Verify via a staging soak + an approval-gated read-only advisor re-check. | `docs/auth-connection-cap-runbook.md`; `docs/process-hardening.md` (Known follow-up debts) | 2026-07-21 | | #012 | P3 | rec | Slim the lazy cross-mode differentials chunk | `cross-mode-differentials.ts` is dynamically imported (correctly code-split **out** of the initial/dashboard bundle — verified), but it pulls the full ~860 KB differentials snapshot (~125 KB gzip lazy chunk) just to build a tiny `{slug,title,clinicalHinge}` + presentations + aliases catalog. A precomputed lightweight index (generator + drift check, like the `specifiers-content` split / medications `fields=index`) would cut that lazy chunk ~5–10×. Not a bundle leak — an M-effort slim. | `src/lib/cross-mode-differentials.ts`; `src/components/clinical-dashboard/cross-mode-links.tsx:150`; session 2026-07-21 (build:analyze) | 2026-07-21 | | #013 | P3 | rec | Route-chunk + mockup catalogue JSON weight | `build:analyze`: `/specifiers` ships `specifiers-search-index.json` (~180 KB parsed), `/forms` ships `forms-catalog.json` (~132 KB), `/formulation` ships `formulation-content.json` (~52 KB, client-side local search — needs index/full split or a search endpoint, architectural). All route-scoped (not initial bundle). Also `*-mockups.tsx` (~100 KB across chunks) build though `/mockups` 404s in prod — exclude from the prod artifact. | session 2026-07-21 (build:analyze) | 2026-07-21 | -| #014 | P3 | rec | Realize the `next/image` win on signed previews | `next.config` `images` (AVIF + `*.supabase.co` `remotePatterns` pinned to the project host, from #1024) is currently inert — signed document/image previews still render as raw ``. Route them through `next/image` to actually get AVIF + lazy optimization. | #1024; `src/components/clinical-dashboard/signed-image.tsx`; session 2026-07-21 | 2026-07-21 | | #016 | P3 | rec | "Big but not easy" structural + motion perf | Deferred larger levers: (a) nonce-CSP forces every product route to `ƒ Dynamic` (zero static generation) — evaluate Partial Prerendering / static shells for the static clinical catalogues (DSM/differentials/therapy/specifiers/formulation); (b) sidebar expand/collapse animates `grid-template-columns` (biggest smoothness cost, motion-gated — needs a transform-overlay rethink); (c) Therapy Compass fetches 692 KB / 2.5 MB JSON client-side (defer until interaction + confirm brotli); (d) settings/setup/admin dialogs static-imported into the home chunk (`next/dynamic` them). | session 2026-07-21 (build route table + design audit) | 2026-07-21 | | #017 | P3 | task | Field Web-Vitals baseline via live Lighthouse | In-sandbox runtime vitals were blocked (prod server hard-requires Supabase secrets; dev-mode CLS measured excellent at 0.00–0.04, content-first pages 0.000). Run Lighthouse against `psychiatry.tools` for real LCP/INP/CLS to prioritize #012–#016 by measured impact rather than reasoning. | session 2026-07-21 (measurement pass) | 2026-07-21 | | #018 | P2 | task | Split the lithium, ADHD and metabolic residuals by mechanism | Revalidated on current main 2026-07-23: these are not one composer defect. Lithium reproduced an unrelated-table retrieval fast-path defect; ADHD retrieves a relevant chart-heavy CAMHS source but exhausts the extractive route budget; metabolic retrieves the correct AKG source but selects schedule-free prose. The narrow lithium subject-evidence guard improved targeting from 0 to 1 with golden recall 1.0 and no reciprocal-rank regressions, but it was reverted because the full canary failed. After #051 stabilises the canary, add independent current-main reproducers and assess each mechanism separately. Do not widen the matcher or combine these into a broad ranking/composer change. | runs `30007833352` and `30009207429`; PR #1093; session 2026-07-23 | 2026-07-21 | @@ -56,14 +123,13 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | #023 | P3 | task | Read Sunday 2026-07-26 scheduled-run artifacts | The 18:00 UTC scheduled runs deliver three free datapoints at once: first full-44 weekly canary (validates the #1044 ANSWER_CASE_LIMIT raise), browser-matrix flake second datapoint (webkit ui-route-coverage now reproduced + root-caused 2026-07-22 → see #024; firefox ui-formulation:91 still awaits a datapoint), and the irrelevant@10 labeling-audit artifact (§3.1 human-decision class). Read all three, then disposition. | sessions 2026-07-20/21; branch-review-ledger convergence notes | 2026-07-21 | | #024 | P3 | issue | WebKit e2e `_rsc`-prefetch access-control-checks errors | verify:release:offline on `main` ce32fe170 (2026-07-22) reproduced #023's webkit clause: **6/6 deterministic** failures in `tests/ui-route-coverage.spec.ts` (Therapy Compass; DSM home/comparison; Specifier comparison/map; Differential stream), each a `pageerror … ?_rsc=… due to access control checks` on Next.js RSC prefetch — Chromium + Firefox clean. Not merge-blocking (required gate `test:e2e:pr` is chromium-only; the full webkit matrix is advisory/release-time). Most likely a Playwright route-interception × WebKit interaction, not a Safari user defect. Next: decide (a) allow/mock the `_rsc` routes for the `webkit` e2e project, or (b) confirm real Safari impact — before trusting the full-matrix webkit gate at release. NB the 2 other webkit fails (`ui-stress:412`, `ui-universal-search:210`) passed on isolated re-run = true flake. | session 2026-07-22 (verify:release:offline, `main` ce32fe170); refines #023 | 2026-07-22 | | #025 | P2 | task | Activate the three webhooks (operator secrets) | Merged (#968) + deployed but inert — verified live: `POST /api/webhooks/railway` returns `503 webhook_not_configured`. To turn on: (1) Railway → set `RAILWAY_WEBHOOK_SECRET` + add the `?token=…` webhook URL; (2) the chat URLs `SLACK_WEBHOOK_URL`/`DISCORD_WEBHOOK_URL` must be set in BOTH places — the Railway **app/server env** (the receiver forwards deploy alerts via `postChatNotification`, which reads server env, so repo-secret-only leaves the Railway webhook authenticated but returning `delivered:false`) AND as **GitHub repo secrets** (the CI-failure workflow reads `secrets.*`); (3) `SUPABASE_INGESTION_WEBHOOK_SECRET`. Each fails closed until set, so this is pure ops. See docs/webhooks.md. | session 2026-07-22; PR #968; docs/webhooks.md | 2026-07-22 | -| #026 | P2 | task | Wire the Supabase document-change trigger | Implementation is complete on `codex/supabase-document-change-trigger`: forward migration `20260723150000_document_change_ingestion_webhook.sql`, schema mirror, update-only/minimal/fail-safe contract test and rollback docs. On 2026-07-24, `drift:manifest` replayed the full schema successfully in disposable Supabase Postgres 17.6.1.127, regenerated the manifest and removed the container; focused schema/drift tests passed 79/79 plus migration-role, function-grant and owner-scope guards. Next: protected-main PR and hosted migration-chain replay, then apply the committed migration through the normal Supabase path before configuring the Vault secret and base-URL GUC. The trigger remains inert until all three live steps are complete. | local branch/worktree; `docs/webhooks.md` section 3; session 2026-07-24 | 2026-07-22 | +| #026 | P2 | task | Activate the merged Supabase document-change trigger safely | PR #1100 merged the forward migration, schema mirror, contract test, drift evidence and rollback docs to `main`. Before activation, add or prove idempotent recovery for sufficiently old `queued` documents with no open ingestion job. Then, only with explicit approval, verify hosted migration state, apply the committed migration through the normal Supabase path, configure the Vault secret and base-URL GUC, and verify one controlled event. The trigger remains inert until those steps complete. | PR #1100; `docs/webhooks.md` section 3; session 2026-07-24 | 2026-07-22 | | #027 | P3 | rec | External uptime monitor independent of GitHub/Railway | `live-domain-monitor.yml` runs on GitHub's cron, so it won't run in exactly the outage it should catch (Actions or the deploy itself down). Add an off-platform synthetic monitor (UptimeRobot / Better Stack / Checkly) hitting `/api/health` with a webhook alert. Provider setup, not code. | session 2026-07-22 webhook review | 2026-07-22 | | #028 | P3 | rec | Runtime error tracking (Sentry or similar) | No error tracking in the repo — production exceptions on `psychiatry.tools`, including how often `RAG_PROVIDER_MODE=auto` silently degrades to source-only, are invisible. Weigh adding `@sentry/nextjs` (dependency + DSN secret + instrumentation) vs cost; alert → chat/issue. Provider-backed; needs explicit sign-off before adding the dependency. | session 2026-07-22 webhook review | 2026-07-22 | | #029 | P2 | issue | 12 of 30 answer-quality cases return the fallback stub | run #61 --dump-answers: 12/30 quality cases emit the source_backed_review_fallback boilerplate with answer_sections: [], all grounded with 4-6 citations. Some still PASS targeting because the stub echoes query keywords (the contraindication/document_lookup matchers need only a keyword), so the targeting metric MASKS the problem for those intents. Superset of #018 — fix in the extractive composer, validate with the provider-backed answer eval. | run #61 dump artifact; session 2026-07-22 | 2026-07-22 | | #030 | P3 | issue | Wide-tier alias lets one doc satisfy both comparison slots | In src/lib/eval-document-matching.ts, "Admission to Discharge for Mental Health Inpatients" appears in BOTH the AdmissionCommunityPts and Discharge alias lists, so a single document can satisfy both expectedFiles slots and make allHit true — a latent false-pass on admission-discharge cases. Not firing today (that doc is not in the failing top-5) but it would mask a real miss. Tighten the tables so one doc cannot fill both sides. | src/lib/eval-document-matching.ts:32-65; session 2026-07-22 | 2026-07-22 | | #032 | P3 | rec | Governance ranking weighting: REFUTED, not debt | The source-governance audit (PR #1051) flagged three "gaps": `review_due` carries no ranking penalty, `unknownCurrentnessPenalty` ships at 0, and `selectBestSourceRecommendation` ignores governance metadata. **These are deliberate, measured decisions — do NOT implement them as written.** Blanket metadata boosts/penalties in selection ordering were measured on 2026-07-02 to regress the golden retrieval eval to 16/23 (doc-recall@5 1.0→0.76, mrr 0.75→0.64). Two corpus facts make it unsafe: scores saturate at the clamp so stacked boosts fully override lexical relevance, and the corpus is only partially metadata-enriched while `normalizeSourceMetadata` coerces unenriched docs to `unknown`/`unverified` — so "unknown" ≠ "bad" and blanket weighting swings ranking approx. 0.35 for reasons unrelated to relevance. Even governance-as-tiebreak buried correct unenriched docs (3 designs bisected). Next action: none — treat as a guardrail. If ever revisited, RC8 (source-strength as a _filter_) is the tracked path, gated on `eval:retrieval:quality` 36/36 plus a live canary pair. | PR #118; `docs/rag-behaviour/refuted-approaches.md`; PR #1051 items 4/5/6 | 2026-07-22 | | #033 | P3 | rec | Source governance metadata absent from the LLM prompt | `buildRagSourceBlock` omits `document_status`, `clinical_validation_status`, and `extraction_quality`, so the model cannot self-caveat during generation and governance is enforced only post-hoc. Generation-surface change: needs `eval:rag` plus `eval:quality --rag-only` (grounded-supported must not drop, citation-failure 0) and explicit approval. Carries the same "unknown ≠ bad" hazard as #032 — on a partially-enriched corpus the model would likely over-caveat correct sources, so design the prompt wording before spending an eval. | `src/lib/rag/rag-source-block.ts:126-198`; PR #1051 audit item 8 | 2026-07-22 | -| #034 | P3 | issue | Answer cache can serve stale governance metadata | `cacheIndexingVersion` derives the version from `updated_at` / `indexed_at` / `index_generation_id`, so a metadata-only `document_status` flip that bumps none of those is invisible to the passive guard. **Already mitigated**: every known status-write path calls `invalidateRagCachesForOwner` or `invalidateRagCachesForDocumentMutation`. Residual risk only — a future write path that omits the invalidator would serve stale governance until TTL. Next action: add a regression test pinning the invalidator call on status-mutating routes (cheaper and safer than touching the protected cache key). | `src/lib/rag/rag-cache.ts:382-438`; PR #1051 audit item 10 | 2026-07-22 | | #035 | P3 | rec | Threshold-conflict detection covers only 3 params | `detectThresholdDisagreements` checks only ANC, WBC, and platelets paired with withholding verbs, so cross-source conflicts on medication doses, lithium/thyroid levels, or vital signs go undetected. Deliberately narrow (see the comment at `:469-474`). Broadening changes when an answer is classified `conflicting` and adds warnings — real false-positive risk. Needs new fixtures plus a behaviour review before any change. | `src/lib/evidence.ts:469-574`; PR #1051 audit item 7 | 2026-07-22 | | #036 | P3 | rec | No explicit `is_public` visibility flag on documents | Public-corpus visibility is implicit: `owner_id IS NULL` on an `indexed` document (`resolveSearchScope`). The `metadata.public_corpus` marker is written by the promotion migrations but never used as a retrieval filter. Promotion is unconditional on `clinical_validation_status`, so unverified documents are publicly searchable — compensated by keeping `unverified_source` in the frontend-visible warning set. A hard schema flag touches RLS and the clinical-risk-gated retrieval RPCs; weigh against the existing compensating control before acting. | `supabase/schema.sql:61-108`; `src/lib/search-scope.ts:181-236`; PR #1051 audit item 3 | 2026-07-22 | | #037 | P3 | rec | D5 trust-cap-all-claims flag parked OFF | `NEXT_PUBLIC_RAG_TRUST_CAP_ALL_CLAIMS` extends authority gating from high-risk claims to **all** supported claims (`deriveTrust`). Ships OFF by design; flipping it caps trust to `medium` for routine claims across the board — a product/clinical-UX decision, not a defect. Both states are test-pinned. Next action: product decision, then flip and re-baseline the UI expectations. | `src/lib/answer-render-policy.ts:159-177`; PR #1051 audit item 11 | 2026-07-22 | @@ -78,6 +144,8 @@ Move resolved rows here with the resolution date and a one-line outcome. Keep th | ID | Type | Summary | Outcome | Resolved | | ---- | ----- | ---------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | +| #034 | issue | Answer cache can serve stale governance metadata | Current-source verification found direct route coverage already asserts RAG-cache invalidation on document PATCH, source review, label, bulk and reindex mutation paths. The residual test recommendation is therefore already met; changing the protected cache key is unnecessary. | 2026-07-24 | +| #014 | rec | Realize the `next/image` win on signed previews | Superseded: `SignedImage` uses `next/image` for layout/sizing but deliberately sets `unoptimized`, preventing bearer signed URLs from entering the unauthenticated optimizer cache where cached content could outlive the token. No optimization task remains unless private-image delivery changes. | 2026-07-24 | | #031 | issue | Populate canary Source Governance table | The answer-quality step now consumes the preceding `golden-retrieval.json` only for source-governance reporting. Offline replay of run `30018289898` populated 338 top results, including 202 review-required entries, while retaining zero retrieval cases and no additional threshold failures. Retrieval and ranking behavior are unchanged. | 2026-07-24 | | #020 | task | Validate eval:quality cost readout post-fix | Confirmed on merged-main canary run `30018289898`: Answer Metrics reported 9 nonzero-cost cases and an estimated answer cost of `$0.234736`; the structured report retained the same value. The PR #1050 estimator fix is operationally proven. | 2026-07-23 | | #003 | task | Staging tenancy release evidence outstanding | Ran GitHub Action and validated isolation | 2026-07-21 |