diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index 0c7ba502..ac089035 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -49,7 +49,7 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | #019 | P2 | task | Admission doc dropped by answer-stage re-rank, not retrieval | DIAGNOSED 2026-07-22 from run #61 artifacts: retrieval coverage is PERFECT — golden documentRecallAt5=1, reciprocalRankAt10=1, missingDocumentSubstrings=[], with "Admission of Community Patients (AKG).pdf" at raw rank 4-5 (matches the strict alias AdmissionCommunityPts). The answer stage then DROPS it from answer.sources top-5 and promotes "Falls Prevention and Management (AKG).pdf" (raw rank 8, off-topic) in its place. So the defect is post-retrieval evidence balancing (cross-document-synthesis.ts balanceCrossDocumentResults / rankAnswerEvidence), NOT comparison-class retrieval coverage. Separately: nothing guarantees BOTH sides of a comparison are represented — comparison_multi_document_gate (rag.ts:2043) is satisfied by any 2 distinct docs. Fix = protected surface (canary pair + ~$2-4 approval). | run #61 artifacts; session 2026-07-22 | 2026-07-21 | | #020 | P3 | task | Validate eval:quality cost readout post-fix | The estimator fix landed in PR #1050 (zero-usage cases count $0 with rates configured; provider-attempted cases with missing usage stay n/a). Remaining action: confirm the next canary run prints a real dollar figure in Answer Metrics, then archive this row. | PR #1050; run #60/#61 Answer Metrics tables | 2026-07-21 | | #021 | P3 | rec | E-3d H2 residual: strong/comparison generation discards | approx. 6 generation attempts per full 44-case run still fail the final quality gate and fall to extractive on strong-route comparison/complex shapes (the designed-conservative outcome). PARKED: weakest cost/benefit on the queue — a wave (approx. $2-4 pair + reviewer cycle) to shave seconds off a few hard cases. Revisit only if latency/waste complaints or a cheaper lever appears. | E-3c design record; runs #59-#61 diagnostics | 2026-07-21 | -| #022 | P2 | task | Source-governance metadata refresh (operator) | Governance warning rate ~0.84 across canaries: stale/review-required/unknown source metadata on most top results. Operator work (document review-status attestation in the app), not a code defect. Next action: generate the prioritized refresh worklist ($0 read-only) and schedule a metadata pass. | runs #57-#61 Source Governance tables; docs/observability-slos.md | 2026-07-21 | +| #022 | P2 | task | Source-governance metadata refresh (operator) | **Worklist generated 2026-07-22 ($0, read-only): `docs/source-governance-refresh-worklist-2026-07-22.md`.** Reframed - this is NOT 59 clinical reviews. Of the 124 documents surfacing in canary top results, 59 are review-required, and **38 (64 pct) are the BMJ published-reference tier all sitting at `clinical_validation_status: unverified`** - one attestation-policy decision, not 38 reviews. The remaining 21 are genuine local WA health-service reviews (FSH 7, NMHS 4, CAMHS 3, AKG 2, KEMH 2, RPBG 2, RKPG 1), mostly `document_status: review_due`. Burn-down: top-10 documents clear 44 pct of flagged slots, top-20 clear 66 pct. Next: decide the BMJ attestation policy, then attest local docs by visibility (start `Clozapine Management by GP (NMHS)`, 22 slots at rank 1). | runs #61/#57 Source Governance data; `docs/source-governance-refresh-worklist-2026-07-22.md` | 2026-07-21 | | #023 | P3 | task | Read Sunday 2026-07-26 scheduled-run artifacts | The 18:00 UTC scheduled runs deliver three free datapoints at once: first full-44 weekly canary (validates the #1044 ANSWER_CASE_LIMIT raise), browser-matrix flake second datapoint (webkit ui-route-coverage now reproduced + root-caused 2026-07-22 → see #024; firefox ui-formulation:91 still awaits a datapoint), and the irrelevant@10 labeling-audit artifact (§3.1 human-decision class). Read all three, then disposition. | sessions 2026-07-20/21; branch-review-ledger convergence notes | 2026-07-21 | | #024 | P3 | issue | WebKit e2e `_rsc`-prefetch access-control-checks errors | verify:release:offline on `main` ce32fe170 (2026-07-22) reproduced #023's webkit clause: **6/6 deterministic** failures in `tests/ui-route-coverage.spec.ts` (Therapy Compass; DSM home/comparison; Specifier comparison/map; Differential stream), each a `pageerror … ?_rsc=… due to access control checks` on Next.js RSC prefetch — Chromium + Firefox clean. Not merge-blocking (required gate `test:e2e:pr` is chromium-only; the full webkit matrix is advisory/release-time). Most likely a Playwright route-interception × WebKit interaction, not a Safari user defect. Next: decide (a) allow/mock the `_rsc` routes for the `webkit` e2e project, or (b) confirm real Safari impact — before trusting the full-matrix webkit gate at release. NB the 2 other webkit fails (`ui-stress:412`, `ui-universal-search:210`) passed on isolated re-run = true flake. | session 2026-07-22 (verify:release:offline, `main` ce32fe170); refines #023 | 2026-07-22 | | #025 | P2 | task | Activate the three webhooks (operator secrets) | Merged (#968) + deployed but inert — verified live: `POST /api/webhooks/railway` returns `503 webhook_not_configured`. To turn on: (1) Railway → set `RAILWAY_WEBHOOK_SECRET` + add the `?token=…` webhook URL; (2) the chat URLs `SLACK_WEBHOOK_URL`/`DISCORD_WEBHOOK_URL` must be set in BOTH places — the Railway **app/server env** (the receiver forwards deploy alerts via `postChatNotification`, which reads server env, so repo-secret-only leaves the Railway webhook authenticated but returning `delivered:false`) AND as **GitHub repo secrets** (the CI-failure workflow reads `secrets.*`); (3) `SUPABASE_INGESTION_WEBHOOK_SECRET`. Each fails closed until set, so this is pure ops. See docs/webhooks.md. | session 2026-07-22; PR #968; docs/webhooks.md | 2026-07-22 | @@ -58,6 +58,7 @@ Durable, cross-session memory of everything still outstanding for this repo: ope | #028 | P3 | rec | Runtime error tracking (Sentry or similar) | No error tracking in the repo — production exceptions on `psychiatry.tools`, including how often `RAG_PROVIDER_MODE=auto` silently degrades to source-only, are invisible. Weigh adding `@sentry/nextjs` (dependency + DSN secret + instrumentation) vs cost; alert → chat/issue. Provider-backed; needs explicit sign-off before adding the dependency. | session 2026-07-22 webhook review | 2026-07-22 | | #029 | P2 | issue | 12 of 30 answer-quality cases return the fallback stub | run #61 --dump-answers: 12/30 quality cases emit the source_backed_review_fallback boilerplate with answer_sections: [], all grounded with 4-6 citations. Some still PASS targeting because the stub echoes query keywords (the contraindication/document_lookup matchers need only a keyword), so the targeting metric MASKS the problem for those intents. Superset of #018 — fix in the extractive composer, validate with the provider-backed answer eval. | run #61 dump artifact; session 2026-07-22 | 2026-07-22 | | #030 | P3 | issue | Wide-tier alias lets one doc satisfy both comparison slots | In src/lib/eval-document-matching.ts, "Admission to Discharge for Mental Health Inpatients" appears in BOTH the AdmissionCommunityPts and Discharge alias lists, so a single document can satisfy both expectedFiles slots and make allHit true — a latent false-pass on admission-discharge cases. Not firing today (that doc is not in the failing top-5) but it would mask a real miss. Tighten the tables so one doc cannot fill both sides. | src/lib/eval-document-matching.ts:32-65; session 2026-07-22 | 2026-07-22 | +| #031 | P3 | issue | Canary Source Governance table reports all zeros | Run #61 `answer-quality.log` prints `## Source Governance` with `Top results \| 0` and every rate 0, even though the run's `golden-retrieval.json` `topResults` carry full per-result governance metadata (`document_status`, `clinical_validation_status`, `extraction_quality`). The operator-facing table is therefore not populated and the #022 worklist had to be derived from raw JSON. Fix the table's data wiring so governance is visible from the log itself. | run #61 eval-canary artifacts; session 2026-07-22 | 2026-07-22 | | #032 | P3 | rec | Governance ranking weighting: REFUTED, not debt | The source-governance audit (PR #1051) flagged three "gaps": `review_due` carries no ranking penalty, `unknownCurrentnessPenalty` ships at 0, and `selectBestSourceRecommendation` ignores governance metadata. **These are deliberate, measured decisions — do NOT implement them as written.** Blanket metadata boosts/penalties in selection ordering were measured on 2026-07-02 to regress the golden retrieval eval to 16/23 (doc-recall@5 1.0→0.76, mrr 0.75→0.64). Two corpus facts make it unsafe: scores saturate at the clamp so stacked boosts fully override lexical relevance, and the corpus is only partially metadata-enriched while `normalizeSourceMetadata` coerces unenriched docs to `unknown`/`unverified` — so "unknown" ≠ "bad" and blanket weighting swings ranking approx. 0.35 for reasons unrelated to relevance. Even governance-as-tiebreak buried correct unenriched docs (3 designs bisected). Next action: none — treat as a guardrail. If ever revisited, RC8 (source-strength as a _filter_) is the tracked path, gated on `eval:retrieval:quality` 36/36 plus a live canary pair. | PR #118; `docs/rag-behaviour/refuted-approaches.md`; PR #1051 items 4/5/6 | 2026-07-22 | | #033 | P3 | rec | Source governance metadata absent from the LLM prompt | `buildRagSourceBlock` omits `document_status`, `clinical_validation_status`, and `extraction_quality`, so the model cannot self-caveat during generation and governance is enforced only post-hoc. Generation-surface change: needs `eval:rag` plus `eval:quality --rag-only` (grounded-supported must not drop, citation-failure 0) and explicit approval. Carries the same "unknown ≠ bad" hazard as #032 — on a partially-enriched corpus the model would likely over-caveat correct sources, so design the prompt wording before spending an eval. | `src/lib/rag/rag-source-block.ts:126-198`; PR #1051 audit item 8 | 2026-07-22 | | #034 | P3 | issue | Answer cache can serve stale governance metadata | `cacheIndexingVersion` derives the version from `updated_at` / `indexed_at` / `index_generation_id`, so a metadata-only `document_status` flip that bumps none of those is invisible to the passive guard. **Already mitigated**: every known status-write path calls `invalidateRagCachesForOwner` or `invalidateRagCachesForDocumentMutation`. Residual risk only — a future write path that omits the invalidator would serve stale governance until TTL. Next action: add a regression test pinning the invalidator call on status-mutating routes (cheaper and safer than touching the protected cache key). | `src/lib/rag/rag-cache.ts:382-438`; PR #1051 audit item 10 | 2026-07-22 | diff --git a/docs/source-governance-refresh-worklist-2026-07-22.md b/docs/source-governance-refresh-worklist-2026-07-22.md new file mode 100644 index 00000000..811c8032 --- /dev/null +++ b/docs/source-governance-refresh-worklist-2026-07-22.md @@ -0,0 +1,101 @@ +# Source-governance refresh worklist — 2026-07-22 + +Successor to [`source-review-priority-2026-07-02.md`](source-review-priority-2026-07-02.md), regenerated +from live canary artifacts. Ledger item: **#022**. Produced read-only at **$0** — no provider calls, no +live queries; everything below is derived from the Eval Canary artifacts for runs **#61** and **#57**. + +## What the two governance numbers actually mean + +They are different denominators and are often conflated: + +| Number | Source | Denominator | Meaning | +| -------------------- | -------------------------------------------------------------------------------- | ----------------------------------------------------------------- | ----------------------------------------------------------- | +| **0.8409** | run #61 `answer-quality.log` → Answer Metrics → "Source governance warning rate" | **44 answer-quality cases** | ~37 of 44 cases raised at least one governance warning | +| **0.5976** (404/676) | runs #61 + #57 `golden-retrieval.json` → `topResults` | **676 individual top-result slots** (36 retrieval cases × 2 runs) | 60% of surfaced result slots carry review-required metadata | + +Policy (verbatim from the canary log): _"unknown, unverified, review_due, outdated, unknown extraction, +and poor extraction metadata are treated as review-required; do not silently default them to current or +approved."_ + +> **Reporting gap worth noting:** run #61's own `## Source Governance` table reports `Top results | 0` and +> all-zero rates, even though the underlying `topResults` records carry full governance metadata. The +> operator-facing table in the log is therefore **not** populated — this worklist had to be derived from +> the raw JSON. Worth fixing so the canary log surfaces this directly. + +## The reframing: this is not 59 document reviews + +**59 of the 124 distinct documents** appearing in top results are review-required. But they fall into two +very different classes: + +| Class | Docs | Flag | What it actually needs | +| ------------------------------------- | -------------------------------------------------------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| **BMJ published reference tier** | **38 (64%)** | `clinical_validation_status: unverified` | **One policy decision**, not 38 clinical reviews — decide how third-party BMJ Best Practice content is attested (bulk attestation via an explicit, auditable metadata update / a dedicated "third-party published" validation state). Do **not** silently default ingestion metadata to `current` or `approved`; preserve the third-party/unverified distinction until the approved policy is applied. | +| **Local WA health-service documents** | **21** (FSH 7, NMHS 4, CAMHS 3, AKG 2, KEMH 2, RPBG 2, RKPG 1) | mostly `document_status: review_due` | **Genuine periodic review** — real attestation work, but a tractable ~21 documents | + +By flag combination: `val:unverified` 30 · `doc:review_due` 19 · both 8 · `doc:unknown` 1 · +`val:unverified + doc:unknown` 1. + +## Burn-down math + +Ranked by number of review-required top-result slots (both runs combined; 404 flagged slots total): + +| Top N documents | Slots cleared | Share of all flagged slots | +| --------------- | ------------- | -------------------------- | +| 5 | 110 / 404 | **27%** | +| 10 | 176 / 404 | **44%** | +| 15 | 226 / 404 | **56%** | +| 20 | 266 / 404 | **66%** | +| 30 | 332 / 404 | **82%** | + +Because the BMJ tier dominates, resolving the BMJ attestation policy alone would clear the large majority +of these slots in one action. + +## Prioritized worklist (top 25 by surfaced slots) + +`Slots` = review-required top-result appearances across runs #61 + #57. `Best` = highest rank achieved +(rank 1 = most user-visible). `Cum.` = cumulative share of all 404 flagged slots. + +| # | Document | Slots | Best | Publisher | Needs | Cum. | +| --: | ---------------------------------------------------------------------- | ----: | ---: | --------- | ----------------------------------------------------- | ---: | +| 1 | Bipolar disorder in adults.pdf | 32 | 1 | BMJ | validation (unverified) | 8% | +| 2 | Clozapine Management by GP (NMHS).pdf | 22 | 1 | NMHS | document status (review_due) | 13% | +| 3 | Alcohol withdrawal.pdf | 22 | 1 | BMJ | document status (review_due); validation (unverified) | 19% | +| 4 | Opioid use disorder.pdf | 18 | 1 | BMJ | validation (unverified) | 23% | +| 5 | Postnatal depression.pdf | 16 | 1 | BMJ | validation (unverified) | 27% | +| 6 | Anorexia nervosa.pdf | 16 | 1 | BMJ | validation (unverified) | 31% | +| 7 | Generalised anxiety disorder.pdf | 14 | 2 | BMJ | validation (unverified) | 35% | +| 8 | Alcohol and Other Drugs - Addiction, Toxicity and Withdrawal (FSH).pdf | 12 | 1 | FSH | document status (review_due) | 38% | +| 9 | Attention deficit hyperactivity disorder in adults.pdf | 12 | 1 | BMJ | validation (unverified) | 41% | +| 10 | Alcohol use disorder.pdf | 12 | 5 | BMJ | validation (unverified) | 44% | +| 11 | Insomnia.pdf | 10 | 1 | BMJ | validation (unverified) | 46% | +| 12 | Schizophrenia.pdf | 10 | 1 | BMJ | validation (unverified) | 49% | +| 13 | Panic disorders.pdf | 10 | 1 | BMJ | validation (unverified) | 51% | +| 14 | Depression in adults.pdf | 10 | 2 | BMJ | validation (unverified) | 53% | +| 15 | Attention deficit hyperactivity disorder in children.pdf | 10 | 4 | BMJ | validation (unverified) | 56% | +| 16 | Clozapine Coordinator and Clozapine Clinic (NMHS).pdf | 8 | 1 | NMHS | document status (review_due) | 58% | +| 17 | Suicide risk mitigation.pdf | 8 | 1 | BMJ | validation (unverified) | 60% | +| 18 | Post-traumatic stress disorder.pdf | 8 | 1 | BMJ | validation (unverified) | 62% | +| 19 | Obsessive-compulsive disorder.pdf | 8 | 1 | BMJ | validation (unverified) | 64% | +| 20 | Tourette's syndrome.pdf | 8 | 1 | BMJ | document status (review_due); validation (unverified) | 66% | +| 21 | Functional neurological and somatic symptom disorders.pdf | 8 | 3 | BMJ | validation (unverified) | 68% | +| 22 | MHATT Assessment and Treatment Process (AKG).pdf | 8 | 4 | AKG | document status (review_due) | 70% | +| 23 | Social anxiety disorder.pdf | 8 | 5 | BMJ | validation (unverified) | 72% | +| 24 | Personality disorders.pdf | 8 | 6 | BMJ | validation (unverified) | 74% | +| 25 | Depression in children.pdf | 6 | 1 | BMJ | validation (unverified) | 75% | + +## Suggested order of work + +1. **Decide the BMJ attestation policy** (clears ~64% of review-required documents in one action) — any chosen option must be an explicit, auditable metadata update that never silently defaults to `current`/`approved` and preserves the third-party/unverified distinction; this remains open debt (ledger #022) until the approved policy is implemented, not merely decided. +2. **Attest the local documents by visibility** — start with `Clozapine Management by GP (NMHS)` (22 slots, + rank 1), then the FSH addiction/withdrawal document, then the remaining NMHS/AKG/CAMHS/KEMH items. +3. **Re-read the warning rate** on the next canary to confirm the burn-down. + +## Scope and limits + +- This is **operator work** (document review-status attestation in the app), not a code defect. No code + change is proposed here. +- The list reflects **only documents that surfaced in golden-case top results**, so it is a + visibility-weighted worklist, not a full corpus audit. The 2026-07-02 predecessor notes the live corpus + is far larger (~2,065 indexed documents at that time). +- Counts combine two runs (#61, #57); a document appearing in both runs counts twice, which is intentional + — it weights persistently-surfaced documents higher.