From d834dec0906099549a74f2285b35f41a1abdb2d0 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 04:41:13 +0000 Subject: [PATCH 1/9] =?UTF-8?q?docs(ledger):=20E-4=20pair=20verdict=20?= =?UTF-8?q?=E2=80=94=20E-3=20wave=20adopted=20(run=20#60=20vs=20#57/#58/#5?= =?UTF-8?q?9)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- docs/branch-review-ledger.md | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index 2925fa99..b0e28108 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -672,3 +672,4 @@ Use this ledger to prevent repeated branch and PR reviews when the reviewed HEAD | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (PR: E-3c PR-B short-circuit) | f7e6cbb + hardening commit | E-3c PR-B: pre-generation validated-extractive short-circuit for gate-passed routine procedural "What...process/include/required" queries (marker validated_routine_extractive_first), generalizing the LAI + blocked-recovery precedents via new rag-extractive-first.ts (3 predicates moved byte-verbatim, machine-verified; rag.ts 5029→4908 vs 5030 budget). Kills the run-#57 6x wasted-generation class. REVIEWS (both pre-push): rag-retrieval-reviewer APPROVE-WITH-NITS — move fidelity brace-diff verified byte-identical; confidence-gate skip PROVEN safe (passed-gate is a no-op in applyConfidenceGate; markers pairwise mutually exclusive); comparison false-positives blocked by unchanged classifier precedence; offline/source-only idempotent; zero retrieval/ranking/selection/threshold change; P2 = eval-only assertions (intent_coverage/artifact_leaks/expected-file) unverifiable offline for flip candidates (quality-nocc-document-support, quality-form-required-documentation, quality-discharge-documentation, quality-duress-pathway + rag-set siblings) → pre-merge BRANCH canary recommended and ADOPTED (offline-green + review-approved proven insufficient for this surface, 2026-07-20). clinical-governance-reviewer APPROVE-WITH-NITS — full gate-stack trace: nothing bypassed (same finalizeRagAnswerQuality, same citation scoping, numeric verification not fail-open, ungrounded-finalize defense at rag.ts:3694); P2 = pre-existing bare-cross-reference-with-overlap gap, NOT materially widened (new trigger anti-correlates), hardening recommended → APPLIED this PR: !isBareCrossReferenceAnswer screen in hasValidatedExtractiveCandidate (closes all three short-circuit paths; discriminating test added; disclosed post-review delta, strictly narrows shipping). MERGE GATE: draft until the branch canary pair (baseline #57/#58 vs branch run with answer_case_limit=44 + answer_quality_eval=true, est $3-6 of authorized envelope) is green — zero per-case regressions, recalls 1.0, quality/targeting rates >= baseline. | Red-proof + 3 negative guards; focused 61/61 + fallback/offline/contract suites; full suite 3061 passed / 1 known container artifact (pre-hardening tree; hardening re-verified focused); typecheck+lint+prettier clean; maintainability budget passed (4908/5030) | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (pair verdict — no code change; main stays at 22b6a2e) | canary run 29794759627 (#59, branch head 7310cb3 = merged PR-B content) | E-3b + E-3c PR-B LIVE-VALIDATED (pair vs banked #57/#58): route_ceiling_failures 2→0 (E-3b proven: agitation timeout now fits inside budget; clozapine retrieval-exhausted ceiling honestly excused via the triple-condition cross-region carve-out — exactly one "retrieval-exhausted" audit cell in the report); p95 17.4s→15.47s (-11%); golden retrieval SUCCESS 36/36 (stop-ship criterion held); validated_routine_extractive_first fired on 4 of the 6 target cases (patient-safety-plan 3.3s, treatment-team-process 2.6s, ect-procedure 2.2s, illegal-substances 2.2s — all pure extractive, zero generation, was 6-9s each with a discarded attempt), the other 2 (community-home-visits, best-practice-prescribing) stayed on generation+fallback because their extractive candidates legitimately fail validation gates = the designed-conservative outcome; ZERO new failing cases (list 5→2, both known residuals: neuroleptic citation red = Option A territory, admission-comparison expected-doc = non-blocking labeling residual); grounded 1.0 + unsupported_correct 1.0 held; targeting 0.5909→0.619, fail_closed 0.9→0.9333, readability/artifact_leaks/intent_coverage unchanged; relevance 0.6→0.5667 = single-case wobble on n=30, WATCH in E-4, not a gate. Discarded-generation rate materially down (4 conversions; residual = the H2 strong-route slice named as E-3d candidate, per design's 20-33% expectation band). MERGE STANDS (user had armed auto-merge pre-verdict; revert drill not triggered). Spend +~$3-6 → Phase E total ~$6-12 of ≤$20. | Pair evidence: run #59 job log (Blocking failures = citation only; Answer Case Diagnostics markers; metric_rates + targeting blocks); dump artifact populated for PR-C diagnosis | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (PR: E-3c PR-C figure-aware selection) | 043b030 + P2-fix commit | E-3c PR-C: dose/monitoring extractive answers now carry the asked-for figure/schedule when the cited chunk verbatim supports it — lead-slot promotion (dose swaps last of 2 slots, monitoring appends 2nd sentence; no-op when a lead already carries a figure) guarded by the claim-support atom corpus (sourceEvidenceText exported, promotionAtomKey byte-identical to claim-support's atomKey), plus the dose/threshold generation-fallback preferring the safe figure-carrying candidate (safety gate unchanged, filter-order-stable). Fallback helpers extracted to rag-extractive-answer (cycle-check verified); rag.ts 4908→4901. Six discriminating tests each verified red-on-prior-code incl. proving the nuke-guard load-bearing by disabling it. REVIEWS (both pre-push on 043b030): rag-retrieval-reviewer APPROVE-WITH-NITS — no-op path byte-identical verified, atom-key identity verified, filter-vs-find proven side-effect-free, 2-sentence append gate-safe, intent double-gated, zero retrieval/ordering change, imputation contract green; P2 = zero-atom monitoring figures ("every 6 weeks" yields no value atom) pass the guard trivially and can be nuked by claim support if sourced only from adjacent context (fails SAFE — evidence gap, never a wrong figure). clinical-governance-reviewer APPROVE-WITH-NITS — all six clinical concerns CLEARED end-to-end (verbatim-support guarantee, citation binding preserved, conservative failure test-proven, unsafe candidates impossible, wrong-drug risk controlled by pre-existing entity/multi-drug guards, no PHI); same zero-atom finding as P3 + one comment-precision nit. P2 FIXED post-review (disclosed): zero-atom figures now require the matched figure substring verbatim in sourceEvidenceText (intentFigureMatchText); proven both directions by 3 new tests (promotes from content, refuses from adjacent-context-only); comment-precision nit folded in. | Focused post-fix: extractive-formatting 32/32 + fallback/eval-cases/offline/contract/extractive-first 98/98 incl. imputation contract; typecheck+prettier+budgets clean; full-suite 3068-passed baseline pre-P2-fix (fix re-verified focused). Live proof = E-4 pair next | +| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (E-4 pair verdict — no code change; #1039 merged as 9b655fa) | canary run 29800029819 (#60, main 9b655fa = E-3b+PR-B+PR-C) | PHASE E-3 WAVE CLOSED — E-4 VERDICT: ADOPT (no revert). 44-case: the ONLY blocking red is the KNOWN persisting neuroleptic citation case (0.0227, identical #57 signature: generation quality-failed → extractive fallback 1 citation — the pre-declared Option A carve-out, retrieval-side, untouched by answer waves); route_ceiling_failures 0 CONFIRMED ON MAIN (E-3b: agitation 23.1s < 25s after 20.3s provider timeout; clozapine 14.5s with exactly one retrieval-exhausted audit cell); grounded 1.0 / unsupported_correct 1.0 / numeric 0 / governance-danger 0 all held; expected_source_hit 0.6136→0.6364 (#1020 widen); generation attempts 19→9 across 44 cases (10+ cases short-circuit via validated_routine_extractive_first at 2-6s, zero generation spend — the absolute wasted-generation seconds collapse; residual 6 discarded attempts are the named E-3d H2 strong/comparison slice). Targeting vs #58 baseline: rate 0.5909→0.6667, dose 1/5→2/4 (sertraline + quetiapine still miss), document_lookup 5/5→6/6, contraindication 2/2, red_result 2/2, pathway 1/2; readability/artifact_leaks 1.0. NOT met: monitoring_schedule flat 1/5 — per-miss lens shows answers of 73-232 chars with NO schedule token available to promote (olanzapine-lai 79ch, metabolic 73ch = single-fact extractive answers; the PR-C promotion is a no-op when no figure-bearing fact is extracted) → root is fact-extraction/retrieval depth on monitoring shapes, queued as the Option-A-wave companion diagnosis (dump artifact 8483731630, 30d retention). WATCH escalated: relevance 0.6 (#58) → 0.5667 (#59) → 0.5333 (#60) — two single-case steps coinciding with more terse extractive answers; fail_closed 0.9 = exactly the #58 main baseline (#59's 0.9333 was the outlier), safety texture flat. Adoption per plan criteria: targeting ≥ baseline ✓, grounded/refusal 1.0 ✓, golden 36/36 ✓ (in-run), ceilings 0 ✓, citation red = carved known case ✓. CodeRabbit post-review follow-up landed pre-merge (02b5c78): interval-regex full-match reorder (atom path proven to intercept the claimed exploit; reorder = drift hardening), clinicalValueAtomKey exported (mirror deleted), guard tests made honestly discriminating + genuine zero-atom "annually" coverage both directions. Instrument note: cost rates live on the targeting step env but eval:quality still reports cost n/a (estimator not consuming them in the 44-case path) — minor tooling residual. Spend +~$2-4 → Phase E total ~$8-16 of ≤$20. | Evidence: run #60 job log read in full (Threshold Status: citation-only; Answer Metrics + 44-row diagnostics; targeting metric_rates + 7-miss list); artifact eval-canary-output 8483731630 sha256 d5c7006e… (download blocked in-session — GitHub App scope; log tee carried the targeting output) | From 1aebf02c0e9c8347efa50b8c9255e1ee73242480 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:09:08 +0000 Subject: [PATCH 2/9] fix(rag): monitoring evidence gate parity with fact filter (run-#60 miss class) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The monitoring-schedule answer gate was narrower than the fact filter and the eval judge: inflected schedules (reviewed annually, monitored for 3 hours), bare digit durations, and the level-range/metabolic-panel vocabulary passed fact filtering but died at the gate or at kind classification, so answers for those cases shipped without the asked-for schedule (run-#60 targeting misses). One shared monitoringScheduleEvidencePattern now feeds both the gate and the filter (lockstep by construction); the kind arm gains inflection tolerance only (panel/range vocab would steal sentences from the dose arm — pinned by the restored adjacent-context refusal test); and both the result-level and sentence-level coverage checks gain a figure escape mirroring the existing concrete-dose escape, tightly scoped to sentences/chunks that carry the asked-for interval or unit range themselves. Four red-proof tests (each verified failing pre-change) plus a negative guard pin the widen. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- src/lib/rag/rag-extractive-answer.ts | 47 +++++++++++--- tests/extractive-answer-formatting.test.ts | 73 ++++++++++++++++++++-- 2 files changed, 108 insertions(+), 12 deletions(-) diff --git a/src/lib/rag/rag-extractive-answer.ts b/src/lib/rag/rag-extractive-answer.ts index 1755b98f..7ef6971e 100644 --- a/src/lib/rag/rag-extractive-answer.ts +++ b/src/lib/rag/rag-extractive-answer.ts @@ -332,6 +332,17 @@ function queryIntentTokens(query: string, intent: AnswerIntent) { return tokens; } +// Shared monitoring-schedule evidence vocabulary. The answer-intent gate +// (answerIntentEvidencePattern) and the fact filter (factSupportsAnswerIntent) +// must accept the same schedule/parameter language: the run-#60 targeting +// misses came from range, metabolic-panel, and inflected-schedule sentences +// ("reviewed annually", "monitored for 3 hours", "Maintenance range +// 0.6-0.8 mmol/L") passing the filter but dying at the narrower gate. +// Inflections match by prefix (monitor\w*), acronyms accept plurals, and bare +// digit durations ("at 12 weeks") count as schedule evidence. +const monitoringScheduleEvidenceSource = String.raw`monitor\w*|follow[-\s]?up|baseline|weekly|monthly|annual(?:ly)?|yearly|every|several\s+times\s+a\s+year|screen(?:ing|ed)?|levels?|blood tests?|bloods|fbcs?|ancs?|wbcs?|ecgs?|lfts?|renal|thyroid|metabolic|glucose|bsl|lipids|cholesterol|triglycerides|blood pressure|bp|pulse|weight|bmi|mmol\/l|mcg\/l|ng\/ml|range|target|therapeutic|maintenance|review\w*|\d+\s*(?:week|month|day|hour|year)s?`; +const monitoringScheduleEvidencePattern = new RegExp(String.raw`\b(?:${monitoringScheduleEvidenceSource})\b`, "i"); + /** Answer intent evidence pattern. */ function answerIntentEvidencePattern(intent: AnswerIntent) { switch (intent) { @@ -340,7 +351,7 @@ function answerIntentEvidencePattern(intent: AnswerIntent) { case "contraindication": return /\b(?:contraindicat\w*|avoid|must not|do not|should not|not use|opioid[-\s]?free|withdrawal|precipitat\w*)\b/i; case "monitoring_schedule": - return /\b(?:monitor|monitoring|baseline|weekly|monthly|annual|every|level|levels|blood test|fbc|anc|ecg|lft|renal|review)\b/i; + return monitoringScheduleEvidencePattern; case "red_result_action": return /\b(?:red|amber|green|threshold|withhold|cease|stop|discontinue|discontinued|urgent|contact|repeat|review|anc|fbc|wbc|neutrophil|toxic\w*|action|patholog\w*|haematolog\w*|hematolog\w*)\b/i; case "pathway_referral": @@ -463,12 +474,21 @@ function resultCoversAnswerIntent(result: SearchResult, query: string, intent: A if (intent === "general") return true; const asksForMaximumDose = intent === "dose" && /\bmax(?:imum)?\b/i.test(query); const maximumDoseCoverage = asksForMaximumDose && hasMaximumDoseEvidence(text); - const intentCoverage = answerIntentEvidencePattern(intent).test(text) || maximumDoseCoverage; + // Result-level twin of the sentence-level figure escape: a chunk that carries + // the asked-for schedule/interval or unit range (a bare range table, "reviewed + // annually" prose) is monitoring evidence even when it never says + // monitor/level — the run-#60 miss class rejected such chunks wholesale here. + const monitoringFigureCoverage = + intent === "monitoring_schedule" && + (monitoringIntervalFigurePattern.test(text) || monitoringUnitRangeFigurePattern.test(text)); + const intentCoverage = + answerIntentEvidencePattern(intent).test(text) || maximumDoseCoverage || monitoringFigureCoverage; if (!intentCoverage) return false; if ( intentTokens.length > 0 && !intentTokens.some((token) => queryTokenMatchesText(token, text)) && - !maximumDoseCoverage + !maximumDoseCoverage && + !monitoringFigureCoverage ) { return false; } @@ -726,8 +746,14 @@ function factKindForSentence(sentence: string, query: string, intent: AnswerInte if (/\b(?:pathway|refer|referral|criteria|indicat\w*|ect|electroconvulsive|specialist|psychiat\w*)\b/i.test(text)) { return "pathway_referral"; } + // Inflection-tolerant only (monitored/annually/LFTs): the wider panel/range + // vocabulary must NOT move here — it would steal sentences like "Maintenance + // range 0.6-0.8 mmol/L" from the dose arm below and change fact priorities + // for non-monitoring intents. "review" stays exact for the same reason: + // review\w* would reclassify dose-review sentences ("doses should be + // reviewed daily") away from the dose arm. if ( - /\b(?:monitor|monitoring|baseline|weekly|monthly|annual|every|level|levels|blood test|ecg|lft|review)\b/i.test(text) + /\b(?:monitor\w*|baseline|weekly|monthly|annual(?:ly)?|every|levels?|blood tests?|ecgs?|lfts?|review)\b/i.test(text) ) { return "monitoring"; } @@ -786,9 +812,7 @@ function factSupportsAnswerIntent( // are classified as renal_limit (renal check triggers before monitoring), but are directly relevant // to monitoring schedule answers. if (kind !== "monitoring" && kind !== "dose" && kind !== "renal_limit") return false; - return /\b(?:monitor|monitoring|follow[-\s]?up|baseline|weekly|monthly|annual|every|several\s+times\s+a\s+year|level|levels|blood test|fbc|anc|wbc|ecg|lft|renal|thyroid|metabolic|glucose|bsl|lipids|cholesterol|triglycerides|blood pressure|bp|pulse|weight|bmi|mmol\/l|range)\b/i.test( - text, - ); + return monitoringScheduleEvidencePattern.test(text); case "red_result_action": if (kind !== "threshold_action" && kind !== "caveat") return false; if (requiresBloodCountEvidence(query) && !hasBloodCountEvidence(text)) return false; @@ -861,10 +885,17 @@ function factSentenceMatchesQueryFromResult( const normalized = normalizeSectionText(sentence).toLowerCase(); const intentTokens = queryIntentTokens(query, intent); + // Mirrors the dose escape below: a sentence that carries the asked-for + // schedule/interval or unit range itself ("reviewed annually", "monitored + // for 3 hours", "Maintenance range 0.6-0.8 mmol/L") is monitoring evidence + // even when it names no query token — the run-#60 miss class. Figure-bearing + // sentences only, so plain schedule-free prose still needs token coverage. const intentCovered = intentTokens.length === 0 || intentTokens.some((token) => queryTokenMatchesText(token, normalized)) || - (intent === "dose" && extractiveConcreteDosePattern.test(normalized)); + (intent === "dose" && extractiveConcreteDosePattern.test(normalized)) || + (intent === "monitoring_schedule" && + (monitoringIntervalFigurePattern.test(normalized) || monitoringUnitRangeFigurePattern.test(normalized))); return answerIntentEvidencePattern(intent).test(normalized) && intentCovered; } diff --git a/tests/extractive-answer-formatting.test.ts b/tests/extractive-answer-formatting.test.ts index df575f72..4c22eb72 100644 --- a/tests/extractive-answer-formatting.test.ts +++ b/tests/extractive-answer-formatting.test.ts @@ -462,12 +462,14 @@ describe("zero-atom figure promotion guard (reviewer P2)", () => { const answer = extractiveAnswerFor(question, [ figureChunk({ id: "monitoring-guard-1", - section_heading: "Metabolic monitoring", - // The synopsis shares the bare number but not the unit, and carries no - // intent tokens so it can never be admitted as a competing fact. The + // The synopsis shares the bare number but not the unit. The // "every 6 weeks" figure yields a 6/week quantity atom, so atom // identity must refuse the 6/month corpus — and the interval pattern's // full-match ordering keeps the verbatim fallback equally honest. + // (The synopsis's own "every 6 months" is corpus-verbatim and may be + // legitimately admitted via the figure coverage escape; the pin here is + // only that the mismatched adjacent-context figure never ships.) + section_heading: "Metabolic monitoring", retrieval_synopsis: "Reviewed every 6 months", content: leadOnly, adjacent_context: intervalSentence, @@ -475,7 +477,6 @@ describe("zero-atom figure promotion guard (reviewer P2)", () => { ]); const plain = (answer.answer ?? "").replace(/\*\*/g, ""); expect(plain).not.toMatch(/every 6 weeks/i); - expect(plain).not.toMatch(/every 6 months/i); }); // Longer than leadOnly so it stays a promotion candidate instead of taking @@ -533,3 +534,67 @@ describe("zero-atom figure promotion guard (reviewer P2)", () => { expect(plain).not.toMatch(/every 6 weeks/i); }); }); + +// Run-#60 miss class: the monitoring-schedule evidence GATE was narrower than +// the fact filter and the eval judge, so schedule/range/panel sentences passed +// fact filtering but died at the gate (or at kind classification) — "reviewed", +// "annually", "monitored", bare durations, and the level-range/metabolic-panel +// vocabulary were all rejected. These pin the widened shared vocabulary. +describe("monitoring evidence gate parity (run-#60 miss class)", () => { + it("admits a level-range fact phrased without monitor/level tokens", () => { + const answer = extractiveAnswerFor("What lithium level range is used for maintenance monitoring?", [ + figureChunk({ + id: "lithium-range-chunk-1", + title: "Lithium (AKG)", + file_name: "Lithium (AKG).pdf", + section_heading: "Lithium", + content: "Lithium is prescribed for bipolar maintenance. Maintenance range is 0.6-0.8 mmol/L.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(plain).toMatch(/0\.6-0\.8 mmol\/L/i); + }); + + it("still refuses a schedule-free monitoring sentence with no query-token coverage", () => { + const answer = extractiveAnswerFor("What metabolic monitoring is required for antipsychotics?", [ + figureChunk({ + id: "metabolic-uncovered-chunk-1", + section_heading: "Metabolic monitoring", + content: + "Metabolic monitoring is required for all patients prescribed antipsychotics. Community staff arrange the reviews with the general practitioner.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + // "reviews" passes the widened evidence gate but carries no schedule figure + // and names no query token, so the coverage escape must NOT admit it. + expect(plain).toBe("Metabolic monitoring is required for all patients prescribed antipsychotics."); + }); + + it("admits a metabolic-panel fact whose schedule uses inflected tokens", () => { + const answer = extractiveAnswerFor("What metabolic monitoring is required for antipsychotics?", [ + figureChunk({ + id: "metabolic-annual-chunk-1", + section_heading: "Metabolic monitoring", + content: + "Metabolic monitoring is required for all patients prescribed antipsychotics. Weight, BMI, fasting glucose and lipids are reviewed annually.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(plain).toMatch(/annually/i); + }); + + it("admits a duration-figure fact phrased with 'monitored for N hours'", () => { + const answer = extractiveAnswerFor("What monitoring is required after olanzapine LAI?", [ + figureChunk({ + id: "olanzapine-lai-chunk-1", + title: "Olanzapine LAI (AKG)", + file_name: "Olanzapine LAI (AKG).pdf", + section_heading: "Post-injection observation", + content: + "Olanzapine LAI requires post-injection observation. The patient is monitored for 3 hours after each injection.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(plain).toMatch(/3 hours/i); + }); +}); From 0abf3c9c4287eb5796848e8c9404050c69d67fac Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:15:33 +0000 Subject: [PATCH 3/9] feat(rag): title-supported escalation rescue via the S3 document-lookup layer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The neuroleptic-side-effect-escalation eval case failed the citation gate on every canary (#57-#60): the query misclassifies as medication_dose_risk, its honest lexical pool fails the dose fast-path floor (measured 0.184/0.0125 vs 0.66/0.055), the S3 document-lookup title layer that would surface the title-named SOP was excluded by the class allowlist, and the vector leg then injects the wrong sibling doc. shouldAttemptDocumentLookupFastPath now also engages the layer for medication_dose_risk when the classifier's existing deterministic signals say the query is escalation-shaped (intent escalation_risk — only assigned when drug_dosing wording did NOT match, so pure dose/route/frequency questions structurally cannot fire) AND a curated title alias phrase is present for the alias tier to rescue with (documentTitleTerms > 0). Rescue semantics by construction: the S3 block sits after both fast-path returns, so a medication_dose_risk query only reaches it once the floor already rejected the pool. Executed blast-radius sweep: exactly one eval case fires across the 44-case answer canary, the 36-case golden retrieval fixture, and the 30-case answer-quality harness — every non-firing case executes byte-identical code. No imputation formula, released-comparator key, selection clamp, or threshold is touched. Red-proof: the predicate test and the end-to-end rescue ordering test both fail with the gate reverted (verified); negative guards pin pure-dose, no-title, and non-escalation shapes plus the four allowlisted classes. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- src/lib/rag/rag.ts | 24 +- ...-document-lookup-escalation-rescue.test.ts | 279 ++++++++++++++++++ 2 files changed, 300 insertions(+), 3 deletions(-) create mode 100644 tests/rag-document-lookup-escalation-rescue.test.ts diff --git a/src/lib/rag/rag.ts b/src/lib/rag/rag.ts index 50b99848..fae60829 100644 --- a/src/lib/rag/rag.ts +++ b/src/lib/rag/rag.ts @@ -2122,12 +2122,30 @@ function markEmbeddingSkippedByTextFastPath(telemetry: SearchTelemetry, reason: } /** Should attempt document lookup fast path. */ -function shouldAttemptDocumentLookupFastPath(queryClass: RagQueryClass) { - return ( +export function shouldAttemptDocumentLookupFastPath( + queryClass: RagQueryClass, + analysis?: Pick, +) { + if ( queryClass === "document_lookup" || queryClass === "broad_summary" || queryClass === "table_threshold" || queryClass === "comparison" + ) { + return true; + } + // Title-supported escalation rescue: a medication_dose_risk query only + // reaches the S3 block after its lexical pool failed the dose fast-path + // floor (decideTextFastPath), so this is rescue semantics by construction. + // Both predicates are existing deterministic classifier signals — the + // escalation_risk intent is only assigned when drug_dosing wording did NOT + // match, so pure dose/route/frequency questions can never engage this layer, + // and documentTitleTerms > 0 means a curated title alias phrase is present + // for the alias tier to rescue with. + return ( + queryClass === "medication_dose_risk" && + analysis?.intent === "escalation_risk" && + (analysis?.documentTitleTerms.length ?? 0) > 0 ); } @@ -2532,7 +2550,7 @@ export async function searchChunksWithTelemetry( } } - if (shouldAttemptDocumentLookupFastPath(queryClassification.queryClass)) { + if (shouldAttemptDocumentLookupFastPath(queryClassification.queryClass, queryAnalysis)) { const documentLookupStartedAt = Date.now(); const documentLookupData = await searchDocumentLookupFastPath({ supabase, diff --git a/tests/rag-document-lookup-escalation-rescue.test.ts b/tests/rag-document-lookup-escalation-rescue.test.ts new file mode 100644 index 00000000..7d5fa457 --- /dev/null +++ b/tests/rag-document-lookup-escalation-rescue.test.ts @@ -0,0 +1,279 @@ +import { describe, expect, it } from "vitest"; +import { + applySecondStageRerankIfNeeded, + decideTextFastPath, + shouldAttemptDocumentLookupFastPath, +} from "../src/lib/rag/rag"; +import { analyzeClinicalQuery, classifyRagQuery, rankClinicalResults } from "../src/lib/clinical-search"; +import { selectRetrievalEvidence } from "../src/lib/retrieval-selection"; +import { resultsHaveReleaseRankScore, stabilizeReleasedSearchOrder } from "../src/lib/released-search-order"; +import type { SearchTelemetry } from "../src/lib/rag/rag-contracts"; +import type { SearchResult } from "../src/lib/types"; + +// Option A (2026-07-21): title-supported escalation rescue. The live case +// neuroleptic-side-effect-escalation failed the citation gate on every canary +// (#57-#60): the query misclassifies as medication_dose_risk, its honest +// lexical pool fails the dose fast-path floor (decideTextFastPath 0.66/0.055), +// and the S3 document-lookup title layer that would surface the title-named +// SOP never ran because shouldAttemptDocumentLookupFastPath allowlisted only +// document_lookup | broad_summary | table_threshold | comparison. The rescue +// engages the layer for medication_dose_risk ONLY when the classifier's +// existing deterministic signals say the query is escalation-shaped +// (intent === "escalation_risk", assigned only when drug_dosing wording did +// NOT match) AND a curated title alias phrase is present for the alias tier to +// rescue with (documentTitleTerms > 0). + +type QueryClass = ReturnType["queryClass"]; + +const escalationQuery = "When should neuroleptic side effects be escalated?"; + +// Mirrors the production pipeline exactly as tests/rag-fast-path-ordering.test.ts +// does: rankClinicalResults → selectRetrievalEvidence → applySecondStageRerankIfNeeded +// → stabilizeReleasedSearchOrder. +function runOrderingPipeline(args: { + query: string; + queryClass: QueryClass; + candidates: SearchResult[]; + topK: number; +}): SearchResult[] { + const selection = selectRetrievalEvidence({ + query: args.query, + queryClass: args.queryClass, + results: rankClinicalResults(args.query, args.candidates), + topK: args.topK, + maxResultsPerDocument: 4, + }); + const reranked = applySecondStageRerankIfNeeded({ + queryClass: args.queryClass, + results: selection.results, + telemetry: {} as SearchTelemetry, + topK: args.topK, + }); + stabilizeReleasedSearchOrder(reranked, resultsHaveReleaseRankScore(reranked)); + return reranked; +} + +// Honest S1 text-fast-path shape for the wrong sibling doc, pinned to the +// live-measured pool the dose floor rejected (strongest 0.184 / topTextRank +// 0.0125 vs the 0.66/0.055 floor). +function zuclopenthixolS1(id: string, chunkIndex: number, content: string): SearchResult { + return { + id, + document_id: "zuclopenthixol-doc", + title: "Zuclopenthixol Acuphase (AKG)", + file_name: "Zuclopenthixol Acuphase (AKG).pdf", + page_number: 2, + chunk_index: chunkIndex, + section_heading: "Administration", + content, + image_ids: [], + images: [], + similarity: 0, + text_rank: 0.0125, + hybrid_score: 0.184, + }; +} + +// Vector-leg shape for the injected sibling (the live top source). +function zuclopenthixolVector(id: string, content: string): SearchResult { + return { + id, + document_id: "zuclopenthixol-doc", + title: "Zuclopenthixol Acuphase (AKG)", + file_name: "Zuclopenthixol Acuphase (AKG).pdf", + page_number: 3, + chunk_index: 9, + section_heading: "Adverse effects", + content, + image_ids: [], + images: [], + similarity: 0.62, + text_rank: 0.02, + hybrid_score: 0.78, + similarity_origin: "cosine", + }; +} + +// S3 document-lookup title-alias chunk shape, values pinned to production +// (rag-candidate-sources.ts searchDocumentLookupFastPath: alias tier text_rank +// 0.34, similarity min(0.92, 0.58 + 0.34 + chunk bonus), hybrid 0.94, +// similarity_origin "synthetic_text"). +function neurolepticS3(id: string, chunkIndex: number, heading: string, content: string): SearchResult { + return { + id, + document_id: "neuroleptic-doc", + title: "Neuroleptic Side Effects (AKG)", + file_name: "Neuroleptic Side Effects (AKG).pdf", + page_number: 1, + chunk_index: chunkIndex, + section_heading: heading, + content, + image_ids: [], + images: [], + similarity: 0.92, + text_rank: 0.34, + hybrid_score: 0.94, + similarity_origin: "synthetic_text", + }; +} + +function neurolepticPool(): SearchResult[] { + return [ + neurolepticS3( + "neuroleptic-1", + 0, + "Escalation", + "Escalate neuroleptic side effects to the treating psychiatrist when symptoms are severe, progressive, or accompanied by fever or rigidity.", + ), + neurolepticS3( + "neuroleptic-2", + 1, + "Monitoring and escalation", + "Contact the duty medical officer urgently if neuroleptic malignant syndrome is suspected; cease the antipsychotic and escalate care.", + ), + neurolepticS3( + "neuroleptic-3", + 2, + "Side effect review", + "Document neuroleptic side effects at each review and escalate persistent extrapyramidal symptoms for specialist assessment.", + ), + ]; +} + +function zuclopenthixolLexicalPool(): SearchResult[] { + return [ + zuclopenthixolS1("zuclo-1", 0, "Zuclopenthixol acuphase doses must be prescribed by a consultant psychiatrist."), + zuclopenthixolS1("zuclo-2", 1, "Observe the patient after each zuclopenthixol acuphase dose is administered."), + ]; +} + +function zuclopenthixolMergedPool(): SearchResult[] { + return [ + ...zuclopenthixolLexicalPool(), + zuclopenthixolVector( + "zuclo-3", + "Common adverse effects of zuclopenthixol include sedation and extrapyramidal symptoms.", + ), + ]; +} + +describe("escalation-rescue gate predicate", () => { + it("fires for the title-supported escalation-shaped medication_dose_risk query", () => { + const analysis = analyzeClinicalQuery(escalationQuery); + // Pin the mechanism-chain premises so classifier drift is loud. + expect(analysis.queryClass).toBe("medication_dose_risk"); + expect(analysis.intent).toBe("escalation_risk"); + expect(analysis.documentTitleTerms.length).toBeGreaterThan(0); + expect(shouldAttemptDocumentLookupFastPath(analysis.queryClass, analysis)).toBe(true); + }); + + it("never fires for pure dose/route/frequency questions", () => { + for (const query of [ + "What is the maximum sertraline dose?", + "Show the clozapine missed-dose monitoring table guidance.", + "What agitation medication can be given IM?", + ]) { + const analysis = analyzeClinicalQuery(query); + expect(shouldAttemptDocumentLookupFastPath("medication_dose_risk", analysis)).toBe(false); + } + }); + + it("never fires for escalation wording without a curated title alias", () => { + const analysis = analyzeClinicalQuery("What are naltrexone contraindications?"); + expect(analysis.documentTitleTerms.length).toBe(0); + expect(shouldAttemptDocumentLookupFastPath("medication_dose_risk", analysis)).toBe(false); + }); + + it("never fires for title-supported non-escalation questions", () => { + const analysis = analyzeClinicalQuery( + "Which observations and blood monitoring are needed while a patient is taking clozapine?", + ); + expect(analysis.intent).not.toBe("escalation_risk"); + expect(shouldAttemptDocumentLookupFastPath("medication_dose_risk", analysis)).toBe(false); + }); + + it("keeps the four allowlisted classes engaged with and without analysis", () => { + for (const queryClass of ["document_lookup", "broad_summary", "table_threshold", "comparison"] as const) { + expect(shouldAttemptDocumentLookupFastPath(queryClass)).toBe(true); + expect(shouldAttemptDocumentLookupFastPath(queryClass, analyzeClinicalQuery("random query"))).toBe(true); + } + expect(shouldAttemptDocumentLookupFastPath("unsupported_or_general")).toBe(false); + }); +}); + +describe("escalation rescue end-to-end ordering", () => { + it("surfaces the title-named SOP above the vector-injected sibling once the gate admits S3 candidates", () => { + const analysis = analyzeClinicalQuery(escalationQuery); + // The pool is constructed THROUGH the production gate, so this test is red + // while the gate excludes medication_dose_risk: without the S3 candidates + // the sibling doc tops the release order and the rescue assertions fail. + const pool = shouldAttemptDocumentLookupFastPath(analysis.queryClass, analysis) + ? [...zuclopenthixolMergedPool(), ...neurolepticPool()] + : [...zuclopenthixolMergedPool()]; + + const released = runOrderingPipeline({ + query: escalationQuery, + queryClass: analysis.queryClass, + candidates: pool, + topK: 8, + }); + + expect(released[0]?.document_id).toBe("neuroleptic-doc"); + const topFiveNeuroleptic = released.slice(0, 5).filter((result) => result.document_id === "neuroleptic-doc"); + // minCitations 2 for the live case: at least two citable chunks of the + // rescued doc must reach the released top-5. + expect(topFiveNeuroleptic.length).toBeGreaterThanOrEqual(2); + // Conservative availability: the sibling stays retrievable, just not on top. + expect(released.some((result) => result.document_id === "zuclopenthixol-doc")).toBe(true); + // The dose fast-path floor now passes on the rescued pool (0.94 >= 0.66). + expect(decideTextFastPath(escalationQuery, released, analysis.queryClass).returnFastPath).toBe(true); + + // Arrival-order invariance: byte-identical S3 primaries must not depend on + // candidate insertion order. + const reversed = runOrderingPipeline({ + query: escalationQuery, + queryClass: analysis.queryClass, + candidates: [...pool].reverse(), + topK: 8, + }); + expect(reversed.map((result) => result.id)).toEqual(released.map((result) => result.id)); + }); + + it("documents the pre-rescue floor rejection of the honest lexical pool", () => { + // Chain link 2 as executable documentation: without the S3 candidates the + // sibling-only pool fails the medication_dose_risk fast-path floor. + const decision = decideTextFastPath(escalationQuery, zuclopenthixolLexicalPool(), "medication_dose_risk"); + expect(decision.returnFastPath).toBe(false); + }); + + it("keeps non-firing dose-shaped pools byte-identical", () => { + const doseQuery = "What is the usual lithium dose for maintenance?"; + const analysis = analyzeClinicalQuery(doseQuery); + expect(shouldAttemptDocumentLookupFastPath("medication_dose_risk", analysis)).toBe(false); + + const pool = [ + zuclopenthixolS1("lithium-1", 0, "Lithium maintenance dosing is 250 mg twice daily for most adults."), + zuclopenthixolS1("lithium-2", 1, "Adjust the lithium dose according to trough levels."), + ].map((result, index) => ({ + ...result, + id: `lithium-${index + 1}`, + document_id: "lithium-doc", + title: "Lithium (AKG)", + file_name: "Lithium (AKG).pdf", + })); + + const released = runOrderingPipeline({ + query: doseQuery, + queryClass: "medication_dose_risk", + candidates: pool, + topK: 8, + }); + const releasedAgain = runOrderingPipeline({ + query: doseQuery, + queryClass: "medication_dose_risk", + candidates: pool, + topK: 8, + }); + expect(released.map((result) => result.id)).toEqual(releasedAgain.map((result) => result.id)); + }); +}); From a3b9a541b34c6401afb5fba5667e3280f95048d1 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:24:36 +0000 Subject: [PATCH 4/9] fix(rag): dose-value sentences resist inflected monitoring-kind steal (review P2/P3) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit rag-retrieval-reviewer reproduced a dose-intent regression from the parity commit: the inflected monitoring kind tokens (monitor\w*, annually, LFTs) claimed sentences carrying a concrete dose value ('Quetiapine is monitored at a dose of 200 mg daily'), and dose-intent answers then rejected the monitoring-kind fact — a source-gap where the figure previously shipped. The kind arm now classifies the legacy tokens byte-identically and lets a sentence claimed ONLY via the new inflections fall through to the dose arm when it carries a clinicalDoseValuePattern value; both reviewer repro sentences are pinned as dose-intent tests. Also extends the multi-drug bare-row guard to monitoring_schedule (reviewer P3 symmetry): the figure coverage escape made bare interval rows admissible without a query token, so a row naming no drug inside a multi-drug chunk must not inherit the queried drug — red-proof verified both directions. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- src/lib/rag/rag-extractive-answer.ts | 20 +++++++++---- tests/extractive-answer-formatting.test.ts | 34 ++++++++++++++++++++++ 2 files changed, 49 insertions(+), 5 deletions(-) diff --git a/src/lib/rag/rag-extractive-answer.ts b/src/lib/rag/rag-extractive-answer.ts index 7ef6971e..6cfaa4fc 100644 --- a/src/lib/rag/rag-extractive-answer.ts +++ b/src/lib/rag/rag-extractive-answer.ts @@ -751,9 +751,17 @@ function factKindForSentence(sentence: string, query: string, intent: AnswerInte // range 0.6-0.8 mmol/L" from the dose arm below and change fact priorities // for non-monitoring intents. "review" stays exact for the same reason: // review\w* would reclassify dose-review sentences ("doses should be - // reviewed daily") away from the dose arm. + // reviewed daily") away from the dose arm. The legacy tokens classify + // byte-identically; a sentence claimed ONLY via the new inflections that + // also carries a concrete dose value ("Quetiapine is monitored at a dose of + // 200 mg daily") falls through to the dose arm — otherwise a sole + // dose-bearing fact would be dropped by dose-intent answers (reviewer P2). + const legacyMonitoringKindPattern = + /\b(?:monitor|monitoring|baseline|weekly|monthly|annual|every|level|levels|blood test|ecg|lft|review)\b/i; + const inflectedMonitoringKindPattern = /\b(?:monitor\w*|annual(?:ly)?|levels?|blood tests?|ecgs?|lfts?)\b/i; if ( - /\b(?:monitor\w*|baseline|weekly|monthly|annual(?:ly)?|every|levels?|blood tests?|ecgs?|lfts?|review)\b/i.test(text) + legacyMonitoringKindPattern.test(text) || + (inflectedMonitoringKindPattern.test(text) && !clinicalDoseValuePattern.test(text)) ) { return "monitoring"; } @@ -870,13 +878,15 @@ function factSentenceMatchesQueryFromResult( const sentenceMedicationEntities = medicationEntitiesInText(sentence); const resultMedicationEntities = medicationEntitiesInText(resultText); if ( - intent === "dose" && + (intent === "dose" || intent === "monitoring_schedule") && queryMedicationEntities.length > 0 && sentenceMedicationEntities.length === 0 && resultMedicationEntities.length > 1 ) { - // A bare dose row from a multi-drug table cannot safely inherit the query's - // medication/class label. Require the row itself to name its medication. + // A bare dose or schedule row from a multi-drug table cannot safely inherit + // the query's medication/class label. Require the row itself to name its + // medication. (Monitoring joined dose here when the figure coverage escape + // below made bare interval/range rows admissible without a query token.) return false; } const entityCoveredByResult = diff --git a/tests/extractive-answer-formatting.test.ts b/tests/extractive-answer-formatting.test.ts index 4c22eb72..d19b1fac 100644 --- a/tests/extractive-answer-formatting.test.ts +++ b/tests/extractive-answer-formatting.test.ts @@ -555,6 +555,40 @@ describe("monitoring evidence gate parity (run-#60 miss class)", () => { expect(plain).toMatch(/0\.6-0\.8 mmol\/L/i); }); + it("keeps a sole dose figure whose sentence uses inflected monitoring tokens (reviewer P2)", () => { + // The inflected monitoring kind tokens must not steal a dose-value sentence + // from the dose arm: with these as the only dose-bearing facts, dose-intent + // answers previously flipped to a source-gap after the inflection widen. + for (const sentence of [ + "Quetiapine is monitored at a dose of 200 mg daily.", + "Quetiapine 200 mg daily is reviewed annually.", + ]) { + const answer = extractiveAnswerFor("What is the quetiapine dose?", [ + figureChunk({ id: "dose-steal-guard-1", content: sentence }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(answer.grounded).toBe(true); + expect(plain).toContain("200 mg"); + } + }); + + it("refuses a bare schedule row from a multi-drug chunk (reviewer P3)", () => { + const answer = extractiveAnswerFor("What monitoring is required for clozapine?", [ + figureChunk({ + id: "multi-drug-monitoring-guard-1", + title: "Antipsychotic Monitoring Summary", + file_name: "Antipsychotic Monitoring Summary.pdf", + section_heading: "Monitoring", + // The bare interval row names no drug; the chunk names two. The figure + // coverage escape must not let the row inherit the queried drug. + content: + "Clozapine and olanzapine monitoring requirements are listed below. Reviewed every 6 months with fasting bloods.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(plain).not.toMatch(/every 6 months/i); + }); + it("still refuses a schedule-free monitoring sentence with no query-token coverage", () => { const answer = extractiveAnswerFor("What metabolic monitoring is required for antipsychotics?", [ figureChunk({ From 55a6bc30477482ad5ad50ac3ea286365c8455d47 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:27:27 +0000 Subject: [PATCH 5/9] docs(ledger): rescue-retrieval review row + take redundant-optional-chain nit Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- docs/branch-review-ledger.md | 1 + src/lib/rag/rag.ts | 2 +- 2 files changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index b0e28108..fc7eefba 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -673,3 +673,4 @@ Use this ledger to prevent repeated branch and PR reviews when the reviewed HEAD | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (pair verdict — no code change; main stays at 22b6a2e) | canary run 29794759627 (#59, branch head 7310cb3 = merged PR-B content) | E-3b + E-3c PR-B LIVE-VALIDATED (pair vs banked #57/#58): route_ceiling_failures 2→0 (E-3b proven: agitation timeout now fits inside budget; clozapine retrieval-exhausted ceiling honestly excused via the triple-condition cross-region carve-out — exactly one "retrieval-exhausted" audit cell in the report); p95 17.4s→15.47s (-11%); golden retrieval SUCCESS 36/36 (stop-ship criterion held); validated_routine_extractive_first fired on 4 of the 6 target cases (patient-safety-plan 3.3s, treatment-team-process 2.6s, ect-procedure 2.2s, illegal-substances 2.2s — all pure extractive, zero generation, was 6-9s each with a discarded attempt), the other 2 (community-home-visits, best-practice-prescribing) stayed on generation+fallback because their extractive candidates legitimately fail validation gates = the designed-conservative outcome; ZERO new failing cases (list 5→2, both known residuals: neuroleptic citation red = Option A territory, admission-comparison expected-doc = non-blocking labeling residual); grounded 1.0 + unsupported_correct 1.0 held; targeting 0.5909→0.619, fail_closed 0.9→0.9333, readability/artifact_leaks/intent_coverage unchanged; relevance 0.6→0.5667 = single-case wobble on n=30, WATCH in E-4, not a gate. Discarded-generation rate materially down (4 conversions; residual = the H2 strong-route slice named as E-3d candidate, per design's 20-33% expectation band). MERGE STANDS (user had armed auto-merge pre-verdict; revert drill not triggered). Spend +~$3-6 → Phase E total ~$6-12 of ≤$20. | Pair evidence: run #59 job log (Blocking failures = citation only; Answer Case Diagnostics markers; metric_rates + targeting blocks); dump artifact populated for PR-C diagnosis | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (PR: E-3c PR-C figure-aware selection) | 043b030 + P2-fix commit | E-3c PR-C: dose/monitoring extractive answers now carry the asked-for figure/schedule when the cited chunk verbatim supports it — lead-slot promotion (dose swaps last of 2 slots, monitoring appends 2nd sentence; no-op when a lead already carries a figure) guarded by the claim-support atom corpus (sourceEvidenceText exported, promotionAtomKey byte-identical to claim-support's atomKey), plus the dose/threshold generation-fallback preferring the safe figure-carrying candidate (safety gate unchanged, filter-order-stable). Fallback helpers extracted to rag-extractive-answer (cycle-check verified); rag.ts 4908→4901. Six discriminating tests each verified red-on-prior-code incl. proving the nuke-guard load-bearing by disabling it. REVIEWS (both pre-push on 043b030): rag-retrieval-reviewer APPROVE-WITH-NITS — no-op path byte-identical verified, atom-key identity verified, filter-vs-find proven side-effect-free, 2-sentence append gate-safe, intent double-gated, zero retrieval/ordering change, imputation contract green; P2 = zero-atom monitoring figures ("every 6 weeks" yields no value atom) pass the guard trivially and can be nuked by claim support if sourced only from adjacent context (fails SAFE — evidence gap, never a wrong figure). clinical-governance-reviewer APPROVE-WITH-NITS — all six clinical concerns CLEARED end-to-end (verbatim-support guarantee, citation binding preserved, conservative failure test-proven, unsafe candidates impossible, wrong-drug risk controlled by pre-existing entity/multi-drug guards, no PHI); same zero-atom finding as P3 + one comment-precision nit. P2 FIXED post-review (disclosed): zero-atom figures now require the matched figure substring verbatim in sourceEvidenceText (intentFigureMatchText); proven both directions by 3 new tests (promotes from content, refuses from adjacent-context-only); comment-precision nit folded in. | Focused post-fix: extractive-formatting 32/32 + fallback/eval-cases/offline/contract/extractive-first 98/98 incl. imputation contract; typecheck+prettier+budgets clean; full-suite 3068-passed baseline pre-P2-fix (fix re-verified focused). Live proof = E-4 pair next | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (E-4 pair verdict — no code change; #1039 merged as 9b655fa) | canary run 29800029819 (#60, main 9b655fa = E-3b+PR-B+PR-C) | PHASE E-3 WAVE CLOSED — E-4 VERDICT: ADOPT (no revert). 44-case: the ONLY blocking red is the KNOWN persisting neuroleptic citation case (0.0227, identical #57 signature: generation quality-failed → extractive fallback 1 citation — the pre-declared Option A carve-out, retrieval-side, untouched by answer waves); route_ceiling_failures 0 CONFIRMED ON MAIN (E-3b: agitation 23.1s < 25s after 20.3s provider timeout; clozapine 14.5s with exactly one retrieval-exhausted audit cell); grounded 1.0 / unsupported_correct 1.0 / numeric 0 / governance-danger 0 all held; expected_source_hit 0.6136→0.6364 (#1020 widen); generation attempts 19→9 across 44 cases (10+ cases short-circuit via validated_routine_extractive_first at 2-6s, zero generation spend — the absolute wasted-generation seconds collapse; residual 6 discarded attempts are the named E-3d H2 strong/comparison slice). Targeting vs #58 baseline: rate 0.5909→0.6667, dose 1/5→2/4 (sertraline + quetiapine still miss), document_lookup 5/5→6/6, contraindication 2/2, red_result 2/2, pathway 1/2; readability/artifact_leaks 1.0. NOT met: monitoring_schedule flat 1/5 — per-miss lens shows answers of 73-232 chars with NO schedule token available to promote (olanzapine-lai 79ch, metabolic 73ch = single-fact extractive answers; the PR-C promotion is a no-op when no figure-bearing fact is extracted) → root is fact-extraction/retrieval depth on monitoring shapes, queued as the Option-A-wave companion diagnosis (dump artifact 8483731630, 30d retention). WATCH escalated: relevance 0.6 (#58) → 0.5667 (#59) → 0.5333 (#60) — two single-case steps coinciding with more terse extractive answers; fail_closed 0.9 = exactly the #58 main baseline (#59's 0.9333 was the outlier), safety texture flat. Adoption per plan criteria: targeting ≥ baseline ✓, grounded/refusal 1.0 ✓, golden 36/36 ✓ (in-run), ceilings 0 ✓, citation red = carved known case ✓. CodeRabbit post-review follow-up landed pre-merge (02b5c78): interval-regex full-match reorder (atom path proven to intercept the claimed exploit; reorder = drift hardening), clinicalValueAtomKey exported (mirror deleted), guard tests made honestly discriminating + genuine zero-atom "annually" coverage both directions. Instrument note: cost rates live on the targeting step env but eval:quality still reports cost n/a (estimator not consuming them in the 44-case path) — minor tooling residual. Spend +~$2-4 → Phase E total ~$8-16 of ≤$20. | Evidence: run #60 job log read in full (Threshold Status: citation-only; Answer Metrics + 44-row diagnostics; targeting metric_rates + 7-miss list); artifact eval-canary-output 8483731630 sha256 d5c7006e… (download blocked in-session — GitHub App scope; log tee carried the targeting output) | +| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (Option A: title-supported escalation rescue) | 0abf3c9 (parent 1aebf02) | rag-retrieval-reviewer PROTECTED-surface review of the S3 document-lookup escalation rescue: shouldAttemptDocumentLookupFastPath exported + gains medication_dose_risk branch firing ONLY when analysis.intent==="escalation_risk" && documentTitleTerms.length>0; call site passes queryAnalysis; new tests/rag-document-lookup-escalation-rescue.test.ts. | APPROVE-WITH-NITS. No P0/P1/P2. Binding constraints held: released-search-order.ts + retrieval-selection.ts clamp + rag-candidate-sources.ts imputation all byte-identical to parent (git diff empty); imputation-contract test green; 0.66/0.055 floor (rag.ts:1765) unmodified and still gates rescued pools (S3 block sits after the 2513 fast-path return + re-runs decideTextFastPath at 2631). BLAST RADIUS empirically proven via 110-case offline probe (44 ragEvalCases + 30 answerQualityEvalCases + 36 golden): EXACTLY 1 fires (neuroleptic-side-effect-escalation, intent=escalation_risk tt=3); all 8 named dose cases non-firing (clozapine-monitoring general, paraphrase general, agitation-pharm general, im-po drug_dosing, typo-dosing drug_dosing, missed-dose-table drug_dosing, LAI general/tt0, prompt-injection-forge intent=protocol). intentFromSignals precedence (clinical-search.ts:575-585) returns drug_dosing before escalation_risk so pure-dose structurally cannot fire — confirmed by construction AND empirically. ADVERSARIAL: prompt-injection-forge intent=protocol => cannot fire (empirical); unsupported short-circuit (rag.ts:2371) precedes S3 block. DOWNSTREAM: buildRetrievalIntent for the escalation query yields EMPTY requiredTermSignals => demote/promote arms (retrieval-selection.ts:354-355/529/555) + wrong-medication cap (rag.ts:730-737, gated on clinical_subject) all inert; end-to-end test proves neuroleptic-doc rank#1 with >=2 citations, sibling retained, arrival-order invariant. golden vector-* probes = broad_summary (already allowlisted) => predicate byte-identical => unaffected. TEST HONESTY: S3 fixture (0.92/0.34/0.94/synthetic_text) faithful to searchDocumentLookupFastPath (rag-candidate-sources.ts:570 caps alias documentScore at 0.34 => sim=min(0.92,0.58+0.34+bonus)=0.92, hybrid=min(0.94,0.94)); differently-relevant fixtures (not identical-content); red-proof structurally airtight (e2e gates its own fixture pool on the predicate). EFFORT: one S3 RPC (same call allowlisted classes issue), rescue-only in the non-forceEmbedding path (floor already rejected => query was headed to embedding anyway; successful rescue short-circuits at 2646 pre-embedding = net-neutral/positive). NITS (P3, non-blocking, no code change pre-canary): (1) redundant analysis?. on the documentTitleTerms clause (&& short-circuit already guarantees analysis defined there); (2) a queryMode forcing medication_dose_risk over an originally-escalation-shaped table_threshold could newly fire S3, but table_threshold already ran S3 so behavior-consistent, not a regression; (3) "forceEmbedding discards S3 merges" is imprecise — line 2649 merges S3 even under forceEmbedding; the accurate unaffected-reason for vector-* probes is broad_summary allowlisting. TRUST GATE = the mandated live canary pair (correctly deferred by the commit). | Offline only, no provider calls: new test 10/10 + imputation-contract 2/2; protected-surface batch 31/31 (escalation-rescue + imputation-contract + fast-path-ordering + released-search-order + retrieval-selection); 110-case classifier probe (temp test, removed); npm run typecheck exit 0. Not run (provider-gated): eval:retrieval:quality, eval:rag, the live canary pair. | diff --git a/src/lib/rag/rag.ts b/src/lib/rag/rag.ts index fae60829..6f3836ed 100644 --- a/src/lib/rag/rag.ts +++ b/src/lib/rag/rag.ts @@ -2145,7 +2145,7 @@ export function shouldAttemptDocumentLookupFastPath( return ( queryClass === "medication_dose_risk" && analysis?.intent === "escalation_risk" && - (analysis?.documentTitleTerms.length ?? 0) > 0 + analysis.documentTitleTerms.length > 0 ); } From a8ed7a44aecbd456ca2f0009cf40d58300430908 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:33:19 +0000 Subject: [PATCH 6/9] docs(ledger): all four Option-A-wave reviewer verdicts (parity + rescue) Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- docs/branch-review-ledger.md | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index fc7eefba..84497173 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -674,3 +674,6 @@ Use this ledger to prevent repeated branch and PR reviews when the reviewed HEAD | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (PR: E-3c PR-C figure-aware selection) | 043b030 + P2-fix commit | E-3c PR-C: dose/monitoring extractive answers now carry the asked-for figure/schedule when the cited chunk verbatim supports it — lead-slot promotion (dose swaps last of 2 slots, monitoring appends 2nd sentence; no-op when a lead already carries a figure) guarded by the claim-support atom corpus (sourceEvidenceText exported, promotionAtomKey byte-identical to claim-support's atomKey), plus the dose/threshold generation-fallback preferring the safe figure-carrying candidate (safety gate unchanged, filter-order-stable). Fallback helpers extracted to rag-extractive-answer (cycle-check verified); rag.ts 4908→4901. Six discriminating tests each verified red-on-prior-code incl. proving the nuke-guard load-bearing by disabling it. REVIEWS (both pre-push on 043b030): rag-retrieval-reviewer APPROVE-WITH-NITS — no-op path byte-identical verified, atom-key identity verified, filter-vs-find proven side-effect-free, 2-sentence append gate-safe, intent double-gated, zero retrieval/ordering change, imputation contract green; P2 = zero-atom monitoring figures ("every 6 weeks" yields no value atom) pass the guard trivially and can be nuked by claim support if sourced only from adjacent context (fails SAFE — evidence gap, never a wrong figure). clinical-governance-reviewer APPROVE-WITH-NITS — all six clinical concerns CLEARED end-to-end (verbatim-support guarantee, citation binding preserved, conservative failure test-proven, unsafe candidates impossible, wrong-drug risk controlled by pre-existing entity/multi-drug guards, no PHI); same zero-atom finding as P3 + one comment-precision nit. P2 FIXED post-review (disclosed): zero-atom figures now require the matched figure substring verbatim in sourceEvidenceText (intentFigureMatchText); proven both directions by 3 new tests (promotes from content, refuses from adjacent-context-only); comment-precision nit folded in. | Focused post-fix: extractive-formatting 32/32 + fallback/eval-cases/offline/contract/extractive-first 98/98 incl. imputation contract; typecheck+prettier+budgets clean; full-suite 3068-passed baseline pre-P2-fix (fix re-verified focused). Live proof = E-4 pair next | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (E-4 pair verdict — no code change; #1039 merged as 9b655fa) | canary run 29800029819 (#60, main 9b655fa = E-3b+PR-B+PR-C) | PHASE E-3 WAVE CLOSED — E-4 VERDICT: ADOPT (no revert). 44-case: the ONLY blocking red is the KNOWN persisting neuroleptic citation case (0.0227, identical #57 signature: generation quality-failed → extractive fallback 1 citation — the pre-declared Option A carve-out, retrieval-side, untouched by answer waves); route_ceiling_failures 0 CONFIRMED ON MAIN (E-3b: agitation 23.1s < 25s after 20.3s provider timeout; clozapine 14.5s with exactly one retrieval-exhausted audit cell); grounded 1.0 / unsupported_correct 1.0 / numeric 0 / governance-danger 0 all held; expected_source_hit 0.6136→0.6364 (#1020 widen); generation attempts 19→9 across 44 cases (10+ cases short-circuit via validated_routine_extractive_first at 2-6s, zero generation spend — the absolute wasted-generation seconds collapse; residual 6 discarded attempts are the named E-3d H2 strong/comparison slice). Targeting vs #58 baseline: rate 0.5909→0.6667, dose 1/5→2/4 (sertraline + quetiapine still miss), document_lookup 5/5→6/6, contraindication 2/2, red_result 2/2, pathway 1/2; readability/artifact_leaks 1.0. NOT met: monitoring_schedule flat 1/5 — per-miss lens shows answers of 73-232 chars with NO schedule token available to promote (olanzapine-lai 79ch, metabolic 73ch = single-fact extractive answers; the PR-C promotion is a no-op when no figure-bearing fact is extracted) → root is fact-extraction/retrieval depth on monitoring shapes, queued as the Option-A-wave companion diagnosis (dump artifact 8483731630, 30d retention). WATCH escalated: relevance 0.6 (#58) → 0.5667 (#59) → 0.5333 (#60) — two single-case steps coinciding with more terse extractive answers; fail_closed 0.9 = exactly the #58 main baseline (#59's 0.9333 was the outlier), safety texture flat. Adoption per plan criteria: targeting ≥ baseline ✓, grounded/refusal 1.0 ✓, golden 36/36 ✓ (in-run), ceilings 0 ✓, citation red = carved known case ✓. CodeRabbit post-review follow-up landed pre-merge (02b5c78): interval-regex full-match reorder (atom path proven to intercept the claimed exploit; reorder = drift hardening), clinicalValueAtomKey exported (mirror deleted), guard tests made honestly discriminating + genuine zero-atom "annually" coverage both directions. Instrument note: cost rates live on the targeting step env but eval:quality still reports cost n/a (estimator not consuming them in the 44-case path) — minor tooling residual. Spend +~$2-4 → Phase E total ~$8-16 of ≤$20. | Evidence: run #60 job log read in full (Threshold Status: citation-only; Answer Metrics + 44-row diagnostics; targeting metric_rates + 7-miss list); artifact eval-canary-output 8483731630 sha256 d5c7006e… (download blocked in-session — GitHub App scope; log tee carried the targeting output) | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (Option A: title-supported escalation rescue) | 0abf3c9 (parent 1aebf02) | rag-retrieval-reviewer PROTECTED-surface review of the S3 document-lookup escalation rescue: shouldAttemptDocumentLookupFastPath exported + gains medication_dose_risk branch firing ONLY when analysis.intent==="escalation_risk" && documentTitleTerms.length>0; call site passes queryAnalysis; new tests/rag-document-lookup-escalation-rescue.test.ts. | APPROVE-WITH-NITS. No P0/P1/P2. Binding constraints held: released-search-order.ts + retrieval-selection.ts clamp + rag-candidate-sources.ts imputation all byte-identical to parent (git diff empty); imputation-contract test green; 0.66/0.055 floor (rag.ts:1765) unmodified and still gates rescued pools (S3 block sits after the 2513 fast-path return + re-runs decideTextFastPath at 2631). BLAST RADIUS empirically proven via 110-case offline probe (44 ragEvalCases + 30 answerQualityEvalCases + 36 golden): EXACTLY 1 fires (neuroleptic-side-effect-escalation, intent=escalation_risk tt=3); all 8 named dose cases non-firing (clozapine-monitoring general, paraphrase general, agitation-pharm general, im-po drug_dosing, typo-dosing drug_dosing, missed-dose-table drug_dosing, LAI general/tt0, prompt-injection-forge intent=protocol). intentFromSignals precedence (clinical-search.ts:575-585) returns drug_dosing before escalation_risk so pure-dose structurally cannot fire — confirmed by construction AND empirically. ADVERSARIAL: prompt-injection-forge intent=protocol => cannot fire (empirical); unsupported short-circuit (rag.ts:2371) precedes S3 block. DOWNSTREAM: buildRetrievalIntent for the escalation query yields EMPTY requiredTermSignals => demote/promote arms (retrieval-selection.ts:354-355/529/555) + wrong-medication cap (rag.ts:730-737, gated on clinical_subject) all inert; end-to-end test proves neuroleptic-doc rank#1 with >=2 citations, sibling retained, arrival-order invariant. golden vector-* probes = broad_summary (already allowlisted) => predicate byte-identical => unaffected. TEST HONESTY: S3 fixture (0.92/0.34/0.94/synthetic_text) faithful to searchDocumentLookupFastPath (rag-candidate-sources.ts:570 caps alias documentScore at 0.34 => sim=min(0.92,0.58+0.34+bonus)=0.92, hybrid=min(0.94,0.94)); differently-relevant fixtures (not identical-content); red-proof structurally airtight (e2e gates its own fixture pool on the predicate). EFFORT: one S3 RPC (same call allowlisted classes issue), rescue-only in the non-forceEmbedding path (floor already rejected => query was headed to embedding anyway; successful rescue short-circuits at 2646 pre-embedding = net-neutral/positive). NITS (P3, non-blocking, no code change pre-canary): (1) redundant analysis?. on the documentTitleTerms clause (&& short-circuit already guarantees analysis defined there); (2) a queryMode forcing medication_dose_risk over an originally-escalation-shaped table_threshold could newly fire S3, but table_threshold already ran S3 so behavior-consistent, not a regression; (3) "forceEmbedding discards S3 merges" is imprecise — line 2649 merges S3 even under forceEmbedding; the accurate unaffected-reason for vector-* probes is broad_summary allowlisting. TRUST GATE = the mandated live canary pair (correctly deferred by the commit). | Offline only, no provider calls: new test 10/10 + imputation-contract 2/2; protected-surface batch 31/31 (escalation-rescue + imputation-contract + fast-path-ordering + released-search-order + retrieval-selection); 110-case classifier probe (temp test, removed); npm run typecheck exit 0. Not run (provider-gated): eval:retrieval:quality, eval:rag, the live canary pair. | +| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (parity commit review) | 1aebf02 (fix landed a3b9a54) | rag-retrieval-reviewer on the monitoring evidence-gate parity commit: REQUEST-CHANGES (soft) — P2 reproduced: inflected monitoring kind tokens (monitor\w*/annual(?:ly)?/blood tests?/ecgs?/lfts?) steal sole-dose-value sentences from the dose arm; dose-intent answers then reject the monitoring-kind fact ("Quetiapine is monitored at a dose of 200 mg daily" flipped grounded true→false, source-gap — fails CLOSED, never a wrong dose). P3: monitoring figure escape lacked the dose escape's multi-drug bare-row guard. Clean: over-admission bounded (broad vocab lives in gate/filter only, promotion still corpus-guarded, claim-support unchanged); regex cost negligible; mismatched-unit test relaxation legitimate (synopsis is corpus-verbatim; weeks pin enforced by atom identity + adjacent_context exclusion from both gate and claim corpora). | BOTH FINDINGS FIXED in a3b9a54: kind arm classifies legacy tokens byte-identically and new-inflection-only sentences fall through to the dose arm when they carry a clinicalDoseValuePattern value (both repro sentences pinned as dose-intent tests); multi-drug bare-row guard extended to monitoring_schedule with a discriminating test — red-proven both directions. Reviewer checks: formatting 38/38, focused 644/644; targeting eval deferred to the wave's live canary pair. | +| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (parity commit review) | 1aebf02 (fix landed a3b9a54) | clinical-governance-reviewer on the same commit: APPROVE-WITH-NITS. P2 (independently converged with the retrieval reviewer's P3): monitoring figure-escape lacked the dose-path multi-drug cross-entity guard — a bare wrong-drug schedule/level row in a multi-drug chunk could be entity-prefixed for a named-drug monitoring query; downstream gates verify text-vs-source presence, never attribution (worked lithium/valproate LFT path traced through finalize). FIXED in a3b9a54 exactly as its smallest-fix prescribed (guard at the :872-881 site now fires for monitoring_schedule; negative multi-drug test added, red-proven). Clean: unsupported figures impossible (admission-only change; promotion corpus guard + numeric verification + claim support all byte-unchanged); conservative failure intact (figure-bearing-only escape, schedule-free refusal pinned); adjacent-context safety held (sourceEvidenceText excludes adjacent_context; weeks refusal confirmed by probe); no PHI/provider/ranking surface. P3s: RAG impact line (present in the PR body — behaviour-change form, correct since the PR also carries the Option A retrieval change); multi-drug negative test (landed in a3b9a54). | Offline guard-chain trace + targeted vitest probes (named-drug guard, bare-figure admission, conservative gap, adjacent refusal). Provider/release gates deferred per confirmation boundary. | +| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (Option A rescue review) | 0abf3c9 | clinical-governance-reviewer on the S3 escalation rescue: APPROVE-WITH-NITS, no P0/P1. P2 = the mandated live canary pair itself (process gate, declared in the PR body; offline-green + review-approved proven insufficient for this surface 2026-07-20). P3s: multi-drug escalation-query recall edge (titled drug + untitled drug — fast-path return can skip the vector leg; recall limitation, not misattribution, mirrors the pre-existing allowlisted-class tradeoff); reviewer probe files must stay uncommitted (relocated to scratchpad). All six clinical concerns verified safe: wrong-document impossible (alias phrases must appear in the query; per-document alias groups, no cross-drug conflation), conservative availability (S3 purely additive via keyed-union merge; sibling retention test-pinned), live expansion acceptable (title-named correct-entity SOP in every firing shape), fail-closed double layer (adversarial short-circuit precedes the predicate; injection-forge case intent=protocol cannot fire — executed), governance metadata unbypassed (same attachDocumentRankingMetadata + status=indexed + access-scope filters), no PHI/provider/schema surface. | Reviewer checks: escalation-rescue suite 8/8, injection-forge intent derivation executed, static trace of the full S3 chain. Live canary pair = the trust gate, dispatched post-merge. | From 57a898fe87a544f8c85d3aa1bc4a360194a4eb7e Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:46:19 +0000 Subject: [PATCH 7/9] fix(rag): dose-kind monitoring evidence requires a cadence or level signal MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit CodeRabbit major on the parity widen: kind===dose facts passed the shared monitoring vocabulary via generic dose wording (range/therapeutic/maintenance around a mg figure), so 'The therapeutic dose range is 300-450 mg daily.' counted as monitoring-schedule evidence with no cadence/level signal. Dose-kind facts now additionally require a cadence or level token (monitoringCadenceOrLevelPattern — the strict schedule/level subset of the shared vocabulary); level ranges like 'Maintenance range is 0.6-0.8 mmol/L' stay admissible through their level units (red-proof pinned both directions). Also escapes the literal vector-* asterisks in the ledger row flagged MD037. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- docs/branch-review-ledger.md | 2 +- src/lib/rag/rag-extractive-answer.ts | 15 +++++++++++++++ tests/extractive-answer-formatting.test.ts | 16 ++++++++++++++++ 3 files changed, 32 insertions(+), 1 deletion(-) diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index 84497173..10a512d9 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -673,7 +673,7 @@ Use this ledger to prevent repeated branch and PR reviews when the reviewed HEAD | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (pair verdict — no code change; main stays at 22b6a2e) | canary run 29794759627 (#59, branch head 7310cb3 = merged PR-B content) | E-3b + E-3c PR-B LIVE-VALIDATED (pair vs banked #57/#58): route_ceiling_failures 2→0 (E-3b proven: agitation timeout now fits inside budget; clozapine retrieval-exhausted ceiling honestly excused via the triple-condition cross-region carve-out — exactly one "retrieval-exhausted" audit cell in the report); p95 17.4s→15.47s (-11%); golden retrieval SUCCESS 36/36 (stop-ship criterion held); validated_routine_extractive_first fired on 4 of the 6 target cases (patient-safety-plan 3.3s, treatment-team-process 2.6s, ect-procedure 2.2s, illegal-substances 2.2s — all pure extractive, zero generation, was 6-9s each with a discarded attempt), the other 2 (community-home-visits, best-practice-prescribing) stayed on generation+fallback because their extractive candidates legitimately fail validation gates = the designed-conservative outcome; ZERO new failing cases (list 5→2, both known residuals: neuroleptic citation red = Option A territory, admission-comparison expected-doc = non-blocking labeling residual); grounded 1.0 + unsupported_correct 1.0 held; targeting 0.5909→0.619, fail_closed 0.9→0.9333, readability/artifact_leaks/intent_coverage unchanged; relevance 0.6→0.5667 = single-case wobble on n=30, WATCH in E-4, not a gate. Discarded-generation rate materially down (4 conversions; residual = the H2 strong-route slice named as E-3d candidate, per design's 20-33% expectation band). MERGE STANDS (user had armed auto-merge pre-verdict; revert drill not triggered). Spend +~$3-6 → Phase E total ~$6-12 of ≤$20. | Pair evidence: run #59 job log (Blocking failures = citation only; Answer Case Diagnostics markers; metric_rates + targeting blocks); dump artifact populated for PR-C diagnosis | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (PR: E-3c PR-C figure-aware selection) | 043b030 + P2-fix commit | E-3c PR-C: dose/monitoring extractive answers now carry the asked-for figure/schedule when the cited chunk verbatim supports it — lead-slot promotion (dose swaps last of 2 slots, monitoring appends 2nd sentence; no-op when a lead already carries a figure) guarded by the claim-support atom corpus (sourceEvidenceText exported, promotionAtomKey byte-identical to claim-support's atomKey), plus the dose/threshold generation-fallback preferring the safe figure-carrying candidate (safety gate unchanged, filter-order-stable). Fallback helpers extracted to rag-extractive-answer (cycle-check verified); rag.ts 4908→4901. Six discriminating tests each verified red-on-prior-code incl. proving the nuke-guard load-bearing by disabling it. REVIEWS (both pre-push on 043b030): rag-retrieval-reviewer APPROVE-WITH-NITS — no-op path byte-identical verified, atom-key identity verified, filter-vs-find proven side-effect-free, 2-sentence append gate-safe, intent double-gated, zero retrieval/ordering change, imputation contract green; P2 = zero-atom monitoring figures ("every 6 weeks" yields no value atom) pass the guard trivially and can be nuked by claim support if sourced only from adjacent context (fails SAFE — evidence gap, never a wrong figure). clinical-governance-reviewer APPROVE-WITH-NITS — all six clinical concerns CLEARED end-to-end (verbatim-support guarantee, citation binding preserved, conservative failure test-proven, unsafe candidates impossible, wrong-drug risk controlled by pre-existing entity/multi-drug guards, no PHI); same zero-atom finding as P3 + one comment-precision nit. P2 FIXED post-review (disclosed): zero-atom figures now require the matched figure substring verbatim in sourceEvidenceText (intentFigureMatchText); proven both directions by 3 new tests (promotes from content, refuses from adjacent-context-only); comment-precision nit folded in. | Focused post-fix: extractive-formatting 32/32 + fallback/eval-cases/offline/contract/extractive-first 98/98 incl. imputation contract; typecheck+prettier+budgets clean; full-suite 3068-passed baseline pre-P2-fix (fix re-verified focused). Live proof = E-4 pair next | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (E-4 pair verdict — no code change; #1039 merged as 9b655fa) | canary run 29800029819 (#60, main 9b655fa = E-3b+PR-B+PR-C) | PHASE E-3 WAVE CLOSED — E-4 VERDICT: ADOPT (no revert). 44-case: the ONLY blocking red is the KNOWN persisting neuroleptic citation case (0.0227, identical #57 signature: generation quality-failed → extractive fallback 1 citation — the pre-declared Option A carve-out, retrieval-side, untouched by answer waves); route_ceiling_failures 0 CONFIRMED ON MAIN (E-3b: agitation 23.1s < 25s after 20.3s provider timeout; clozapine 14.5s with exactly one retrieval-exhausted audit cell); grounded 1.0 / unsupported_correct 1.0 / numeric 0 / governance-danger 0 all held; expected_source_hit 0.6136→0.6364 (#1020 widen); generation attempts 19→9 across 44 cases (10+ cases short-circuit via validated_routine_extractive_first at 2-6s, zero generation spend — the absolute wasted-generation seconds collapse; residual 6 discarded attempts are the named E-3d H2 strong/comparison slice). Targeting vs #58 baseline: rate 0.5909→0.6667, dose 1/5→2/4 (sertraline + quetiapine still miss), document_lookup 5/5→6/6, contraindication 2/2, red_result 2/2, pathway 1/2; readability/artifact_leaks 1.0. NOT met: monitoring_schedule flat 1/5 — per-miss lens shows answers of 73-232 chars with NO schedule token available to promote (olanzapine-lai 79ch, metabolic 73ch = single-fact extractive answers; the PR-C promotion is a no-op when no figure-bearing fact is extracted) → root is fact-extraction/retrieval depth on monitoring shapes, queued as the Option-A-wave companion diagnosis (dump artifact 8483731630, 30d retention). WATCH escalated: relevance 0.6 (#58) → 0.5667 (#59) → 0.5333 (#60) — two single-case steps coinciding with more terse extractive answers; fail_closed 0.9 = exactly the #58 main baseline (#59's 0.9333 was the outlier), safety texture flat. Adoption per plan criteria: targeting ≥ baseline ✓, grounded/refusal 1.0 ✓, golden 36/36 ✓ (in-run), ceilings 0 ✓, citation red = carved known case ✓. CodeRabbit post-review follow-up landed pre-merge (02b5c78): interval-regex full-match reorder (atom path proven to intercept the claimed exploit; reorder = drift hardening), clinicalValueAtomKey exported (mirror deleted), guard tests made honestly discriminating + genuine zero-atom "annually" coverage both directions. Instrument note: cost rates live on the targeting step env but eval:quality still reports cost n/a (estimator not consuming them in the 44-case path) — minor tooling residual. Spend +~$2-4 → Phase E total ~$8-16 of ≤$20. | Evidence: run #60 job log read in full (Threshold Status: citation-only; Answer Metrics + 44-row diagnostics; targeting metric_rates + 7-miss list); artifact eval-canary-output 8483731630 sha256 d5c7006e… (download blocked in-session — GitHub App scope; log tee carried the targeting output) | -| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (Option A: title-supported escalation rescue) | 0abf3c9 (parent 1aebf02) | rag-retrieval-reviewer PROTECTED-surface review of the S3 document-lookup escalation rescue: shouldAttemptDocumentLookupFastPath exported + gains medication_dose_risk branch firing ONLY when analysis.intent==="escalation_risk" && documentTitleTerms.length>0; call site passes queryAnalysis; new tests/rag-document-lookup-escalation-rescue.test.ts. | APPROVE-WITH-NITS. No P0/P1/P2. Binding constraints held: released-search-order.ts + retrieval-selection.ts clamp + rag-candidate-sources.ts imputation all byte-identical to parent (git diff empty); imputation-contract test green; 0.66/0.055 floor (rag.ts:1765) unmodified and still gates rescued pools (S3 block sits after the 2513 fast-path return + re-runs decideTextFastPath at 2631). BLAST RADIUS empirically proven via 110-case offline probe (44 ragEvalCases + 30 answerQualityEvalCases + 36 golden): EXACTLY 1 fires (neuroleptic-side-effect-escalation, intent=escalation_risk tt=3); all 8 named dose cases non-firing (clozapine-monitoring general, paraphrase general, agitation-pharm general, im-po drug_dosing, typo-dosing drug_dosing, missed-dose-table drug_dosing, LAI general/tt0, prompt-injection-forge intent=protocol). intentFromSignals precedence (clinical-search.ts:575-585) returns drug_dosing before escalation_risk so pure-dose structurally cannot fire — confirmed by construction AND empirically. ADVERSARIAL: prompt-injection-forge intent=protocol => cannot fire (empirical); unsupported short-circuit (rag.ts:2371) precedes S3 block. DOWNSTREAM: buildRetrievalIntent for the escalation query yields EMPTY requiredTermSignals => demote/promote arms (retrieval-selection.ts:354-355/529/555) + wrong-medication cap (rag.ts:730-737, gated on clinical_subject) all inert; end-to-end test proves neuroleptic-doc rank#1 with >=2 citations, sibling retained, arrival-order invariant. golden vector-* probes = broad_summary (already allowlisted) => predicate byte-identical => unaffected. TEST HONESTY: S3 fixture (0.92/0.34/0.94/synthetic_text) faithful to searchDocumentLookupFastPath (rag-candidate-sources.ts:570 caps alias documentScore at 0.34 => sim=min(0.92,0.58+0.34+bonus)=0.92, hybrid=min(0.94,0.94)); differently-relevant fixtures (not identical-content); red-proof structurally airtight (e2e gates its own fixture pool on the predicate). EFFORT: one S3 RPC (same call allowlisted classes issue), rescue-only in the non-forceEmbedding path (floor already rejected => query was headed to embedding anyway; successful rescue short-circuits at 2646 pre-embedding = net-neutral/positive). NITS (P3, non-blocking, no code change pre-canary): (1) redundant analysis?. on the documentTitleTerms clause (&& short-circuit already guarantees analysis defined there); (2) a queryMode forcing medication_dose_risk over an originally-escalation-shaped table_threshold could newly fire S3, but table_threshold already ran S3 so behavior-consistent, not a regression; (3) "forceEmbedding discards S3 merges" is imprecise — line 2649 merges S3 even under forceEmbedding; the accurate unaffected-reason for vector-* probes is broad_summary allowlisting. TRUST GATE = the mandated live canary pair (correctly deferred by the commit). | Offline only, no provider calls: new test 10/10 + imputation-contract 2/2; protected-surface batch 31/31 (escalation-rescue + imputation-contract + fast-path-ordering + released-search-order + retrieval-selection); 110-case classifier probe (temp test, removed); npm run typecheck exit 0. Not run (provider-gated): eval:retrieval:quality, eval:rag, the live canary pair. | +| 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (Option A: title-supported escalation rescue) | 0abf3c9 (parent 1aebf02) | rag-retrieval-reviewer PROTECTED-surface review of the S3 document-lookup escalation rescue: shouldAttemptDocumentLookupFastPath exported + gains medication_dose_risk branch firing ONLY when analysis.intent==="escalation_risk" && documentTitleTerms.length>0; call site passes queryAnalysis; new tests/rag-document-lookup-escalation-rescue.test.ts. | APPROVE-WITH-NITS. No P0/P1/P2. Binding constraints held: released-search-order.ts + retrieval-selection.ts clamp + rag-candidate-sources.ts imputation all byte-identical to parent (git diff empty); imputation-contract test green; 0.66/0.055 floor (rag.ts:1765) unmodified and still gates rescued pools (S3 block sits after the 2513 fast-path return + re-runs decideTextFastPath at 2631). BLAST RADIUS empirically proven via 110-case offline probe (44 ragEvalCases + 30 answerQualityEvalCases + 36 golden): EXACTLY 1 fires (neuroleptic-side-effect-escalation, intent=escalation_risk tt=3); all 8 named dose cases non-firing (clozapine-monitoring general, paraphrase general, agitation-pharm general, im-po drug_dosing, typo-dosing drug_dosing, missed-dose-table drug_dosing, LAI general/tt0, prompt-injection-forge intent=protocol). intentFromSignals precedence (clinical-search.ts:575-585) returns drug_dosing before escalation_risk so pure-dose structurally cannot fire — confirmed by construction AND empirically. ADVERSARIAL: prompt-injection-forge intent=protocol => cannot fire (empirical); unsupported short-circuit (rag.ts:2371) precedes S3 block. DOWNSTREAM: buildRetrievalIntent for the escalation query yields EMPTY requiredTermSignals => demote/promote arms (retrieval-selection.ts:354-355/529/555) + wrong-medication cap (rag.ts:730-737, gated on clinical_subject) all inert; end-to-end test proves neuroleptic-doc rank#1 with >=2 citations, sibling retained, arrival-order invariant. golden vector-\* probes = broad_summary (already allowlisted) => predicate byte-identical => unaffected. TEST HONESTY: S3 fixture (0.92/0.34/0.94/synthetic_text) faithful to searchDocumentLookupFastPath (rag-candidate-sources.ts:570 caps alias documentScore at 0.34 => sim=min(0.92,0.58+0.34+bonus)=0.92, hybrid=min(0.94,0.94)); differently-relevant fixtures (not identical-content); red-proof structurally airtight (e2e gates its own fixture pool on the predicate). EFFORT: one S3 RPC (same call allowlisted classes issue), rescue-only in the non-forceEmbedding path (floor already rejected => query was headed to embedding anyway; successful rescue short-circuits at 2646 pre-embedding = net-neutral/positive). NITS (P3, non-blocking, no code change pre-canary): (1) redundant analysis?. on the documentTitleTerms clause (&& short-circuit already guarantees analysis defined there); (2) a queryMode forcing medication_dose_risk over an originally-escalation-shaped table_threshold could newly fire S3, but table_threshold already ran S3 so behavior-consistent, not a regression; (3) "forceEmbedding discards S3 merges" is imprecise — line 2649 merges S3 even under forceEmbedding; the accurate unaffected-reason for vector-\* probes is broad_summary allowlisting. TRUST GATE = the mandated live canary pair (correctly deferred by the commit). | Offline only, no provider calls: new test 10/10 + imputation-contract 2/2; protected-surface batch 31/31 (escalation-rescue + imputation-contract + fast-path-ordering + released-search-order + retrieval-selection); 110-case classifier probe (temp test, removed); npm run typecheck exit 0. Not run (provider-gated): eval:retrieval:quality, eval:rag, the live canary pair. | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (parity commit review) | 1aebf02 (fix landed a3b9a54) | rag-retrieval-reviewer on the monitoring evidence-gate parity commit: REQUEST-CHANGES (soft) — P2 reproduced: inflected monitoring kind tokens (monitor\w*/annual(?:ly)?/blood tests?/ecgs?/lfts?) steal sole-dose-value sentences from the dose arm; dose-intent answers then reject the monitoring-kind fact ("Quetiapine is monitored at a dose of 200 mg daily" flipped grounded true→false, source-gap — fails CLOSED, never a wrong dose). P3: monitoring figure escape lacked the dose escape's multi-drug bare-row guard. Clean: over-admission bounded (broad vocab lives in gate/filter only, promotion still corpus-guarded, claim-support unchanged); regex cost negligible; mismatched-unit test relaxation legitimate (synopsis is corpus-verbatim; weeks pin enforced by atom identity + adjacent_context exclusion from both gate and claim corpora). | BOTH FINDINGS FIXED in a3b9a54: kind arm classifies legacy tokens byte-identically and new-inflection-only sentences fall through to the dose arm when they carry a clinicalDoseValuePattern value (both repro sentences pinned as dose-intent tests); multi-drug bare-row guard extended to monitoring_schedule with a discriminating test — red-proven both directions. Reviewer checks: formatting 38/38, focused 644/644; targeting eval deferred to the wave's live canary pair. | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (parity commit review) | 1aebf02 (fix landed a3b9a54) | clinical-governance-reviewer on the same commit: APPROVE-WITH-NITS. P2 (independently converged with the retrieval reviewer's P3): monitoring figure-escape lacked the dose-path multi-drug cross-entity guard — a bare wrong-drug schedule/level row in a multi-drug chunk could be entity-prefixed for a named-drug monitoring query; downstream gates verify text-vs-source presence, never attribution (worked lithium/valproate LFT path traced through finalize). FIXED in a3b9a54 exactly as its smallest-fix prescribed (guard at the :872-881 site now fires for monitoring_schedule; negative multi-drug test added, red-proven). Clean: unsupported figures impossible (admission-only change; promotion corpus guard + numeric verification + claim support all byte-unchanged); conservative failure intact (figure-bearing-only escape, schedule-free refusal pinned); adjacent-context safety held (sourceEvidenceText excludes adjacent_context; weeks refusal confirmed by probe); no PHI/provider/ranking surface. P3s: RAG impact line (present in the PR body — behaviour-change form, correct since the PR also carries the Option A retrieval change); multi-drug negative test (landed in a3b9a54). | Offline guard-chain trace + targeted vitest probes (named-drug guard, bare-figure admission, conservative gap, adjacent refusal). Provider/release gates deferred per confirmation boundary. | | 2026-07-21 | claude/clinical-kb-pwa-review-asi3wb (Option A rescue review) | 0abf3c9 | clinical-governance-reviewer on the S3 escalation rescue: APPROVE-WITH-NITS, no P0/P1. P2 = the mandated live canary pair itself (process gate, declared in the PR body; offline-green + review-approved proven insufficient for this surface 2026-07-20). P3s: multi-drug escalation-query recall edge (titled drug + untitled drug — fast-path return can skip the vector leg; recall limitation, not misattribution, mirrors the pre-existing allowlisted-class tradeoff); reviewer probe files must stay uncommitted (relocated to scratchpad). All six clinical concerns verified safe: wrong-document impossible (alias phrases must appear in the query; per-document alias groups, no cross-drug conflation), conservative availability (S3 purely additive via keyed-union merge; sibling retention test-pinned), live expansion acceptable (title-named correct-entity SOP in every firing shape), fail-closed double layer (adversarial short-circuit precedes the predicate; injection-forge case intent=protocol cannot fire — executed), governance metadata unbypassed (same attachDocumentRankingMetadata + status=indexed + access-scope filters), no PHI/provider/schema surface. | Reviewer checks: escalation-rescue suite 8/8, injection-forge intent derivation executed, static trace of the full S3 chain. Live canary pair = the trust gate, dispatched post-merge. | diff --git a/src/lib/rag/rag-extractive-answer.ts b/src/lib/rag/rag-extractive-answer.ts index 6cfaa4fc..ecd4ec51 100644 --- a/src/lib/rag/rag-extractive-answer.ts +++ b/src/lib/rag/rag-extractive-answer.ts @@ -342,6 +342,14 @@ function queryIntentTokens(query: string, intent: AnswerIntent) { // digit durations ("at 12 weeks") count as schedule evidence. const monitoringScheduleEvidenceSource = String.raw`monitor\w*|follow[-\s]?up|baseline|weekly|monthly|annual(?:ly)?|yearly|every|several\s+times\s+a\s+year|screen(?:ing|ed)?|levels?|blood tests?|bloods|fbcs?|ancs?|wbcs?|ecgs?|lfts?|renal|thyroid|metabolic|glucose|bsl|lipids|cholesterol|triglycerides|blood pressure|bp|pulse|weight|bmi|mmol\/l|mcg\/l|ng\/ml|range|target|therapeutic|maintenance|review\w*|\d+\s*(?:week|month|day|hour|year)s?`; const monitoringScheduleEvidencePattern = new RegExp(String.raw`\b(?:${monitoringScheduleEvidenceSource})\b`, "i"); +// The strict subset of the monitoring vocabulary that is an actual cadence or +// level signal: dose-kind facts must carry one of these before they count as +// monitoring evidence, so generic dose wording (range/therapeutic/maintenance +// around a mg figure) cannot masquerade as a schedule. +const monitoringCadenceOrLevelPattern = new RegExp( + String.raw`\b(?:monitor\w*|follow[-\s]?up|baseline|weekly|monthly|annual(?:ly)?|yearly|every|several\s+times\s+a\s+year|screen(?:ing|ed)?|review\w*|levels?|serum|trough|plasma|mmol\/l|mcg\/l|ng\/ml|\d+\s*(?:week|month|day|hour|year)s?)\b`, + "i", +); /** Answer intent evidence pattern. */ function answerIntentEvidencePattern(intent: AnswerIntent) { @@ -820,6 +828,13 @@ function factSupportsAnswerIntent( // are classified as renal_limit (renal check triggers before monitoring), but are directly relevant // to monitoring schedule answers. if (kind !== "monitoring" && kind !== "dose" && kind !== "renal_limit") return false; + // A dose-kind fact must carry an actual cadence or level signal — the + // shared vocabulary alone would admit generic dose language + // ("The therapeutic dose range is 300-450 mg daily.") as monitoring + // evidence via range/therapeutic/maintenance tokens (CodeRabbit major). + // Level ranges like "Maintenance range is 0.6-0.8 mmol/L" stay admissible + // through their level units. + if (kind === "dose" && !monitoringCadenceOrLevelPattern.test(text)) return false; return monitoringScheduleEvidencePattern.test(text); case "red_result_action": if (kind !== "threshold_action" && kind !== "caveat") return false; diff --git a/tests/extractive-answer-formatting.test.ts b/tests/extractive-answer-formatting.test.ts index d19b1fac..c3c7fe0f 100644 --- a/tests/extractive-answer-formatting.test.ts +++ b/tests/extractive-answer-formatting.test.ts @@ -572,6 +572,22 @@ describe("monitoring evidence gate parity (run-#60 miss class)", () => { } }); + it("refuses generic dose-range prose as monitoring evidence (CodeRabbit major)", () => { + const answer = extractiveAnswerFor("What monitoring is required for quetiapine?", [ + figureChunk({ + id: "dose-range-not-schedule-1", + section_heading: "Quetiapine", + // The dose range carries range/therapeutic/maintenance vocabulary and a + // mg unit-range figure, but no cadence or level signal — it must not be + // admitted as monitoring-schedule evidence. + content: + "Quetiapine monitoring requirements are described in the prescribing guideline. The therapeutic dose range is 300-450 mg daily.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(plain).not.toContain("300-450 mg"); + }); + it("refuses a bare schedule row from a multi-drug chunk (reviewer P3)", () => { const answer = extractiveAnswerFor("What monitoring is required for clozapine?", [ figureChunk({ From d6c218aaaa1ba22a2f5ea34c3c6499a4cd438e43 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:52:32 +0000 Subject: [PATCH 8/9] fix(rag): monitoring coverage escapes accept level ranges only (CodeRabbit major 2) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The result-level and sentence-level monitoring coverage escapes reused monitoringUnitRangeFigurePattern, whose unit list includes dose-amount units (mg/micrograms/mcg) — so a therapeutic dose range could still grant monitoring intent coverage even though the dose-kind fact guard rejects the sentence itself. Both escapes now use monitoringLevelRangeCoveragePattern (mmol/L | nmol/L | mcg/L | ng/mL only); the wider pattern stays reserved for the corpus-guarded figure-promotion checks. Regression test pins a token-free chunk whose only candidate coverage grantor is a mg dose range. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- src/lib/rag/rag-extractive-answer.ts | 9 +++++++-- tests/extractive-answer-formatting.test.ts | 15 +++++++++++++++ 2 files changed, 22 insertions(+), 2 deletions(-) diff --git a/src/lib/rag/rag-extractive-answer.ts b/src/lib/rag/rag-extractive-answer.ts index ecd4ec51..a14911ca 100644 --- a/src/lib/rag/rag-extractive-answer.ts +++ b/src/lib/rag/rag-extractive-answer.ts @@ -350,6 +350,11 @@ const monitoringCadenceOrLevelPattern = new RegExp( String.raw`\b(?:monitor\w*|follow[-\s]?up|baseline|weekly|monthly|annual(?:ly)?|yearly|every|several\s+times\s+a\s+year|screen(?:ing|ed)?|review\w*|levels?|serum|trough|plasma|mmol\/l|mcg\/l|ng\/ml|\d+\s*(?:week|month|day|hour|year)s?)\b`, "i", ); +// Level/concentration ranges only, for the monitoring intent COVERAGE escapes: +// dose-amount unit ranges ("300-450 mg") must not grant coverage — the wider +// monitoringUnitRangeFigurePattern (which also matches mg/mcg dose ranges) +// stays reserved for the corpus-guarded figure-promotion checks. +const monitoringLevelRangeCoveragePattern = /\d+(?:\.\d+)?\s*[-–]\s*\d+(?:\.\d+)?\s*(?:mmol\/L|nmol\/L|mcg\/L|ng\/mL)/i; /** Answer intent evidence pattern. */ function answerIntentEvidencePattern(intent: AnswerIntent) { @@ -488,7 +493,7 @@ function resultCoversAnswerIntent(result: SearchResult, query: string, intent: A // monitor/level — the run-#60 miss class rejected such chunks wholesale here. const monitoringFigureCoverage = intent === "monitoring_schedule" && - (monitoringIntervalFigurePattern.test(text) || monitoringUnitRangeFigurePattern.test(text)); + (monitoringIntervalFigurePattern.test(text) || monitoringLevelRangeCoveragePattern.test(text)); const intentCoverage = answerIntentEvidencePattern(intent).test(text) || maximumDoseCoverage || monitoringFigureCoverage; if (!intentCoverage) return false; @@ -920,7 +925,7 @@ function factSentenceMatchesQueryFromResult( intentTokens.some((token) => queryTokenMatchesText(token, normalized)) || (intent === "dose" && extractiveConcreteDosePattern.test(normalized)) || (intent === "monitoring_schedule" && - (monitoringIntervalFigurePattern.test(normalized) || monitoringUnitRangeFigurePattern.test(normalized))); + (monitoringIntervalFigurePattern.test(normalized) || monitoringLevelRangeCoveragePattern.test(normalized))); return answerIntentEvidencePattern(intent).test(normalized) && intentCovered; } diff --git a/tests/extractive-answer-formatting.test.ts b/tests/extractive-answer-formatting.test.ts index c3c7fe0f..41f41b91 100644 --- a/tests/extractive-answer-formatting.test.ts +++ b/tests/extractive-answer-formatting.test.ts @@ -572,6 +572,21 @@ describe("monitoring evidence gate parity (run-#60 miss class)", () => { } }); + it("denies monitoring intent coverage to a dose-amount range in a token-free chunk (CodeRabbit major #2)", () => { + const answer = extractiveAnswerFor("What monitoring is required for quetiapine?", [ + figureChunk({ + id: "dose-range-coverage-guard-1", + section_heading: "Quetiapine", + // No monitoring token anywhere: the ONLY thing that could grant intent + // coverage is the mg dose range, which must not count as a monitoring + // figure at either the result level or the sentence level. + content: "Quetiapine prescribing information follows. The therapeutic dose range is 300-450 mg daily.", + }), + ]); + const plain = (answer.answer ?? "").replace(/\*\*/g, ""); + expect(plain).not.toContain("300-450 mg"); + }); + it("refuses generic dose-range prose as monitoring evidence (CodeRabbit major)", () => { const answer = extractiveAnswerFor("What monitoring is required for quetiapine?", [ figureChunk({ From e2a39962db189c87354ca543cbd45ae15c51d600 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 21 Jul 2026 05:57:20 +0000 Subject: [PATCH 9/9] test(rag): make the dose-range coverage fixture genuinely token-free (CodeRabbit) Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01EXsJcLrbZUXwnBeG91cVo9 --- tests/extractive-answer-formatting.test.ts | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/tests/extractive-answer-formatting.test.ts b/tests/extractive-answer-formatting.test.ts index 41f41b91..fc710ae8 100644 --- a/tests/extractive-answer-formatting.test.ts +++ b/tests/extractive-answer-formatting.test.ts @@ -577,10 +577,13 @@ describe("monitoring evidence gate parity (run-#60 miss class)", () => { figureChunk({ id: "dose-range-coverage-guard-1", section_heading: "Quetiapine", - // No monitoring token anywhere: the ONLY thing that could grant intent - // coverage is the mg dose range, which must not count as a monitoring - // figure at either the result level or the sentence level. - content: "Quetiapine prescribing information follows. The therapeutic dose range is 300-450 mg daily.", + // Genuinely token-free: no monitoring vocabulary at all, so the ONLY + // candidate coverage grantor is the mg dose range, which must not + // count as a monitoring figure at either the result level or the + // sentence level. (End-to-end this is defense-in-depth layered with + // the gate miss and the dose-kind cadence/level guard — the + // generic-vocabulary path is pinned by the neighbouring test.) + content: "Quetiapine prescribing information follows. 300-450 mg is listed for adults.", }), ]); const plain = (answer.answer ?? "").replace(/\*\*/g, "");