P3-8: CSI500 independent generalization check for value/lowvol signals - #21
Merged
Merged
Conversation
…aming) Ask whether the P3-7 sign-level conclusion (value_ep/value_bp positive, volatility_20 negative on the independent 2024-07..2026-05 window) generalizes outside the screened universes, specifically to CSI500 (000905.SH) — a cell independent in BOTH universe and time. Machinery is the unchanged P3-7 independent-validation layer: same factor groups, same cost scenarios, no new factors, no tuning, no alpha/portfolio/ execution/OOS-slicing change. - config/phase3_real_csi500_generalization.yaml: screened anchor SSE50|2022-2024 (must reproduce P3-6/P3-7), independent anchor SSE50|2024-2026 (must reproduce the P3-7 verdict numbers), and the NEW independent cell 000905.SH|2024-2026; CSI500|2022-2024 skipped + disclosed (runtime budget). Data feasibility probed: monthly 500-name index_weight snapshots through 2026-05-29 incl. a pre-start 2024-06-28 snapshot; 645 distinct in-window constituents. - output.subset_report_name (baseline_report_name precedent): each subset-validation study owns its report filename; None keeps the historical default bitwise (locked by tests) — a P3-8 run no longer clobbers the accepted P3-7 artifact. - tests: +4 (CSI500 config cells/groups/scenarios/hypotheses, sample classes, default report filename preserved, P3-8 owns its filename) -> 379 passed; ruff clean; 11 configs validate; demo run-phase0 unchanged (ic 0.96 / annual 0.8408)
- CLAUDE.md / AGENTS.md (verbatim copy): P3-8 progress entry — both anchors reconciled exactly (screened raw ICs 22/22 == P3-5; the independent SSE50 verdict ICs identical to P3-7, a second reproducibility confirmation); CSI500|2024-2026 verdict SUPPORTED (value_ep +0.0083/+0.0145, value_bp +0.0230/+0.0127, volatility_20 -0.0350/-0.0272; 21 settled rebalances vs minimum 8) with LESS attenuation than the SSE50/CSI300 holdouts -> the P3-7 sign-level conclusion GENERALIZES outside the screened universes; CSI500 portfolio results positive across all groups and the full cost ladder (honest caveat: legacy_trio +17.8% is another small-sample/regime ranking flip; not a return claim); gates updated to 379 passed - TEST_REPORT.md: P3-8 per-file breakdown (4) + real-data validation entry (anchor reconciliation, verdict, preserved P3-7 artifact, secret scan 0) - RUNBOOK.md: Phase 3-8 section (cell roles, subset_report_name, report-type title note, skip disclosure)
…(review HIGH) The P3-8 CSI500 report rendered with the hardcoded "Phase 3-7 — Independent-Sample Validation" H1 (qt/reports.py only branched the title on has_independent), contradicting the config and docs that identify it as the P3-8 CSI500 generalization study — the same class of stale report wording flagged on P3-7. - config: OutputCfg gains `subset_report_title` (the `subset_report_name` precedent — config owns the study identity). None keeps the renderer's sample-aware default (P3-7 independent / P3-6 post-hoc). - render_subset_validation: the H1 is config-driven when output.subset_report_title is set, else the existing sample-aware default. The body framing (screened-vs-independent mechanics) stays sample-aware — correct for any such study, so only the H1 changes. - config/phase3_real_csi500_generalization.yaml: sets the title to "Phase 3-8 — CSI500 Independent Generalization Check ...". - tests: +4 (config sets title; P3-6/P3-7 leave it unset; rendered CSI500 H1 starts with Phase 3-8 + contains CSI500 + is NOT Phase 3-7; default title preserved without the override) -> 383 passed. - artifact: regenerated deterministically — only the H1 line changes with this fix, so the accepted real numbers are untouched (the 3.5h matrix is NOT rerun); diff is exactly the 1 title line; secret scan 0. - docs: RUNBOOK/TEST_REPORT/CLAUDE.md/AGENTS.md note the config-driven title and the corrected "title carries report type" wording.
StackOverFlow11
added a commit
that referenced
this pull request
Jun 13, 2026
docs: mark PR #21 (P3-8) merged in progress docs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
P3-8: does the P3-7 sign-level conclusion (value_ep/value_bp positive, volatility_20 negative on the independent 2024-07..2026-05 window) generalize outside the screened universes, specifically to CSI500 (000905.SH, mid-caps) — a cell independent in BOTH universe and time? Report-only; the unchanged P3-7 independent-validation machinery (same factor groups, same cost scenarios, same hypotheses); no new factors, no tuning, no alpha/portfolio/execution/OOS-slicing change.
config/phase3_real_csi500_generalization.yaml: screened anchorSSE50|2022-2024(must reproduce P3-6/P3-7), independent anchorSSE50|2024-2026(must reproduce the P3-7 verdict numbers), and the NEW independent cell000905.SH|2024-2026;CSI500|2022-2024skipped + disclosed (runtime budget). Data feasibility probed first (monthly 500-nameindex_weightsnapshots through 2026-05-29 incl. a pre-start 2024-06-28 snapshot).output.subset_report_name(thebaseline_report_nameprecedent): each subset-validation study owns its report filename, so the P3-8 run does not clobber the accepted P3-7 artifact. Configs without the key keep the historical filename bitwise (locked by tests).output.subset_report_title(review HIGH fix): the report H1 is config-driven so the CSI500 study names itself ("Phase 3-8 — CSI500 Independent Generalization Check") instead of the machinery's default P3-7 phase label; the body framing stays sample-aware. P3-6/P3-7 configs leave it unset and keep the renderer's sample-aware default (regression-locked).Real run (3 cells, ~3.55h; the 735-name CSI500 cell dominates)
legacy_trioposts +17.8%→+10.6% — another small-sample/regime ranking flip across cells; ~21 rebalances, single window. NOT a return claim.Test plan
pytest -p no:cacheprovider— 383 passed (+8: CSI500 config cells/groups/scenarios/hypotheses, sample classes, default + own report filename, config-driven title incl. "first line is NOT Phase 3-7", default-title regression)ruff check .— cleanvalidate-configon all 11 configs — OKrun-phase0demo regression — ic 0.9600 / annual 0.8408 unchangedgit diff --check+ conflict marker scan — clean