Restore and expand Chronicle cross-dataset evaluation - #159
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
added 26 commits
August 12, 2026 21:08
This reverts commit 55493c8.
anth-volk
force-pushed
the
restore-microcosm-comparisons
branch
from
August 12, 2026 18:17
bf7262f to
7b59b75
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #157
Summary
Why
The prior Cross-dataset page had been removed, and the available comparison logic did not provide a complete, auditable classification of Chronicle facts across heterogeneous datasets and model/dataset pairs. Period-aligned facts also require source-specific handling so 2022 and 2023 observations are not incorrectly compared directly with 2024 estimates.
The restored page now uses explicit adapter capabilities, immutable evaluation artifacts, approved Microcosm aging alignments, and a responsive layout that switches before horizontal scrolling begins. The model-backed Python verification is an explicit periodic maintenance operation; lightweight frontend checks remain in CI.
User impact
Users can compare error and coverage across all four current sources, group results by Chronicle source, period, or geography, and drill down to individual facts. The matrix keeps all four sources visible on standard displays and becomes a 2×2 or 1×4 source grid when its own content area narrows.
Future maintainers can rerun the locked harness checks with
uv run python scripts/verify_evaluation_harness.py, optionally supplying pinned Microcosm and ACS inputs for their numerical gates.Validation
bun test— 82 passedbun run lint— TypeScript passedevaluation-02ba5db8abd8414c22eb6625