TODO review — 2026-09-18
Reviewed EXECUTION-PLAN.md and reports/kaggle-community-review-findings-2026-09-17.md in full. Updated TODO.md, retaining existing completed tasks and historical experiment rows. Neither input document nor running evaluation code was changed.
Decisions
Confirmed current-code issues
probe-current-contract.py calls the current evaluator with mocked tiny arrays and no training/index data. Its result semantic-probes.json reproduces Regime-A query expansion/self-comparison, the conditional-MRR denominator error, full-pool coverage being truncated to top-1000 recall, and strict-output validation gaps. These are findings to fix, not passing G1 tests. Source hashes are attached to the probes.
Evidence and scope
TODO.before-review.mdreview-check.jsonprobe-current-contract.py, semantic-probes.json