← Back

📑 Contents

Enveda CASMI 2026 TodoNext execution orderEvidence rules and current snapshotMorning findings — disposition after the evening auditImmediate audit follow-ups — historical local fixes and remaining acceptance workPhase 0 — Ground truth and infrastructurePhase 1 — Spectral retrieval baselineControlled retrieval ablations — after the corrected baseline is lockedPhase 2 — Learned fingerprint ranking and domain adaptationPhase 3 — Known structures without reference spectraPhase 3A — Pinned structure-only universe (promoted; first after G1)September 19 audit repair gates (before promotion)Phase 3B — Formula/mass-gated candidates and non-neural baselinePhase 3C — Mass-shifted analogue propagation (separate route)Later representations and forward checkingPhase 4 — Novel structures and de novo generationPhase 5 — Unified ensemblePhase 6 — Kaggle packaging and final submissionActive experiment logBlockers and decisions

Enveda CASMI 2026 Todo


Legend: [ ] pending · [-] in progress · [x] complete · [!] blocked


Reviewed: 2026-09-20 (MYT), after Phase 3A index completion. Current findings: [audit report](reports/review-20260919-1817/AUDIT.md). This review supersedes stale active-process and completion statements below; historical experiment entries remain historical.


Current audit snapshot: 45 maintained tests passed across the five focused suites. G1 saved-result arithmetic and placeholder output were independently audited previously, but full G1 remains OPEN. The shadow is a seven-row repair copy, not a completed source rebuild. Production is not promoted or changed.


Material corrections: the scaffold classifier compared identity overlap in its category logic: 1,741 scaffold-overlap changes, not zero (1,295 lost / 446 gained; 7,853 other scaffold-string differences). The registered readout is not certified untouched: 164 selected identities belong to the previously inspected fold 4. Analogue combined_mrr is untruncated reciprocal rank, not MRR@25, and its failure/stratum counters need correction. Keep all Phase 3A results diagnostic-only.


Checklist count after reconciliation: counts below include the G0 closure evidence; gate states are separate and this is not an effort-weighted or competition-readiness percentage. Four explicit shadow/analogue repair tasks remain tracked.


Evening audit: 2026-09-18, current source + isolated synthetic probes + saved artifacts. Evidence and pre-audit backup: reports/audit-20260918-1923/. The audit changes documentation only; the active evaluator and its inputs remain untouched. Earlier completion claims are qualified below.


Audit report: [AUDIT.md](reports/audit-20260918-1923/AUDIT.md). Fresh lightweight suite: 24 passed, 2 real-index tests deliberately deselected; additional independent synthetic evaluator/entrypoint/failure-injection evidence is saved alongside it. The earlier 26-pass result is not a claim that all gate scenarios are covered.


The priority order is deliberate: establish trustworthy evaluation, expand candidate coverage, then learn to rank. Community methods are hypotheses, not validated improvements. This review updates the backlog; it does not mark unrun experiments complete.


Next execution order


1. P0 — Close reproduced audit defects before promotion. Correct scaffold classification, analogue scorer-identity/query accounting, MRR@25, failure categories and report-set exposure tracking. Resolve shadow manifest/progress and reproducible repair evidence. Preserve existing outputs; do not overwrite production or rerun the full spectral evaluation as part of this review.

2. Lock a qualified local retrieval baseline, then run bounded mass-window, candidate-cap and multi-spectrum ablations on development folds. Keep an untouched report partition.

3. Promote Phase 3A/3B ahead of learned ranking: pinned/eligible structure sources, mass/formula indexing, candidate recall, reference-availability accounting, and a non-neural structure-aware baseline. A ranker cannot recover an answer absent from its pool, and a reference-spectrum matcher cannot rank an identity whose eligible reference spectra were removed.

4. Separate spectral-library diagnostics from unseen-identity ranking: retain library-ceiling reranking as an optimistic implementation control; report identity-held-out cross-identity decoy ordering/reference availability separately, without treating an absent target reference as a conventional zero-quality rank. Then proceed to structure-aware fingerprint/analogue methods before learned ranking.

5. Phase 2 + Phase 3C: leakage-safe spectrum-to-fingerprint/ranking experiments and a separate mass-shifted analogue-propagation benchmark. The analogue experiment may run alongside bounded ranking once both use the same frozen candidate pool and folds; it is not a prerequisite for starting every ranker experiment.

6. Keep de novo generation and cross-route fusion deferred. Revisit Phase 4 when database/analogue coverage failures justify it; validate route fusion in Phase 5, not by copying public weights.


Phase numbers are retained for traceability. These dependencies take precedence over the original plan's provisional Phase 2-before-Phase 3 calendar; no new dates or compute commitments are implied.


Evidence rules and current snapshot


  • Placeholder warning: downloadable test.parquet contains training examples. Its 400-molecule/1,213-spectrum output is a schema/runtime diagnostic only, not CV, public-leaderboard evidence or a G1 quality score. Production IDs and counts must come from the runtime test file, never be hardcoded from the sample.
  • Authoritative namespace: use artifacts/rebuild-20260917/ and manifests/rebuild-20260917/ together. The root retrieval symlink currently resolves there. Historical 274,288-identity and raw 275,810-identity artifacts are not interchangeable with the rebuilt 274,195 scorer identities.
  • Regimes are different questions: A deliberately permits same-identity references only under the documented acquisition-exclusion proxy; B must exclude all held-out-identity spectra across every library, analogue prototype and derived spectral feature. A library-only holdout must never be presented as clean B/novelty validation. C additionally excludes structures/structure targets and requires a frozen scaffold-aware proxy.
  • Current data count: use the hashed release's 1,184 NP-example rows / 250 identities. The published 1,151-row count is a documented external discrepancy, not a reason to trim the local cohort or block its evaluation.
  • Infrastructure evidence: reports/phase0-qualified-g0-20260917.json records local guards/tests and hashes. Official scorer source acquisition and pinned-runtime fixture certification are now complete; acquisition-lineage certification remains unavailable.
  • Earlier G1 evidence: reports/phase1-g1-molecule-full-20260917.json uses one valid spectrum per identity, not all-spectrum validation. Preserve it as a bounded diagnostic; audit metric semantics before comparisons. Library-ceiling scores are optimistic diagnostics, not held-out generalization estimates.
  • Completed G1, not an active run: reports/g1-audit-20260919-1005/AUDIT.md reconciles the later 20:02 launch, saved ranks and 400-row validated placeholder submission. Evaluation took 9h04m20s / ~4.90 GiB peak RSS; placeholder inference 1m49.03s / ~1.44 GiB. Saved arithmetic passed; this is not immutable launch provenance or a full hidden-workload timing test. No matching evaluator/build/pytest process was active at this review's initial inspection.
  • Training release rechecked: the CLI redownload matched size, SHA-256, schema, 2,539,608 rows and 21 row groups. Evidence: reports/todo-review-20260918/kaggle-redownload-comparison-20260918.json. The verified staged duplicate was permanently deleted at the user's request; authoritative training data and the comparison report were retained. This proves equality to the download checked on September 18, not perpetual freshness.
  • Existing rebuilt submission: submissions/phase1-retrieval-molecule-rebuilt-20260917.csv exists and its generation report records 400 rows and 1,213 spectra. A local schema check passed earlier in this session; this does not establish strict candidate validation, hidden-test accuracy or current-run provenance. The legacy submissions/phase1-retrieval.csv remains unsuitable as current evidence.
  • Source packaging partly repaired: correctly named manifests/pubchem-structure-only-1-500000.tsv and source manifest exist; original misnamed .tsv.gz is preserved. Source hash is 6006275eb9fde9101ec1f342eacb1240fad10201c3e8182eaf584cf31ea91152; 240,536 rows / 231,755 declared identities. Verifier checks suffix/magic, but the builder still accepts misleading plain-TSV suffixes and maintained format-rejection tests are missing. Upstream snapshot/licensing remains uncertified.

  • Morning findings — disposition after the evening audit


    The morning evidence below describes the pre-fix implementation, not current source. Preserve [semantic-probes.json](reports/todo-review-20260918/semantic-probes.json) and its source hashes as historical evidence; do not rerun its old defect-expecting script over that report. Fresh evening probes are in reports/audit-20260918-1923/evaluator-edge-probes.json and submission-wrapper-probes.json.


  • Fixed in the all-spectrum evaluator: A now receives the frozen row query mask. A fresh fixture verifies that the eligible reference is not ingested as a query and that no query/reference self-comparison occurs.
  • Fixed in the all-spectrum evaluator: conditional MRR divides by all pool-covered molecules; pool coverage is computed before top-k truncation. Fresh fixtures cover ranks 1/25/26/100/1,000/1,001/absent, an explicit tie and an empty pool, including the coverage × conditional-MRR identity. This does not validate every legacy evaluator or full-run result.
  • Historical validator defects, superseded by A3: raw-field/token/ID checks were subsequently fixed and the current local output audited. Robust precursor coercion and explicit validator path/hash binding remain separate open issues.

  • Current consequence: the old run has ended and wrapper fixes are verified. Saved per-molecule target ranks exist, but complete candidate/stratum evidence, immutable provenance, runtime hardening and C remain incomplete. Preserve the useful results without full G1 promotion.


    Immediate audit follow-ups — historical local fixes and remaining acceptance work


  • [x] P0 · A1 · locally fixed and failure-injection verified Wrapper now uses set -Eeuo pipefail, stage labels and an ERR trap; failed evaluator/submission/validation stages stop before later DONE markers. The exact copied-wrapper probe now returns 41/42/43 respectively with no false success. The active long-running invocation was not restarted and therefore still used the pre-fix shell.
  • [x] P0 · A2 · locally fixed and cold-invocation verified Runtime validator inserts the repository root for clean direct-script execution and writes phase1-retrieval-rebuilt-20260918.validation.json, preserving generation metadata in the original .json. Synthetic clean-environment and metadata-preservation probes pass. The active invocation's already-running process is not retroactively changed.
  • [x] P0 · A3 · local contract fixed and probed Validator now rejects missing fields, null/blank runtime IDs before string coercion, extra CSV fields and interior empty candidate tokens; an entirely blank smiles field is the explicit no-candidate policy. Added contract tests and clean helper probes. This is local validation, not official scorer parity.
  • [x] P0 · A4 · post-hoc provenance snapshot verified 2026-09-20 Created manifests/a4-provenance-snapshot-20260920.json with SHA-256 hashes for current Python source and frozen inputs/artifacts, runtime/platform metadata, dependency freeze, Git state, and explicit missing/limitation fields. This is a reproducibility snapshot for future launch discipline, not immutable launch-time provenance for earlier runs; G0/G1 remain open.
  • [x] P0 · official scorer source and pinned-runtime certification · verified 2026-09-20 Replaced src/scorer_contract.py with the exact Kaggle metric/casmi-mean-reciprocal-rank source (SHA-256 b542cd945203cf4305a0a01537a5fa9e93fa4e2d79400b8563286bce0f73ed2d) and ran a deterministic fixture under RDKit 2026.03.3. Expected and actual MRR@25 were both 0.625; rank-26 rejection and invalid-solution host-error probes passed. Evidence: reports/official-scorer-certification-20260920/evidence.json. This certifies local execution parity for the fixture, not hidden-test accuracy or data acquisition lineage.
  • [-] P0 · A5 · partial evidence package verified 2026-09-20 Packaged existing all-spectrum G1 per-molecule ranks/query counts for library-ceiling, Regime B and Regime A, plus available library-ceiling fold/query-count strata, in reports/a5-g1-evidence-pack-20260920.json (12,198 per-molecule records). Explicitly recorded missing complete candidate IDs/scores/source rows, full A/B strata, immutable launch binding, and Regime C (NOT_EVALUATED). This closes evidence reconciliation/packaging only; a clean A5 completion still requires a rerun or persisted candidate-level evidence.
  • [x] P1 · A6 · maintained regression coverage verified 2026-09-20 The focused G1/Phase 3A regression suite ran successfully in .venv-rdkit2026: 26 passed (tests/test_phase3a_regressions.py and tests/test_g1_contract.py). This confirms the maintained local contract tests; it does not close full G1 acceptance.
  • [ ] P1 · A7 Replace coarse progress records with time/query/chunk-level updates and, separately, real resume state. Current JSON files contain counters, not candidate scores or restart state; restarts lose work. REGIME_COMPLETE is written before rank/metric calculation and does not persist those results. Do not infer ETA from rows, stale checkpoint elapsed time, or CPU utilization.
  • [x] P1 · A8 · draft eligibility corrected Phase 3A explicitly retains permitted database structures for B while excluding their spectra/features; C exclusions remain separate. GPU speedup is still unbenchmarked and optimization remains design-only. The historical instruction to preserve the then-active run is not a claim that it is still running.

  • Phase 0 — Ground truth and infrastructure


  • [x] P0 Download the current train.parquet release from Kaggle.
  • [x] P0 · 2026-09-18 recheck Redownload via Kaggle CLI into staging and compare SHA-256, size, schema, rows and row groups. The saved comparison matches in every field. Permanently remove only that verified duplicate when requested; retain authoritative data and comparison evidence.
  • [x] P0 Record file size, SHA-256, Kaggle metadata and download timestamp in manifests/.
  • [x] P0 Inspect the actual training schema, row groups, library counts and identity counts.
  • [x] P0 Add a low-RAM data access path using PyArrow row-group streaming and selected columns.
  • [x] P0 Create a pinned RDKit 2026.03.3 environment for scorer-compatible identity matching.
  • [x] P0 Implement tautomer-canonicalized InChIKey14 identity logic in src/identity.py.
  • [x] P0 Add scorer unit tests for tautomer equivalence, stereochemistry, invalid SMILES and rank cutoff.
  • [x] P0 Implement all ten advertised adduct mass conversions and round-trip tests.
  • [x] P0 · basic alias fix Normalize [M+FA-H]- to [M+CH2O2-H]- in both conversion directions; retain independent formate-mass and round-trip checks. Current tests verify alias use in neutral-mass conversion, reject a non-string adduct and reject the wrong-polarity FA spelling. This is not exhaustive validation of all numeric/polarity inputs.
  • [x] P1 · symmetric adduct/input edge tests · verified 2026-09-20 Added positive finite-mass policy checks for every advertised adduct, canonical/alias formate parity, non-numeric/non-finite/zero/negative precursor and neutral inputs, and reran the focused pinned-runtime suite: 48 passed. Evaluator-level malformed-row handling remains separately covered by G1-04.
  • [-] P0 Freeze molecule-grouped validation manifests for library, database-only and novelty-proxy regimes. The rebuilt scorer-compatible five-fold and strict external manifests are internally consistent; qualified Regime A query/reference masks are now operational and runtime-validated, while Regime B identity-wide exclusion manifests and runtime checks now pass locally, and Regime C remains proxy-only.
  • [x] P0 · cross-library leakage fixture · verified 2026-09-20 Added an identity-wide Regime-B exclusion helper and synthetic regression fixtures proving a held-out identity is excluded from every library, including references, spectral prototypes and derived features; a library-only exclusion is explicitly rejected. Focused suite: 25 passed; maintained six-file suite: 52 passed. This verifies the local mask contract, not external acquisition-lineage certification.
  • [ ] P0 Freeze development/report partitions, seeds and tuning policy before community-inspired sweeps; protect the 250-compound NP report set from repeated tuning with grouped/nested validation.
  • [x] P1 Add a reproducible experiment configuration format under configs/.

  • Resource constraint: the local development host has approximately 11 GiB RAM. Full-data operations must stream or use bounded artifacts; do not load the entire training table into pandas.


    Gate G0: [x] Qualified local infrastructure pass — verified 2026-09-20. The pinned runtime, official scorer, 15-file artifact/source binding, fail-closed runtime/hash preflight, deliberate tamper failure, and 48 maintained tests pass. Evidence: reports/g0-local-closure-20260920/decision.json and reports/g0-local-closure-20260920/manifest.json. This closes qualified local G0 only; acquisition-lineage certification remains externally unavailable, Regime A is an exact-available-field proxy, Regime C is proxy-only, and this is not hidden-test or competition-performance certification.


    Phase 1 — Spectral retrieval baseline


  • [x] P0 Build a streaming Parquet reader and compact peak-array representation.
  • [x] P0 Implement precursor/adduct to neutral-mass filtering.
  • [x] P0 Implement peak cleaning variants: intensity floor, precursor exclusion and top-N/windowed peaks.
  • [x] P0 Implement fragment cosine similarity.
  • [x] P0 Implement spectral entropy similarity.
  • [x] P0 Implement neutral-loss similarity.
  • [-] P0 Build and verify indexed reference search. Rebuilt index metadata and the qualified infrastructure report record 2,539,608 rows, 2,160,035 valid spectra and 274,195 scorer identities, with five fold masks. Preserve the legacy index under artifacts/legacy-20260917/; complete end-to-end artifact binding rather than trusting a symlink or status label alone.
  • [x] P0 · local fix verified Pass the A row query mask into the all-spectrum evaluator. Independent evening synthetic evidence confirms only the permitted query is used against eligible references; full cohort accounting and aggregation acceptance remain G1-01 work.
  • [-] P0 · G1-01 Complete correct all-spectrum molecule-level aggregation on frozen query cohorts. For A, use the actual regime_a_query.npy row mask, not every row sharing one of its 250 identities. Assert eligible query/reference separation and reconcile total, valid, skipped and zero-valid-spectrum molecules. For B/C, retain identity-wide exclusions. Aggregate all permitted spectra within each evaluation molecule; produce exactly one output per runtime molecule_id.
  • [x] P0 · G1-02 · local math verified Independently probe the corrected all-spectrum evaluator at its default top-k=1,000: target ranks 1/25/26/100/1,000/1,001/absent, a deterministic equal-score tie and an empty pool. Full pool coverage is distinct from Recall@1,000; conditional MRR divides by all covered targets; MRR@25 = pool_coverage × conditional_MRR@25 holds, with zero returned for empty coverage. Evidence: reports/audit-20260918-1923/evaluator-edge-probes.json. Promote these probes into maintained tests (A6); the legacy evaluators and full-run evidence are not thereby certified.
  • [-] P0 · G1-03 · bounded audit passed 2026-09-20 Independently audited the existing 400-row molecule-level submission: 9,364 raw candidates, 9,364 unique scorer identities, maximum 25 per row, 0 invalid SMILES, 0 duplicate identity rows, 1,001 sampled index identity→SMILES mappings with 0 stale mappings, and tautomer/stereo equivalence probes passed. Evidence: reports/g1-03-identity-dedup-20260920/evidence.json. Full closure still requires candidate identity/source-row evidence from the active all-spectrum evaluator.
  • [x] P0 · G1-04 · expanded offline packaging acceptance verified 2026-09-20 Expanded local fixture used six renamed/reversed runtime molecule IDs across 24 spectra, including empty peaks, null precursor, unsupported adduct and multiple spectra. Network socket construction was disabled; inference produced 6/6 rows with 100% coverage and strict validation passed. Four negative validator probes rejected duplicate IDs, interior empty candidates, >25 candidates and extra columns. Evidence: reports/g1-04-offline-acceptance-20260920/evidence.json. Fixed null precursor coercion so malformed precursor values are skipped instead of crashing. This closes local packaging/validator acceptance, not hidden-test accuracy, official scorer parity or full G1.
  • [x] P0 · G1-05 · post-run audit verified 2026-09-21 Completed all-spectrum G1 rerun (12h03m / 43,232s) produced complete ranked candidate IDs, scores, source reference rows and query source rows for all three regimes: library_ceiling (212,831 evidence rows), regime_b (204,513), regime_a (7,138). Independent recomputation from evidence reproduced every reported metric with zero drift: library ceiling MRR@25 0.8624 / coverage 90.37%, Regime B 0.0 (identity-wide exclusion, 0 violations), Regime A MRR@25 0.9297 / coverage 100%. Per-molecule ranks: 0 mismatches. Candidate identity dedup: 0 duplicates. Source rows: all within bounds. Evidence: reports/g1-post-run-audit-20260921/evidence.json.
  • [x] P0 · G1-06 · provenance snapshot verified 2026-09-21 Created manifests/g1-06-provenance-snapshot-20260921.json binding 21 files (7 source, 10 input, 4 output) with SHA-256 hashes, plus /usr/bin/time -v stats (peak RSS 5,206,524 KiB ≈ 4.96 GB, user 34,381s), runtime metadata (Python 3.10, RDKit 2026.03.3, aarch64), git state (head 8609672198), pip freeze, and 5 explicit limitations. Self-hash included for tamper detection. This is a post-run binding, not launch-time capture; launch-time immutability remains a stated limitation.
  • [-] P1 · A7 Progress counters are not resumable checkpoints. REGIME_COMPLETE is written before rank/metric calculation and does not persist those results. Do not infer ETA from rows, stale checkpoint elapsed time, or CPU utilization. Less urgent now that G1 completed, but needed before any future rerun.
  • [x] P0 · G1-07 · offline smoke verified 2026-09-20 Ran scripts/run_g1_07_offline_smoke.py in persistent tmux session g1-07-smoke against a four-molecule/15-spectrum fixture with renamed and reversed IDs (g107_3…g107_0) and network socket construction disabled. Inference produced 4/4 rows with 100% coverage; strict local validation passed with zero errors. Evidence: reports/g1-07-offline-smoke-20260920/evidence.json. Production artifacts were not modified. This is an offline packaging smoke test, not the two final clean-room notebook reruns required by G6.
  • [x] P0 · G1-08 · decision written, gate not closed Independent decision is reports/g1-audit-20260919-1005/audit-decision.json: LOCAL_DIAGNOSTIC_PASS_FULL_G1_ACCEPTANCE_OPEN. Saved arithmetic/output pass; provenance permutation guard, malformed precursor handling, legacy metrics, packaging and C remain open. Completing this decision task does not close G1.

  • Gate G1: [!] OPEN — execution and evidence audit in progress. A successful process exit, 400/400 placeholder rows, single-spectrum metrics or B=0 by construction cannot close it. Regime-B zero target coverage for a pure reference-spectrum matcher is structurally expected after identity-wide exclusion: it demonstrates the method's unseen-identity limitation, not structure-only identification quality. Keep candidate coverage, reference availability and ranking quality as separate measurements. Do not begin learned-model training before qualified local acceptance.


    Controlled retrieval ablations — after the corrected baseline is locked


    These are follow-on comparisons, not an excuse to withhold a correctly audited baseline indefinitely. Use the same frozen cohorts/masks and bounded development experiments; promote only with reproducible molecule-level evidence.


  • [ ] P1 Compare max across permitted spectra against mean/top-k evidence, within-molecule reciprocal-rank fusion, collision-energy-aware aggregation and same-adduct/same-polarity policies. Include a matched single-spectrum control, reference-count bias checks and correlated-acquisition controls. Do not merge incompatible ions into one raw peak list. Cross-route ensemble weights remain Phase 5 work.
  • [ ] P1 Add source, adduct, polarity and collision-energy compatibility features, including missingness and original energy units; validate before adopting weights.
  • [ ] P1 Compare mass windows 8.5 / 10 / 20 / 30 ppm, with a calibrated absolute floor; candidate caps none / 100 / 1,000 / 6,000; relative intensity floors 0.002 / 0.005 / 0.01. Retain the original plan's preprocessing variants as further hypotheses, not an obligatory full Cartesian sweep. An uncapped experiment must still stream within the RAM budget.
  • [ ] P1 For every cap/window/preprocessing setting, record candidate rows and unique identities before/after truncation, correct-candidate retention, full coverage, conditional MRR, MRR@25, Top-1/Top-5, runtime and peak memory. Tune on development folds and evaluate the untouched report partition with paired molecule bootstraps.
  • [ ] P1 Profile search overhead and reuse a precomputed sorted mass array rather than repeatedly materializing masses[order] inside every query. Prove candidate membership and ranking equivalence before accepting this optimization; do not alter a running evaluator in place.
  • [x] P1 · design only Record an optimization path that preserves the current evaluator as an oracle, separates query extraction from scoring, validates mass-bin/vectorized candidates against a bounded fixture, and checkpoints at candidate-chunk boundaries. Design: reports/todo-review-20260918/G1-OPTIMIZATION-DESIGN.md. No optimized implementation or full-data rerun is claimed complete.
  • [ ] P2 · optional, unbenchmarked Evaluate a separate Kaggle-GPU implementation only after CPU oracle/fixture equivalence is established, including the existing tolerance-based greedy peak matching and deterministic ranks. No GPU notebook, GPU timings or measured speedup exists; previous numerical speedup/ETA estimates are hypotheses, not planning evidence.

  • Phase 2 — Learned fingerprint ranking and domain adaptation


    Entry dependency: qualified local G1 plus a frozen Phase 3A/3B candidate pool and non-neural baseline. Do not train against an empty B target pool and call the result a ranking experiment.


  • [ ] P0 Generate out-of-fold candidate features without molecule leakage. Under B masks, remove held-out spectra from supervised training, retrieval, analogue prototypes and all derived features; A/C must follow their own frozen contracts. Candidate-generation models must also be out of fold.
  • [ ] P0 Establish a bounded spectrum-to-fingerprint baseline: spectrum -> fingerprint logits/embedding -> candidate fingerprint similarity -> molecule-level ranking. Start with a documented 2,048-bit Morgan target as a hypothesis; compare single-spectrum and permitted multi-spectrum pooling, likelihood/cosine scoring and calibration without copying public constants.
  • [ ] P0 Train a first molecule-grouped LightGBM/CatBoost ranker on eligible OOF features; benchmark simple fingerprint transfer before a larger encoder. Keep model versions, targets, folds, seeds and compute budgets reproducible.
  • [ ] P0 Compare learned ranking against the locked heuristic retrieval/database baselines on the same candidate pools; separate coverage gains from reranking gains. Use exact scorer identity and MRR@25 rather than fingerprint similarity as the promotion criterion.
  • [ ] P1 Balance or reweight training sources to reduce synthetic-library dominance.
  • [ ] P1 Add reference-count bias correction.
  • [ ] P1 Add confidence margins and uncertainty diagnostics.
  • [ ] P1 Run ablations for fragment, entropy, neutral-loss and multi-spectrum evidence.

  • Gate G2: learned ranking improves grouped validation with paired molecule-level evidence, no unexplained severe regression in any intended regime, and acceptable measured resources. Random-row/library-only splits and public leaderboard claims do not satisfy the gate.


    Phase 3 — Known structures without reference spectra


    Phase 3A — Pinned structure-only universe (promoted; first after G1)


  • [x] P1 · specification only Structure-index acceptance specification exists at manifests/phase3a-structure-index-spec-20260918.json. B eligibility text is corrected; this does not certify all sources or satisfy acceptance.
  • [x] P0 · specification corrected The draft retains permitted held-out database structures for B while excluding their spectra/features; C structural/scaffold exclusions remain separate. Runtime leakage certification remains a distinct task.
  • [ ] P0 Acquire and document an eligible COCONUT snapshot: source/version/date, download hash, license, public accessibility and offline asset path. Normalize/deduplicate with the pinned scorer identity and count rejected structures. Keep structures explicitly separate from spectral references.
  • [x] P0 · Phase 3A rederived structure index · verified 2026-09-20 Built artifacts/phase3a-rederived-20260920/pubchem-structure-index.sqlite from the consolidated Phase 3A source. Terminal manifest status is COMPLETE; 444,792 source rows processed, 424,291 valid/unique scorer identities, 0 invalid SMILES, 0 missing identities and 0 identity mismatches. The 131 MB SQLite artifact and progress manifest are present; no active builder process remains. This completes the bounded structure index build, not source provenance/licensing certification or production promotion.
  • [x] P0 · Phase 3A rederived pool consolidation · verified 2026-09-20 Consolidated the completed chunk outputs into manifests/rederived-20260919/pubchem-rederived-consolidated.tsv: 444,792 valid rows, 444,792 unique CIDs, 424,291 scorer identities, 19,801 training-overlap rows, 424,991 structure-only rows and 179,581 novelty-proxy rows. Independently checked schema and SMILES parseability; no duplicate CIDs. Preserved 24 chunk-reported rejected records in pubchem-rederived-exceptions.jsonl. Manifest: reports/phase3a-rederived-20260919/consolidated-manifest.json. This is a bounded consolidated structure pool, not yet a mass/formula index or promoted production source.
  • [x] P1 · format/count audit only Record the local PubChem text file's physical format, hash, stored row/identity-field counts and malformed-row count in reports/todo-review-20260918/pubchem-local-audit.json. This did not recompute chemical identities, training-overlap or novelty labels against the rebuilt namespace, and did not physically relocate the file.
  • [x] P1 · identity namespace reconciliation Compared source-declared and index-recomputed identity sets read-only: 240,536 source rows; 231,755 declared identities; 231,749 indexed/recomputed identities; 578 declared-only and 572 computed-only identities, net difference 6. The index correctly follows recomputed scorer identity14; the net six is not six catalogue misses. Full sets and hashes: reports/phase3a-pubchem-identity-reconciliation-20260919.json and .md.
  • [-] P0 · source normalization/provenance The Phase 3A bounded PubChem-derived structure index is now built and terminally verified, but source normalization/provenance and licensing acceptance remain open. Identity-overlap labels match (12,331), but 9,594 scaffold-string differences include 1,741 changed scaffold-overlap memberships; prior zero-change classification used the wrong variable. Production has seven invalid representatives (six isotope cases plus phosphorus CID 14858); the separate shadow repairs their SMILES only. Keep scaffolds quarantined for C/novelty. See reports/review-20260919-1817/AUDIT.md.
  • [ ] P1 Audit and, if storage/licensing permit, add a broader pinned PubChem-derived structure source. The existing bounded PubChem archive used for validation cohort construction is not evidence of a complete production database.

  • September 19 audit repair gates (before promotion)


  • [x] P0 · S1 Corrected the scaffold classifier to compare scaffold membership separately from identity overlap and string serialization. Read-only rerun: 222,148 exact scaffold/membership rows, 7,853 string-only differences, 1,741 scaffold-membership changes, 0 identity-overlap changes, and 7 invalid production representatives. Added executable regression coverage; novelty metadata remains quarantined. Evidence: reports/phase3a-pubchem-scaffold-discrepancy-20260919.json and reports/review-20260919-1817/AUDIT.md.
  • [x] P0 · S2 · guarded repair verified 2026-09-20 Guarded repair path was executed against the exact base artifact and manifest; the seven explicit repairs passed parseability/identity checks, row-count and no-overwrite guards, and produced artifacts/phase3a/pubchem-structure-index-repaired.sqlite with PASS evidence in artifacts/phase3a/pubchem-structure-index-repaired.repair-verification.json. This closes the repair execution task only; full source provenance/licensing and promotion remain open.
  • [-] P0 · S3 Cache-aligned analogue output persists candidate IDs, ranked IDs, reference IDs, rank, route, source row, adduct, precursor and fallback state for 1,440 policy/query records. Same-query production/shadow comparison and fingerprint availability comparison both pass: 11,106/11,106 candidate identities have fingerprints; 420/2,095 reference identities do; zero availability changes occur between artifacts. Remaining: upstream provenance/licensing before promotion.
  • [-] P0 · S4 Metric contract and cache-aligned 500-identity/480-query rerun completed: enforced top-1000 ranking, true MRR@25, separate route/ranked denominators, mutually exclusive reference/pool failure categories, and operative P2 candidate/reference fallback fields. The rerun selects identities from identity-cache.sqlite via normalized_smiles, never raw Parquet inchikey14. Cache-aligned P0/P2 combined MRR@25 are 0.3003/0.3024 and Recall@25 81.67%/82.08%. Identity audit still finds 1,559 registered-derived rows across 44 stored→derived mappings and 164 identities overlapping prior fold 4. Do not label the partition untouched.

  • Phase 3B — Formula/mass-gated candidates and non-neural baseline


  • [ ] P0 Evaluate formula inference and neutral-mass consistency, retaining alternative formulas when uncertain. Measure true-formula recall and answer retention before tightening gates; report both naturally database-covered molecules and the full B cohort.
  • [x] P0 · reference availability audit Separate PubChem structure coverage from local spectral-reference availability on the post-freeze bounded readout. P0/P1/P2 structure-pool recall was 94.11% on the selected 500-query stability sample; 95.58% had any local candidate reference and 58.11%/58.53% had any fold-eligible candidate reference. All targets had full-library references but 0% had fold-eligible target references by identity-wide exclusion. Evidence: reports/phase3a-reference-availability-20260919.json and .md. This is a post-freeze stability diagnostic, not prospective validation.
  • [-] P0 Bounded reference-availability evidence exists (preceding completed diagnostic task); extend to every frozen pool with candidate/reference-row accounting. Correct route labels and preserve unavailable-reference versus wrong-ranking failures.
  • [ ] P0 Retain a bounded library-ceiling spectral reranking control with query-row exclusion only. Label it optimistic and do not use it as unseen-identity evidence.
  • [ ] P0 Add an identity-held-out cross-identity spectral diagnostic that measures decoy ordering and candidate/reference coverage, but reports target rank only when an eligible target reference exists. A pure spectral-library matcher has structurally undefined target rank when all target references are excluded.
  • [x] P0 · fallback/category audit Independently verified deterministic fallback ordering with zero equivalence mismatches. On 475 valid audit queries: 447 targets were in the structure pool, 28 were absent; 260 had an eligible spectral route and 215 used fallback. Fallback recovered 90.23% in-pool and Recall@25 88.84%. Oracle formula gating was identical (90.23% pool / 88.84% @25), but is explicitly non-deployable because it uses visible-row formula metadata. Evidence: reports/phase3a-analogue-fallback-audit-20260919.json and .md.
  • [-] P0 Registered readout now has a cache-backed identity/exposure audit: 739/739 registered identities are present; 100 direct identity14 recomputations match; 1,559 registered-derived rows across 44 stored→derived mappings disagree with Parquet inchikey14; 164 identities overlap prior fold 4. Query extraction is cache-aligned, per-query provenance is persisted, and regression-tested, but the report remains diagnostic because prior exposure is unresolved.

  • [-] P0 Superseded by cache-aligned report reports/phase3a-analogue-transfer-preregistered-corrected-cache-aligned-20260919.json: 500 identities, 480 valid queries, 1,440 policy/query evidence records. P0 combined MRR@25 is 0.3003 / Recall@25 81.67%; P2 is 0.3024 / 82.08%. These remain diagnostic because 164 identities overlap prior fold 4 and the partition is not untouched.

  • [-] P0 · provenance Bounded PubChem source audit is explicit in reports/phase3a-pubchem-provenance-audit-20260919.json: local SHA-256 6006275e…91152, 240,536 rows and 231,755 declared identities are reproducible, but source_url, upstream version, download date, license and derivation commit are absent. Do not promote or claim complete PubChem coverage/novelty certification until those fields are independently recovered and verified.

  • Phase 3C — Mass-shifted analogue propagation (separate route)


  • [ ] P1 Retrieve eligible spectral neighbours -> search a mass-shifted structural neighbourhood -> apply neutral-mass/formula consistency -> deduplicate scorer identities -> rank with spectral/fingerprint/mass evidence. Keep this route separate from exact-library retrieval and from de novo analogue editing.
  • [ ] P0 Persist source spectrum/source identity, candidate database snapshot/hash, mass-shift and structural-similarity rules, generation/retention counts and deduplication method. Prove no held-out-identity spectral source enters the route or its derived features.
  • [ ] P1 Ablate analogue propagation against the frozen direct-library + structure-only baseline, on identical folds and pools where appropriate. Select weights locally; do not import public leaderboard-tuned weights or arbitrary candidate quotas.

  • Later representations and forward checking


  • [ ] P1 Test pretrained spectrum embeddings such as DreaMS/MIST after license, training-provenance and offline checks. Qualify contamination risk rather than claiming clean novelty performance when overlap is unknown.
  • [ ] P1 Integrate the trained spectrum-to-fingerprint model from Phase 2 into structure retrieval; do not maintain a second conflicting training task here.
  • [ ] P2 Add a forward-spectrum reranker for shortlisted structures.

  • Gate G3: measurable structure-known/no-spectrum improvement over the non-neural baseline, with leakage-safe spectral inputs, licensed/pinned structure provenance and separate pool coverage/conditional MRR. Database absence and ranking failure must remain distinguishable.


    Phase 4 — Novel structures and de novo generation


    Deferred: no immediate autoregressive SMILES or large-transformer training. First diagnose residual failures after database/analogue/fingerprint work and approve a bounded experiment budget. A failed de novo gate need not prevent shipping a validated retrieval/database solution, but forbids claiming novelty coverage.


    The C-proxy split/leakage audit needed by G1 is validation infrastructure and is not deferred with generation.


  • [ ] P1 Reuse and extend the audited G1 novelty-proxy split with structure, spectral and generator-training-target exclusions before training a generator.
  • [ ] P1 Build a formula-, adduct- and collision-energy-conditioned spectrum encoder.
  • [ ] P1 Implement a bounded SMILES decoder/generator.
  • [ ] P1 Add beam search and controlled sampling with valid-SMILES filtering.
  • [ ] P1 Add de novo analogue editing from eligible close spectral neighbours; distinguish newly generated structures here from database analogue propagation in Phase 3C.
  • [ ] P1 Deduplicate generated outputs by scorer-compatible identity.
  • [ ] P2 Add forward-spectrum and formula consistency reranking.
  • [ ] P2 Measure unique candidate recall and actual MRR contribution, not validity alone.

  • Gate G4: de novo candidates add measurable clean-validation MRR or candidate coverage within the offline runtime budget.


    Phase 5 — Unified ensemble


  • [ ] P0 Union validated retrieval, database/analogue and, only if G4 supports it, de novo candidate pools.
  • [ ] P0 Train a final OOF molecule-level ensemble ranker.
  • [ ] P0 Produce exactly one row per runtime molecule with at most 25 unique candidates.
  • [ ] P1 Calibrate route confidence and preserve exploration for uncertain molecules.
  • [ ] P1 Stress-test rare adducts, high masses, one-spectrum molecules and noisy spectra.
  • [ ] P1 Compare library-heavy, balanced and novelty-heavy validation mixtures.
  • [ ] P1 Perform paired-bootstrap comparisons and record promotion decisions.
  • [ ] P1 Compare cross-route reciprocal-rank/static fusion and calibrated learned fusion only on common frozen folds with common scorer-identity deduplication. Choose weights without report/hidden-test leakage; require paired improvement over the best individual route, not a public notebook's claimed score.

  • Gate G5: robust integrated model beats the retrieval baseline across the selected validation scenarios.


    Phase 6 — Kaggle packaging and final submission


  • [ ] P0 Package model weights, indexes, tokenizers and dependencies for offline notebook execution.
  • [ ] P0 Run with internet disabled.
  • [ ] P0 Test against dynamically loaded test IDs; do not rely on placeholder IDs.
  • [ ] P0 Verify submission.csv columns, nulls, duplicate IDs and maximum candidate count.
  • [ ] P0 Stress-test cold offline execution, unseen/reordered IDs, changed molecule counts, 16+ spectra, rare adducts, high masses, empty/noisy spectra and missing metadata. Derive IDs from runtime Parquet; sample submission is a format check only. Fail clearly on globally corrupted artifacts.
  • [ ] P0 Complete two clean-room notebook reruns within the 9-hour limit.
  • [ ] P1 Prepare two complementary final candidates only if validation supports both.
  • [ ] P1 Record final code, configuration, hashes, licenses and reproduction instructions.
  • [ ] P0 Select final submissions explicitly before the deadline.

  • Gate G6: two successful offline reruns, reproducible output and final compliance audit.


    Active experiment log


    Historical chronology is retained below. Earlier PASS/closure statements and namespace counts were superseded by the 2026-09-17 audit/rebuild; the current gate sections above govern decisions.


    DateExperimentHypothesisResultDecision
    2026-09-16Competition review and workspace setupEstablish a complete execution baselinePlan and project scaffold created; no model runStart Phase 0
    2026-09-16Phase 0 data acquisitionVerify current Kaggle release and characterize it under low RAM3,033,286,496-byte file; SHA-256 recorded; 2,539,608 rows, 21 row groups, 275,810 unique InChIKey14 identities; audit completeProceed to grouped split manifests
    2026-09-16Molecule-grouped splitPrevent spectral/structure leakage in local validationDeterministic 5-fold manifest validated: 275,810 groups and 2,539,608 rows; fold assignment is group-exclusiveUse fold 0 as initial validation; keep novelty labels separate
    2026-09-16G0 data/infrastructure gateConfirm reproducible Phase 0 inputs and assumptionsAll eight gate checks passed; config and split hashes recorded in manifests/phase0-gate.jsonClose data/infrastructure work; proceed to retrieval baseline
    2026-09-16Phase 0 auditVerify that validation grouping uses scorer identityFailed: 12/100 sampled normalized SMILES differed between stored and scorer-derived InChIKey14; full probe hit pathological tautomer enumerationReopen G0; rebuild identity-safe split before retrieval evaluation
    2026-09-16Phase 0 corrective auditVerify corrected scorer-identity split and reproducibility gatePassed: 8 tests; compilation passed; 2,539,608 rows and 274,288 scorer groups validated; cache/checkpoint agree; full G0 gate passed; 100-row identity probe had 0 invalid structuresProceed to retrieval baseline; retain external novelty regimes as pending
    2026-09-16Phase 1 retrieval baselineBuild a low-memory, identity-safe spectral retrieval floorIndex built over all 2,539,608 rows; 2,160,035 valid spectra; fold-0 mask excludes all 503,842 held-out rows; molecule-level fixture submission has 400/400 covered molecules and passes schema/count checksKeep baseline; run grouped MRR/Top-1/Recall evaluation before closing G1
    2026-09-17Qualified rebuilt infrastructureSeparate corrected identity namespace from legacy evidencephase0-qualified-g0-20260917.json records 274,195 identities, 22 passing local tests, A/B guards and source/mask hashesQualified local infrastructure only; preserve official parity/acquisition caveats
    2026-09-17Bounded molecule-level retrievalEstablish a diagnostic floor on the rebuilt namespace5,974-identity library/B cohorts plus 250 A identities; one valid query spectrum per identity; saved report and exit-0 timing logNot all-spectrum G1; audit metric definitions and query selection before promotion
    2026-09-18All-spectrum G1 executionAggregate permitted spectra for each validation moleculeStarted 06:26 MYT; still active with no result artifact at the 08:21 checkpointKeep G1 open; require semantic, provenance, strict-output and resource audit, not just process completion
    2026-09-18Community findings + execution-plan reviewImprove candidate coverage before heavier modelsPrioritized structure-only snapshots/mass-formula candidates, OOF fingerprints and separate analogue propagation; added leakage, adduct, cap and aggregation checksAdopt revised backlog order; no new model training, submission or G1 closure claimed
    2026-09-18Evaluator/validator correctionsRepair the morning audit's concrete defectsA row-mask forwarding, uncapped coverage, covered-query conditional MRR and basic invalid/duplicate/over-25 checks implementedLocal fixes verified; not end-to-end G1 acceptance
    2026-09-18Checkpointed rerun at 12:39Expose progress without waiting for the aggregate reportPID 198260: library scan reached 621 batches / 113,376,230 comparisons; B active; final report absent at evening auditPreserve ongoing work; counters are not resumable state, saved metrics or a reliable ETA
    2026-09-18Kaggle CLI release verificationCheck authoritative training data against a fresh downloadSize/hash/schema/rows/row groups all match; verified staged duplicate permanently deleted on requestRetain original data and comparison report; no rebuild justified by release mismatch
    2026-09-18Non-competing preparationImprove small contracts and document next stepsBasic formate alias tested; PubChem physical-format/count audit and Phase 3A/optimization drafts recordedDrafts/source audit have eligibility/provenance limitations; no new index or GPU runner built
    2026-09-18Evening progress auditReconcile source, tests, claims and TODO24 lightweight tests pass with 2 integration tests excluded; independent metric fixes verified; wrapper failures, entrypoint failure, metadata loss and strict-output gaps reproducedKeep G1 open; prioritize A1-A8, narrow completed claims and retain historical evidence
    2026-09-19Phase 3A fallback/category auditVerify fallback equivalence and separate structure/reference failure classes475 valid audit queries: 28 structure-pool misses; 260 had a fold-eligible local identity in the mass window; 215 used the audit fallback route; fallback pool recall 90.23%, @25 88.84%; 0 mass-order equivalence mismatches; oracle formula gate showed no improvementKeep formula gate non-deployable; do not call the 260 count a positive spectral-neighbour count; preserve categories; add maintained tests and repeat on a genuinely pre-registered untouched partition

    Blockers and decisions


  • Actionable local blockers: G1-01 through G1-08 are not all satisfied. The core A query-mask and all-spectrum metric fixes are verified, and A1-A3 are now fixed locally with isolated tests; outstanding priorities are A4-A5 (source/input binding, per-molecule/stratified/C evidence), A6-A7 (maintained branch coverage and meaningful progress/resume), plus broader adduct/input coverage. A successful current process exit does not satisfy these automatically.
  • Current execution decision: previous runs have ended; no stale PID is a current instruction. Preserve production and historical evidence. The repaired shadow is unpromoted; no full rebuild, learned training, GPU run or external submission was performed by this audit.
  • External qualifications, not reasons to fabricate evidence: exact acquisition lineage remains unavailable; the official scorer source and pinned-runtime fixture certification are recorded in reports/official-scorer-certification-20260920/evidence.json. C is a proxy, not true hidden novelty. The 1,184-versus-1,151 documentation discrepancy is pinned locally and is not itself a remaining implementation blocker.
  • Do not start: learned-model training before qualified local G1; large encoders/de novo before smaller database/fingerprint/analogue baselines justify them. Phase 3A/3B precedes Phase 2 in the revised dependency order. Do not compare the library-ceiling spectral MRR with identity-held-out B as if they were the same task; the latter has no target spectrum for a pure library matcher after identity-wide exclusion.
  • Do not submit: until a regenerated molecule-level output passes strict local checks, caveats are documented, and an offline rebuilt-artifact run succeeds. This TODO review does not authorize a Kaggle submission, paid compute, public redistribution or execution of every newly listed experiment.
  • Completion rule: code present != run completed != independently valid evidence != gate closed. Preserve all four distinctions in status updates and in the goal state.