← Back

📑 Contents

Enveda CASMI 2026 TodoCurrent execution snapshot — 24 September 2026 (MYT)Execution tasks — one authoritative backlogShortlist-to-experiment mapNext execution orderEvidence rules and current snapshotMorning findings — disposition after the evening auditImmediate audit follow-ups — historical local fixes and remaining acceptance workPhase 0 — Ground truth and infrastructurePhase 1 — Spectral retrieval baselineControlled retrieval ablations — after the corrected baseline is lockedPhase 2 — Learned fingerprint ranking and domain adaptationPhase 3 — Known structures without reference spectraPhase 3A — Bounded pinned structure universe (supporting prerequisite, not bulk expansion)September 19 audit repair gates (before promotion)Phase 3B — Formula/mass-gated candidates and non-neural baselinePhase 3C — Mass-shifted analogue propagation (parallel complementary channel)Later representations and forward checkingPhase 4 — Novel structures and de novo generationLater generation experiments — gated, not a submission dependencyPhase 5 — Unified ensemblePhase 6 — Kaggle packaging and final submissionValidation contract for the revised routeActive experiment logBlockers and decisions24 September shortlist replan

Enveda CASMI 2026 Todo


Legend: [ ] pending · [-] in progress · [x] complete · [!] blocked


Reviewed: 2026-09-24 (MYT), against the actual [prvsiyan analog-propagation notebook](https://www.kaggle.com/code/prvsiyan/analog-propagation-casmi-2026-baseline) and the [MS/MS-to-2D-structure arXiv shortlist](research/arxiv-msms-2d-shortlist-20260923.md). New decision/evidence: [plan review](reports/plan-review-20260922-analog/REVIEW.md). Earlier [audit findings](reports/review-20260919-1817/AUDIT.md) and completed experiment records remain historical evidence; they are not erased by this revision.


Current direction: prioritize learned spectrum-to-fingerprint ranking and complementary mass-shifted analogue propagation on a bounded, eligible candidate pool. No open-ended PubChem slice expansion. Retain the already-authorized Phase 4 experiment as an unqualified feasibility branch; do not scale it or reserve fixed output slots for it without measured benefit.


Current execution snapshot — 24 September 2026 (MYT)


  • Pool: competition training structures only. PubChem remains quarantined and its rebuild deferred by CC; no new catalogue ingestion.
  • Reusable evidence: official-scorer fixture certification, qualified G0, AUDIT-01 and BASE-01/02 evidence exist. Preserve their scope; do not restart them as unchecked setup tasks.
  • FP-01: independent kernel exists at kaggle-kernel/fp-01-ranking/main.py; saved implementation report records local compile/schema checks. Training/OOF outputs and model quality remain unverified. See reports/fp-01-implementation-20260924.json.
  • Latest launch evidence: reports/fp-01-launch-check-20260924-0625.json records HTTP 409 and status permission denial; execution is unverified. Prior GPU-quota exhaustion is historical, not reconfirmed by that check.
  • Validation: ISO-01 formula-grouped panel is a development diagnostic; the 164 prior fold-4 overlaps and earlier fold-0 exposure forbid an untouched-report claim. Split membership and model-specific training leakage must still be checked independently of historical exposure labels.
  • Scope of this revision: documentation only; no kernel push, training, submission, asset download, process or automation change.

  • Execution tasks — one authoritative backlog


    The phase checklist below owns scientific acceptance. This queue references those IDs rather than creating a second TOP5 completion ledger. Historical gates retain their original names; task implementation, run completion and promotion are separate.


  • [!] P0 · RUN-01 Resolve exact FP-01 owner/slug, account permissions, remote state and available approved GPU quota. Latest saved result: HTTP 409 / permission denied; no confirmed execution. Retry only on new evidence, not blind hourly pushes. Existing scheduler unchanged by this edit.
  • [-] P0 · EVAL-01 Reuse existing scorer/BASE/AUDIT evidence; bind competition-only pool, B-safe supervised exclusions and diagnostic exposure labels. Freeze nested/grouped selection/report rules and common controls before model promotion. No clean claim from merely reseeding exposed identities.
  • [ ] P0 · SUB-01 After RUN-01 and applicable evaluation/package checks, produce the first evidence-qualified offline candidate from the best validated route. Record source/checkpoint/pool hashes, dynamic-ID validation, measured end-to-end runtime, notebook version and official submission result. Final G6 two-rerun acceptance remains separate; novel generation is not required for a development submission.

  • Shortlist-to-experiment map


    This mapping uses the supplied local shortlist (abstract-level screening), not a fresh full-paper replication or asset licence certification. Before implementing a new paper-specific mechanism, record its version, exact method section and independently written design under LIT-01. Published benchmark numbers are not expected CASMI scores.


    Technical referenceIndependent experimentAcceptance / limit
    Retrieval objectives, arXiv:2602.16507FP-01/02: BCE + f · z versus hard-negative rankingSame pool/cohort; molecule MRR@25 and paired ranks decide, not bit accuracy
    MassSpecGym / In the Wild, 2410.23326 / 2606.19624EVAL-01, VAL-01/02, AUDIT-01Scorer parity, acquisition/identity exclusions, exposure and shortcut audits; no blanket indictment of other papers
    JESTR, 2411.14464JESTR-01: joint spectrum/structure embeddings from scratchSame frozen pools/folds and comparable budget versus FP-01
    GLACIER, 2606.29161RERANK-01: forward-spectrum agreementRetained-answer recall, no-reranker control, OOF features and measured latency
    MS-BART, 2510.20615GEN-03A: structure/fingerprint decoding from scratch, adapted to OOF predicted fingerprintsNo MIST weights; structure targets follow B/C exclusions; compare noisy versus perfect/oracle conditioning
    MARLIN, 2607.04774GEN-03B: fingerprint + mass, formula-free generation conceptBounded challenger; no downloaded weights; measure exact new identities and runtime
    DiffMS, 2502.09571GEN-03C: predicted-formula-constrained graph generationGround-truth formula is oracle-only; count formula misses in full-cohort score
    MADGEN, 2501.01950GEN-03D: retrieved scaffold and bounded analogue editingPredicted scaffold, never held-out truth; separate retrieval failure from generation failure
    MSFlow, 2602.19912Excluded assets; no scheduled experimentShortlist describes non-commercial trained version. Withdrawal is not established by this source; remove the earlier unsupported withdrawal claim

    Evidence status: qualified local G0 and later G1 diagnostic audit records exist; full G1/G4 promotion is not established by them. The September 20 five-suite/45-test result is historical, not a test run performed by this review. A seven-row repair shadow and larger rebuilt structure artifacts are different objects. This revision changes documentation only: no training, notebook execution/push, new dataset ingestion, production promotion or automation changes.


    Binding implementation policy: papers/notebooks are technical references only; independently implement adopted concepts. Do not copy their implementations or download/use external pretrained weights, checkpoints, tokenizers or training archives—even if their licences permit it. Commercially compatible general-purpose libraries remain allowed after review. Our own checkpoints trained from scratch on authorized inputs may be saved, transferred from Kaggle and packaged. Competition data are used under the competition-specific permission; this does not make the dataset commercially licensed. External data/libraries/tools require documented compatible rights and provenance. PubChem rebuilding and COCONUT acquisition are deferred, not prerequisites to the competition-only route.


    Historical material corrections: the scaffold classifier compared identity overlap in its category logic: 1,741 scaffold-overlap changes, not zero (1,295 lost / 446 gained; 7,853 other scaffold-string differences). The registered readout is not certified untouched: 164 selected identities belong to the previously inspected fold 4. Earlier analogue combined_mrr was untruncated reciprocal rank; corrected S4 evidence below supersedes that metric but not the exposure qualification. Keep those Phase 3A readouts diagnostic-only.


    Checklist interpretation: task states include dated historical evidence; gate states are separate and are not an effort-weighted competition-readiness percentage. September 22 additions remain pending unless explicitly identified as this completed documentation review.


    Historical evening audit: 2026-09-18 source + isolated synthetic probes + saved artifacts. Evidence and pre-audit backup: reports/audit-20260918-1923/. That review left the then-active evaluator and its inputs untouched. Earlier completion claims are qualified below; its process status is not current.


    Audit report: [AUDIT.md](reports/audit-20260918-1923/AUDIT.md). Fresh lightweight suite: 24 passed, 2 real-index tests deliberately deselected; additional independent synthetic evaluator/entrypoint/failure-injection evidence is saved alongside it. The earlier 26-pass result is not a claim that all gate scenarios are covered.


    The priority order is deliberate: establish trustworthy evaluation and a bounded high-recall pool, then measure learned ranking and complementary evidence. Community performance is author-reported, not our result. This review updates the backlog; it does not mark unrun experiments complete.


    Next execution order


    1. Recover execution access (RUN-01). Inspect the exact FP-01 owner/slug, account access and remote state before any retry. The latest saved check is HTTP 409 / permission denied, not renewed proof of quota exhaustion. Earlier attempts did confirm the weekly GPU quota. Do not blindly push hourly; resume when access/quota evidence changes. No paid compute or local training is implied.

    2. Finish the bounded comparison contract (EVAL-01). Reuse the certified scorer and corrected baseline; bind the competition-only pool, identity-wide training exclusions, candidate membership and predeclared development cohorts. Existing exposed panels support diagnostics, not untouched claims. Freeze a nested/grouped selection/report protocol with prior exposure disclosed before promotion.

    3. Run the existing fingerprint kernel (FP-01/02/03). Start with a bounded Kaggle GPU sanity run, then BCE-only versus BCE plus same-mass negatives on identical folds/pools; add same-formula negatives and multi-spectrum variants one change at a time. Evaluate exact molecule MRR@25, not fingerprint accuracy alone. Persist logits, candidate scores/IDs, ranks, checkpoint and run provenance; code presence is not execution evidence.

    4. Package the first qualified candidate (SUB-01, PKG-01). Compare against matched retrieval/transfer controls, select using local evidence, verify offline dynamic-ID inference and strict output validation, then submit an evidence-qualified version under existing execution authority. Do not wait for every research branch or the final two clean-room finalist runs before a bounded development submission. Record remote notebook version and official score; a valid submission is not proof of improvement.

    5. Challenge the baseline (JESTR-01, AN-01/02, FUSE-01). Independently implement joint spectrum/structure contrastive ranking after the FP comparison; measure analogue transfer as complementary evidence on the same folds/pool. Compare single routes, static fusion and OOF learned fusion. Retain only attributable improvements.

    6. Bounded forward reranking (RERANK-01). Use GLACIER as a conceptual reference for an independently trained predictor, starting with cheap fragment-explanation features. Only proceed when failure analysis shows ranking, rather than absent candidates, is limiting performance. Measure shortlist retention and whole-pipeline runtime; no fixed top-K guarantees nine hours.

    7. Conditional generation (GEN-01/02/03/04). After useful fingerprint ranking, test predicted/noisy-fingerprint-conditioned decoding, then bounded scaffold editing and formula-free/formula-conditioned challengers. Require deployable predicted formulas/scaffolds, full-cohort failure accounting and marginal gains after shared reranking. No fixed generation quotas. Sampling/diffusion are optional experiments, not prerequisites to a leaderboard submission.


    Phase numbers remain historical identifiers, not a mandatory execution sequence.


    Interpretation guardrails: the author's 0.151/0.93 is an assumption-dependent estimate on the public test, not a known 16%/84% Class-1/Class-2 split. The remainder includes Class 3. His 99.6% COCONUT coverage is for a particular NP sample, not all natural products; 100% in-pool validation recall cannot establish hidden-test pool coverage. Reported 0.517/0.52/0.63 scores use different controls/cohorts from our library-ceiling result and have not been reproduced here. External DreaMS weights/representations are excluded, and no such result is established for our system.


    Evidence rules and current snapshot


  • Placeholder warning: downloadable test.parquet contains training examples. Its 400-molecule/1,213-spectrum output is a schema/runtime diagnostic only, not CV, public-leaderboard evidence or a G1 quality score. Production IDs and counts must come from the runtime test file, never be hardcoded from the sample.
  • Authoritative namespace: use artifacts/rebuild-20260917/ and manifests/rebuild-20260917/ together. The root retrieval symlink currently resolves there. Historical 274,288-identity and raw 275,810-identity artifacts are not interchangeable with the rebuilt 274,195 scorer identities.
  • Regimes are different questions: A deliberately permits same-identity references only under the documented acquisition-exclusion proxy; B must exclude all held-out-identity spectra across every library, analogue prototype and derived spectral feature. A library-only holdout must never be presented as clean B/novelty validation. C additionally excludes structures/structure targets and requires a frozen scaffold-aware proxy.
  • Current data count: use the hashed release's 1,184 NP-example rows / 250 identities. The published 1,151-row count is a documented external discrepancy, not a reason to trim the local cohort or block its evaluation.
  • Infrastructure evidence: reports/phase0-qualified-g0-20260917.json records local guards/tests and hashes. Official scorer source acquisition and pinned-runtime fixture certification are now complete; acquisition-lineage certification remains unavailable.
  • Earlier G1 evidence: reports/phase1-g1-molecule-full-20260917.json uses one valid spectrum per identity, not all-spectrum validation. Preserve it as a bounded diagnostic; audit metric semantics before comparisons. Library-ceiling scores are optimistic diagnostics, not held-out generalization estimates.
  • Completed G1, not an active run: reports/g1-audit-20260919-1005/AUDIT.md reconciles the later 20:02 launch, saved ranks and 400-row validated placeholder submission. Evaluation took 9h04m20s / ~4.90 GiB peak RSS; placeholder inference 1m49.03s / ~1.44 GiB. Saved arithmetic passed; this is not immutable launch provenance or a full hidden-workload timing test. No matching evaluator/build/pytest process was active at this review's initial inspection.
  • Training release rechecked: the CLI redownload matched size, SHA-256, schema, 2,539,608 rows and 21 row groups. Evidence: reports/todo-review-20260918/kaggle-redownload-comparison-20260918.json. The verified staged duplicate was permanently deleted at the user's request; authoritative training data and the comparison report were retained. This proves equality to the download checked on September 18, not perpetual freshness.
  • Existing rebuilt submission: submissions/phase1-retrieval-molecule-rebuilt-20260917.csv exists and its generation report records 400 rows and 1,213 spectra. A local schema check passed earlier in this session; this does not establish strict candidate validation, hidden-test accuracy or current-run provenance. The legacy submissions/phase1-retrieval.csv remains unsuitable as current evidence.
  • Source packaging partly repaired: correctly named manifests/pubchem-structure-only-1-500000.tsv and source manifest exist; original misnamed .tsv.gz is preserved. Source hash is 6006275eb9fde9101ec1f342eacb1240fad10201c3e8182eaf584cf31ea91152; 240,536 rows / 231,755 declared identities. Verifier checks suffix/magic, but the builder still accepts misleading plain-TSV suffixes and maintained format-rejection tests are missing. Upstream snapshot/licensing remains uncertified.

  • Morning findings — disposition after the evening audit


    The morning evidence below describes the pre-fix implementation, not current source. Preserve [semantic-probes.json](reports/todo-review-20260918/semantic-probes.json) and its source hashes as historical evidence; do not rerun its old defect-expecting script over that report. Fresh evening probes are in reports/audit-20260918-1923/evaluator-edge-probes.json and submission-wrapper-probes.json.


  • Fixed in the all-spectrum evaluator: A now receives the frozen row query mask. A fresh fixture verifies that the eligible reference is not ingested as a query and that no query/reference self-comparison occurs.
  • Fixed in the all-spectrum evaluator: conditional MRR divides by all pool-covered molecules; pool coverage is computed before top-k truncation. Fresh fixtures cover ranks 1/25/26/100/1,000/1,001/absent, an explicit tie and an empty pool, including the coverage × conditional-MRR identity. This does not validate every legacy evaluator or full-run result.
  • Historical validator defects, superseded by A3 and G1-04: raw-field/token/ID checks and null-precursor coercion were subsequently fixed and locally tested. Explicit launch-time validator/path/hash binding and any additional untested malformed-input cases remain separate acceptance work; do not reopen the already-tested null-precursor fix.

  • Historical consequence and later disposition: that run ended and wrapper fixes were verified. Later G1-05/06 records add candidate evidence and post-run binding, while G1-04/07 record local malformed-input/packaging acceptance. Full independent evidence/mask certification, immutable launch provenance, remaining stress cases and C remain open; those later completions do not close full G1.


    Immediate audit follow-ups — historical local fixes and remaining acceptance work


  • [x] P0 · A1 · locally fixed and failure-injection verified Wrapper now uses set -Eeuo pipefail, stage labels and an ERR trap; failed evaluator/submission/validation stages stop before later DONE markers. The exact copied-wrapper probe now returns 41/42/43 respectively with no false success. The active long-running invocation was not restarted and therefore still used the pre-fix shell.
  • [x] P0 · A2 · locally fixed and cold-invocation verified Runtime validator inserts the repository root for clean direct-script execution and writes phase1-retrieval-rebuilt-20260918.validation.json, preserving generation metadata in the original .json. Synthetic clean-environment and metadata-preservation probes pass. The active invocation's already-running process is not retroactively changed.
  • [x] P0 · A3 · local contract fixed and probed Validator now rejects missing fields, null/blank runtime IDs before string coercion, extra CSV fields and interior empty candidate tokens; an entirely blank smiles field is the explicit no-candidate policy. Added contract tests and clean helper probes. This is local validation, not official scorer parity.
  • [x] P0 · A4 · post-hoc provenance snapshot verified 2026-09-20 Created manifests/a4-provenance-snapshot-20260920.json with SHA-256 hashes for current Python source and frozen inputs/artifacts, runtime/platform metadata, dependency freeze, Git state, and explicit missing/limitation fields. This is a reproducibility snapshot for future launch discipline, not immutable launch-time provenance for earlier runs; G0/G1 remain open.
  • [x] P0 · official scorer source and pinned-runtime certification · verified 2026-09-20 Replaced src/scorer_contract.py with the exact Kaggle metric/casmi-mean-reciprocal-rank source (SHA-256 b542cd945203cf4305a0a01537a5fa9e93fa4e2d79400b8563286bce0f73ed2d) and ran a deterministic fixture under RDKit 2026.03.3. Expected and actual MRR@25 were both 0.625; rank-26 rejection and invalid-solution host-error probes passed. Evidence: reports/official-scorer-certification-20260920/evidence.json. This certifies local execution parity for the fixture, not hidden-test accuracy or data acquisition lineage.
  • [-] P0 · A5 · partial evidence package verified 2026-09-20 Packaged the historical per-molecule G1 evidence. Later G1-05 candidate-level evidence and AUDIT-01 exact identity/mask validation supersede the missing-evidence limitation for the saved rerun, but A5 remains a historical partial package and does not establish launch-time binding or Regime C.
  • [x] P1 · A6 · maintained regression coverage verified 2026-09-20 The focused G1/Phase 3A regression suite ran successfully in .venv-rdkit2026: 26 passed (tests/test_phase3a_regressions.py and tests/test_g1_contract.py). This confirms the maintained local contract tests; it does not close full G1 acceptance.
  • A7 reconciliation: the current all-spectrum evaluator now writes its terminal checkpoint after metrics/evidence; do not list the historical ordering defect as still unfixed. Genuine resumable scoring state and reliable progress acceptance remain tracked once under Phase 1 A7 below.
  • [x] P1 · A8 · draft eligibility corrected Phase 3A explicitly retains permitted database structures for B while excluding their spectra/features; C exclusions remain separate. GPU speedup is still unbenchmarked and optimization remains design-only. The historical instruction to preserve the then-active run is not a claim that it is still running.

  • Phase 0 — Ground truth and infrastructure


  • [x] P0 Download the current train.parquet release from Kaggle.
  • [x] P0 · 2026-09-18 recheck Redownload via Kaggle CLI into staging and compare SHA-256, size, schema, rows and row groups. The saved comparison matches in every field. Permanently remove only that verified duplicate when requested; retain authoritative data and comparison evidence.
  • [x] P0 Record file size, SHA-256, Kaggle metadata and download timestamp in manifests/.
  • [x] P0 Inspect the actual training schema, row groups, library counts and identity counts.
  • [x] P0 Add a low-RAM data access path using PyArrow row-group streaming and selected columns.
  • [x] P0 Create a pinned RDKit 2026.03.3 environment for scorer-compatible identity matching.
  • [x] P0 Implement tautomer-canonicalized InChIKey14 identity logic in src/identity.py.
  • [x] P0 Add scorer unit tests for tautomer equivalence, stereochemistry, invalid SMILES and rank cutoff.
  • [x] P0 Implement all ten advertised adduct mass conversions and round-trip tests.
  • [x] P0 · basic alias fix Normalize [M+FA-H]- to [M+CH2O2-H]- in both conversion directions; retain independent formate-mass and round-trip checks. Current tests verify alias use in neutral-mass conversion, reject a non-string adduct and reject the wrong-polarity FA spelling. This is not exhaustive validation of all numeric/polarity inputs.
  • [x] P1 · symmetric adduct/input edge tests · verified 2026-09-20 Added positive finite-mass policy checks for every advertised adduct, canonical/alias formate parity, non-numeric/non-finite/zero/negative precursor and neutral inputs, and reran the focused pinned-runtime suite: 48 passed. Evaluator-level malformed-row handling remains separately covered by G1-04.
  • [-] P0 Freeze molecule-grouped validation manifests for library, database-only and novelty-proxy regimes. The rebuilt scorer-compatible five-fold and strict external manifests are internally consistent; qualified Regime A query/reference masks are now operational and runtime-validated, while Regime B identity-wide exclusion manifests and runtime checks now pass locally, and Regime C remains proxy-only.
  • [x] P0 · cross-library leakage fixture · verified 2026-09-20 Added an identity-wide Regime-B exclusion helper and synthetic regression fixtures proving a held-out identity is excluded from every library, including references, spectral prototypes and derived features; a library-only exclusion is explicitly rejected. Focused suite: 25 passed; maintained six-file suite: 52 passed. This verifies the local mask contract, not external acquisition-lineage certification.
  • [ ] P0 Freeze development/report partitions, seeds and tuning policy before community-inspired sweeps; protect the 250-compound NP report set from repeated tuning with grouped/nested validation.
  • [-] P0 · VAL-01 · diagnostic audit complete 2026-09-23 Authoritative cache alignment and bounded scorer-identity recomputation passed: 739/739 registered identities are in the cache/training-derived records and 832/832 direct identity14 recomputations matched. Saved evidence: reports/val-01-02-exposure-20260923.json. Raw Parquet stored InChIKey14 differs from the scorer-derived namespace for many training rows; do not treat raw-hash parity as scorer-identity parity. Full cross-artifact VAL-01 binding remains open.
  • [-] P0 · VAL-02 · exposure ledger diagnostic complete 2026-09-23 The registered prospective partition contains 164 identities overlapping prior fold 4; it is not an untouched report set. Phase 4 also inspected/trained against raw-key fold 0, so the prior blanket fold-0-untouched claim is withdrawn. Saved evidence: reports/val-01-02-exposure-20260923.json. Next requirement is a genuinely unexposed report partition or nested/grouped evaluation with explicit prior-exposure labels; changing the seed does not erase exposure.
  • [x] P0 · ISO-01 · training diagnostic panel built 2026-09-23 Built manifests/iso-01-training-diagnostic-panel-20260923.json from data/train.parquet using the authoritative identity cache and deterministic formula grouping, without model outputs: 24,858 same-formula groups and 247,777 scorer identities. All 247,777 panel identities are training-exposed, so this is a development diagnostic only; it is not a clean report cohort and cannot close ISO-01. The original selection-free specification remains at manifests/iso-01-panel-spec-20260923.json.
  • [x] P1 Add a reproducible experiment configuration format under configs/.

  • Resource constraint: the local development host has approximately 11 GiB RAM. Full-data operations must stream or use bounded artifacts; do not load the entire training table into pandas.


    Gate G0: [x] Qualified local infrastructure pass — verified 2026-09-20. The pinned runtime, official scorer, 15-file artifact/source binding, fail-closed runtime/hash preflight, deliberate tamper failure, and 48 maintained tests pass. Evidence: reports/g0-local-closure-20260920/decision.json and reports/g0-local-closure-20260920/manifest.json. This closes qualified local G0 only; acquisition-lineage certification remains externally unavailable, Regime A is an exact-available-field proxy, Regime C is proxy-only, and this is not hidden-test or competition-performance certification.


    Phase 1 — Spectral retrieval baseline


  • [x] P0 Build a streaming Parquet reader and compact peak-array representation.
  • [x] P0 Implement precursor/adduct to neutral-mass filtering.
  • [x] P0 Implement peak cleaning variants: intensity floor, precursor exclusion and top-N/windowed peaks.
  • [x] P0 Implement fragment cosine similarity.
  • [x] P0 Implement spectral entropy similarity.
  • [x] P0 Implement neutral-loss similarity.
  • [-] P0 Build and verify indexed reference search. Rebuilt index metadata and the qualified infrastructure report record 2,539,608 rows, 2,160,035 valid spectra and 274,195 scorer identities, with five fold masks. Preserve the legacy index under artifacts/legacy-20260917/; complete end-to-end artifact binding rather than trusting a symlink or status label alone.
  • [x] P0 · local fix verified Pass the A row query mask into the all-spectrum evaluator. Independent evening synthetic evidence confirms only the permitted query is used against eligible references; full cohort accounting and aggregation acceptance remain G1-01 work.
  • [-] P0 · G1-01 Complete correct all-spectrum molecule-level aggregation on frozen query cohorts. For A, use the actual regime_a_query.npy row mask, not every row sharing one of its 250 identities. Assert eligible query/reference separation and reconcile total, valid, skipped and zero-valid-spectrum molecules. For B/C, retain identity-wide exclusions. Aggregate all permitted spectra within each evaluation molecule; produce exactly one output per runtime molecule_id.
  • [x] P0 · G1-02 · local math verified Independently probe the corrected all-spectrum evaluator at its default top-k=1,000: target ranks 1/25/26/100/1,000/1,001/absent, a deterministic equal-score tie and an empty pool. Full pool coverage is distinct from Recall@1,000; conditional MRR divides by all covered targets; MRR@25 = pool_coverage × conditional_MRR@25 holds, with zero returned for empty coverage. Evidence: reports/audit-20260918-1923/evaluator-edge-probes.json. A6 records subsequent maintained regression coverage; these local checks do not certify legacy evaluators or full-run evidence. Add only demonstrably missing cases under AUDIT-01 rather than repeating the completed A6 task.
  • [x] P0 · G1-03 · bounded identity audit passed 2026-09-20; full gate caveat retained The 400-row submission audit passed: 9,364 raw candidates, 9,364 unique scorer identities, 0 invalid SMILES, 0 duplicate identity rows, 1,001 sampled mappings with 0 stale mappings, and identity probes passed. Later candidate/source-row evidence and AUDIT-01 passed for the saved all-spectrum run. Evidence: reports/g1-03-identity-dedup-20260920/evidence.json and reports/audit-01-exact-identity-mask-20260923/evidence.json. This does not close full G1, launch-time immutability or official hidden-test parity.
  • [x] P0 · G1-04 · expanded offline packaging acceptance verified 2026-09-20 Expanded local fixture used six renamed/reversed runtime molecule IDs across 24 spectra, including empty peaks, null precursor, unsupported adduct and multiple spectra. Network socket construction was disabled; inference produced 6/6 rows with 100% coverage and strict validation passed. Four negative validator probes rejected duplicate IDs, interior empty candidates, >25 candidates and extra columns. Evidence: reports/g1-04-offline-acceptance-20260920/evidence.json. Fixed null precursor coercion so malformed precursor values are skipped instead of crashing. This closes local packaging/validator acceptance, not hidden-test accuracy, official scorer parity or full G1.
  • [x] P0 · G1-05 · post-run audit verified 2026-09-21 Completed all-spectrum G1 rerun (12h03m / 43,232s) produced complete ranked candidate IDs, scores, source reference rows and query source rows for all three regimes: library_ceiling (212,831 evidence rows), regime_b (204,513), regime_a (7,138). Independent recomputation from evidence reproduced every reported metric with zero drift: library ceiling MRR@25 0.8624 / coverage 90.37%, Regime B 0.0 (0 target matches in this reference-only B route; later AUDIT-01 separately verified saved identity/mask joins), Regime A MRR@25 0.9297 / coverage 100%. Per-molecule ranks: 0 mismatches. Candidate identity dedup: 0 duplicates. Source rows: all within bounds. Evidence: reports/g1-post-run-audit-20260921/evidence.json.
  • [x] P0 · G1-06 · provenance snapshot verified 2026-09-21 Created manifests/g1-06-provenance-snapshot-20260921.json binding 21 files (7 source, 10 input, 4 output) with SHA-256 hashes, plus /usr/bin/time -v stats (peak RSS 5,206,524 KiB ≈ 4.96 GB, user 34,381s), runtime metadata (Python 3.10, RDKit 2026.03.3, aarch64), git state (head 8609672198), pip freeze, and 5 explicit limitations. Self-hash included for tamper detection. This is a post-run binding, not launch-time capture; launch-time immutability remains a stated limitation.
  • [-] P1 · A7 Terminal metrics/evidence now precede the completion checkpoint in src/g1_molecule_eval_all.py; this ordering fix is present. Complete real resumable per-query/chunk score state, safe recovery and stage-aware progress separately. Counters and a metrics-bearing terminal JSON are not resume state. Validate recovery before a future rerun; do not infer ETA from stale totals or CPU utilization.
  • [x] P0 · G1-07 · offline smoke verified 2026-09-20 Ran scripts/run_g1_07_offline_smoke.py in persistent tmux session g1-07-smoke against a four-molecule/15-spectrum fixture with renamed and reversed IDs (g107_3…g107_0) and network socket construction disabled. Inference produced 4/4 rows with 100% coverage; strict local validation passed with zero errors. Evidence: reports/g1-07-offline-smoke-20260920/evidence.json. Production artifacts were not modified. This is an offline packaging smoke test, not the two final clean-room notebook reruns required by G6.
  • [x] P0 · G1-08 · historical decision reconciled, gate not closed The 19 September decision remains LOCAL_DIAGNOSTIC_PASS_FULL_G1_ACCEPTANCE_OPEN. Its then-open packaging items have later dispositions in G1-04/G1-07, maintained coverage in A6, candidate evidence in G1-05 and exact identity/mask validation in AUDIT-01. Immutable launch binding, acquisition lineage, official hidden-test parity, Regime C and final clean-room runs remain open.

  • [x] P0 · AUDIT-01 · exact identity/mask audit verified 2026-09-23 Rejoined every saved G1 query/reference evidence row to the authoritative scorer identity cache and frozen masks for library ceiling, Regime A and Regime B; exact target-cohort checks, held-out-fold exclusion, self-exclusion, malformed-row and metric/evidence joins passed with zero errors. Evidence: reports/audit-01-exact-identity-mask-20260923/evidence.json. This validates the saved local qualified run only; launch-time immutability, acquisition lineage and official hidden-test parity remain open.

  • Gate G1: [!] OPEN — full acceptance not established. A successful process exit, 400/400 placeholder rows, single-spectrum metrics or B=0 by construction cannot close it. Regime-B zero target coverage for a pure reference-spectrum matcher is structurally expected after identity-wide exclusion: it demonstrates the method's unseen-identity limitation, not structure-only identification quality. Keep candidate coverage, reference availability and ranking quality separate. Further learned experiments require the task-relevant qualified scorer/split/pool contract; full G1 closure and the unrun C proxy are not to be asserted merely because Phase 4 has already executed.


    Controlled retrieval ablations — after the corrected baseline is locked


    These are follow-on comparisons, not an excuse to withhold a correctly audited baseline indefinitely. Use the same frozen cohorts/masks and bounded development experiments; promote only with reproducible molecule-level evidence.


  • [x] P0 · BASE-01 · resolved 2026-09-23 Reconciled the executable baseline: effective 30-ppm window, uncapped identity pool with top-1000 reporting, 64-peak index, scorer-compatible split, and RDKit 2026.03.3. manifests/retrieval-baseline-resolved-20260923.json supersedes the disputed historical lock; this is a qualified local contract, not hidden-test certification.
  • [x] P0 · BASE-02 · corrected bounded replay closed 2026-09-23 Fixed eligibility-before-collapse, cap ordering, restored similarity ranking and pool-coverage semantics; focused fixtures passed and the corrected 300-identity replay completed across folds 1–4 with fold 0 held as report partition. Best uncapped result was 30 ppm: MRR@25 0.8797, Top-1 0.8567, coverage 0.9100. The 20-vs-10 ppm paired difference was +0.0067, 95% CI [-0.0017, 0.0167], randomization p 0.314: inconclusive. Persisted per-molecule paired outcomes and uncertainty evidence in reports/base02-ablation-corrected-20260923/closure-results.json; do not present this bounded library-ceiling result as generalization or hidden-test performance.

  • [ ] P1 Compare max across permitted spectra against mean/top-k evidence, within-molecule reciprocal-rank fusion, collision-energy-aware aggregation and same-adduct/same-polarity policies. Include a matched single-spectrum control, reference-count bias checks and correlated-acquisition controls. Do not merge incompatible ions into one raw peak list. Cross-route ensemble weights remain Phase 5 work.
  • [ ] P1 Add source, adduct, polarity and collision-energy compatibility features, including missingness and original energy units; validate before adopting weights.
  • [ ] P1 · CHEM-01 Add independent ionic/electron-mass round-trip fixtures (including sodium/potassium/dehydration) and precursor/isotope-peak exclusion ablations. Track true-candidate retention after every filter; don't borrow the notebook's ppm window or precursor cleaning without instrument/adduct-tail evidence.
  • [ ] P1 · CHEM-02 Compare reference-index masses derived from eligible known structure/formula metadata against precursor-derived reference masses, preserving raw residuals and provenance. Keep this separate from query mass/formula inference: never use hidden-query truth formulas. Test formula parser coverage, charge/multimer conventions and unknown/adduct fallback, then paired A/B retention and ranking.
  • [ ] P1 Compare mass windows 8.5 / 10 / 20 / 30 ppm, with a calibrated absolute floor; candidate caps none / 100 / 1,000 / 6,000; relative intensity floors 0.002 / 0.005 / 0.01. Retain the original plan's preprocessing variants as further hypotheses, not an obligatory full Cartesian sweep. An uncapped experiment must still stream within the RAM budget.
  • [ ] P1 For every cap/window/preprocessing setting, record candidate rows and unique identities before/after truncation, correct-candidate retention, full coverage, conditional MRR, MRR@25, Top-1/Top-5, runtime and peak memory. Tune on inner development folds and evaluate the predeclared exposure-labelled report protocol with paired molecule bootstraps; do not assume any existing fold is untouched.
  • [ ] P1 Profile search overhead and reuse a precomputed sorted mass array rather than repeatedly materializing masses[order] inside every query. Prove candidate membership and ranking equivalence before accepting this optimization; do not alter a running evaluator in place.
  • [x] P1 · design only Record an optimization path that preserves the current evaluator as an oracle, separates query extraction from scoring, validates mass-bin/vectorized candidates against a bounded fixture, and checkpoints at candidate-chunk boundaries. Design: reports/todo-review-20260918/G1-OPTIMIZATION-DESIGN.md. No optimized implementation or full-data rerun is claimed complete.
  • [ ] P2 · optional, unbenchmarked Evaluate a separate Kaggle-GPU implementation only after CPU oracle/fixture equivalence is established, including the existing tolerance-based greedy peak matching and deterministic ranks. No GPU notebook, GPU timings or measured speedup exists; previous numerical speedup/ETA estimates are hypotheses, not planning evidence.

  • Phase 2 — Learned fingerprint ranking and domain adaptation


    Entry dependency: qualified task-relevant scorer/split/exposure contracts, a frozen eligible Phase 3A/3B pool and non-neural controls. Full PubChem coverage is not required. Do not train against an empty B target pool and call it a ranking experiment. All model training remains on approved Kaggle GPU resources, never this local host.


  • [ ] P0 Generate out-of-fold candidate features without molecule leakage. Under B masks, remove held-out spectra from supervised training, retrieval, analogue prototypes and all derived features; A/C must follow their own frozen contracts. Candidate-generation models must also be out of fold.
  • [ ] P0 · LIT-01 Use the arXiv shortlist as reference-only design input. Record paper/version/claim-to-test, implement independently, and reject copied code, checkpoints, tokenizers, datasets and ambiguous/non-commercial assets. Evidence: research/arxiv-msms-2d-shortlist-20260923.md; no experiment is licensed or approved by citation alone.
  • [!] P0 · FP-01 Establish spectrum → binary fingerprint logits → candidate ranking, starting with documented Morgan targets and a multi-family union challenger. Freeze fingerprint/bit selection on permitted training data, save ordering and candidate-ID mappings, and test candidate-bit alignment. Compare BCE-only training and f · z ranking with neighbour transfer and simple cosine controls; the independent-Bernoulli score is not a calibrated exact-match probability. Independently implement the concept; do not use paper code, weights or training data. Contract frozen at manifests/fp-01-training-pool-contract-20260923.json; Kaggle handoff at manifests/fp-01-kaggle-handoff-20260923.json; PubChem is excluded and no local training is authorized. Local implementation and compile/schema evidence exists (2026-09-24). Earlier launches hit the weekly GPU quota; the latest 06:25 check returned HTTP 409 and status permission denied without reconfirming quota. Execution remains unverified; no verified model metrics or OOF outputs exist. RUN-01 owns launch recovery.
  • [ ] P0 · FP-02 Add same-ppm hard-negative contrastive ranking to BCE; separately test same-formula hard negatives and false-negative removal by scorer identity. Tune loss weight/temperature on development data. Persist negative-pool hashes and OOF logits; ablate the loss on identical candidate pools and report both easy and hard-isomer strata.
  • [ ] P1 · FP-03 Compare per-spectrum encoding, compatible multi-spectrum training and late molecule pooling. Condition on adduct/polarity, precursor mass, instrument and collision-energy/missingness; align train/inference input distributions. Keep cleaned surviving spectra paired with their own metadata, and handle mixed instruments per query rather than a global test-mode assumption. Do not adopt the notebook's first-adduct/median-precursor/fixed-energy merge for incompatible ions. Ablate source-balanced inclusion/downweighting/exclusion of enveda-180 rather than labelling it uniformly low quality.
  • [ ] P1 · JESTR-01 Independently implement a joint spectrum/structure contrastive embedding challenger motivated by JESTR. Use only permitted structures/spectra and commercially compatible libraries; compare against FP-01 on identical candidate pools, folds, scorer identity and molecule-level MRR@25. Do not download/use JESTR code, weights or data.
  • [ ] P0 Train a first molecule-grouped LightGBM/CatBoost ranker on eligible OOF features; benchmark simple fingerprint transfer before a larger encoder. Keep model versions, targets, folds, seeds and compute budgets reproducible.
  • [ ] P0 Compare learned ranking against the locked heuristic retrieval/database baselines on the same candidate pools; separate coverage gains from reranking gains. Use exact scorer identity and MRR@25 rather than fingerprint similarity as the promotion criterion.
  • [ ] P1 Balance or reweight training sources to reduce synthetic-library dominance.
  • [ ] P1 Add reference-count bias correction.
  • [ ] P1 Add confidence margins and uncertainty diagnostics.
  • [ ] P1 Run ablations for fragment, entropy, neutral-loss and multi-spectrum evidence.

  • Gate G2: learned ranking improves grouped validation with paired molecule-level evidence, no unexplained severe regression in any intended regime, and acceptable measured resources. Random-row/library-only splits and public leaderboard claims do not satisfy the gate.


    Phase 3 — Known structures without reference spectra


    Phase 3A — Bounded pinned structure universe (supporting prerequisite, not bulk expansion)


  • [x] P1 · specification only Structure-index acceptance specification exists at manifests/phase3a-structure-index-spec-20260918.json. B eligibility text is corrected; this does not certify all sources or satisfy acceptance.
  • [x] P0 · specification corrected The draft retains permitted held-out database structures for B while excluding their spectra/features; C structural/scaffold exclusions remain separate. Runtime leakage certification remains a distinct task.
  • [ ] P2 · POOL-01 · DEFERRED Acquire and document an eligible COCONUT snapshot: source/version/date, download hash, exact license/public-access terms and offline path. Compare a bounded union with permitted training structures, deduplicate by pinned scorer identity and count rejected/changed structures. Audit salt/charge/element/mass preprocessing rather than copying the notebook's domain restrictions. Keep structure availability distinct from spectral reference availability.
  • [!] P0 · RIGHTS-01 · fail-closed audit recorded 2026-09-23 reports/rights-01-audit-20260923.json records separate dispositions for papers/notebooks, the bounded PubChem source/index, competition inputs, local runtime and future external models. The recovered PubChem upstream is conditionally usable under recorded CC0 terms, but the CURRENT-Full source is not an immutable release. The exact SDF→TSV→index transformation is now bound to commit 8020971feff895b962a8520f01130c6e05fe9abc; existing derived outputs have not been reproduced under that chain in this turn and remain quarantined. The fail-closed offline package inventory is manifests/offline-package-manifest-20260923.json; it blocks adoption of that external-pool package until rebuild verification and complete dependency/license records; the later competition-only FP-01 contract is a separate route, not blocked by the deferred PubChem rebuild. No paper code, checkpoint, tokenizer or training archive was reused.
  • [x] P0 · Phase 3A rederived structure index · verified 2026-09-20 Built artifacts/phase3a-rederived-20260920/pubchem-structure-index.sqlite from the consolidated Phase 3A source. Terminal manifest status is COMPLETE; 444,792 source rows processed, 424,291 valid/unique scorer identities, 0 invalid SMILES, 0 missing identities and 0 identity mismatches. The 131 MB SQLite artifact and progress manifest are present; no active builder process remains. This completes the bounded structure index build, not source provenance/licensing certification or production promotion.
  • [x] P0 · Phase 3A rederived pool consolidation · verified 2026-09-20 Consolidated the completed chunk outputs into manifests/rederived-20260919/pubchem-rederived-consolidated.tsv: 444,792 valid rows, 444,792 unique CIDs, 424,291 scorer identities, 19,801 training-overlap rows, 424,991 structure-only rows and 179,581 novelty-proxy rows. Independently checked schema and SMILES parseability; no duplicate CIDs. Preserved 24 chunk-reported rejected records in pubchem-rederived-exceptions.jsonl. Manifest: reports/phase3a-rederived-20260919/consolidated-manifest.json. This is a bounded consolidated structure pool, not yet a mass/formula index or promoted production source.
  • [x] P1 · format/count audit only Record the local PubChem text file's physical format, hash, stored row/identity-field counts and malformed-row count in reports/todo-review-20260918/pubchem-local-audit.json. This did not recompute chemical identities, training-overlap or novelty labels against the rebuilt namespace, and did not physically relocate the file.
  • [x] P1 · identity namespace reconciliation Compared source-declared and index-recomputed identity sets read-only: 240,536 source rows; 231,755 declared identities; 231,749 indexed/recomputed identities; 578 declared-only and 572 computed-only identities, net difference 6. The index correctly follows recomputed scorer identity14; the net six is not six catalogue misses. Full sets and hashes: reports/phase3a-pubchem-identity-reconciliation-20260919.json and .md.
  • [-] P0 · source normalization/provenance The Phase 3A derivation chain is now commit-bound at 8020971feff895b962a8520f01130c6e05fe9abc, and the content-pinned archive hash is recorded. Existing TSV/index outputs still require a reproducible rebuild/verification under that exact chain; the upstream CURRENT-Full label is not immutable. Identity-overlap labels match (12,331), but 9,594 scaffold-string differences include 1,741 changed scaffold-overlap memberships; production has seven invalid representatives. Keep the pool quarantined for C/novelty and do not promote until rebuild and RIGHTS-01/offline manifest gates pass.
  • [ ] P2 · DEFERRED · POOL-02 Reconsider broader PubChem-derived coverage only after independent full-cohort queries demonstrate material missing-structure failures and a bounded proposal passes licensing/storage/runtime review. Compare new coverage against same-query ranking dilution and retained candidate recall. No additional CID-slice download is authorized by this plan update.

  • September 19 audit repair gates (before promotion)


  • [x] P0 · S1 Corrected the scaffold classifier to compare scaffold membership separately from identity overlap and string serialization. Read-only rerun: 222,148 exact scaffold/membership rows, 7,853 string-only differences, 1,741 scaffold-membership changes, 0 identity-overlap changes, and 7 invalid production representatives. Added executable regression coverage; novelty metadata remains quarantined. Evidence: reports/phase3a-pubchem-scaffold-discrepancy-20260919.json and reports/review-20260919-1817/AUDIT.md.
  • [x] P0 · S2 · guarded repair verified 2026-09-20 Guarded repair path was executed against the exact base artifact and manifest; the seven explicit repairs passed parseability/identity checks, row-count and no-overwrite guards, and produced artifacts/phase3a/pubchem-structure-index-repaired.sqlite with PASS evidence in artifacts/phase3a/pubchem-structure-index-repaired.repair-verification.json. This closes the repair execution task only; full source provenance/licensing and promotion remain open.
  • [-] P0 · S3 Cache-aligned analogue output persists candidate IDs, ranked IDs, reference IDs, rank, route, source row, adduct, precursor and fallback state for 1,440 policy/query records. Same-query production/shadow comparison and fingerprint availability comparison both pass: 11,106/11,106 candidate identities have fingerprints; 420/2,095 reference identities do; zero availability changes occur between artifacts. Remaining: upstream provenance/licensing before promotion.
  • [-] P0 · S4 Metric contract and cache-aligned 500-identity/480-query rerun completed: enforced top-1000 ranking, true MRR@25, separate route/ranked denominators, mutually exclusive reference/pool failure categories, and operative P2 candidate/reference fallback fields. The rerun selects identities from identity-cache.sqlite via normalized_smiles, never raw Parquet inchikey14. Cache-aligned P0/P2 combined MRR@25 are 0.3003/0.3024 and Recall@25 81.67%/82.08%. Identity audit still finds 1,559 registered-derived rows across 44 stored→derived mappings and 164 identities overlapping prior fold 4. Do not label the partition untouched.

  • Phase 3B — Formula/mass-gated candidates and non-neural baseline


  • [ ] P0 Evaluate formula inference and neutral-mass consistency, retaining alternative formulas when uncertain. Measure true-formula recall and answer retention before tightening gates; report both naturally database-covered molecules and the full B cohort.
  • [x] P0 · reference availability audit Separate PubChem structure coverage from local spectral-reference availability on the post-freeze bounded readout. P0/P1/P2 structure-pool recall was 94.11% on the selected 500-query stability sample; 95.58% had any local candidate reference and 58.11%/58.53% had any fold-eligible candidate reference. All targets had full-library references but 0% had fold-eligible target references by identity-wide exclusion. Evidence: reports/phase3a-reference-availability-20260919.json and .md. This is a post-freeze stability diagnostic, not prospective validation.
  • [-] P0 Bounded reference-availability evidence exists (preceding completed diagnostic task); extend to every frozen pool with candidate/reference-row accounting. Correct route labels and preserve unavailable-reference versus wrong-ranking failures.
  • [ ] P0 Retain a bounded library-ceiling spectral reranking control with query-row exclusion only. Label it optimistic and do not use it as unseen-identity evidence.
  • [ ] P0 Add an identity-held-out cross-identity spectral diagnostic that measures decoy ordering and candidate/reference coverage, but reports target rank only when an eligible target reference exists. A pure spectral-library matcher has structurally undefined target rank when all target references are excluded.
  • [x] P0 · fallback/category audit Independently verified deterministic fallback ordering with zero equivalence mismatches. On 475 valid audit queries: 447 targets were in the structure pool, 28 were absent; 260 had an eligible spectral route and 215 used fallback. Fallback recovered 90.23% in-pool and Recall@25 88.84%. Oracle formula gating was identical (90.23% pool / 88.84% @25), but is explicitly non-deployable because it uses visible-row formula metadata. Evidence: reports/phase3a-analogue-fallback-audit-20260919.json and .md.
  • [-] P0 Registered readout now has a cache-backed identity/exposure audit: 739/739 registered identities are present; 100 direct identity14 recomputations match; 1,559 registered-derived rows across 44 stored→derived mappings disagree with Parquet inchikey14; 164 identities overlap prior fold 4. Query extraction is cache-aligned, per-query provenance is persisted, and regression-tested, but the report remains diagnostic because prior exposure is unresolved.

  • [-] P0 Superseded by cache-aligned report reports/phase3a-analogue-transfer-preregistered-corrected-cache-aligned-20260919.json: 500 identities, 480 valid queries, 1,440 policy/query evidence records. P0 combined MRR@25 is 0.3003 / Recall@25 81.67%; P2 is 0.3024 / 82.08%. These remain diagnostic because 164 identities overlap prior fold 4 and the partition is not untouched.

  • [-] P0 · provenance The September 19 audit recorded missing upstream/derivation certification. The derivation scripts are now present in commit 8020971feff895b962a8520f01130c6e05fe9abc, and manifests/offline-dependency-inventory-20260923.json records 14 packages from the pinned RDKit environment. The historical TSV/index outputs have not yet been reproduced and verified under that commit; the offline package manifest remains fail-closed and inventory-only for licences/wheels. Do not promote this external pool until those fields are closed. Its quarantine does not block competition-only FP-01/AN-01; their own dependency, rights and evaluation checks still apply.

  • Phase 3C — Mass-shifted analogue propagation (parallel complementary channel)


  • [ ] P1 · AN-01 Retrieve eligible reference neighbours by direct and precursor-mass-shifted fragment matching, then transfer evidence through candidate↔neighbour fingerprint similarity. Existing 10/30-ppm direct-cosine transfer is not the notebook's wide-window shifted route. Compare max(direct, uniformly shifted) with a properly one-to-one mixed-match challenger; do not describe the first as the second. The neighbour structure is not automatically the answer. Select mass-shift window, neighbour count and weighting on development data; keep this separate from exact-library retrieval and de novo editing.
  • [ ] P0 Persist source spectrum/source identity, candidate database snapshot/hash, mass-shift and structural-similarity rules, generation/retention counts and deduplication method. Prove no held-out-identity spectral source enters the route or its derived features.
  • [ ] P1 · AN-02 Ablate analogue propagation against frozen direct-library and structure-only controls, on the same folds/pool as FP-01/02. Test adduct/charge/polarity, instrument/energy compatibility and duplicate-reference bias. Report unavailable-neighbour and unavailable-target-reference cases separately; select weights locally, without public leaderboard-tuned constants or fixed candidate quotas.

  • Later representations and forward checking


  • [!] P2 · EMB-01 · excluded by user policy Do not adopt external DreaMS/MIST encoders, pretrained embeddings or checkpoints. Papers may guide independent from-scratch designs; use JESTR-01 for the scheduled challenger. Licence approval does not override the no-external-weights instruction.
  • [ ] P1 Integrate the trained spectrum-to-fingerprint model from Phase 2 into structure retrieval; do not maintain a second conflicting training task here.
  • [ ] P1 · RERANK-01 Test bounded candidate bond-fragment explanation as an independent feature, then an independently implemented spectrum–candidate cross-encoder or from-scratch forward-spectrum model inspired by GLACIER on a high-recall shortlist if ISO-01 confirms a same-formula ranking bottleneck. Measure shortlist recall, complete molecule MRR and stage latency; retain a no-reranker control and explicit chemistry/timeout fallbacks. Do not copy or download paper code, weights or data.

  • Gate G3: measurable structure-known/no-spectrum improvement over the non-neural baseline, with leakage-safe spectral inputs, licensed/pinned structure provenance and separate pool coverage/conditional MRR. Database absence and ranking failure must remain distinguishable.


    Phase 4 — Novel structures and de novo generation


    Feasibility branch, not promoted: a spectrum-to-SMILES implementation and Kaggle experiment already exist. Preserve their source/checkpoint/output history; this review neither stops the existing run nor launches another. Further scale-up is deferred until the gates below pass and a bounded experiment budget is approved. A failed G4 does not prevent shipping a validated fingerprint/analogue/ranking solution, but forbids claiming proven novelty coverage.


    The C-proxy split/leakage audit needed by G1 is validation infrastructure and is not deferred with generation.


  • [ ] P1 Reuse and extend the audited G1 novelty-proxy split with structure, spectral and generator-training-target exclusions before training a generator.
  • [ ] P0 · GEN-01 Bind the exact remotely executed source, checkpoint, tokenizer, package versions and training/split hashes. Shifted labels and unconditional causal masking are already present in current local source; verify behavioral prefix-logit parity, positional encoding, padding/EOS/BOS handling and accumulation remainder with synthetic contracts, then a bounded Kaggle GPU sanity run before long training. A compile pass or low teacher-forced loss is insufficient; include actual generation and spectrum-conditioning checks. No local training, including "smoke" training.
  • [ ] P0 · GEN-02 Replace approximate validation with pinned RDKit 2026.03.3 and official scorer-identity matching; fail clearly if a required dependency is unavailable offline. Audit VAL-01/02 before clean claims. Persist molecule-level unique top-25 outputs across permitted spectra, per-query ranks, validity/duplicate/timeout counts and exact Recall@25/MRR@25. Greedy top-1, string equality and syntax-only validity are not full acceptance metrics.
  • [ ] P1 Build a formula-, adduct- and collision-energy-conditioned spectrum encoder.
  • [-] P1 Bounded SMILES decoder code exists in kaggle-kernel/phase4-denovo/; qualification remains GEN-01/02. No useful generalization result is asserted by code presence or a completed job.
  • [-] P1 Beam-search code exists; complete valid-unique-candidate, special-token and search-budget tests, and compare controlled sampling. Beam width 25 neither guarantees 25 valid scoring identities nor directly optimizes MRR@25.
  • [ ] P2 · GEN-03 After FP-01/02, compare independently implemented predicted/noisy-fingerprint-conditioned generation with the unconstrained baseline, plus soft predicted-formula and valence constraints. MS-BART, MARLIN, DiffMS and MADGEN are technical references only; eligible scaffold conditioning must use deployable inputs, not held-out ground truth. Audit all external assets separately under RIGHTS-01; a paper recommendation is not permission to use its weights/code/data. MSFlow trained assets are excluded because the shortlist describes non-commercial availability; withdrawal is not established by this source.
  • [ ] P1 Add de novo analogue editing from eligible close spectral neighbours; distinguish newly generated structures here from database analogue propagation in Phase 3C.
  • [ ] P1 Deduplicate generated outputs by scorer-compatible identity.
  • [ ] P2 Add forward-spectrum and formula consistency reranking.
  • [ ] P1 · GEN-04 Measure new correct identities, displacement of strong ranked candidates, conditional pool-missing performance and full-cohort MRR after common deduplication/reranking. Keep generation only where its marginal contribution survives resource and uncertainty checks; never hardcode five generator slots.

  • Gate G4: de novo candidates add measurable clean-validation MRR or candidate coverage within the offline runtime budget.



    Later generation experiments — gated, not a submission dependency


  • [ ] P2 · GEN-03A MS-BART-inspired fingerprint-to-structure decoder from scratch on eligible structures; adapt with OOF predicted/noisy fingerprints and preserve B/C exclusions.
  • [ ] P2 · GEN-03B MARLIN-inspired mass/fingerprint-conditioned formula-free challenger only after GEN-03A establishes a useful baseline; bounded sample budget.
  • [ ] P2 · GEN-03C DiffMS-inspired predicted-formula graph challenger; report predicted-formula recall and end-to-end misses, not oracle-formula scores as deployment evidence.
  • [ ] P2 · GEN-03D MADGEN-inspired retrieved-scaffold completion; measure scaffold retrieval failure separately and newly correct identities after common reranking.
  • [ ] P2 · UNC-01 Optional calibrated fingerprint sampling/diversity ablation only after FP-01; independent bit samples need not be chemically realizable. Compare deterministic logits, do not displace plausible isomers with unvalidated scaffold quotas, and retain only measured MRR/recall gains.

  • Phase 5 — Unified ensemble


  • [ ] P0 Union validated retrieval, database/analogue and, only if G4 supports it, de novo candidate pools.
  • [ ] P0 · FUSE-01 Fit a seeded OOF molecule-level evidence ranker across library, analogue, fingerprint and qualified fragmentation/generator features. Exclude pool provenance/source/NP-likeness features by default when simulated positives are all training-derived; audit any shipped feature/weight file. Compare single channels, static fusion and calibrated fusion; select simulation mixture/feature transformations using inner development folds only.
  • [ ] P0 · FUSE-02 Bind each advertised configuration knob to its active code path; test that changing it changes the intended feature/output. The notebook's CFG.SIM_POWER/CFG.W1 are unused, and its candidate cap affects all downstream channels. When caps/pools/channels change, recompute and refit set-relative rank/z-score features or prove invariance; reject stale feature-schema/checkpoint combinations.
  • [ ] P0 Produce exactly one row per runtime molecule with at most 25 unique candidates.
  • [ ] P1 Calibrate route confidence and preserve exploration for uncertain molecules.
  • [ ] P1 Stress-test rare adducts, high masses, one-spectrum molecules and noisy spectra.
  • [ ] P1 Compare library-heavy, balanced and novelty-heavy validation mixtures.
  • [ ] P1 · REPRO-01 Repeat the exact source/config/data/seed pipeline and selected multi-seed controls, storing per-molecule predictions and component scores. Pair molecule bootstrap intervals with our measured numerical/algorithmic noise; do not adopt the author's 0.006 leaderboard spread as a universal cutoff. Promotion requires attribution to an isolated change.
  • [ ] P1 Compare cross-route reciprocal-rank/static fusion and calibrated learned fusion only on common frozen folds with common scorer-identity deduplication. Choose weights without report/hidden-test leakage; require paired improvement over the best individual route, not a public notebook's claimed score.

  • Gate G5: robust integrated model beats the retrieval baseline across the selected validation scenarios.


    Phase 6 — Kaggle packaging and final submission


  • [ ] P0 Package our own from-scratch model weights, indexes, locally built tokenizers and rights-cleared dependencies for offline notebook execution; no external pretrained assets.
  • [ ] P0 · PKG-01 Attach version-pinned offline wheels (including compatible RDKit/NumPy/Python), prove required imports with network access disabled before training/inference, and bind code/weights/bit-order/tokenizer/pool as one run manifest. Preserve downloaded outputs in durable, versioned project storage, not only /tmp; write metrics/checkpoints before potentially failing post-processing. No package-registry fallback in the final offline path.
  • [ ] P0 Run with internet disabled.
  • [ ] P0 Test against dynamically loaded test IDs; do not rely on placeholder IDs.
  • [ ] P0 Verify submission.csv columns, nulls, duplicate IDs and maximum candidate count.
  • [ ] P0 Stress-test cold offline execution, unseen/reordered IDs, changed molecule counts, 16+ spectra, rare adducts, high masses, empty/noisy spectra and missing metadata. Derive IDs from runtime Parquet; sample submission is a format check only. Fail clearly on globally corrupted artifacts.
  • [ ] P0 Complete two clean-room notebook reruns within the 9-hour limit.
  • [ ] P1 Prepare two complementary final candidates only if validation supports both.
  • [ ] P1 Record final code, configuration, hashes, licenses and reproduction instructions.
  • [ ] P0 Select final submissions explicitly before the deadline.

  • Gate G6: two successful offline reruns, reproducible output and final compliance audit.


    Validation contract for the revised route


  • A: held-out acquisitions with independently eligible same-identity references; preserve the documented acquisition-lineage proxy limitation.
  • B: candidate structures and deterministic candidate fingerprints may remain in the frozen database. Exclude all held-out-identity spectra, supervised target pairs, spectral prototypes and learned/in-sample features from model training and reference retrieval. Do not remove the answer from B's structure pool by blanket rule.
  • C: additionally exclude held-out structures/structure targets and apply the declared scaffold proxy. Never present C-proxy results as true hidden novelty.
  • Every experiment: full-cohort and in-pool MRR@25, Top-1/5, Recall@25, pool/shortlist retention, runtime/RAM/VRAM, per-molecule scores and ranks, failure taxonomy, seeds and source/input/output hashes. Candidate source/NP priors stay off by default; formula/scaffold truth is oracle-only.
  • Decision: use paired molecule bootstrap and repeated-seed evidence on the same pool/cohort. Missing structures, formula/scaffold prediction failures, pruning loss and ranking error remain separate. Small uncertain gains stay provisional; a public-score bump alone does not establish generalization.

  • Active experiment log


    Historical chronology is retained below. Earlier PASS/closure statements and namespace counts were superseded by the 2026-09-17 audit/rebuild; the current gate sections above govern decisions.


    DateExperimentHypothesisResultDecision
    2026-09-16Competition review and workspace setupEstablish a complete execution baselinePlan and project scaffold created; no model runStart Phase 0
    2026-09-16Phase 0 data acquisitionVerify current Kaggle release and characterize it under low RAM3,033,286,496-byte file; SHA-256 recorded; 2,539,608 rows, 21 row groups, 275,810 unique InChIKey14 identities; audit completeProceed to grouped split manifests
    2026-09-16Molecule-grouped splitPrevent spectral/structure leakage in local validationDeterministic 5-fold manifest validated: 275,810 groups and 2,539,608 rows; fold assignment is group-exclusiveUse fold 0 as initial validation; keep novelty labels separate
    2026-09-16G0 data/infrastructure gateConfirm reproducible Phase 0 inputs and assumptionsAll eight gate checks passed; config and split hashes recorded in manifests/phase0-gate.jsonClose data/infrastructure work; proceed to retrieval baseline
    2026-09-16Phase 0 auditVerify that validation grouping uses scorer identityFailed: 12/100 sampled normalized SMILES differed between stored and scorer-derived InChIKey14; full probe hit pathological tautomer enumerationReopen G0; rebuild identity-safe split before retrieval evaluation
    2026-09-16Phase 0 corrective auditVerify corrected scorer-identity split and reproducibility gatePassed: 8 tests; compilation passed; 2,539,608 rows and 274,288 scorer groups validated; cache/checkpoint agree; full G0 gate passed; 100-row identity probe had 0 invalid structuresProceed to retrieval baseline; retain external novelty regimes as pending
    2026-09-16Phase 1 retrieval baselineBuild a low-memory, identity-safe spectral retrieval floorIndex built over all 2,539,608 rows; 2,160,035 valid spectra; fold-0 mask excludes all 503,842 held-out rows; molecule-level fixture submission has 400/400 covered molecules and passes schema/count checksKeep baseline; run grouped MRR/Top-1/Recall evaluation before closing G1
    2026-09-17Qualified rebuilt infrastructureSeparate corrected identity namespace from legacy evidencephase0-qualified-g0-20260917.json records 274,195 identities, 22 passing local tests, A/B guards and source/mask hashesQualified local infrastructure only; preserve official parity/acquisition caveats
    2026-09-17Bounded molecule-level retrievalEstablish a diagnostic floor on the rebuilt namespace5,974-identity library/B cohorts plus 250 A identities; one valid query spectrum per identity; saved report and exit-0 timing logNot all-spectrum G1; audit metric definitions and query selection before promotion
    2026-09-18All-spectrum G1 executionAggregate permitted spectra for each validation moleculeStarted 06:26 MYT; still active with no result artifact at the 08:21 checkpointKeep G1 open; require semantic, provenance, strict-output and resource audit, not just process completion
    2026-09-18Community findings + execution-plan reviewImprove candidate coverage before heavier modelsPrioritized structure-only snapshots/mass-formula candidates, OOF fingerprints and separate analogue propagation; added leakage, adduct, cap and aggregation checksAdopt revised backlog order; no new model training, submission or G1 closure claimed
    2026-09-18Evaluator/validator correctionsRepair the morning audit's concrete defectsA row-mask forwarding, uncapped coverage, covered-query conditional MRR and basic invalid/duplicate/over-25 checks implementedLocal fixes verified; not end-to-end G1 acceptance
    2026-09-18Checkpointed rerun at 12:39Expose progress without waiting for the aggregate reportPID 198260: library scan reached 621 batches / 113,376,230 comparisons; B active; final report absent at evening auditPreserve ongoing work; counters are not resumable state, saved metrics or a reliable ETA
    2026-09-18Kaggle CLI release verificationCheck authoritative training data against a fresh downloadSize/hash/schema/rows/row groups all match; verified staged duplicate permanently deleted on requestRetain original data and comparison report; no rebuild justified by release mismatch
    2026-09-18Non-competing preparationImprove small contracts and document next stepsBasic formate alias tested; PubChem physical-format/count audit and Phase 3A/optimization drafts recordedDrafts/source audit have eligibility/provenance limitations; no new index or GPU runner built
    2026-09-18Evening progress auditReconcile source, tests, claims and TODO24 lightweight tests pass with 2 integration tests excluded; independent metric fixes verified; wrapper failures, entrypoint failure, metadata loss and strict-output gaps reproducedKeep G1 open; prioritize A1-A8, narrow completed claims and retain historical evidence
    2026-09-19Phase 3A fallback/category auditVerify fallback equivalence and separate structure/reference failure classes475 valid audit queries: 28 structure-pool misses; 260 had a fold-eligible local identity in the mass window; 215 used the audit fallback route; fallback pool recall 90.23%, @25 88.84%; 0 mass-order equivalence mismatches; oracle formula gate showed no improvementKeep formula gate non-deployable; do not call the 260 count a positive spectral-neighbour count; preserve categories; add maintained tests and repeat on a genuinely pre-registered untouched partition
    2026-09-23FP-01 Kaggle handoffPrepare the smallest independent fingerprint-ranking implementation without local training or external assetsHandoff manifest binds the FP-01 contract, diagnostic ISO-01 panel, baseline and no-submission/no-external-asset policy; implementation and Kaggle run remain pendingImplement and run a bounded Kaggle smoke/ablation: BCE-only vs BCE+hard-negative on approved compute; do not submit until molecule-level evidence and package checks pass


    Blockers and decisions


  • Actionable blockers (24 September): RUN-01 access/remote-state verification and EVAL-01/OOF/package qualification for the competition-only FP route. Latest check does not reconfirm quota. PubChem rebuild and COCONUT acquisition are deferred; diagnostic ISO-01 work is reusable, clean report claims remain gated.
  • Current execution decision: carry forward authorized Kaggle execution toward leaderboard progress once RUN-01 and task-specific checks pass; no local training or paid compute. This replan itself changes documentation only and does not launch or alter jobs.
  • External qualifications, not reasons to fabricate evidence: exact acquisition lineage remains unavailable; the official scorer source and pinned-runtime fixture certification are recorded in reports/official-scorer-certification-20260920/evidence.json. C is a proxy, not true hidden novelty. The 1,184-versus-1,151 documentation discrepancy is pinned locally and is not itself a remaining implementation blocker.
  • Do not expand: unbounded PubChem acquisition, larger encoders or additional long generator runs without the relevant validation/rights gate and approved bounded GPU budget. A frozen bounded pool/non-neural control enables Phase 2/3C; full catalogue coverage is not a prerequisite. Do not compare library-ceiling MRR with clean identity-held-out B as the same task, or assert clean generator folds merely because raw-ID hashes match.
  • Submission gate: SUB-01 requires recorded local evidence, qualified package/runtime and strict dynamic-ID output checks. Development submissions need not wait for deferred generation/external-pool work; final finalists still require G6 clean-room acceptance. This documentation edit submits nothing.
  • Completion rule: code present != run completed != independently valid evidence != gate closed. Preserve all four distinctions in status updates and in the goal state.

  • 24 September shortlist replan


    Reconciled the supplied shortlist with the competition-only decision and saved FP-01 launch evidence. Removed the duplicate TOP5 status ledger, excluded external pretrained routes, corrected B/C eligibility and unsupported MSFlow withdrawal wording, and scheduled first qualified submission before optional heavy research. Existing completed evidence retained; no experiment promoted by this edit.