← Back

πŸ“‘ Contents

Enveda CASMI 2026 β€” MRR@25 / Top-5 Execution TODO1. Current position2. P0 β€” Unlock the first measured candidate2A. Ready now β€” non-GPU work, including arXiv ideas3. P1 β€” Improve measured MRR@254. P2 β€” Generation and uncertainty, only when justified5. Final leaderboard selection6. Non-negotiable experiment and promotion contract7. Constraints and deferred work8. Reuse completed work; preserve history

Enveda CASMI 2026 β€” MRR@25 / Top-5 Execution TODO


Updated: 24 September 2026 (MYT)

Goal: maximize official final private-leaderboard MRR@25 and finish top 5. No fixed score guarantees that rank.

Source of strategy: [latest execution plan](EXECUTION-PLAN.md), authoritative 24 September execution order; [paper shortlist](research/arxiv-msms-2d-shortlist-20260923.md).

Status: [ ] pending Β· [-] partial Β· [x] completed within stated scope Β· [!] blocked. Deferred/excluded work is labelled separately.


Success is measured by ranking quality and official resultsβ€”not checklist completion, training loss, fingerprint accuracy or the number of models built. Promote useful improvements; stop branches that fail to add value.


1. Current position


AreaVerified scope / next gap
Scorer and infrastructureOfficial-scorer fixture and qualified local G0 evidence exist; 64 scoped regression tests pass in .venv-rdkit2026; Phase 0 chemistry/manifest validation passes after repairing the archived PubChem audit path; not proof of hidden-test performance
RetrievalBaseline semantics repaired and 64 scoped tests pass; fresh bounded ablation still required; reference-only B failure is not a learned structure-ranking result
FP-01Independent kernel implemented with saved local compile/schema checks; static handoff audit passes (kernel/metadata/contract/baseline hashes, GPU/no-internet metadata, scorer identity, hard-negative identity exclusion, checkpoint/OOF provenance); no verified trained checkpoint, OOF results or model-quality metrics
Remote launchSaved quota check: 0.00h remaining; refresh 26 September 2026 00:00 UTC = Saturday 26 September, 08:00 MYT, not next week. Recheck availability at refresh; no fresh remote check in this review
ValidationExisting ISO-01 panel is diagnostic. Prior fold-0 exposure and 164 fold-4 overlaps prohibit an untouched-report claim; historical G1-05 evidence reproduces its reported arithmetic and found 0 current-target candidate matches under its limited audit, not full identity-wide fold/mask exclusion; AUDIT-01 already passed exact identity/mask joins for the saved G1 run; new learned routes still require their own exclusion verification
Candidate universeCompetition-training structures only; PubChem quarantined/rebuild deferred; COCONUT acquisition deferred
LeaderboardCurrent official score/rank not verified in this cleanup; establish an account/submission snapshot under RUN-01

Evidence: [local EVAL/VAL readiness](reports/local-eval-val-readiness-20260924.json), [Phase 0 validation](reports/phase0-validation-20260924.json), [FP implementation](reports/fp-01-implementation-20260924.json), [local validation](reports/fp-01-local-validation-20260924.json), [local tests](reports/fp-01-local-preparation-tests-20260924.json), [package provenance](reports/fp-01-package-provenance-20260924.json), [latest launch check](reports/fp-01-launch-check-20260924-0625.json), [exposure audit](reports/val-01-02-exposure-20260923.json).


2. P0 β€” Unlock the first measured candidate


Do these first. RUN-01 blocks remote execution, not useful local preparation. Do not wait for every research branch before a qualified development submission.


  • [!] RUN-01 β€” Recover Kaggle execution at the recorded quota refresh. GPU quota is confirmed exhausted today: 40.89h used, 0.00h remaining, refresh 2026-09-26T00:00:00Z; no push or launch attempted this turn. The cache-bound Phase 4 sanity kernel is statically ready (reports/phase4-sanity-launch-audit-20260924.json) but remains unexecuted. After refresh, verify account access and exact owner/slug, push only the bounded sanity kernel, inspect its checkpoint/metrics, then decide whether to scale. Avoid duplicate launches and blind retries. Exit: an accessible, identified run and actionable compute status.
  • [-] EVAL-01 β€” Freeze the comparison contract. Reuse certified scorer and BASE/AUDIT evidence. Bind competition-only candidate IDs, fingerprints/bit order, identity namespace, pool/fold hashes, all-spectrum query aggregation and non-neural controls. Freeze nested/grouped development/report rules with exposure disclosed; audit model-specific training exclusions. Completed this run: validated the persisted ISO-01 panel contract (24,858 formula groups / 247,777 exposed identities; deterministic group membership and required metrics; panel SHA-256 unchanged) and passed the read-only cross-artifact consistency audit in reports/eval-val-consistency-audit-20260924.json. Exit: reproducible same-query/same-pool comparisons, with diagnostics distinguished from promotion evidence.
  • [-] VAL-01/02 β€” Close relevant identity/exposure gaps. Verify cross-artifact joins and all-library held-out-identity exclusion in every learned route. Maintain an exposure ledger; no reseeding/relabeling makes old queries untouched. Use a predeclared nested/grouped evaluation when a genuinely unexposed report set is unavailable. Completed this run: reran the read-only authoritative-cache audit: 739/739 registered identities aligned, 100/100 direct identity recomputations matched, 0 uncached rows; prior fold-4 overlap remains 164, so no clean-untouched claim. The consistency audit confirms these counts and the clean-claim guard; new learned-route cross-library fold/mask exclusion remains open; the saved G1 AUDIT-01 is already PASS.
  • [-] ISO-01 β€” Qualify diagnostic strata. Reuse existing formula-grouped panel, but freeze independent full-cohort coverage and in-pool/hard-isomer strata without selecting on model failures. Keep pool-absent, zero-candidate and unusable-spectrum cases in the end-to-end denominator. Diagnostic construction does not close clean evaluation.
  • [x] BASE-02 β€” Repair retrieval-ablation semantics and complete bounded replay. Mass-window eligibility now uses each source query spectrum mass; candidate caps apply per query spectrum before identity aggregation; candidate rows retain query provenance; rank >25 remains pool coverage. Synthetic regression fixtures and the full scoped suite pass (51 tests). Evidence: reports/base02-retrieval-semantics-repair-20260924.json. A fresh paired CPU diagnostic on the frozen 246-identity cohort is recorded in reports/base02-analogue-paired-20260924.json: corrected 10-ppm BASE-02 MRR@25 0.95935, Top-1 0.93902, coverage 0.98374. Analogue policies are materially worse and not promoted. These are diagnostic, not official or clean-generalization results.
  • [!] FP-01 β€” Execute the implemented BCE baseline. Verify code/input/dependency hashes and failure handling; run a bounded approved Kaggle GPU sanity job before scaling. Pool spectrum logits once per query identity, score the frozen candidate universe, and persist checkpoint, logits, candidate IDs/scores, ranks, timing and memory. Local scorer-identity binding was corrected after audit; pinned-runtime semantic checks pass. Same-mass hard-negative pairs now exclude same-identity spectra; 64-test suite passes. Static pre-launch handoff audit passes in reports/fp-01-handoff-audit-20260924.json; remote execution remains blocked by current Kaggle quota. See reports/fp-01-local-validation-20260924.json. The 19 September G1-08 decision is historical; its then-open malformed-precursor/local-packaging items have later local dispositions in G1-04/G1-07, with maintained regression coverage in A6. Remaining work is immutable launch binding, GPU execution, checkpoint inspection and full G1/C qualification. Exit: completed, inspectable runβ€”not compilation or an accepted push.
  • [ ] FP-02 β€” Test ranking loss. Compare BCE-only with BCE + same-mass hard-negative ranking on identical cohorts/pools/budgets; separately test same-formula negatives. Remove scorer-equivalent false negatives, save negative-pool hashes, and tune weights/temperature only inside development folds. Exit: paired molecule MRR@25 comparison against FP-01 and matched controls.
  • [ ] PKG-01 β€” Build offline inference. Existing retrieval packaging passes network-guarded G1-04/G1-07 fixtures directly from the repository root; dynamic IDs/order, malformed CSV rejection and offline validation are proven in reports/pkg-01-local-readiness-20260924.json. The static FP-01 handoff audit also passes. A checkpoint-pending package manifest now binds the FP-01 kernel, metadata, training-pool contract, baseline and handoff audit in manifests/fp-01-package-pending-20260924.json; the origin-aware dependency inventory confirms NumPy/PyArrow/RDKit are from .venv-rdkit2026, while Torch is intentionally absent locally and wheel hashes/license records remain unbound (reports/fp-01-dependency-inventory-20260924.json). Bundle our own trained checkpoint, code, pool/index, bit ordering and rights-cleared pinned dependencies after GPU execution. Exit: reproducible inference artifact, not placeholder accuracy.
  • [ ] SUB-01 β€” Submit the best qualified development candidate. Select the strongest validated eligible route, including the retrieval control if learned ranking does not improve it. Record hypothesis, local evidence, immutable notebook/source/asset version, successful offline check and official submission result. Respect submission limits; no probing hidden labels. Exit: confirmed scored submission and a recorded baseline to beat. Finalist two-clean-room-run acceptance remains FINAL-01, not a reason to postpone every development submission.

  • Critical path: RUN-01 at quota refresh + EVAL-01 β†’ bounded FP-01/02 comparison β†’ PKG-01 β†’ SUB-01. During the quota pause, only safe local preparation/documentation is allowed; no local training or paid compute. Pandas is installed; the full scoped suite now passes 64 tests after retrieval, chemistry, analogue-route, candidate-propagation and metadata-eligibility contract work. Existing code and diagnostics are reusable; external-pool rebuilding is not on this path.


    2A. Ready now β€” non-GPU work, including arXiv ideas


    GPU quota blocks training, not reading papers, independent design, CPU chemistry/metrics or interface preparation. Do not repeatedly run the same 64-test suite or rewrite completed audits as substitute progress. These tasks remain pending until their deliverables exist.


    Priority / IDTask available nowConcrete completion evidence
  • [x] LIT-02 β€” Complete primary-method extraction. Primary arXiv HTML methods for JESTR (2411.14464), retrieval objectives (2602.16507), MassSpecGym (2410.23326) and MassSpecGym in the Wild (2606.19624) were inspected and implementation-relevant notes saved in research/lit-02-method-extraction-20260924.md; provenance and decision record: reports/lit-02-method-extraction-20260924.json. No paper code, checkpoint, tokenizer, pretrained representation or external dataset was adopted. JESTR remains an independent design-only challenger; retrieval MRR@25β€”not fingerprint similarityβ€”remains the promotion metric.
  • 2 Β· JESTR-PREPPrepared and synthetically verified in research/jestr-prep-20260924.md, manifests/jestr-prep-contract-20260924.json and reports/jestr-prep-validation-20260924.json. The contract binds scorer identity, competition-only pool, grouped folds, molecule-level scoring, negative exclusions, deterministic ties and MRR@25 metrics. No training or paper reproduction is claimed.
    3 Β· COMP-01Scoped audit complete in reports/comp-01-analogue-scope-audit-20260924.json. The historical 246-record P0 report is internally coherent: 5,890 candidate rows, zero candidate IDs missing from the referenced index, zero duplicate candidate rows, 25/246 targets absent from candidate lists (10.16%), and route target ranks present for direct 121, shifted 116 and combined 121. This separates candidate absence from route ranking/abstention; it does not certify the quarantined index, hidden-test generalization or the full 250-identity denominator.
    4 Β· BASE-02-CPUCompleted bounded corrected CPU paired replay in reports/base02-analogue-paired-20260924.json on frozen 246-identity cohort: BASE-02 10-ppm control MRR@25 0.95935, Top-1 0.93902, coverage 0.98374; analogue policies 0.226–0.229 MRR@25 and 0.610–0.618 coverage, all ~0.73 MRR below BASE-02. Diagnostic only; no promotion.
    5 Β· PKG-LICENSEPartial local inventory complete in reports/pkg-license-local-inventory-20260924.json. NumPy 2.2.6, PyArrow 25.0.1 and local RDKit 2026.3.3 metadata/notice state are recorded; Torch is not installed in the project venv. Kaggle runtime/ABI, RDKit wheel provenance, CUDA Torch and transitive notices remain unverified.
    6 Β· GLACIER-PREPPrepared and synthetically verified in research/glacier-prep-20260924.md and reports/arxiv-prep-validation-20260924.json. The contract separates candidate recall from reranking, preserves upstream scores on forward-model failure, binds failure strata, and requires paired molecule-level MRR@25 before promotion. No predictor or paper reproduction is claimed.
    7 Β· GEN-PREPPrepared and synthetically verified in research/gen-prep-20260924.md and reports/arxiv-prep-validation-20260924.json. The shared contract covers MS-BART/MARLIN/DiffMS/MADGEN-inspired routes, deployable-vs-oracle inputs, scorer-identity deduplication, failure strata and marginal-MRR gates. No architecture selection, weights or training claim is made.

    Immediate choice: LIT-02 β†’ JESTR-PREP, while COMP-01 corrects the evidence boundary. Reading/design may proceed before FP-01; training, measured comparison and promotion remain sequenced behind the first qualified FP baseline. GLACIER/generation preparation does not authorize expensive training branches.


    3. P1 β€” Improve measured MRR@25


    After the initial comparison, choose the next experiment from the observed failure breakdown. One isolated change per comparison; no obligatory Cartesian sweep.


  • [ ] FP-03 β€” Multi-spectrum/domain ablation. Compare single-spectrum, compatible pooling and late fusion; preserve each spectrum's adduct, polarity, instrument, collision energy and missingness. Test source balance/weighting rather than assuming all synthetic data are harmful. Compare Morgan with a fixed multi-family fingerprint challenger only after baseline evidence.
  • [-] AN-01/02 β€” Scoped paired diagnostic complete; deprioritized. The frozen 246-identity replay found analogue MRR@25 0.226–0.229 versus BASE-02 0.95935, with materially lower coverage; no analogue policy is promoted. The historical scope limits remain: 25/246 target-absent candidate lists and route abstentions, diagnostic-only cohort, no official score or clean generalization claim. Evidence: reports/base02-analogue-paired-20260924.json. Reopen only with a justified eligible pool and new hypothesis.
  • [x] COMP-01 β€” Scope-audit the historical analogue comparison. scripts/audit_comp01_analogue_scope.py and reports/comp-01-analogue-scope-audit-20260924.json verify the saved report’s 246-record P0 cohort, 5,890 candidate rows, candidate/index membership, duplicate status, explicit 25/246 target-absence stratum, route abstentions and source hashes. This is a scoped evidence audit, not a fresh run or promotion; the quarantined-index rights and full 250-identity denominator remain unresolved.
  • [ ] FUSE-01/02 β€” OOF evidence fusion. Compare best single route, simple static/rank fusion and a seeded molecule-level learned ranker. Train only on fold-safe route features; exclude candidate-source/natural-product priors by default. Bind every configuration knob and recompute set-relative features after pool/pruning changes. Require marginal gain over the strongest individual route.
  • [-] CHEM-01/02 β€” Targeted candidate-retention fixes. Investigate mass/adduct/formula failures before tightening filters. Validate ionic masses and unknown/multicharge handling; compare eligible structure-derived reference masses with precursor-derived masses. Keep query formula inference deployable and quantify every filter's answer loss. Completed local chemistry gate: all ten advertised adducts round-trip; formate physical delta and alias are independently tested; malformed, unknown, nonfinite and nonpositive inputs fail closed. Phase 0 validator now resolves the archived immutable PubChem audit path and passes all checks. No new adduct semantics were invented; multicharge remains unsupported by the declared singly charged contract.
  • [x] GLACIER-PREP β€” Freeze the bounded forward-reranker contract. research/glacier-prep-20260924.md defines deployable inputs, explicit failure statuses, upstream-score preservation, feature boundaries, shortlist-retention gates and molecule-level promotion metrics. reports/arxiv-prep-validation-20260924.json passes its synthetic contract checks. This is design-only and not a paper reproduction or trained predictor.
  • [ ] REPRO-01 β€” Confirm gains. Repeat selected frozen pipelines/seeds, store per-molecule deltas and paired bootstrap intervals, and test hard-isomer/domain regressions. Small uncertain improvements remain provisional. Preserve the last reproducible best candidate while testing challengers.

  • 4. P2 β€” Generation and uncertainty, only when justified


    Not prerequisites to SUB-01. Proceed when useful fingerprints and failure analysis justify the compute; measure incremental correct identities and final MRR after shared ranking.


  • [ ] GEN-01/02 β€” Qualify existing generator engineering. Bind source/tokenizer/checkpoint/runtime; test shifted targets, causal/padding/position behavior, BOS/EOS, accumulation and actual spectrum-conditioned generation. Use the pinned official identity, molecule-level unique top-25 outputs and B/C-safe splits. No local smoke training; existing feasibility code is not a validated model. Current preparation: cache-only identity mapping now fails closed on any missing SMILES; validation emits pool coverage, Recall@25, MRR@25, Top-1, Top-5, validity/uniqueness/duplicate and invalid-output diagnostics; static sanity launch audit passes. This extends the existing precursor-m/z conditioning path; adduct, collision-energy/missingness and formula conditioning remain future work. GPU evidence and checkpoint quality remain unverified.
  • [x] GEN-PREP β€” Freeze the shared generation contract. research/gen-prep-20260924.md defines deployable inputs, oracle-only boundaries, MS-BART/MARLIN/DiffMS/MADGEN-inspired route contracts, scorer-identity deduplication, failure strata and marginal-MRR acceptance gates. reports/arxiv-prep-validation-20260924.json passes synthetic checks. Training and architecture selection remain gated by FP-01/02 evidence; no external assets are approved.
  • [ ] GEN-03B β€” MARLIN-inspired formula-free challenger. Test bounded fingerprint-plus-mass generation only after a useful simpler baseline; measure new exact identities, not merely valid SMILES.
  • [ ] GEN-03C β€” DiffMS-inspired formula-conditioned challenger. Use predicted formula hypotheses; report true-formula retention and full-cohort misses. Ground-truth formula results are oracle diagnostics only.
  • [ ] GEN-03D β€” MADGEN-inspired scaffold completion. Use retrieved/predicted scaffolds and bounded analogue edits. Separate scaffold retrieval failure, invalid generation and wrong ranking; never condition deployment on hidden structural truth.
  • [ ] UNC-01 β€” Optional uncertainty sampling. Compare deterministic logits with bounded calibrated fingerprint sampling. Independent bit samples need not describe realizable molecules. Test diversity policies without removing plausible constitutional isomers; no fixed route/scaffold quotas.
  • [ ] GEN-04 β€” Keep only incremental value. Count newly correct scorer identities, candidate displacement, pool-missing gains and final full-cohort MRR under a fixed runtime budget. Drop ineffective branches rather than reserving generation slots.

  • 5. Final leaderboard selection


  • [ ] FINAL-01 β€” Two clean-room reruns per finalist. Pin all sources/assets, disable internet and run within the competition limit twice. Verify dynamic/unseen/reordered IDs, variable spectra counts, empty/noisy spectra, rare adducts/high masses, strict output schema and deterministic outputs or explained bounded variation.
  • [ ] FINAL-02 β€” Select final entries. Maximize expected private-leaderboard MRR using local generalization evidence and observed official results, not repeated public-score tuning. Select up to two complementary candidates only if evidence supports both; otherwise retain the strongest reproducible system. Confirm final selection before the recorded competition deadline in the execution plan.
  • [ ] FINAL-03 β€” Reproduction package and result. Save code, our own checkpoints, dependency/data permissions, input/output hashes, runtime, experiment/submission ledger and limitations. Record official final MRR@25 and rank; claim top five only when officially confirmed.

  • 6. Non-negotiable experiment and promotion contract


    Optimize the actual score: MRR@25 = mean(1 / first_correct_rank); missing answers or ranks beyond 25 score zero. Aim to move correct identities toward rank 1 while preserving useful alternatives through rank 25.


    RegimeEligible evidence
    A β€” reference availableExclude query acquisitions/duplicates; independently eligible same-identity reference spectra may remain. Disclose acquisition-lineage proxy limits
    B β€” structure known, spectra unavailableCandidate structures and deterministic fingerprints may remain. Exclude held-out spectra, supervised target pairs, spectral prototypes and learned/in-sample features from training/reference retrieval across every library
    C β€” novelty proxyAdditionally exclude held-out structures/structure targets; enforce the declared scaffold proxy. Do not claim this is true hidden novelty

    Every comparison must record:


  • Same frozen query cohort/pool, identity deduplication, masks, seeds and source/input/output hashes; training and selection exposure disclosed.
  • Full-cohort MRR@25, covered-pool conditional MRR, Top-1/5, Recall@25, candidate Recall@100/1000, pool/shortlist retention, runtime and peak RAM/VRAM.
  • Per-molecule IDs, candidate membership/scores/ranks, OOF provenance and hard-isomer, source/instrument, adduct, mass and spectrum-count strata.
  • Separate unavailable structure, unavailable reference, wrong formula/scaffold, pruning loss, ranking failure, invalid/duplicate output and timeout.
  • Paired molecule-level uncertainty, repeated-run evidence where needed, marginal gain over the best control, and no unexplained severe intended-regime regression.

  • Training loss, validity, bit accuracy, in-pool recall and public leaderboard bumps do not individually establish improvement. Placeholder test data are schema/runtime fixtures, not generalization evidence. Do not claim full historical G1/G4 closure from these scoped tasks.


    7. Constraints and deferred work


  • Papers are technical references only. Independently implement; no copied paper/notebook implementations or external pretrained weights/checkpoints/tokenizers/training archives. Our own from-scratch checkpoints may be retained and packaged.
  • Use commercially compatible general-purpose libraries/tools and documented eligible data. Competition inputs retain their competition-specific restrictions; do not claim they are commercially licensed or publicly redistribute derived archives.
  • RIGHTS-01 / LIT-01: before adopting an input/dependency or a new paper mechanism, record rights/provenance and paper version/method-to-test. Citation alone approves no asset. Review actual method details before implementation; the shortlist is abstract-level screening.
  • Train only on approved Kaggle GPU resources; no local training or newly paid compute. Local data preparation must stream within host memory; long jobs need persistent supervised execution and recoverable checkpoints (historical A7 resume work applies before a long local rerun).
  • POOL-01/02 β€” DEFERRED: COCONUT acquisition, PubChem rebuild and broader catalogue expansion. Reconsider only after independently measured missing-structure failures justify a bounded rights/resource-qualified proposal. Existing PubChem artifacts remain quarantined; they do not block the competition-only route.
  • EMB-01 β€” EXCLUDED: external pretrained DreaMS/MIST representations. MSFlow trained assets are likewise excluded; the shortlist does not establish paper withdrawal.
  • Current source/NP priors, fixed 10/5/5/5 allocations, mandatory diffusion and unvalidated diversity quotas are not default production policies.

  • 8. Reuse completed work; preserve history


    These are recorded scoped completions, not fresh reruns or blanket promotion:


    Reusable resultEvidence
    Official scorer fixture, RDKit 2026.03.3[Certification](reports/official-scorer-certification-20260920/evidence.json)
    Qualified local G0[Decision](reports/g0-local-closure-20260920/decision.json)
    Saved exact identity/mask audit (AUDIT-01)[Evidence](reports/audit-01-exact-identity-mask-20260923/evidence.json)
    Executable retrieval contract (BASE-01)[Manifest](manifests/retrieval-baseline-resolved-20260923.json)
    Historical BASE-02 replay β€” superseded by 24 September semantics repair; not current control metrics[Results](reports/base02-ablation-corrected-20260923/closure-results.json)
    Local offline packaging/validator fixtures, not final Kaggle acceptance[G1-04](reports/g1-04-offline-acceptance-20260920/evidence.json), [G1-07](reports/g1-07-offline-smoke-20260920/evidence.json)
    FP contract and diagnostic panel[Contract](manifests/fp-01-training-pool-contract-20260923.json), [Panel](manifests/iso-01-training-diagnostic-panel-20260923.json)

    The complete prior checklist, all 49 completed entries, historical experiments and unresolved legacy caveats are preserved in [archived TODO](reports/todo-cleanup-20260924/TODO.before.md). This file replaces duplicate status ledgers; the archive is evidence, not the current execution order. Old G0–G6 gate names in the execution plan retain their historical scope.


    Next action: LIT-02 method extraction and JESTR-PREP are now complete at the documented design scope; qualify COMP-01 and the bounded BASE-02 CPU replay are complete. A private Kaggle dataset jamesl8/casmi26-base02-retrieval-index-20260924 was created and verified ready on 24 September, containing the frozen corrected index for the 10-ppm BASE-02 submission. Do not retry GPU before the recorded 26 September 08:00 MYT refresh. Recheck quota/access then; a bounded Phase 4 engineering sanity check is not evidence that generation should displace FP-01/02. Preserve the objective: higher official MRR@25 and a top-five finish.