Enveda CASMI 2026 — Winning-Solution Execution Plan
Prepared: 16 September 2026 (Malaysia time)
Competition: https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra
Objective: maximize final private-leaderboard MRR@25 with a compliant, reproducible, offline solution capable of competing for first place. Winning is an objective, not a guarantee.
Last revised: 24 September 2026 (MYT), following the [analog-propagation notebook review](reports/plan-review-20260922-analog/REVIEW.md) and the [MS/MS-to-2D-structure arXiv shortlist](research/arxiv-msms-2d-shortlist-20260923.md).
Status: retrieval diagnostics, bounded structure indexes and a Phase 4 feasibility implementation exist; experimental execution is not evidence of clean generalization or gate closure. TODO.md owns task status. This revision changes research priorities and acceptance requirements only; it does not launch training, modify a running kernel or promote artifacts.
1. Executive judgment
This is a molecular structure identification and ranking competition, not a conventional tabular prediction problem. Given several mass spectra of a molecule, return up to 25 candidate chemical structures ordered by probability of being correct.
The recommended strategy is a three-route candidate system with one calibrated molecule-level ranker:
1. Spectral-library retrieval: win the known-compound cases with accurate mass filtering, fragment/neutral-loss matching, and evidence across collision energies and adducts.
2. Structure-database retrieval: recover molecules with no reference spectrum using formula inference, spectrum-to-fingerprint/embedding models, and forward-spectrum reranking.
3. De novo generation: propose genuinely new structures, including local edits of spectral analogues, then verify/rerank them using the same chemistry and spectral evidence.
Current priority: after repairing evaluation and provenance gaps, build spectrum-to-fingerprint learning plus calibrated candidate ranking, with mass-shifted analogue propagation as a complementary channel. Keep exact-library retrieval as a Class-1 control. These channels can rank a permitted structure even when that structure has no eligible reference spectrum; they are not merely exact-spectrum lookup.
Bound the catalogue work: respect CC's decision not to pursue open-ended PubChem downloads. Use the competition-training structure pool now; defer COCONUT acquisition and PubChem rebuilding. Future external-pool comparisons require independently selected queries and a measured coverage gap. A pool is supporting infrastructure, not the main learning strategy. Reconsider expansion only after a measured coverage gap and resource/licensing review, not because a larger CID range sounds better.
Keep generation as a bounded complementary experiment, not an assumed winner or a discarded route. Exact-structure scoring does not make generation intrinsically wrong. It can recover in-pool or out-of-pool answers, but must add correct unique candidates and/or improve final MRR. Do not allocate fixed retrieval/generation slot quotas. NotebookLM's fingerprint-to-constrained-generator proposal fits Track C after a useful fingerprint predictor and a valid evaluation contract exist.
Execution order: execution access + task-specific evaluation contract → existing FP-01 loss comparison → qualified first submission → JESTR/analogue challengers and fusion → bounded forward reranking → conditional generation. Preserve existing experiment outputs; do not expand training before the applicable gate passes. Phase numbering is retained for traceability, not a requirement to follow the original calendar.
Binding implementation policy: papers/notebooks are technical references only; independently implement adopted concepts. Do not copy their implementations or download/use external pretrained weights, checkpoints, tokenizers or training archives—even if their licences permit it. Commercially compatible general-purpose libraries remain allowed after review. Our own checkpoints trained from scratch on authorized inputs may be saved, transferred from Kaggle and packaged. Competition data are used under the competition-specific permission; this does not make the dataset commercially licensed. External data/libraries/tools require documented compatible rights and provenance. PubChem rebuilding and COCONUT acquisition are deferred, not prerequisites to the competition-only route.
Current execution snapshot — 24 September 2026 (MYT)
kaggle-kernel/fp-01-ranking/main.py; saved implementation report records local compile/schema checks. Training/OOF outputs and model quality remain unverified. See reports/fp-01-implementation-20260924.json.reports/fp-01-launch-check-20260924-0625.json records HTTP 409 and status permission denial; execution is unverified. CC has now confirmed the weekly Kaggle GPU quota is exhausted, so RUN-01 is deferred until next week; do not retry this week.Authoritative execution order
1. Recover execution access next week (RUN-01). Deferred until the weekly Kaggle GPU quota resets by CC's explicit decision. Do not retry this week. Inspect the exact FP-01 owner/slug, account access and remote state before any retry; distinguish permission, conflict and quota failures. No paid compute or local training is implied.
2. Finish the bounded comparison contract (EVAL-01). Reuse the certified scorer and corrected baseline; bind the competition-only pool, identity-wide training exclusions, candidate membership and predeclared development cohorts. Existing exposed panels support diagnostics, not untouched claims. Freeze a nested/grouped selection/report protocol with prior exposure disclosed before promotion.
3. Run the existing fingerprint kernel (FP-01/02/03). Start with a bounded Kaggle GPU sanity run, then BCE-only versus BCE plus same-mass negatives on identical folds/pools; add same-formula negatives and multi-spectrum variants one change at a time. Evaluate exact molecule MRR@25, not fingerprint accuracy alone. Persist logits, candidate scores/IDs, ranks, checkpoint and run provenance; code presence is not execution evidence.
4. Package the first qualified candidate (SUB-01, PKG-01). Compare against matched retrieval/transfer controls, select using local evidence, verify offline dynamic-ID inference and strict output validation, then submit an evidence-qualified version under existing execution authority. Do not wait for every research branch or the final two clean-room finalist runs before a bounded development submission. Record remote notebook version and official score; a valid submission is not proof of improvement.
5. Challenge the baseline (JESTR-01, AN-01/02, FUSE-01). Independently implement joint spectrum/structure contrastive ranking after the FP comparison; measure analogue transfer as complementary evidence on the same folds/pool. Compare single routes, static fusion and OOF learned fusion. Retain only attributable improvements.
6. Bounded forward reranking (RERANK-01). Use GLACIER as a conceptual reference for an independently trained predictor, starting with cheap fragment-explanation features. Only proceed when failure analysis shows ranking, rather than absent candidates, is limiting performance. Measure shortlist retention and whole-pipeline runtime; no fixed top-K guarantees nine hours.
7. Conditional generation (GEN-01/02/03/04). After useful fingerprint ranking, test predicted/noisy-fingerprint-conditioned decoding, then bounded scaffold editing and formula-free/formula-conditioned challengers. Require deployable predicted formulas/scaffolds, full-cohort failure accounting and marginal gains after shared reranking. No fixed generation quotas. Sampling/diffusion are optional experiments, not prerequisites to a leaderboard submission.
Shortlist-to-experiment map
This mapping uses the supplied local shortlist (abstract-level screening), not a fresh full-paper replication or asset licence certification. Before implementing a new paper-specific mechanism, record its version, exact method section and independently written design under LIT-01. Published benchmark numbers are not expected CASMI scores.
| Technical reference | Independent experiment | Acceptance / limit |
|---|---|---|
| Retrieval objectives, arXiv:2602.16507 | FP-01/02: BCE + f · z versus hard-negative ranking | Same pool/cohort; molecule MRR@25 and paired ranks decide, not bit accuracy |
| MassSpecGym / In the Wild, 2410.23326 / 2606.19624 | EVAL-01, VAL-01/02, AUDIT-01 | Scorer parity, acquisition/identity exclusions, exposure and shortcut audits; no blanket indictment of other papers |
| JESTR, 2411.14464 | JESTR-01: joint spectrum/structure embeddings from scratch | Same frozen pools/folds and comparable budget versus FP-01 |
| GLACIER, 2606.29161 | RERANK-01: forward-spectrum agreement | Retained-answer recall, no-reranker control, OOF features and measured latency |
| MS-BART, 2510.20615 | GEN-03A: structure/fingerprint decoding from scratch, adapted to OOF predicted fingerprints | No MIST weights; structure targets follow B/C exclusions; compare noisy versus perfect/oracle conditioning |
| MARLIN, 2607.04774 | GEN-03B: fingerprint + mass, formula-free generation concept | Bounded challenger; no downloaded weights; measure exact new identities and runtime |
| DiffMS, 2502.09571 | GEN-03C: predicted-formula-constrained graph generation | Ground-truth formula is oracle-only; count formula misses in full-cohort score |
| MADGEN, 2501.01950 | GEN-03D: retrieved scaffold and bounded analogue editing | Predicted scaffold, never held-out truth; separate retrieval failure from generation failure |
| MSFlow, 2602.19912 | Excluded assets; no scheduled experiment | Shortlist describes non-commercial trained version. Withdrawal is not established by this source; remove the earlier unsupported withdrawal claim |
---
2. Verified competition facts
| Item | Verified requirement / observation |
|---|---|
| Target | 2D atom connectivity, submitted as SMILES |
| Prediction unit | One row per molecule_id, aggregating all its spectra |
| Metric | Mean Reciprocal Rank at 25, higher is better |
| Correctness | RDKit tautomer canonicalization, then first InChIKey block (InChIKey14); RDKit pinned to 2026.03.3 |
| Stereochemistry | Not required for scoring; tautomer-equivalent guesses match after canonicalization |
| Output | submission.csv, columns molecule_id,smiles; guesses separated by semicolons |
| Guess count | Up to 25; wrong guesses only cost their occupied rank |
| Execution | Notebook submissions; CPU or GPU runtime ≤9 hours, internet disabled |
| External assets | Freely/publicly accessible external data and pretrained models allowed, subject to rules, accessibility and licensing |
| Submission limit | 5 per day; select at most 2 final submissions |
| Team limit | 5 people |
| Prize pool | US$50,000; first prize US$16,000; top five receive prizes |
| Start | 14 September 2026 |
| Entry / team merger deadline | 7 December 2026, 23:59 UTC = 8 December, 07:59 MYT |
| Final deadline | 14 December 2026, 23:59 UTC = 15 December, 07:59 MYT |
| Winning code | MIT winner-license type; training/inference code and reproducible documentation required |
| Competition data | Competition use and non-commercial/academic research; CC BY-NC 4.0; do not redistribute to nonparticipants |
Historical leaderboard/account snapshot — 16 September 2026
At the original 16 September review, the CLI listed 147 teams. The leading public score was 0.332, second 0.285, and tenth 0.254. This is an early leaderboard, not evidence of the score needed to win in December. Our account was already entered, with 0 lifetime submissions and all 5 daily submissions remaining.
Three hidden novelty classes
Class proportions and per-molecule class assignments are hidden. "No public reference spectrum" does not mean nobody has ever measured the compound; public availability, our local library and a fold-eligible library are different sets. The classes concern reference/structure availability, not a guarantee that cosine retrieval will identify every Class-1 query.
The public notebook estimates Class 1 as 0.151 / 0.93 ≈ 16%. This assumes its local Class-1 quality transfers to the public test and other classes contribute negligibly to its library-only score. The remaining approximately 84% is not a measured Class-2 fraction: it includes other classes, and public/private mixtures may differ. Use the estimate only for sensitivity scenarios; do not use it as a known training prior or hidden label.
---
3. What was actually inspected
Official sources retrieved
Authenticated Kaggle CLI access successfully retrieved the description, data description, evaluation, timeline, code requirements, competition-specific and foundational rules, file listing, public leaderboard, discussions listing and public notebook listing. Public notebook source files were downloaded and inspected; their claimed performance was not independently reproduced.
Original 16 September inspection snapshot — historical, not current execution status
| Artifact | Observed size / contents |
|---|---|
train.parquet | Listed at 3,033,286,496 bytes; not downloaded or fully profiled in this review |
test.parquet | Downloaded; 4,848,729 bytes, 1,213 spectra, 400 molecules, 12 columns |
sample_submission.csv | Downloaded; 43,619 bytes, 400 rows; IDs exactly match visible test IDs |
| Visible spectra per molecule | Minimum 1, median 3, maximum 9 |
| Visible instrument | All timsTOF |
| Visible adducts | 7 observed; the hidden-test description lists 10 |
| Local RDKit | 2022.09.5, not the scoring version |
Critical distinction: the downloadable test consists of training examples and is replaced by hidden data during scoring. It is useful for schema/runtime tests, not a generalization benchmark. Never precompute predictions or build the final candidate index only around its IDs, formulas or mass windows.
The official hidden-test description is approximately 1,500 spectra / 400 molecules, 1–16 spectra per molecule, neutral masses 157–1,159 Da. These are descriptions, not permission to hardcode dimensions.
Data-source implications
Official training description: approximately 2.5 million spectra and 275,810 unique structures. Major sources include:
enveda-180: 1,153,785 spectra, 182,941 structures; instrument-matched, but mostly synthetic drug-like compounds.pluskal_ms2: 527,581 spectra; Orbitrap, different collision-energy conventions.riken, gnps, massbank, mona, spectraverse, msdial: important natural-product coverage, heterogeneous acquisition/curation.enveda-np-examples: 1,151 spectra / 250 common natural products, measured with the target instrument/pipeline. Particularly valuable for domain-matched validation.Data release follow-up: the 18 September redownload comparison matched the retained training file in SHA-256, bytes, schema, 2,539,608 rows and 21 row groups (reports/todo-review-20260918/kaggle-redownload-comparison-20260918.json). The verified local NP-example count is 1,184 rows / 250 identities, not the original description's 1,151 rows. The authoritative rebuilt namespace contains 274,195 scorer identities; raw 275,810-InChIKey grouping is not interchangeable. Pin these artifacts for each experiment; this is not a perpetual freshness claim.
Documentation inconsistency: the data page says “train.parquet (18 columns)” but its listed shared/additional fields imply a different count; it also says “seven” in one adduct description while listing ten elsewhere. Inspect the actual train schema and support all ten advertised hidden-test adducts.
---
4. Optimize the actual metric
For U molecules:
MRR@25 = mean(1 / first_correct_rank)
A missing answer or answer below rank 25 contributes zero.
| Change for one molecule | Gain in reciprocal rank |
|---|---|
| Rank 2 → rank 1 | +0.500 |
| Rank 5 → rank 1 | +0.800 |
| Not found → rank 25 | +0.040 |
| Not found → rank 1 | +1.000 |
Therefore:
1. Prioritize reliable top-1 decisions while expanding candidate coverage.
2. Keep up to 25 distinct scoring identities; stereoisomers and tautomer spellings should not waste slots.
3. Track candidate recall separately from ranking quality. A ranker cannot rescue a missing answer.
4. Do not use Tanimoto similarity, SMILES validity, or formula correctness as substitutes for the official exact-connectivity score. They are diagnostics and training auxiliaries only.
5. A calibrated estimate of candidate correctness should drive final order. Do not impose arbitrary quotas such as “10 retrieved, 10 database, 5 generated” irrespective of evidence.
6. Treat spectra as correlated observations of one molecule. Six energies do not create six independent labels or six votes of equal reliability.
Diagnostic decomposition: if candidate-pool coverage is R and MRR conditional on coverage is Q, total MRR is R × Q. Report both, by novelty regime, to decide whether to improve generation or ranking. In the end-to-end metric, an absent answer contributes zero; in diagnostics, distinguish missing structure, missing eligible reference, filtering loss and wrong rank.
Beam width is a search budget, not the number of valid distinct scoring identities returned and not a direct MRR objective. Token likelihood, fingerprint loss, Tanimoto and ranker classification loss are training/ranking surrogates; promote them only through the official molecule-level score.
---
5. Validation design — the most important investment
5.1 Reproduce the scorer first
Create a dedicated environment with RDKit 2026.03.3 (package version spelling may appear as 2026.3.3). Do not change the global installation used by other projects.
Implement the documented matching operation: parse SMILES → tautomer canonicalize → InChIKey14. Verify any available official implementation before freezing it. Do not silently add salt stripping, uncharging, fragment selection or other standardization steps not specified by the evaluator.
Unit tests must cover:
Use one identity function for labels, split grouping, candidates, evaluator and submission deduplication.
5.2 Three explicit evaluation regimes
A — spectral-library available
Use held-out target-domain spectra, particularly enveda-np-examples, as queries. Remove these exact query acquisitions and any duplicate/reprocessed versions from every learned training corpus and reference index. Allow independently acquired reference spectra of the same compound in other libraries: this deliberately models Class 1.
B — structure available, spectra unavailable
For validation identities, remove all their spectra across all libraries from supervised training, spectral retrieval, analogue prototypes and derived model features. Their structures may remain in the frozen permitted structure database. Report results both for naturally database-covered molecules and for the full cohort; do not inject validation answers into the pool.
C — structure unavailable / de novo proxy
Remove validation identities from spectral sources, candidate databases and supervised structure-target training. Add a scaffold-disjoint split. Audit our own structure-pretraining and checkpoint provenance; if overlap cannot be ruled out, label the result as potentially contaminated rather than a clean novelty estimate. External pretrained models are excluded.
A withheld common compound is only a proxy for genuinely novel natural-product chemistry. Do not claim the proxy exactly represents Class 3.
5.3 Split protocol
5.4 Required experiment report
Every candidate change must report:
MRR@25 | Top-1 | Top-5 | Recall@25 | Candidate Recall@100/1000 | Pool coverage | MRR conditional on coverage | runtime | peak RAM/VRAM
Break down by A/B/C, molecular mass, adduct/polarity, spectrum count, spectrum quality, natural-product family/scaffold, and closest-training similarity. Record unparseable structures and wrong/missing formula cases.
Promotion rule: require reproducible grouped-validation improvement, no unexplained severe regression in another regime, and acceptable runtime. Prefer paired-bootstrap evidence; small gains with wide uncertainty remain provisional. A leaderboard bump alone is insufficient. Preserve per-molecule ranks, unique candidate membership and stage-level attrition so statistical comparisons can be recomputed rather than inferred from rounded aggregate logs.
5.5 Current evidence qualifications identified on 22 September
The [local evidence review](reports/plan-review-20260922-analog/local-evidence-review.md) found semantic defects despite matching file hashes:
Keep the correct all-spectrum evaluator's historical results distinct from the later ablation defects. Strengthen the post-run auditor's exact cohort/mask/row-identity and metric-drift checks before fresh G1 promotion; an audit-scope gap is not proof that a saved run leaked or had score drift.
---
6. Data and chemistry foundation
6.1 Stream, do not materialize everything
Read Parquet metadata and selected columns first; stream row groups. Store compact peak arrays and offsets, canonical identity tables, sorted neutral-mass indexes and reproducible molecule-level split manifests.
The local host has roughly 11 GiB RAM and 36 GiB free disk at review time. A 3 GB Parquet file can expand far beyond 3 GB in pandas. Full PubChem ingestion is not an appropriate first local step.
6.2 Preserve raw information
Keep original arrays plus named preprocessing variants. Evaluate intensity floors of 0.1%, 1%, 2%; square-root/log transforms; top-64/128/256 peaks; and windowed peak selection. Compare entropy, cosine and unfiltered/soft-filtered variants. Do not discard low-quality spectra outright when they may be the only evidence for a molecule.
Account for training precursor_error_ppm, metadata missingness, source quality and duplicate spectra. Avoid indiscriminately dropping one million instrument-matched spectra; use balanced batches or source weights to counter synthetic-chemistry dominance.
Treat removal of enveda-180 as an ablation, not evidence that all its spectra are low quality. Compare source-balanced inclusion, downweighting and exclusion on identical scorer-identity splits. Preserve acquisition metadata and evaluate per-spectrum prediction, compatible-spectrum fusion and late evidence pooling as distinct approaches; an encoder trained on single spectra should not silently receive merged inputs at inference.
6.3 Adduct-aware mass inference
Support all advertised adducts:
[M+H]+, [M+NH4]+, [M-H2O+H]+, [M-2H2O+H]+, [M+Na]+, [M+K]+, [M-H]-, [M-H2O-H]-, [M+CH2O2-H]-, [M+Cl]-.
Use exact ionic/adduct mass constants and explicit charge/multimer handling for additional training adducts. Test round trips from a known molecular mass through every supported adduct. Dehydrated adducts must add the lost water back when inferring the neutral molecule.
Combine neutral-mass estimates across spectra using robust, quality-weighted inference. Preserve alternative hypotheses when spectra disagree. Calibrate a ppm tolerance with an absolute floor; do not adopt a fixed 0.01 Da window as unquestioned truth.
Test the neutral-atom versus ionic/electron-mass convention explicitly, including sodium, potassium and dehydrated ions. Compare adduct-aware precursor/isotope-peak exclusion against the current cleaning policy before spectral matching. Record answer retention before/after mass filtering, identity deduplication and candidate caps. Filter eligible reference rows before collapsing identities, and restore score ordering after any coarse selection; query-specific mass compatibility cannot be inferred from a single globally best spectrum of that identity.
Separately compare reference masses derived from eligible known structure/formula metadata against reference precursor-derived masses (CHEM-02), retaining raw precursor residuals and source tags. This is a reference-index ablation, not permission to use unavailable query truth formulas. Test formula parser coverage and charge/multimer/unknown-adduct behavior; keep query mass inference based on its permitted observations.
6.4 Formula inference
Enumerate/rank plausible formulas using calibrated precursor mass, allowed elements, valence/DBE checks, fragment subformula evidence and cross-adduct consistency. Retain several formulas when uncertain. Start with permissive chemistry constraints and measure true-formula recall before narrowing.
MIST-CF and SIRIUS are candidates for comparison, not assumed plug-and-play dependencies. Verify license, offline execution, model assets, instrument support and runtime first. Do not assume all SIRIUS services are free or callable offline.
Use collision_energy_ev as the cross-source input, with missingness and original-unit/source indicators. NCE-to-eV conversions are approximate, not interchangeable ground truth.
---
7. Modeling workstreams, in priority order
Track A — High-quality spectral retrieval [P0]
Implementation
1. Retrieve mass-compatible reference structures and their spectra.
2. Compare direct fragment cosine, entropy similarity, neutral-loss similarity and carefully adduct-aware modified cosine.
3. Evaluate polarity/adduct compatibility explicitly; do not merge raw spectra from incompatible ions into a single peak list.
4. Aggregate by canonical structure and query molecule: best/second-best match, support across independent acquisitions, collision-energy agreement, mass residual, explained intensity, unmatched intense peaks, source quality and ambiguity margins.
5. Correct for reference-count bias: a structure with many library spectra should not win merely from more chances to match.
6. Train an OOF molecule-grouped candidate ranker, initially LightGBM/CatBoost or a compact pairwise/listwise model. Use reciprocal-rank-aware evaluation and tune top-rank behavior; NDCG training objectives are surrogates, not the competition metric.
Ablations: max pooling vs consensus; fragment-only vs loss-only vs combined; entropy vs cosine; same-adduct-only vs compatible cross-adduct; learned vs heuristic aggregation; single vs multiple preprocessing views.
Deliverable: deterministic CPU baseline, indexed retrieval artifacts, OOF ranked candidate tables and validated submission builder.
Track B — Structures without measured spectra [P1]
Candidate universe
Begin with permitted competition training structures only. COCONUT acquisition and the quarantined PubChem rebuild are deferred; revisit only after measured missing-structure failures and rights/resource review. Broader PubChem ingestion is deferred, not an automatic next step. COCONUT is attractive for natural products but cannot be assumed to cover all analogues or synthetic-looking test compounds. The author's 99.6% NP-example coverage and 100% validation mass-window recall are dataset/protocol-specific, not measured hidden-test coverage.
Build full mass/formula indexes independent of visible test IDs. Dynamic filtering by the hidden query masses is fine; precomputing only the placeholder's mass windows is not.
For a proposed expansion, measure both membership gain and ranking dilution on independently selected full-cohort queries, together with bytes, build time and inference time. The heuristic new_coverage × retained_ranking_quality > old_coverage is a sensitivity analysis, not a substitute for a validation design that can observe missing answers. Keep preprocessing such as salt stripping or largest-fragment selection separate from the official scorer identity; quantify any answer loss it causes.
Learned retrieval
Compare:
Current source note: precursor-m/z conditioning already exists in the baseline path. Future learned retrieval work should extend it with adduct, collision-energy/missingness and predicted-formula conditioning; it should not treat precursor input as absent.
Literature-derived technical references, not dependencies:
Pool multiple spectra at the molecular level using reliability-aware attention or a learned set aggregator. Model missing metadata explicitly. Train/finetune with source balance and natural-product weighting, selected by clean validation.
First learned experiment — fingerprint ranking, not a second SMILES decoder
1. Establish a binary Morgan fingerprint baseline and compare a frozen, documented multi-family fingerprint union; choose retained bits without report-fold labels. Store bit definitions, ordering, fingerprinting version and candidate identity mappings with the model.
2. Predict fingerprint logits z from the spectrum using multi-label BCE. For a binary candidate fingerprint f, compare the score f · z: under an independent-Bernoulli model, it is the candidate log-likelihood up to a query-only constant. It is not itself a calibrated probability of exact correctness; count fingerprints require a different model.
3. Add same-mass hard-negative ranking loss, then separately test same-formula isomers: BCE + λ × contrastive cross-entropy, with temperature/weight chosen on development data. Deduplicate scorer-equivalent positives before sampling negatives and prevent held-out spectra/targets from entering training. Report BCE-only and combined-loss results on the same frozen pool.
4. Compare single-spectrum, compatible multi-spectrum training and late pooling. Condition on adduct, polarity, precursor/neutral-mass hypotheses, instrument and collision-energy availability; do not equate unknown metadata to a true zero.
5. Score every retained candidate or verify shortlist recall before truncation. Save OOF logits, ranks and per-query runtime so a later ranker cannot unknowingly fit in-sample neural features.
Complementary experiment — mass-shifted analogue propagation
Find eligible reference neighbours by direct and precursor-mass-shifted fragment evidence, then score candidate structures by their fingerprint relationship to those neighbours. This is structural evidence transfer, not copying the neighbour's SMILES as the answer. Test direct-only, shifted-only and combined matching; freeze neighbour count, similarity powers and reference-count handling on development folds. Track reference identity/source row, adduct/charge/energy compatibility, candidate identity and score components. Exclude every held-out identity's spectra from all neighbour/prototype stores in B/C.
DreaMS/MIST-family papers may inform independently written designs, but external encoders, frozen representations and pretrained checkpoints are excluded by user policy. JESTR-01 is the scheduled representation challenger.
Forward consistency reranking
For a manageable top-K shortlist, predict spectra from candidate structures using an independently implemented forward model trained from scratch on authorized inputs; GLACIER is the shortlist reference. ICEBERG/SCARF/CFM-ID-family papers may guide the design, but their code, checkpoints and data are not approved by citation alone. Score across actual adducts/energies. Account for cross-instrument errors; a forward model's low score must not automatically veto a strong measured reference.
First test a bounded bond-fragment explanation score as an independent feature, not a universal veto. Cap fragmentation/search work and preserve isotope, charge, hydrogen-rearrangement and unsupported-element semantics. Then consider a spectrum–candidate cross-encoder over a high-recall shortlist, with same-formula hard negatives and OOF predictions. Measure shortlist recall separately: an expensive second-stage model cannot recover a discarded candidate. Keep a no-cross-encoder control to test whether any gain survives latency and regime stress tests.
Gate: demonstrate B-regime coverage and MRR gains beyond neighbour-fingerprint transfer. If pool coverage is high but MRR is low, improve ranking before enlarging databases further.
Track C — De novo and analogue editing [P1/P2]
Start with the official spectrum-encoder/SMILES-decoder tutorial as an engineering reference, not the final architecture.
The existing Phase 4 experiment is a feasibility branch; a running or completed job does not certify it. Before further long training, bind the exact remote source to checkpoint/tokenizer/runtime and pass the bounded GPU sanity gate: authoritative cache load with zero misses, shifted teacher-forced targets, causal masking in both training and generation, position/padding behavior, EOS termination, BOS/PAD suppression, optimizer accumulation/remainder handling and generated-sample inspection. Test next-token prediction rather than same-token copying. Do not reuse copy-objective checkpoints as evidence of useful chemistry learning.
Evaluate outputs with pinned RDKit and the official tautomer-compatible identity; syntax-only checks, string equality and uncanonicalized InChIKey14 are not interchangeable with that score. Evaluate unique top-25 candidates at the molecule level over all permitted query spectra, not only greedy top-1 from one spectrum. Repair raw-identity split/exposure issues before calling a result clean B/C validation.
Improve in this order:
1. Extend existing precursor-m/z conditioning with adduct, collision-energy/missingness and predicted formula conditioning.
2. Learn from sets of spectra for one molecule, rather than treating each acquisition independently.
3. Train a bounded independently implemented decoder from scratch; only adapt our own rights/provenance-bound checkpoints before scaling.
4. Generate chemically valid candidates with beam search and/or controlled sampling; compare tokenizations and randomized-SMILES augmentation.
5. Add analogue editing: modify close spectral neighbours using plausible mass differences and functional-group transformations; score all outputs automatically.
6. Apply formula/mass consistency, exact scoring-identity deduplication and shared forward/spectral reranking.
7. Use decoder scores normalized/calibrated across length, generation route and spectrum count. Raw log probabilities from different models are not comparable.
After the fingerprint baseline is useful, compare fingerprint-conditioned decoding with the unconstrained baseline. Train the generator with realistic predicted/noisy fingerprints, not only perfect target fingerprints it will not receive at inference. MS-BART, MARLIN, DiffMS and MADGEN are technical references for independently implemented experiments only; do not copy their code or use their checkpoints/data. Test soft formula constraints, valence-aware decoding and eligible scaffold conditioning; ground-truth formula/scaffold inputs are oracle controls unless available through a permitted inference path. MIST/MSNovelist assets and all external weights remain excluded. The shortlist describes MSFlow trained assets as non-commercial; they are excluded. It does not establish withdrawal.
Keep the formula filter soft when formula confidence is low. Do not spend weeks optimizing validity while exact structure recall remains zero. Compare meaningful top-K generation recall on the C proxy and contribution after ensemble reranking.
Compute policy: expand beam width or sampling only when it adds unique correct candidates or improves MRR under the offline budget. Generate more candidates for uncertain queries, but reserve a small exploration budget so an overconfident retrieval gate cannot completely suppress the correct route.
Track D — Unified ranking and uncertainty [P1]
Union all unique candidates, attach source and chemistry evidence, and train a molecule-grouped ranker with OOF features.
Inputs include retrieval agreement, fingerprint/embedding similarity, formula posterior, mass residual, forward-spectrum fit, generator likelihood, spectrum quality and calibrated confidence. Build a calibrated OOF ranker against simple and static-fusion controls rather than importing the notebook's hand weights or simulation mixture.
Leakage guard: exclude candidate source labels and database/natural-product provenance priors by default. If all positive answers in a simulation come from the training-structure half of the pool, these features reveal the sampling procedure instead of chemistry. Candidate structure availability can be allowed in B, while its supervised spectral evidence cannot. Fit feature transformations, fingerprints, model/calibrator choices and mixture weights within the declared training/development contract; reject external ranker archives/pretrained checkpoints; bind our own OOF features and from-scratch checkpoints before reuse.
Version the candidate pool, pruning policy, channel set and feature schema together. Pruning alters candidate-relative ranks/z-scores: recompute/refit these features or prove compatibility instead of reusing an uncapped training archive. Verify advertised knobs reach active code; the reviewed notebook's CFG.SIM_POWER/CFG.W1 are unused and its nominal fragmentation cap actually affects all downstream channels (FUSE-02).
Do not train a hard novelty-class classifier without reliable class labels. Use calibrated evidence-based routing, with novelty proxies tested out of distribution.
Return up to 25 candidates ranked by estimated correctness. Deduplicate before truncation. Prefer fewer credible candidates to filling with repeated CCO; an emergency valid output is a runtime safeguard, not a scientifically useful prediction.
---
8. Public baselines: what to borrow and what not to trust
Official de novo tutorial — inversion/casmi-denovo-tutorial-notebook
Inspected: peak encoder + autoregressive SMILES decoder, grouped raw-InChIKey split, a 200k-spectrum training cap, one spectrum per validation structure, and pooled per-molecule generated candidates.
Useful: working model/data/submission wiring and memory-aware batching.
Limitations to fix: excludes enveda-180 in this inspected version; single-spectrum validation is not the full task; output deduplication uses raw InChIKey14 rather than the stated tautomer-canonicalized identity. Do not inherit its scorer assumptions or placeholder-ID handling blindly.
Fast spectral baseline — haideptry/enveda-casmi-2026-fast-spectral-cosine-baseline
The downloaded version describes a richer pipeline than its title: direct/neutral-loss matching, multi-spectrum consensus, COCONUT chemistry heuristics and confidence gating.
Useful: efficient mass-window retrieval and complementary spectral evidence.
Limitations: fixed hand-tuned blends/thresholds and heuristic candidate priors require clean ablation. Notebook narrative scores are author claims, not reproduced results. Inspect and test active code paths before reuse; popularity is not a correctness test.
Evidence ranking — octaviograu/enveda-molecule-evidence-ranking
Inspected: entropy/cosine, mass-compatible database candidates, neighbour-transferred fingerprints, multi-spectrum evidence, target-domain validation, canonical identity handling and learned-ranking support.
Useful: closest starting design for the recommended baseline. It explicitly holds out NP-example acquisitions and separates spectral-reference evidence from analogue inference.
Limits: fingerprint transfer is not a trained spectral foundation model. Audit fold exclusions and external-asset attachments; downloaded metadata is not proof that every dependency is available offline. The notebook itself states it does not use DreaMS/MIST or forward-spectrum models.
Rule: use architecture descriptions as technical references only after review, independently implement the published concepts without copying source. Retain attribution. Do not paste large unverified notebooks together and call the result an ensemble.
Analog propagation — prvsiyan/analog-propagation-casmi-2026-baseline (22 September review)
Evidence, source provenance and precise code anchors are retained in [the technical notebook review](reports/plan-review-20260922-analog/notebook-review.md). No cells or trained assets from this notebook were executed by this review.
Borrow as testable designs: a bounded training-structure/COCONUT pool; direct and mass-shifted reference evidence; spectrum-to-fingerprint BCE plus same-mass hard-negative ranking; candidate bond-fragment explanation; and seeded, calibrated evidence fusion. Prioritize hard same-formula negatives and molecule-level multi-spectrum support rather than another unconditional expansion of the catalogue.
Author-reported results, not our measurements: the narrative includes approximately 711k pool structures, fingerprint-only MRR about 0.517, analogue MRR about 0.52 and combined MRR about 0.63 on its simulations. It also reports public scores and same-code rerun differences. These numbers do not establish superiority over our system because datasets, exclusions, metric contracts, random seeds and runtime paths differ. In particular, our approximately 0.87 library-ceiling diagnostic is not a competing B-regime score.
Do not inherit blindly: raw identity/dedup assumptions, candidate restrictions, shipped training-feature archives, source-dependent priors, adduct compatibility, inferred test mixtures, 10-ppm settings, model/asset availability or unverified license statements. Inspect active code against the narrative and require independent A/B/C accounting. A perfect in-pool holdout cannot settle the decision to expand an external catalogue; a frozen-embedding DreaMS result is not evidence about all possible fine-tuning uses.
Two specific negative results worth learning from: cell C031 reports that substituting DreaMS embedding cosine for mass-shifted entropy in a subsampled analogue benchmark scored 0.1129 versus 0.1555 at 130 queries; this was not a fingerprint-head fine-tuning test. Cell C032's unsuccessful (predicted logits, candidate fingerprint) interaction heads did not directly see peaks. A genuine candidate–peak cross-encoder is therefore a separate hypothesis, not a reproduced improvement; use OOF upstream predictions to avoid training the head on unrealistically accurate in-sample logits. The source snapshot contains no saved executed outputs, so these numerical comparisons remain author reports.
Changes adopted in this plan: FP-01/02/03, AN-01/02, ISO-01, FUSE-01 and RERANK-01 in TODO.md, gated by VAL-01/02, BASE-01/02, RIGHTS-01 and offline packaging. Generation remains a measured complementary branch; the notebook's rhetorical claim that it is the wrong tool is not adopted as a scientific conclusion.
---
9. Calendar and exit gates
The dated table below is the original 16 September target calendar, retained for traceability, not a current phase launch order or a claim of completion. The 22 September execution ladder below supersedes its sequential Phase 2/3/4 assumptions. Competition deadlines are unchanged.
| Phase | Target window | Deliverables | Exit gate |
|---|---|---|---|
| 0 — Ground truth and infrastructure | Sep 16–18 | Current train manifest; schema/source audit; pinned scorer; split spec; offline dependency smoke test | G0: evaluator tests pass, data version fixed, no unsupported hidden adducts |
| 1 — Credible retrieval baseline | Sep 19–25 | Mass-indexed cosine/entropy/loss retrieval; molecule pooling; validation report; first notebook candidate | G1: reproducible scores and complete offline output; no placeholder shortcuts |
| 2 — Ranking and domain adaptation | Sep 26–Oct 9 | OOF ranker; source-balanced models; uncertainty and failure dashboard | G2: measured improvement over simple retrieval with molecule-level uncertainty |
| 3 — Known structure / no spectrum | Oct 10–30 | COCONUT + broader structure index; fingerprint/embedding retrieval; shortlist forward reranker | G3: B-regime gains; explicit coverage/ranking decomposition |
| 4 — Novel structure capability | Oct 31–Nov 20 | Formula-conditioned generator + analogue editing; clean novelty-proxy audit | G4: nontrivial unique correct candidates and ensemble MRR gain within budget |
| 5 — Integrated ensemble | Nov 21–Dec 4 | Unified ranker; complementary models; ablations and realistic mixture stress tests | G5: selected robust system beats baseline across intended regimes |
| 6 — Freeze and final delivery | Dec 5–12 | Two finalist notebooks; clean-room reruns; license/provenance/reproduction package | G6: two successful offline reruns per finalist, all artifacts fixed and recoverable |
| Buffer | Dec 13–14 UTC | Final verification and explicit final-selection check | No last-minute unvalidated architecture changes |
Historical 22 September ladder — superseded by the authoritative 24 September order above:
| Next deliverable | Existing phases / TODO IDs | Acceptance before expansion |
|---|---|---|
| Trustworthy comparison harness | 0/1: VAL-01/02, ISO-01, BASE-01/02 | Canonical identity/exposure accounting, semantic fixtures, reproducible per-molecule evidence |
| Bounded candidate and rights contract | 3A/3B: POOL-01, RIGHTS-01 | Independently measured coverage; licensed/pinned assets; non-neural baseline |
| First learned and analogue comparison | 2/3C: FP-01/02/03, AN-01/02 | Same-pool OOF MRR/recall improvement, hard-isomer and domain checks, measured compute |
| Calibrated complementary evidence | 5 plus RERANK-01, FUSE-01, REPRO-01 | Improvement over the best single channel; shortlist retention; repeatability and latency |
| Generator qualification/extension | 4: GEN-01/02/03/04 | Correct learning objective, official scorer, clean candidate gains and no fixed slot quotas |
Parallelism: after common validation/pool gates, fingerprint and analogue experiments may run independently; an already-authorized generator experiment can be inspected without blocking them. This review does not cancel or relaunch that experiment. A failed de novo gate prevents claims of validated novelty coverage, not shipment of a stronger validated ranking system.
Original first-72-hours acceptance checklist — status tracked in TODO.md
The items below describe the original setup requirements; unchecked boxes are not a current completion tally. Consult the dated evidence and gates in TODO.md rather than repeating completed acquisition or audits.
---
10. Runtime, resources and spending discipline
Local host
Use the current ARM server for orchestration, code/document review, manifests, bounded metadata work and report generation. CC's standing constraint: no model training on this host, including a "smoke" training loop. Put training and heavy model validation/inference on approved Kaggle GPU resources. Keep this project isolated from existing workloads; do not replace global Python/RDKit packages or consume all free disk. Lightweight contract tests do not justify loading the full Parquet table.
Kaggle / GPU work
Training should normally occur separately, with versioned weights attached to the inference notebook. The official tutorial trains inline for illustration; the final 9-hour window is better used for inference/reranking.
Before committing to hardware, check actual accelerator availability, memory, weekly quota and remaining account capacity. Access to the CLI does not guarantee unlimited GPUs.
Proposed benchmark budgets, not promises:
Instrument wall time by stage. Stop unproductive branches when candidate recall and MRR plateau. Do not spend most of the project on large-scale pretraining; prioritize compact from-scratch models and measured ranking gains. External pretrained representations are excluded.
Offline packaging checklist
molecule_id values from the runtime test file; sample submission is only a format cross-check./kaggle/working/submission.csv.External assets and commercial-use review
Record separate license/accessibility evidence for paper text, code, pretrained weights, input datasets and generated/distributed artifacts. CC requires attention to commercial-use permissions; prefer explicitly commercial-use-permitting code and assets, with attribution and redistribution obligations retained. Public downloadability, a preprint's CC-BY notice or an MIT repository does not automatically authorize every linked model or training corpus. Copyleft is not synonymous with non-commercial, but compatibility/redistribution obligations still require review.
For the 23 September arXiv shortlist, the default disposition is reference-only / no asset reuse. Independent reimplementation is allowed as a research direction, subject to patent, competition-rule, library and data review. No paper listed in research/arxiv-msms-2d-shortlist-20260923.md is pre-approved as a source of code, weights, tokenizer, dataset or checkpoint. A paper or asset described as non-commercial, withdrawn, unavailable, or ambiguously licensed is excluded from a commercial-compatible candidate pipeline.
The captured competition-specific rules (research/pages.json, sections 5–6) name MIT winner licensing, require commercial-use-permitting submission code/model terms, and include separate exceptions for third-party software and incompatible-license input data/pretrained models. Those exceptions are not blanket permission to ignore the third party's actual usage license or the host's accessibility requirements. Competition data itself remains governed by its competition/CC-BY-NC terms. Do not claim that a prize competition is automatically an academic-use exemption or that a published algorithm is unpatented. Ambiguous external assets require documented clarification before adoption.
Keep competition-derived structures, fingerprints and training archives access-controlled; do not redistribute them to nonparticipants in a public asset. A public notebook's attached asset is not automatically reusable. Build permitted training-derived features inside the competition-access environment where necessary, and record exact source/checkpoint hashes for reproducibility.
---
11. Submission and final-selection policy
1. Do not submit automatically as part of this planning task.
2. Once execution is approved, first establish a valid offline notebook submission.
3. Treat the 5/day limit as a ceiling, not a target. Submit only versions with a recorded hypothesis and local evidence.
4. Change one major component at a time and retain immutable source/weights/configs.
5. Do not hand-label hidden cases, infer answers through leaderboard probing or use hardcoded test predictions.
6. Compare public score changes with validation-regime changes. Large disagreement triggers investigation, not immediate retuning to the leaderboard.
7. Preserve two final candidates: the best broadly robust model and, only if justified, a complementary model with different novelty/coverage tradeoffs. Do not sacrifice a clearly stronger second model merely for cosmetic diversity.
8. Explicitly verify final selections before the deadline; do not rely on automatic selection.
---
12. Risks, mitigations and decision triggers
| Risk | Mitigation / decision |
|---|---|
| Same compound leaks across libraries | Canonical-identity exclusions across spectra, labels, analogue indexes and supervised training |
| Raw identity split disagrees with scorer grouping | Audit raw→scorer mapping and exposure; never equate deterministic hash parity with leakage safety |
| A perfect in-pool benchmark hides catalogue misses | Independent full-cohort membership test plus separate conditional-ranking panel |
| Ranker learns which dataset supplied the answer | Exclude provenance/NP-prior features by default and audit OOF feature construction |
| Visible test produces deceptively perfect retrieval | Use it only for plumbing; score frozen held-out molecules |
| 250 NP examples overfit through repeated tuning | Nested/grouped validation plus broader scaffold/source holdouts |
| Class mixture unknown | Report multiple mixture scenarios and regime-level performance |
| Formula filter discards the truth | Measure formula/pool recall; retain alternative formulas and fallback paths |
| Common formulas yield huge isomer pools | Hard-negative fingerprint training and shortlist forward reranking |
| De novo gives valid but wrong SMILES | Gate on exact-connectivity recall and ensemble gain, not validity |
| Trivial teacher-forcing shortcut looks like convergence | Shift labels, test causal/prefix behavior, inspect actual generation before long GPU runs |
| Copy of local source differs from executed kernel | Bind remote version/source hash, weights, tokenizer and runtime; do not overwrite launch evidence |
| Spectral consensus amplifies duplicated evidence | Deduplicate acquisitions and cap correlated support |
| Training sources overwhelm natural-product chemistry | Source-balanced sampling and domain-matched validation |
| RDKit version changes identity or ranking | Pin exact scoring version; cache keys with version metadata |
| External model has hidden training overlap | Audit provenance; qualify novelty claims when overlap is unknown |
| External asset licensing/accessibility is unsuitable | Exclude or seek host clarification before relying on it |
| Runtime fails on hidden replacement data | Dynamic query handling, broad indexes and cold offline stress tests |
| Disk/RAM/GPU pressure | Stream data, bounded candidate pools, stage profiling and resource caps |
| Training data changes | Manifest/hash checks; rebuild derived assets and rerun relevant baselines |
Host questions to resolve when needed: exact scorer edge cases beyond the published matching description; any ambiguous external-tool licensing/accessibility; details of the September 15 data correction. Do not assume answers.
---
13. Proposed repository and evidence contract
enveda-casmi26/ EXECUTION-PLAN.md research/ # captured rules, listing snapshots, notebook review data/ # access-controlled; exclude raw data from public git manifests/ # source version, SHA-256, license, schema, split IDs configs/ # experiment and inference configurations src/ identity.py # exact scorer-compatible structure identity adducts.py preprocess.py retrieval.py formula.py embeddings.py generation.py ranking.py evaluate.py submit.py tests/ experiments/ # metrics, OOF predictions, timing and ablations artifacts/ # versioned indexes and weights; not raw-data git blobs notebooks/ # offline training/inference entry points reports/ # decisions, failure categories, final reproductionThis was the proposed layout at the original review. Source, indexes, reports and experimental kernels now exist; their acceptance is tracked separately in TODO.md. Do not treat a directory or checkpoint as evidence that a phase has passed.
Every experiment must record: code commit, configuration, source/checkpoint hashes, licenses, split manifest, seeds, molecule-level predictions, metrics, hardware, wall time and decision. Claims such as “better”, “novel” or “ready to submit” require their corresponding evidence.
Suggested roles (one person may cover several): chemistry/data; retrieval/ranking; generative modeling; validation; deployment/reproducibility. A teammate with mass-spectrometry expertise may add more value than another generic model sweep, but no team formation is assumed or authorized here.
---
14. Evidence and references
Official pages were retrieved via authenticated CLI on 16 September 2026. Local raw snapshot: research/pages.json.
1. [Competition overview](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/overview)
2. [Evaluation](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/overview/evaluation) — matching, MRR@25 and submission format.
3. [Data](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/data) — novelty classes, placeholder test, sources and schema.
4. [Rules](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/rules) — licenses, data use, team/submission limits, prohibition on hand-labeling evaluation records.
5. [Timeline](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/overview/timeline)
6. [Code requirements](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/overview/code-requirements)
7. [Public leaderboard](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/leaderboard) — snapshot in research/leaderboard.json.
8. [Host data-update discussion](https://www.kaggle.com/competitions/enveda-CASMI26-molecule-id-mass-spectra/discussion/741471) — title/date verified; body not retrieved.
9. [Official de novo tutorial](https://www.kaggle.com/code/inversion/casmi-denovo-tutorial-notebook)
10. [Fast spectral baseline](https://www.kaggle.com/code/haideptry/enveda-casmi-2026-fast-spectral-cosine-baseline)
11. [Molecule evidence ranking](https://www.kaggle.com/code/octaviograu/enveda-molecule-evidence-ranking)
12. [Spectral entropy similarity paper](https://doi.org/10.1038/s41592-021-01331-z) — methodological follow-up referenced by the public evidence-ranking notebook; not independently evaluated here.
13. [COCONUT downloads](https://coconut.naturalproducts.net/download) — proposed external structure source; verify chosen snapshot and license before use.
14. [DreaMS resource cited by the host](https://zenodo.org/records/10997887) — historical reference only; external pretrained representation route excluded by user policy.
15. [Analog propagation baseline](https://www.kaggle.com/code/prvsiyan/analog-propagation-casmi-2026-baseline) — reviewed source/narrative; external performance remains author-reported.
16. [22 September review and decision record](reports/plan-review-20260922-analog/REVIEW.md) — adopted priorities, source snapshot links, evidence qualifications and verification.
Review boundaries
The original 16 September review did not train models or reproduce public performance. This 22 September revision reads source, metadata and saved evidence and edits documentation only. It does not execute public notebook cells, train models locally or remotely, push/stop kernels, expand databases, change automations or submit predictions. Previously completed experiments are not rerun merely to update their description.
Bottom line: establish a trustworthy molecule-level benchmark and bounded eligible pool, then prioritize learned fingerprint ranking, mass-shifted analogue evidence and calibrated same-formula discrimination. Keep generation where it adds verified value. The opportunity is complementary chemistry-aware evidence across unknown novelty mixtures—not unconditional catalogue growth or another unvalidated large decoder.