Enveda CASMI 2026 β MRR@25 / Top-5 Execution TODO
Updated: 24 September 2026 (MYT)
Goal: maximize official final private-leaderboard MRR@25 and finish top 5. No fixed score guarantees that rank.
Source of strategy: [latest execution plan](EXECUTION-PLAN.md), authoritative 24 September execution order; [paper shortlist](research/arxiv-msms-2d-shortlist-20260923.md).
Status: [ ] pending Β· [-] partial Β· [x] completed within stated scope Β· [!] blocked. Deferred/excluded work is labelled separately.
Success is measured by ranking quality and official resultsβnot checklist completion, training loss, fingerprint accuracy or the number of models built. Promote useful improvements; stop branches that fail to add value.
1. Current position
| Area | Verified scope / next gap |
|---|---|
| Scorer and infrastructure | Official-scorer fixture and qualified local G0 evidence exist; 64 scoped regression tests pass in .venv-rdkit2026; Phase 0 chemistry/manifest validation passes after repairing the archived PubChem audit path; not proof of hidden-test performance |
| Retrieval | Baseline semantics repaired and 64 scoped tests pass; fresh bounded ablation still required; reference-only B failure is not a learned structure-ranking result |
| FP-01 | Independent kernel implemented with saved local compile/schema checks; static handoff audit passes (kernel/metadata/contract/baseline hashes, GPU/no-internet metadata, scorer identity, hard-negative identity exclusion, checkpoint/OOF provenance); no verified trained checkpoint, OOF results or model-quality metrics |
| Remote launch | Saved quota check: 0.00h remaining; refresh 26 September 2026 00:00 UTC = Saturday 26 September, 08:00 MYT, not next week. Recheck availability at refresh; no fresh remote check in this review |
| Validation | Existing ISO-01 panel is diagnostic. Prior fold-0 exposure and 164 fold-4 overlaps prohibit an untouched-report claim; historical G1-05 evidence reproduces its reported arithmetic and found 0 current-target candidate matches under its limited audit, not full identity-wide fold/mask exclusion; AUDIT-01 already passed exact identity/mask joins for the saved G1 run; new learned routes still require their own exclusion verification |
| Candidate universe | Competition-training structures only; PubChem quarantined/rebuild deferred; COCONUT acquisition deferred |
| Leaderboard | Current official score/rank not verified in this cleanup; establish an account/submission snapshot under RUN-01 |
Evidence: [local EVAL/VAL readiness](reports/local-eval-val-readiness-20260924.json), [Phase 0 validation](reports/phase0-validation-20260924.json), [FP implementation](reports/fp-01-implementation-20260924.json), [local validation](reports/fp-01-local-validation-20260924.json), [local tests](reports/fp-01-local-preparation-tests-20260924.json), [package provenance](reports/fp-01-package-provenance-20260924.json), [latest launch check](reports/fp-01-launch-check-20260924-0625.json), [exposure audit](reports/val-01-02-exposure-20260923.json).
2. P0 β Unlock the first measured candidate
Do these first. RUN-01 blocks remote execution, not useful local preparation. Do not wait for every research branch before a qualified development submission.
2026-09-26T00:00:00Z; no push or launch attempted this turn. The cache-bound Phase 4 sanity kernel is statically ready (reports/phase4-sanity-launch-audit-20260924.json) but remains unexecuted. After refresh, verify account access and exact owner/slug, push only the bounded sanity kernel, inspect its checkpoint/metrics, then decide whether to scale. Avoid duplicate launches and blind retries. Exit: an accessible, identified run and actionable compute status.reports/eval-val-consistency-audit-20260924.json. Exit: reproducible same-query/same-pool comparisons, with diagnostics distinguished from promotion evidence.reports/base02-retrieval-semantics-repair-20260924.json. A fresh paired CPU diagnostic on the frozen 246-identity cohort is recorded in reports/base02-analogue-paired-20260924.json: corrected 10-ppm BASE-02 MRR@25 0.95935, Top-1 0.93902, coverage 0.98374. Analogue policies are materially worse and not promoted. These are diagnostic, not official or clean-generalization results.reports/fp-01-handoff-audit-20260924.json; remote execution remains blocked by current Kaggle quota. See reports/fp-01-local-validation-20260924.json. The 19 September G1-08 decision is historical; its then-open malformed-precursor/local-packaging items have later local dispositions in G1-04/G1-07, with maintained regression coverage in A6. Remaining work is immutable launch binding, GPU execution, checkpoint inspection and full G1/C qualification. Exit: completed, inspectable runβnot compilation or an accepted push.reports/pkg-01-local-readiness-20260924.json. The static FP-01 handoff audit also passes. A checkpoint-pending package manifest now binds the FP-01 kernel, metadata, training-pool contract, baseline and handoff audit in manifests/fp-01-package-pending-20260924.json; the origin-aware dependency inventory confirms NumPy/PyArrow/RDKit are from .venv-rdkit2026, while Torch is intentionally absent locally and wheel hashes/license records remain unbound (reports/fp-01-dependency-inventory-20260924.json). Bundle our own trained checkpoint, code, pool/index, bit ordering and rights-cleared pinned dependencies after GPU execution. Exit: reproducible inference artifact, not placeholder accuracy.Critical path: RUN-01 at quota refresh + EVAL-01 β bounded FP-01/02 comparison β PKG-01 β SUB-01. During the quota pause, only safe local preparation/documentation is allowed; no local training or paid compute. Pandas is installed; the full scoped suite now passes 64 tests after retrieval, chemistry, analogue-route, candidate-propagation and metadata-eligibility contract work. Existing code and diagnostics are reusable; external-pool rebuilding is not on this path.
2A. Ready now β non-GPU work, including arXiv ideas
GPU quota blocks training, not reading papers, independent design, CPU chemistry/metrics or interface preparation. Do not repeatedly run the same 64-test suite or rewrite completed audits as substitute progress. These tasks remain pending until their deliverables exist.
| Priority / ID | Task available now | Concrete completion evidence |
|---|
research/lit-02-method-extraction-20260924.md; provenance and decision record: reports/lit-02-method-extraction-20260924.json. No paper code, checkpoint, tokenizer, pretrained representation or external dataset was adopted. JESTR remains an independent design-only challenger; retrieval MRR@25βnot fingerprint similarityβremains the promotion metric.| 2 Β· JESTR-PREP | Prepared and synthetically verified in research/jestr-prep-20260924.md, manifests/jestr-prep-contract-20260924.json and reports/jestr-prep-validation-20260924.json. The contract binds scorer identity, competition-only pool, grouped folds, molecule-level scoring, negative exclusions, deterministic ties and MRR@25 metrics. No training or paper reproduction is claimed. |
|---|---|
| 3 Β· COMP-01 | Scoped audit complete in reports/comp-01-analogue-scope-audit-20260924.json. The historical 246-record P0 report is internally coherent: 5,890 candidate rows, zero candidate IDs missing from the referenced index, zero duplicate candidate rows, 25/246 targets absent from candidate lists (10.16%), and route target ranks present for direct 121, shifted 116 and combined 121. This separates candidate absence from route ranking/abstention; it does not certify the quarantined index, hidden-test generalization or the full 250-identity denominator. |
| 4 Β· BASE-02-CPU | Completed bounded corrected CPU paired replay in reports/base02-analogue-paired-20260924.json on frozen 246-identity cohort: BASE-02 10-ppm control MRR@25 0.95935, Top-1 0.93902, coverage 0.98374; analogue policies 0.226β0.229 MRR@25 and 0.610β0.618 coverage, all ~0.73 MRR below BASE-02. Diagnostic only; no promotion. |
| 5 Β· PKG-LICENSE | Partial local inventory complete in reports/pkg-license-local-inventory-20260924.json. NumPy 2.2.6, PyArrow 25.0.1 and local RDKit 2026.3.3 metadata/notice state are recorded; Torch is not installed in the project venv. Kaggle runtime/ABI, RDKit wheel provenance, CUDA Torch and transitive notices remain unverified. |
| 6 Β· GLACIER-PREP | Prepared and synthetically verified in research/glacier-prep-20260924.md and reports/arxiv-prep-validation-20260924.json. The contract separates candidate recall from reranking, preserves upstream scores on forward-model failure, binds failure strata, and requires paired molecule-level MRR@25 before promotion. No predictor or paper reproduction is claimed. |
| 7 Β· GEN-PREP | Prepared and synthetically verified in research/gen-prep-20260924.md and reports/arxiv-prep-validation-20260924.json. The shared contract covers MS-BART/MARLIN/DiffMS/MADGEN-inspired routes, deployable-vs-oracle inputs, scorer-identity deduplication, failure strata and marginal-MRR gates. No architecture selection, weights or training claim is made. |
Immediate choice: LIT-02 β JESTR-PREP, while COMP-01 corrects the evidence boundary. Reading/design may proceed before FP-01; training, measured comparison and promotion remain sequenced behind the first qualified FP baseline. GLACIER/generation preparation does not authorize expensive training branches.
3. P1 β Improve measured MRR@25
After the initial comparison, choose the next experiment from the observed failure breakdown. One isolated change per comparison; no obligatory Cartesian sweep.
reports/base02-analogue-paired-20260924.json. Reopen only with a justified eligible pool and new hypothesis.scripts/audit_comp01_analogue_scope.py and reports/comp-01-analogue-scope-audit-20260924.json verify the saved reportβs 246-record P0 cohort, 5,890 candidate rows, candidate/index membership, duplicate status, explicit 25/246 target-absence stratum, route abstentions and source hashes. This is a scoped evidence audit, not a fresh run or promotion; the quarantined-index rights and full 250-identity denominator remain unresolved.research/glacier-prep-20260924.md defines deployable inputs, explicit failure statuses, upstream-score preservation, feature boundaries, shortlist-retention gates and molecule-level promotion metrics. reports/arxiv-prep-validation-20260924.json passes its synthetic contract checks. This is design-only and not a paper reproduction or trained predictor.4. P2 β Generation and uncertainty, only when justified
Not prerequisites to SUB-01. Proceed when useful fingerprints and failure analysis justify the compute; measure incremental correct identities and final MRR after shared ranking.
research/gen-prep-20260924.md defines deployable inputs, oracle-only boundaries, MS-BART/MARLIN/DiffMS/MADGEN-inspired route contracts, scorer-identity deduplication, failure strata and marginal-MRR acceptance gates. reports/arxiv-prep-validation-20260924.json passes synthetic checks. Training and architecture selection remain gated by FP-01/02 evidence; no external assets are approved.5. Final leaderboard selection
6. Non-negotiable experiment and promotion contract
Optimize the actual score: MRR@25 = mean(1 / first_correct_rank); missing answers or ranks beyond 25 score zero. Aim to move correct identities toward rank 1 while preserving useful alternatives through rank 25.
| Regime | Eligible evidence |
|---|---|
| A β reference available | Exclude query acquisitions/duplicates; independently eligible same-identity reference spectra may remain. Disclose acquisition-lineage proxy limits |
| B β structure known, spectra unavailable | Candidate structures and deterministic fingerprints may remain. Exclude held-out spectra, supervised target pairs, spectral prototypes and learned/in-sample features from training/reference retrieval across every library |
| C β novelty proxy | Additionally exclude held-out structures/structure targets; enforce the declared scaffold proxy. Do not claim this is true hidden novelty |
Every comparison must record:
Training loss, validity, bit accuracy, in-pool recall and public leaderboard bumps do not individually establish improvement. Placeholder test data are schema/runtime fixtures, not generalization evidence. Do not claim full historical G1/G4 closure from these scoped tasks.
7. Constraints and deferred work
8. Reuse completed work; preserve history
These are recorded scoped completions, not fresh reruns or blanket promotion:
| Reusable result | Evidence |
|---|---|
| Official scorer fixture, RDKit 2026.03.3 | [Certification](reports/official-scorer-certification-20260920/evidence.json) |
| Qualified local G0 | [Decision](reports/g0-local-closure-20260920/decision.json) |
| Saved exact identity/mask audit (AUDIT-01) | [Evidence](reports/audit-01-exact-identity-mask-20260923/evidence.json) |
| Executable retrieval contract (BASE-01) | [Manifest](manifests/retrieval-baseline-resolved-20260923.json) |
| Historical BASE-02 replay β superseded by 24 September semantics repair; not current control metrics | [Results](reports/base02-ablation-corrected-20260923/closure-results.json) |
| Local offline packaging/validator fixtures, not final Kaggle acceptance | [G1-04](reports/g1-04-offline-acceptance-20260920/evidence.json), [G1-07](reports/g1-07-offline-smoke-20260920/evidence.json) |
| FP contract and diagnostic panel | [Contract](manifests/fp-01-training-pool-contract-20260923.json), [Panel](manifests/iso-01-training-diagnostic-panel-20260923.json) |
The complete prior checklist, all 49 completed entries, historical experiments and unresolved legacy caveats are preserved in [archived TODO](reports/todo-cleanup-20260924/TODO.before.md). This file replaces duplicate status ledgers; the archive is evidence, not the current execution order. Old G0βG6 gate names in the execution plan retain their historical scope.
Next action: LIT-02 method extraction and JESTR-PREP are now complete at the documented design scope; qualify COMP-01 and the bounded BASE-02 CPU replay are complete. A private Kaggle dataset jamesl8/casmi26-base02-retrieval-index-20260924 was created and verified ready on 24 September, containing the frozen corrected index for the 10-ppm BASE-02 submission. Do not retry GPU before the recorded 26 September 08:00 MYT refresh. Recheck quota/access then; a bounded Phase 4 engineering sanity check is not evidence that generation should displace FP-01/02. Preserve the objective: higher official MRR@25 and a top-five finish.