Enveda CASMI 2026 Todo
Legend: [ ] pending · [-] in progress · [x] complete · [!] blocked
Reviewed: 2026-09-24 (MYT), against the actual [prvsiyan analog-propagation notebook](https://www.kaggle.com/code/prvsiyan/analog-propagation-casmi-2026-baseline) and the [MS/MS-to-2D-structure arXiv shortlist](research/arxiv-msms-2d-shortlist-20260923.md). New decision/evidence: [plan review](reports/plan-review-20260922-analog/REVIEW.md). Earlier [audit findings](reports/review-20260919-1817/AUDIT.md) and completed experiment records remain historical evidence; they are not erased by this revision.
Current direction: prioritize learned spectrum-to-fingerprint ranking and complementary mass-shifted analogue propagation on a bounded, eligible candidate pool. No open-ended PubChem slice expansion. Retain the already-authorized Phase 4 experiment as an unqualified feasibility branch; do not scale it or reserve fixed output slots for it without measured benefit.
Current execution snapshot — 24 September 2026 (MYT)
kaggle-kernel/fp-01-ranking/main.py; saved implementation report records local compile/schema checks. Training/OOF outputs and model quality remain unverified. See reports/fp-01-implementation-20260924.json.reports/fp-01-launch-check-20260924-0625.json records HTTP 409 and status permission denial; execution is unverified. Prior GPU-quota exhaustion is historical, not reconfirmed by that check.Execution tasks — one authoritative backlog
The phase checklist below owns scientific acceptance. This queue references those IDs rather than creating a second TOP5 completion ledger. Historical gates retain their original names; task implementation, run completion and promotion are separate.
Shortlist-to-experiment map
This mapping uses the supplied local shortlist (abstract-level screening), not a fresh full-paper replication or asset licence certification. Before implementing a new paper-specific mechanism, record its version, exact method section and independently written design under LIT-01. Published benchmark numbers are not expected CASMI scores.
| Technical reference | Independent experiment | Acceptance / limit |
|---|---|---|
| Retrieval objectives, arXiv:2602.16507 | FP-01/02: BCE + f · z versus hard-negative ranking | Same pool/cohort; molecule MRR@25 and paired ranks decide, not bit accuracy |
| MassSpecGym / In the Wild, 2410.23326 / 2606.19624 | EVAL-01, VAL-01/02, AUDIT-01 | Scorer parity, acquisition/identity exclusions, exposure and shortcut audits; no blanket indictment of other papers |
| JESTR, 2411.14464 | JESTR-01: joint spectrum/structure embeddings from scratch | Same frozen pools/folds and comparable budget versus FP-01 |
| GLACIER, 2606.29161 | RERANK-01: forward-spectrum agreement | Retained-answer recall, no-reranker control, OOF features and measured latency |
| MS-BART, 2510.20615 | GEN-03A: structure/fingerprint decoding from scratch, adapted to OOF predicted fingerprints | No MIST weights; structure targets follow B/C exclusions; compare noisy versus perfect/oracle conditioning |
| MARLIN, 2607.04774 | GEN-03B: fingerprint + mass, formula-free generation concept | Bounded challenger; no downloaded weights; measure exact new identities and runtime |
| DiffMS, 2502.09571 | GEN-03C: predicted-formula-constrained graph generation | Ground-truth formula is oracle-only; count formula misses in full-cohort score |
| MADGEN, 2501.01950 | GEN-03D: retrieved scaffold and bounded analogue editing | Predicted scaffold, never held-out truth; separate retrieval failure from generation failure |
| MSFlow, 2602.19912 | Excluded assets; no scheduled experiment | Shortlist describes non-commercial trained version. Withdrawal is not established by this source; remove the earlier unsupported withdrawal claim |
Evidence status: qualified local G0 and later G1 diagnostic audit records exist; full G1/G4 promotion is not established by them. The September 20 five-suite/45-test result is historical, not a test run performed by this review. A seven-row repair shadow and larger rebuilt structure artifacts are different objects. This revision changes documentation only: no training, notebook execution/push, new dataset ingestion, production promotion or automation changes.
Binding implementation policy: papers/notebooks are technical references only; independently implement adopted concepts. Do not copy their implementations or download/use external pretrained weights, checkpoints, tokenizers or training archives—even if their licences permit it. Commercially compatible general-purpose libraries remain allowed after review. Our own checkpoints trained from scratch on authorized inputs may be saved, transferred from Kaggle and packaged. Competition data are used under the competition-specific permission; this does not make the dataset commercially licensed. External data/libraries/tools require documented compatible rights and provenance. PubChem rebuilding and COCONUT acquisition are deferred, not prerequisites to the competition-only route.
Historical material corrections: the scaffold classifier compared identity overlap in its category logic: 1,741 scaffold-overlap changes, not zero (1,295 lost / 446 gained; 7,853 other scaffold-string differences). The registered readout is not certified untouched: 164 selected identities belong to the previously inspected fold 4. Earlier analogue combined_mrr was untruncated reciprocal rank; corrected S4 evidence below supersedes that metric but not the exposure qualification. Keep those Phase 3A readouts diagnostic-only.
Checklist interpretation: task states include dated historical evidence; gate states are separate and are not an effort-weighted competition-readiness percentage. September 22 additions remain pending unless explicitly identified as this completed documentation review.
Historical evening audit: 2026-09-18 source + isolated synthetic probes + saved artifacts. Evidence and pre-audit backup: reports/audit-20260918-1923/. That review left the then-active evaluator and its inputs untouched. Earlier completion claims are qualified below; its process status is not current.
Audit report: [AUDIT.md](reports/audit-20260918-1923/AUDIT.md). Fresh lightweight suite: 24 passed, 2 real-index tests deliberately deselected; additional independent synthetic evaluator/entrypoint/failure-injection evidence is saved alongside it. The earlier 26-pass result is not a claim that all gate scenarios are covered.
The priority order is deliberate: establish trustworthy evaluation and a bounded high-recall pool, then measure learned ranking and complementary evidence. Community performance is author-reported, not our result. This review updates the backlog; it does not mark unrun experiments complete.
Next execution order
1. Recover execution access (RUN-01). Inspect the exact FP-01 owner/slug, account access and remote state before any retry. The latest saved check is HTTP 409 / permission denied, not renewed proof of quota exhaustion. Earlier attempts did confirm the weekly GPU quota. Do not blindly push hourly; resume when access/quota evidence changes. No paid compute or local training is implied.
2. Finish the bounded comparison contract (EVAL-01). Reuse the certified scorer and corrected baseline; bind the competition-only pool, identity-wide training exclusions, candidate membership and predeclared development cohorts. Existing exposed panels support diagnostics, not untouched claims. Freeze a nested/grouped selection/report protocol with prior exposure disclosed before promotion.
3. Run the existing fingerprint kernel (FP-01/02/03). Start with a bounded Kaggle GPU sanity run, then BCE-only versus BCE plus same-mass negatives on identical folds/pools; add same-formula negatives and multi-spectrum variants one change at a time. Evaluate exact molecule MRR@25, not fingerprint accuracy alone. Persist logits, candidate scores/IDs, ranks, checkpoint and run provenance; code presence is not execution evidence.
4. Package the first qualified candidate (SUB-01, PKG-01). Compare against matched retrieval/transfer controls, select using local evidence, verify offline dynamic-ID inference and strict output validation, then submit an evidence-qualified version under existing execution authority. Do not wait for every research branch or the final two clean-room finalist runs before a bounded development submission. Record remote notebook version and official score; a valid submission is not proof of improvement.
5. Challenge the baseline (JESTR-01, AN-01/02, FUSE-01). Independently implement joint spectrum/structure contrastive ranking after the FP comparison; measure analogue transfer as complementary evidence on the same folds/pool. Compare single routes, static fusion and OOF learned fusion. Retain only attributable improvements.
6. Bounded forward reranking (RERANK-01). Use GLACIER as a conceptual reference for an independently trained predictor, starting with cheap fragment-explanation features. Only proceed when failure analysis shows ranking, rather than absent candidates, is limiting performance. Measure shortlist retention and whole-pipeline runtime; no fixed top-K guarantees nine hours.
7. Conditional generation (GEN-01/02/03/04). After useful fingerprint ranking, test predicted/noisy-fingerprint-conditioned decoding, then bounded scaffold editing and formula-free/formula-conditioned challengers. Require deployable predicted formulas/scaffolds, full-cohort failure accounting and marginal gains after shared reranking. No fixed generation quotas. Sampling/diffusion are optional experiments, not prerequisites to a leaderboard submission.
Phase numbers remain historical identifiers, not a mandatory execution sequence.
Interpretation guardrails: the author's 0.151/0.93 is an assumption-dependent estimate on the public test, not a known 16%/84% Class-1/Class-2 split. The remainder includes Class 3. His 99.6% COCONUT coverage is for a particular NP sample, not all natural products; 100% in-pool validation recall cannot establish hidden-test pool coverage. Reported 0.517/0.52/0.63 scores use different controls/cohorts from our library-ceiling result and have not been reproduced here. External DreaMS weights/representations are excluded, and no such result is established for our system.
Evidence rules and current snapshot
test.parquet contains training examples. Its 400-molecule/1,213-spectrum output is a schema/runtime diagnostic only, not CV, public-leaderboard evidence or a G1 quality score. Production IDs and counts must come from the runtime test file, never be hardcoded from the sample.artifacts/rebuild-20260917/ and manifests/rebuild-20260917/ together. The root retrieval symlink currently resolves there. Historical 274,288-identity and raw 275,810-identity artifacts are not interchangeable with the rebuilt 274,195 scorer identities.reports/phase0-qualified-g0-20260917.json records local guards/tests and hashes. Official scorer source acquisition and pinned-runtime fixture certification are now complete; acquisition-lineage certification remains unavailable.reports/phase1-g1-molecule-full-20260917.json uses one valid spectrum per identity, not all-spectrum validation. Preserve it as a bounded diagnostic; audit metric semantics before comparisons. Library-ceiling scores are optimistic diagnostics, not held-out generalization estimates.reports/g1-audit-20260919-1005/AUDIT.md reconciles the later 20:02 launch, saved ranks and 400-row validated placeholder submission. Evaluation took 9h04m20s / ~4.90 GiB peak RSS; placeholder inference 1m49.03s / ~1.44 GiB. Saved arithmetic passed; this is not immutable launch provenance or a full hidden-workload timing test. No matching evaluator/build/pytest process was active at this review's initial inspection.reports/todo-review-20260918/kaggle-redownload-comparison-20260918.json. The verified staged duplicate was permanently deleted at the user's request; authoritative training data and the comparison report were retained. This proves equality to the download checked on September 18, not perpetual freshness.submissions/phase1-retrieval-molecule-rebuilt-20260917.csv exists and its generation report records 400 rows and 1,213 spectra. A local schema check passed earlier in this session; this does not establish strict candidate validation, hidden-test accuracy or current-run provenance. The legacy submissions/phase1-retrieval.csv remains unsuitable as current evidence.manifests/pubchem-structure-only-1-500000.tsv and source manifest exist; original misnamed .tsv.gz is preserved. Source hash is 6006275eb9fde9101ec1f342eacb1240fad10201c3e8182eaf584cf31ea91152; 240,536 rows / 231,755 declared identities. Verifier checks suffix/magic, but the builder still accepts misleading plain-TSV suffixes and maintained format-rejection tests are missing. Upstream snapshot/licensing remains uncertified.Morning findings — disposition after the evening audit
The morning evidence below describes the pre-fix implementation, not current source. Preserve [semantic-probes.json](reports/todo-review-20260918/semantic-probes.json) and its source hashes as historical evidence; do not rerun its old defect-expecting script over that report. Fresh evening probes are in reports/audit-20260918-1923/evaluator-edge-probes.json and submission-wrapper-probes.json.
Historical consequence and later disposition: that run ended and wrapper fixes were verified. Later G1-05/06 records add candidate evidence and post-run binding, while G1-04/07 record local malformed-input/packaging acceptance. Full independent evidence/mask certification, immutable launch provenance, remaining stress cases and C remain open; those later completions do not close full G1.
Immediate audit follow-ups — historical local fixes and remaining acceptance work
set -Eeuo pipefail, stage labels and an ERR trap; failed evaluator/submission/validation stages stop before later DONE markers. The exact copied-wrapper probe now returns 41/42/43 respectively with no false success. The active long-running invocation was not restarted and therefore still used the pre-fix shell.phase1-retrieval-rebuilt-20260918.validation.json, preserving generation metadata in the original .json. Synthetic clean-environment and metadata-preservation probes pass. The active invocation's already-running process is not retroactively changed.smiles field is the explicit no-candidate policy. Added contract tests and clean helper probes. This is local validation, not official scorer parity.manifests/a4-provenance-snapshot-20260920.json with SHA-256 hashes for current Python source and frozen inputs/artifacts, runtime/platform metadata, dependency freeze, Git state, and explicit missing/limitation fields. This is a reproducibility snapshot for future launch discipline, not immutable launch-time provenance for earlier runs; G0/G1 remain open.src/scorer_contract.py with the exact Kaggle metric/casmi-mean-reciprocal-rank source (SHA-256 b542cd945203cf4305a0a01537a5fa9e93fa4e2d79400b8563286bce0f73ed2d) and ran a deterministic fixture under RDKit 2026.03.3. Expected and actual MRR@25 were both 0.625; rank-26 rejection and invalid-solution host-error probes passed. Evidence: reports/official-scorer-certification-20260920/evidence.json. This certifies local execution parity for the fixture, not hidden-test accuracy or data acquisition lineage..venv-rdkit2026: 26 passed (tests/test_phase3a_regressions.py and tests/test_g1_contract.py). This confirms the maintained local contract tests; it does not close full G1 acceptance.Phase 0 — Ground truth and infrastructure
train.parquet release from Kaggle.manifests/.src/identity.py.[M+FA-H]- to [M+CH2O2-H]- in both conversion directions; retain independent formate-mass and round-trip checks. Current tests verify alias use in neutral-mass conversion, reject a non-string adduct and reject the wrong-polarity FA spelling. This is not exhaustive validation of all numeric/polarity inputs.identity14 recomputations matched. Saved evidence: reports/val-01-02-exposure-20260923.json. Raw Parquet stored InChIKey14 differs from the scorer-derived namespace for many training rows; do not treat raw-hash parity as scorer-identity parity. Full cross-artifact VAL-01 binding remains open.reports/val-01-02-exposure-20260923.json. Next requirement is a genuinely unexposed report partition or nested/grouped evaluation with explicit prior-exposure labels; changing the seed does not erase exposure.manifests/iso-01-training-diagnostic-panel-20260923.json from data/train.parquet using the authoritative identity cache and deterministic formula grouping, without model outputs: 24,858 same-formula groups and 247,777 scorer identities. All 247,777 panel identities are training-exposed, so this is a development diagnostic only; it is not a clean report cohort and cannot close ISO-01. The original selection-free specification remains at manifests/iso-01-panel-spec-20260923.json.configs/.Resource constraint: the local development host has approximately 11 GiB RAM. Full-data operations must stream or use bounded artifacts; do not load the entire training table into pandas.
Gate G0: [x] Qualified local infrastructure pass — verified 2026-09-20. The pinned runtime, official scorer, 15-file artifact/source binding, fail-closed runtime/hash preflight, deliberate tamper failure, and 48 maintained tests pass. Evidence: reports/g0-local-closure-20260920/decision.json and reports/g0-local-closure-20260920/manifest.json. This closes qualified local G0 only; acquisition-lineage certification remains externally unavailable, Regime A is an exact-available-field proxy, Regime C is proxy-only, and this is not hidden-test or competition-performance certification.
Phase 1 — Spectral retrieval baseline
artifacts/legacy-20260917/; complete end-to-end artifact binding rather than trusting a symlink or status label alone.regime_a_query.npy row mask, not every row sharing one of its 250 identities. Assert eligible query/reference separation and reconcile total, valid, skipped and zero-valid-spectrum molecules. For B/C, retain identity-wide exclusions. Aggregate all permitted spectra within each evaluation molecule; produce exactly one output per runtime molecule_id.MRR@25 = pool_coverage × conditional_MRR@25 holds, with zero returned for empty coverage. Evidence: reports/audit-20260918-1923/evaluator-edge-probes.json. A6 records subsequent maintained regression coverage; these local checks do not certify legacy evaluators or full-run evidence. Add only demonstrably missing cases under AUDIT-01 rather than repeating the completed A6 task.reports/g1-03-identity-dedup-20260920/evidence.json and reports/audit-01-exact-identity-mask-20260923/evidence.json. This does not close full G1, launch-time immutability or official hidden-test parity.reports/g1-04-offline-acceptance-20260920/evidence.json. Fixed null precursor coercion so malformed precursor values are skipped instead of crashing. This closes local packaging/validator acceptance, not hidden-test accuracy, official scorer parity or full G1.reports/g1-post-run-audit-20260921/evidence.json.manifests/g1-06-provenance-snapshot-20260921.json binding 21 files (7 source, 10 input, 4 output) with SHA-256 hashes, plus /usr/bin/time -v stats (peak RSS 5,206,524 KiB ≈ 4.96 GB, user 34,381s), runtime metadata (Python 3.10, RDKit 2026.03.3, aarch64), git state (head 8609672198), pip freeze, and 5 explicit limitations. Self-hash included for tamper detection. This is a post-run binding, not launch-time capture; launch-time immutability remains a stated limitation.src/g1_molecule_eval_all.py; this ordering fix is present. Complete real resumable per-query/chunk score state, safe recovery and stage-aware progress separately. Counters and a metrics-bearing terminal JSON are not resume state. Validate recovery before a future rerun; do not infer ETA from stale totals or CPU utilization.scripts/run_g1_07_offline_smoke.py in persistent tmux session g1-07-smoke against a four-molecule/15-spectrum fixture with renamed and reversed IDs (g107_3…g107_0) and network socket construction disabled. Inference produced 4/4 rows with 100% coverage; strict local validation passed with zero errors. Evidence: reports/g1-07-offline-smoke-20260920/evidence.json. Production artifacts were not modified. This is an offline packaging smoke test, not the two final clean-room notebook reruns required by G6.LOCAL_DIAGNOSTIC_PASS_FULL_G1_ACCEPTANCE_OPEN. Its then-open packaging items have later dispositions in G1-04/G1-07, maintained coverage in A6, candidate evidence in G1-05 and exact identity/mask validation in AUDIT-01. Immutable launch binding, acquisition lineage, official hidden-test parity, Regime C and final clean-room runs remain open.reports/audit-01-exact-identity-mask-20260923/evidence.json. This validates the saved local qualified run only; launch-time immutability, acquisition lineage and official hidden-test parity remain open.Gate G1: [!] OPEN — full acceptance not established. A successful process exit, 400/400 placeholder rows, single-spectrum metrics or B=0 by construction cannot close it. Regime-B zero target coverage for a pure reference-spectrum matcher is structurally expected after identity-wide exclusion: it demonstrates the method's unseen-identity limitation, not structure-only identification quality. Keep candidate coverage, reference availability and ranking quality separate. Further learned experiments require the task-relevant qualified scorer/split/pool contract; full G1 closure and the unrun C proxy are not to be asserted merely because Phase 4 has already executed.
Controlled retrieval ablations — after the corrected baseline is locked
These are follow-on comparisons, not an excuse to withhold a correctly audited baseline indefinitely. Use the same frozen cohorts/masks and bounded development experiments; promote only with reproducible molecule-level evidence.
manifests/retrieval-baseline-resolved-20260923.json supersedes the disputed historical lock; this is a qualified local contract, not hidden-test certification.reports/base02-ablation-corrected-20260923/closure-results.json; do not present this bounded library-ceiling result as generalization or hidden-test performance.max across permitted spectra against mean/top-k evidence, within-molecule reciprocal-rank fusion, collision-energy-aware aggregation and same-adduct/same-polarity policies. Include a matched single-spectrum control, reference-count bias checks and correlated-acquisition controls. Do not merge incompatible ions into one raw peak list. Cross-route ensemble weights remain Phase 5 work.masses[order] inside every query. Prove candidate membership and ranking equivalence before accepting this optimization; do not alter a running evaluator in place.reports/todo-review-20260918/G1-OPTIMIZATION-DESIGN.md. No optimized implementation or full-data rerun is claimed complete.Phase 2 — Learned fingerprint ranking and domain adaptation
Entry dependency: qualified task-relevant scorer/split/exposure contracts, a frozen eligible Phase 3A/3B pool and non-neural controls. Full PubChem coverage is not required. Do not train against an empty B target pool and call it a ranking experiment. All model training remains on approved Kaggle GPU resources, never this local host.
research/arxiv-msms-2d-shortlist-20260923.md; no experiment is licensed or approved by citation alone.f · z ranking with neighbour transfer and simple cosine controls; the independent-Bernoulli score is not a calibrated exact-match probability. Independently implement the concept; do not use paper code, weights or training data. Contract frozen at manifests/fp-01-training-pool-contract-20260923.json; Kaggle handoff at manifests/fp-01-kaggle-handoff-20260923.json; PubChem is excluded and no local training is authorized. Local implementation and compile/schema evidence exists (2026-09-24). Earlier launches hit the weekly GPU quota; the latest 06:25 check returned HTTP 409 and status permission denied without reconfirming quota. Execution remains unverified; no verified model metrics or OOF outputs exist. RUN-01 owns launch recovery.enveda-180 rather than labelling it uniformly low quality.Gate G2: learned ranking improves grouped validation with paired molecule-level evidence, no unexplained severe regression in any intended regime, and acceptable measured resources. Random-row/library-only splits and public leaderboard claims do not satisfy the gate.
Phase 3 — Known structures without reference spectra
Phase 3A — Bounded pinned structure universe (supporting prerequisite, not bulk expansion)
manifests/phase3a-structure-index-spec-20260918.json. B eligibility text is corrected; this does not certify all sources or satisfy acceptance.reports/rights-01-audit-20260923.json records separate dispositions for papers/notebooks, the bounded PubChem source/index, competition inputs, local runtime and future external models. The recovered PubChem upstream is conditionally usable under recorded CC0 terms, but the CURRENT-Full source is not an immutable release. The exact SDF→TSV→index transformation is now bound to commit 8020971feff895b962a8520f01130c6e05fe9abc; existing derived outputs have not been reproduced under that chain in this turn and remain quarantined. The fail-closed offline package inventory is manifests/offline-package-manifest-20260923.json; it blocks adoption of that external-pool package until rebuild verification and complete dependency/license records; the later competition-only FP-01 contract is a separate route, not blocked by the deferred PubChem rebuild. No paper code, checkpoint, tokenizer or training archive was reused.artifacts/phase3a-rederived-20260920/pubchem-structure-index.sqlite from the consolidated Phase 3A source. Terminal manifest status is COMPLETE; 444,792 source rows processed, 424,291 valid/unique scorer identities, 0 invalid SMILES, 0 missing identities and 0 identity mismatches. The 131 MB SQLite artifact and progress manifest are present; no active builder process remains. This completes the bounded structure index build, not source provenance/licensing certification or production promotion.manifests/rederived-20260919/pubchem-rederived-consolidated.tsv: 444,792 valid rows, 444,792 unique CIDs, 424,291 scorer identities, 19,801 training-overlap rows, 424,991 structure-only rows and 179,581 novelty-proxy rows. Independently checked schema and SMILES parseability; no duplicate CIDs. Preserved 24 chunk-reported rejected records in pubchem-rederived-exceptions.jsonl. Manifest: reports/phase3a-rederived-20260919/consolidated-manifest.json. This is a bounded consolidated structure pool, not yet a mass/formula index or promoted production source.reports/todo-review-20260918/pubchem-local-audit.json. This did not recompute chemical identities, training-overlap or novelty labels against the rebuilt namespace, and did not physically relocate the file.reports/phase3a-pubchem-identity-reconciliation-20260919.json and .md.8020971feff895b962a8520f01130c6e05fe9abc, and the content-pinned archive hash is recorded. Existing TSV/index outputs still require a reproducible rebuild/verification under that exact chain; the upstream CURRENT-Full label is not immutable. Identity-overlap labels match (12,331), but 9,594 scaffold-string differences include 1,741 changed scaffold-overlap memberships; production has seven invalid representatives. Keep the pool quarantined for C/novelty and do not promote until rebuild and RIGHTS-01/offline manifest gates pass.September 19 audit repair gates (before promotion)
reports/phase3a-pubchem-scaffold-discrepancy-20260919.json and reports/review-20260919-1817/AUDIT.md.artifacts/phase3a/pubchem-structure-index-repaired.sqlite with PASS evidence in artifacts/phase3a/pubchem-structure-index-repaired.repair-verification.json. This closes the repair execution task only; full source provenance/licensing and promotion remain open.identity-cache.sqlite via normalized_smiles, never raw Parquet inchikey14. Cache-aligned P0/P2 combined MRR@25 are 0.3003/0.3024 and Recall@25 81.67%/82.08%. Identity audit still finds 1,559 registered-derived rows across 44 stored→derived mappings and 164 identities overlapping prior fold 4. Do not label the partition untouched.Phase 3B — Formula/mass-gated candidates and non-neural baseline
reports/phase3a-reference-availability-20260919.json and .md. This is a post-freeze stability diagnostic, not prospective validation.reports/phase3a-analogue-fallback-audit-20260919.json and .md.identity14 recomputations match; 1,559 registered-derived rows across 44 stored→derived mappings disagree with Parquet inchikey14; 164 identities overlap prior fold 4. Query extraction is cache-aligned, per-query provenance is persisted, and regression-tested, but the report remains diagnostic because prior exposure is unresolved.reports/phase3a-analogue-transfer-preregistered-corrected-cache-aligned-20260919.json: 500 identities, 480 valid queries, 1,440 policy/query evidence records. P0 combined MRR@25 is 0.3003 / Recall@25 81.67%; P2 is 0.3024 / 82.08%. These remain diagnostic because 164 identities overlap prior fold 4 and the partition is not untouched.8020971feff895b962a8520f01130c6e05fe9abc, and manifests/offline-dependency-inventory-20260923.json records 14 packages from the pinned RDKit environment. The historical TSV/index outputs have not yet been reproduced and verified under that commit; the offline package manifest remains fail-closed and inventory-only for licences/wheels. Do not promote this external pool until those fields are closed. Its quarantine does not block competition-only FP-01/AN-01; their own dependency, rights and evaluation checks still apply.Phase 3C — Mass-shifted analogue propagation (parallel complementary channel)
Later representations and forward checking
Gate G3: measurable structure-known/no-spectrum improvement over the non-neural baseline, with leakage-safe spectral inputs, licensed/pinned structure provenance and separate pool coverage/conditional MRR. Database absence and ranking failure must remain distinguishable.
Phase 4 — Novel structures and de novo generation
Feasibility branch, not promoted: a spectrum-to-SMILES implementation and Kaggle experiment already exist. Preserve their source/checkpoint/output history; this review neither stops the existing run nor launches another. Further scale-up is deferred until the gates below pass and a bounded experiment budget is approved. A failed G4 does not prevent shipping a validated fingerprint/analogue/ranking solution, but forbids claiming proven novelty coverage.
The C-proxy split/leakage audit needed by G1 is validation infrastructure and is not deferred with generation.
kaggle-kernel/phase4-denovo/; qualification remains GEN-01/02. No useful generalization result is asserted by code presence or a completed job.Gate G4: de novo candidates add measurable clean-validation MRR or candidate coverage within the offline runtime budget.
Later generation experiments — gated, not a submission dependency
Phase 5 — Unified ensemble
CFG.SIM_POWER/CFG.W1 are unused, and its candidate cap affects all downstream channels. When caps/pools/channels change, recompute and refit set-relative rank/z-score features or prove invariance; reject stale feature-schema/checkpoint combinations.Gate G5: robust integrated model beats the retrieval baseline across the selected validation scenarios.
Phase 6 — Kaggle packaging and final submission
/tmp; write metrics/checkpoints before potentially failing post-processing. No package-registry fallback in the final offline path.submission.csv columns, nulls, duplicate IDs and maximum candidate count.Gate G6: two successful offline reruns, reproducible output and final compliance audit.
Validation contract for the revised route
Active experiment log
Historical chronology is retained below. Earlier PASS/closure statements and namespace counts were superseded by the 2026-09-17 audit/rebuild; the current gate sections above govern decisions.
| Date | Experiment | Hypothesis | Result | Decision |
|---|---|---|---|---|
| 2026-09-16 | Competition review and workspace setup | Establish a complete execution baseline | Plan and project scaffold created; no model run | Start Phase 0 |
| 2026-09-16 | Phase 0 data acquisition | Verify current Kaggle release and characterize it under low RAM | 3,033,286,496-byte file; SHA-256 recorded; 2,539,608 rows, 21 row groups, 275,810 unique InChIKey14 identities; audit complete | Proceed to grouped split manifests |
| 2026-09-16 | Molecule-grouped split | Prevent spectral/structure leakage in local validation | Deterministic 5-fold manifest validated: 275,810 groups and 2,539,608 rows; fold assignment is group-exclusive | Use fold 0 as initial validation; keep novelty labels separate |
| 2026-09-16 | G0 data/infrastructure gate | Confirm reproducible Phase 0 inputs and assumptions | All eight gate checks passed; config and split hashes recorded in manifests/phase0-gate.json | Close data/infrastructure work; proceed to retrieval baseline |
| 2026-09-16 | Phase 0 audit | Verify that validation grouping uses scorer identity | Failed: 12/100 sampled normalized SMILES differed between stored and scorer-derived InChIKey14; full probe hit pathological tautomer enumeration | Reopen G0; rebuild identity-safe split before retrieval evaluation |
| 2026-09-16 | Phase 0 corrective audit | Verify corrected scorer-identity split and reproducibility gate | Passed: 8 tests; compilation passed; 2,539,608 rows and 274,288 scorer groups validated; cache/checkpoint agree; full G0 gate passed; 100-row identity probe had 0 invalid structures | Proceed to retrieval baseline; retain external novelty regimes as pending |
| 2026-09-16 | Phase 1 retrieval baseline | Build a low-memory, identity-safe spectral retrieval floor | Index built over all 2,539,608 rows; 2,160,035 valid spectra; fold-0 mask excludes all 503,842 held-out rows; molecule-level fixture submission has 400/400 covered molecules and passes schema/count checks | Keep baseline; run grouped MRR/Top-1/Recall evaluation before closing G1 |
| 2026-09-17 | Qualified rebuilt infrastructure | Separate corrected identity namespace from legacy evidence | phase0-qualified-g0-20260917.json records 274,195 identities, 22 passing local tests, A/B guards and source/mask hashes | Qualified local infrastructure only; preserve official parity/acquisition caveats |
| 2026-09-17 | Bounded molecule-level retrieval | Establish a diagnostic floor on the rebuilt namespace | 5,974-identity library/B cohorts plus 250 A identities; one valid query spectrum per identity; saved report and exit-0 timing log | Not all-spectrum G1; audit metric definitions and query selection before promotion |
| 2026-09-18 | All-spectrum G1 execution | Aggregate permitted spectra for each validation molecule | Started 06:26 MYT; still active with no result artifact at the 08:21 checkpoint | Keep G1 open; require semantic, provenance, strict-output and resource audit, not just process completion |
| 2026-09-18 | Community findings + execution-plan review | Improve candidate coverage before heavier models | Prioritized structure-only snapshots/mass-formula candidates, OOF fingerprints and separate analogue propagation; added leakage, adduct, cap and aggregation checks | Adopt revised backlog order; no new model training, submission or G1 closure claimed |
| 2026-09-18 | Evaluator/validator corrections | Repair the morning audit's concrete defects | A row-mask forwarding, uncapped coverage, covered-query conditional MRR and basic invalid/duplicate/over-25 checks implemented | Local fixes verified; not end-to-end G1 acceptance |
| 2026-09-18 | Checkpointed rerun at 12:39 | Expose progress without waiting for the aggregate report | PID 198260: library scan reached 621 batches / 113,376,230 comparisons; B active; final report absent at evening audit | Preserve ongoing work; counters are not resumable state, saved metrics or a reliable ETA |
| 2026-09-18 | Kaggle CLI release verification | Check authoritative training data against a fresh download | Size/hash/schema/rows/row groups all match; verified staged duplicate permanently deleted on request | Retain original data and comparison report; no rebuild justified by release mismatch |
| 2026-09-18 | Non-competing preparation | Improve small contracts and document next steps | Basic formate alias tested; PubChem physical-format/count audit and Phase 3A/optimization drafts recorded | Drafts/source audit have eligibility/provenance limitations; no new index or GPU runner built |
| 2026-09-18 | Evening progress audit | Reconcile source, tests, claims and TODO | 24 lightweight tests pass with 2 integration tests excluded; independent metric fixes verified; wrapper failures, entrypoint failure, metadata loss and strict-output gaps reproduced | Keep G1 open; prioritize A1-A8, narrow completed claims and retain historical evidence |
| 2026-09-19 | Phase 3A fallback/category audit | Verify fallback equivalence and separate structure/reference failure classes | 475 valid audit queries: 28 structure-pool misses; 260 had a fold-eligible local identity in the mass window; 215 used the audit fallback route; fallback pool recall 90.23%, @25 88.84%; 0 mass-order equivalence mismatches; oracle formula gate showed no improvement | Keep formula gate non-deployable; do not call the 260 count a positive spectral-neighbour count; preserve categories; add maintained tests and repeat on a genuinely pre-registered untouched partition |
| 2026-09-23 | FP-01 Kaggle handoff | Prepare the smallest independent fingerprint-ranking implementation without local training or external assets | Handoff manifest binds the FP-01 contract, diagnostic ISO-01 panel, baseline and no-submission/no-external-asset policy; implementation and Kaggle run remain pending | Implement and run a bounded Kaggle smoke/ablation: BCE-only vs BCE+hard-negative on approved compute; do not submit until molecule-level evidence and package checks pass |
Blockers and decisions
reports/official-scorer-certification-20260920/evidence.json. C is a proxy, not true hidden novelty. The 1,184-versus-1,151 documentation discrepancy is pinned locally and is not itself a remaining implementation blocker.24 September shortlist replan
Reconciled the supplied shortlist with the competition-only decision and saved FP-01 launch evidence. Removed the duplicate TOP5 status ledger, excluded external pretrained routes, corrected B/C eligibility and unsupported MSFlow withdrawal wording, and scheduled first qualified submission before optional heavy research. Existing completed evidence retained; no experiment promoted by this edit.