Enveda CASMI 2026 Todo
Legend: [ ] pending · [-] in progress · [x] complete · [!] blocked
Reviewed: 2026-09-20 (MYT), after Phase 3A index completion. Current findings: [audit report](reports/review-20260919-1817/AUDIT.md). This review supersedes stale active-process and completion statements below; historical experiment entries remain historical.
Current audit snapshot: 45 maintained tests passed across the five focused suites. G1 saved-result arithmetic and placeholder output were independently audited previously, but full G1 remains OPEN. The shadow is a seven-row repair copy, not a completed source rebuild. Production is not promoted or changed.
Material corrections: the scaffold classifier compared identity overlap in its category logic: 1,741 scaffold-overlap changes, not zero (1,295 lost / 446 gained; 7,853 other scaffold-string differences). The registered readout is not certified untouched: 164 selected identities belong to the previously inspected fold 4. Analogue combined_mrr is untruncated reciprocal rank, not MRR@25, and its failure/stratum counters need correction. Keep all Phase 3A results diagnostic-only.
Checklist count after reconciliation: counts below include the G0 closure evidence; gate states are separate and this is not an effort-weighted or competition-readiness percentage. Four explicit shadow/analogue repair tasks remain tracked.
Evening audit: 2026-09-18, current source + isolated synthetic probes + saved artifacts. Evidence and pre-audit backup: reports/audit-20260918-1923/. The audit changes documentation only; the active evaluator and its inputs remain untouched. Earlier completion claims are qualified below.
Audit report: [AUDIT.md](reports/audit-20260918-1923/AUDIT.md). Fresh lightweight suite: 24 passed, 2 real-index tests deliberately deselected; additional independent synthetic evaluator/entrypoint/failure-injection evidence is saved alongside it. The earlier 26-pass result is not a claim that all gate scenarios are covered.
The priority order is deliberate: establish trustworthy evaluation, expand candidate coverage, then learn to rank. Community methods are hypotheses, not validated improvements. This review updates the backlog; it does not mark unrun experiments complete.
Next execution order
1. P0 — Close reproduced audit defects before promotion. Correct scaffold classification, analogue scorer-identity/query accounting, MRR@25, failure categories and report-set exposure tracking. Resolve shadow manifest/progress and reproducible repair evidence. Preserve existing outputs; do not overwrite production or rerun the full spectral evaluation as part of this review.
2. Lock a qualified local retrieval baseline, then run bounded mass-window, candidate-cap and multi-spectrum ablations on development folds. Keep an untouched report partition.
3. Promote Phase 3A/3B ahead of learned ranking: pinned/eligible structure sources, mass/formula indexing, candidate recall, reference-availability accounting, and a non-neural structure-aware baseline. A ranker cannot recover an answer absent from its pool, and a reference-spectrum matcher cannot rank an identity whose eligible reference spectra were removed.
4. Separate spectral-library diagnostics from unseen-identity ranking: retain library-ceiling reranking as an optimistic implementation control; report identity-held-out cross-identity decoy ordering/reference availability separately, without treating an absent target reference as a conventional zero-quality rank. Then proceed to structure-aware fingerprint/analogue methods before learned ranking.
5. Phase 2 + Phase 3C: leakage-safe spectrum-to-fingerprint/ranking experiments and a separate mass-shifted analogue-propagation benchmark. The analogue experiment may run alongside bounded ranking once both use the same frozen candidate pool and folds; it is not a prerequisite for starting every ranker experiment.
6. Keep de novo generation and cross-route fusion deferred. Revisit Phase 4 when database/analogue coverage failures justify it; validate route fusion in Phase 5, not by copying public weights.
Phase numbers are retained for traceability. These dependencies take precedence over the original plan's provisional Phase 2-before-Phase 3 calendar; no new dates or compute commitments are implied.
Evidence rules and current snapshot
test.parquet contains training examples. Its 400-molecule/1,213-spectrum output is a schema/runtime diagnostic only, not CV, public-leaderboard evidence or a G1 quality score. Production IDs and counts must come from the runtime test file, never be hardcoded from the sample.artifacts/rebuild-20260917/ and manifests/rebuild-20260917/ together. The root retrieval symlink currently resolves there. Historical 274,288-identity and raw 275,810-identity artifacts are not interchangeable with the rebuilt 274,195 scorer identities.reports/phase0-qualified-g0-20260917.json records local guards/tests and hashes. Official scorer source acquisition and pinned-runtime fixture certification are now complete; acquisition-lineage certification remains unavailable.reports/phase1-g1-molecule-full-20260917.json uses one valid spectrum per identity, not all-spectrum validation. Preserve it as a bounded diagnostic; audit metric semantics before comparisons. Library-ceiling scores are optimistic diagnostics, not held-out generalization estimates.reports/g1-audit-20260919-1005/AUDIT.md reconciles the later 20:02 launch, saved ranks and 400-row validated placeholder submission. Evaluation took 9h04m20s / ~4.90 GiB peak RSS; placeholder inference 1m49.03s / ~1.44 GiB. Saved arithmetic passed; this is not immutable launch provenance or a full hidden-workload timing test. No matching evaluator/build/pytest process was active at this review's initial inspection.reports/todo-review-20260918/kaggle-redownload-comparison-20260918.json. The verified staged duplicate was permanently deleted at the user's request; authoritative training data and the comparison report were retained. This proves equality to the download checked on September 18, not perpetual freshness.submissions/phase1-retrieval-molecule-rebuilt-20260917.csv exists and its generation report records 400 rows and 1,213 spectra. A local schema check passed earlier in this session; this does not establish strict candidate validation, hidden-test accuracy or current-run provenance. The legacy submissions/phase1-retrieval.csv remains unsuitable as current evidence.manifests/pubchem-structure-only-1-500000.tsv and source manifest exist; original misnamed .tsv.gz is preserved. Source hash is 6006275eb9fde9101ec1f342eacb1240fad10201c3e8182eaf584cf31ea91152; 240,536 rows / 231,755 declared identities. Verifier checks suffix/magic, but the builder still accepts misleading plain-TSV suffixes and maintained format-rejection tests are missing. Upstream snapshot/licensing remains uncertified.Morning findings — disposition after the evening audit
The morning evidence below describes the pre-fix implementation, not current source. Preserve [semantic-probes.json](reports/todo-review-20260918/semantic-probes.json) and its source hashes as historical evidence; do not rerun its old defect-expecting script over that report. Fresh evening probes are in reports/audit-20260918-1923/evaluator-edge-probes.json and submission-wrapper-probes.json.
Current consequence: the old run has ended and wrapper fixes are verified. Saved per-molecule target ranks exist, but complete candidate/stratum evidence, immutable provenance, runtime hardening and C remain incomplete. Preserve the useful results without full G1 promotion.
Immediate audit follow-ups — historical local fixes and remaining acceptance work
set -Eeuo pipefail, stage labels and an ERR trap; failed evaluator/submission/validation stages stop before later DONE markers. The exact copied-wrapper probe now returns 41/42/43 respectively with no false success. The active long-running invocation was not restarted and therefore still used the pre-fix shell.phase1-retrieval-rebuilt-20260918.validation.json, preserving generation metadata in the original .json. Synthetic clean-environment and metadata-preservation probes pass. The active invocation's already-running process is not retroactively changed.smiles field is the explicit no-candidate policy. Added contract tests and clean helper probes. This is local validation, not official scorer parity.manifests/a4-provenance-snapshot-20260920.json with SHA-256 hashes for current Python source and frozen inputs/artifacts, runtime/platform metadata, dependency freeze, Git state, and explicit missing/limitation fields. This is a reproducibility snapshot for future launch discipline, not immutable launch-time provenance for earlier runs; G0/G1 remain open.src/scorer_contract.py with the exact Kaggle metric/casmi-mean-reciprocal-rank source (SHA-256 b542cd945203cf4305a0a01537a5fa9e93fa4e2d79400b8563286bce0f73ed2d) and ran a deterministic fixture under RDKit 2026.03.3. Expected and actual MRR@25 were both 0.625; rank-26 rejection and invalid-solution host-error probes passed. Evidence: reports/official-scorer-certification-20260920/evidence.json. This certifies local execution parity for the fixture, not hidden-test accuracy or data acquisition lineage.reports/a5-g1-evidence-pack-20260920.json (12,198 per-molecule records). Explicitly recorded missing complete candidate IDs/scores/source rows, full A/B strata, immutable launch binding, and Regime C (NOT_EVALUATED). This closes evidence reconciliation/packaging only; a clean A5 completion still requires a rerun or persisted candidate-level evidence..venv-rdkit2026: 26 passed (tests/test_phase3a_regressions.py and tests/test_g1_contract.py). This confirms the maintained local contract tests; it does not close full G1 acceptance.REGIME_COMPLETE is written before rank/metric calculation and does not persist those results. Do not infer ETA from rows, stale checkpoint elapsed time, or CPU utilization.Phase 0 — Ground truth and infrastructure
train.parquet release from Kaggle.manifests/.src/identity.py.[M+FA-H]- to [M+CH2O2-H]- in both conversion directions; retain independent formate-mass and round-trip checks. Current tests verify alias use in neutral-mass conversion, reject a non-string adduct and reject the wrong-polarity FA spelling. This is not exhaustive validation of all numeric/polarity inputs.configs/.Resource constraint: the local development host has approximately 11 GiB RAM. Full-data operations must stream or use bounded artifacts; do not load the entire training table into pandas.
Gate G0: [x] Qualified local infrastructure pass — verified 2026-09-20. The pinned runtime, official scorer, 15-file artifact/source binding, fail-closed runtime/hash preflight, deliberate tamper failure, and 48 maintained tests pass. Evidence: reports/g0-local-closure-20260920/decision.json and reports/g0-local-closure-20260920/manifest.json. This closes qualified local G0 only; acquisition-lineage certification remains externally unavailable, Regime A is an exact-available-field proxy, Regime C is proxy-only, and this is not hidden-test or competition-performance certification.
Phase 1 — Spectral retrieval baseline
artifacts/legacy-20260917/; complete end-to-end artifact binding rather than trusting a symlink or status label alone.regime_a_query.npy row mask, not every row sharing one of its 250 identities. Assert eligible query/reference separation and reconcile total, valid, skipped and zero-valid-spectrum molecules. For B/C, retain identity-wide exclusions. Aggregate all permitted spectra within each evaluation molecule; produce exactly one output per runtime molecule_id.MRR@25 = pool_coverage × conditional_MRR@25 holds, with zero returned for empty coverage. Evidence: reports/audit-20260918-1923/evaluator-edge-probes.json. Promote these probes into maintained tests (A6); the legacy evaluators and full-run evidence are not thereby certified.reports/g1-03-identity-dedup-20260920/evidence.json. Full closure still requires candidate identity/source-row evidence from the active all-spectrum evaluator.reports/g1-04-offline-acceptance-20260920/evidence.json. Fixed null precursor coercion so malformed precursor values are skipped instead of crashing. This closes local packaging/validator acceptance, not hidden-test accuracy, official scorer parity or full G1.reports/g1-post-run-audit-20260921/evidence.json.manifests/g1-06-provenance-snapshot-20260921.json binding 21 files (7 source, 10 input, 4 output) with SHA-256 hashes, plus /usr/bin/time -v stats (peak RSS 5,206,524 KiB ≈ 4.96 GB, user 34,381s), runtime metadata (Python 3.10, RDKit 2026.03.3, aarch64), git state (head 8609672198), pip freeze, and 5 explicit limitations. Self-hash included for tamper detection. This is a post-run binding, not launch-time capture; launch-time immutability remains a stated limitation.REGIME_COMPLETE is written before rank/metric calculation and does not persist those results. Do not infer ETA from rows, stale checkpoint elapsed time, or CPU utilization. Less urgent now that G1 completed, but needed before any future rerun.scripts/run_g1_07_offline_smoke.py in persistent tmux session g1-07-smoke against a four-molecule/15-spectrum fixture with renamed and reversed IDs (g107_3…g107_0) and network socket construction disabled. Inference produced 4/4 rows with 100% coverage; strict local validation passed with zero errors. Evidence: reports/g1-07-offline-smoke-20260920/evidence.json. Production artifacts were not modified. This is an offline packaging smoke test, not the two final clean-room notebook reruns required by G6.reports/g1-audit-20260919-1005/audit-decision.json: LOCAL_DIAGNOSTIC_PASS_FULL_G1_ACCEPTANCE_OPEN. Saved arithmetic/output pass; provenance permutation guard, malformed precursor handling, legacy metrics, packaging and C remain open. Completing this decision task does not close G1.Gate G1: [!] OPEN — execution and evidence audit in progress. A successful process exit, 400/400 placeholder rows, single-spectrum metrics or B=0 by construction cannot close it. Regime-B zero target coverage for a pure reference-spectrum matcher is structurally expected after identity-wide exclusion: it demonstrates the method's unseen-identity limitation, not structure-only identification quality. Keep candidate coverage, reference availability and ranking quality as separate measurements. Do not begin learned-model training before qualified local acceptance.
Controlled retrieval ablations — after the corrected baseline is locked
These are follow-on comparisons, not an excuse to withhold a correctly audited baseline indefinitely. Use the same frozen cohorts/masks and bounded development experiments; promote only with reproducible molecule-level evidence.
max across permitted spectra against mean/top-k evidence, within-molecule reciprocal-rank fusion, collision-energy-aware aggregation and same-adduct/same-polarity policies. Include a matched single-spectrum control, reference-count bias checks and correlated-acquisition controls. Do not merge incompatible ions into one raw peak list. Cross-route ensemble weights remain Phase 5 work.masses[order] inside every query. Prove candidate membership and ranking equivalence before accepting this optimization; do not alter a running evaluator in place.reports/todo-review-20260918/G1-OPTIMIZATION-DESIGN.md. No optimized implementation or full-data rerun is claimed complete.Phase 2 — Learned fingerprint ranking and domain adaptation
Entry dependency: qualified local G1 plus a frozen Phase 3A/3B candidate pool and non-neural baseline. Do not train against an empty B target pool and call the result a ranking experiment.
Gate G2: learned ranking improves grouped validation with paired molecule-level evidence, no unexplained severe regression in any intended regime, and acceptable measured resources. Random-row/library-only splits and public leaderboard claims do not satisfy the gate.
Phase 3 — Known structures without reference spectra
Phase 3A — Pinned structure-only universe (promoted; first after G1)
manifests/phase3a-structure-index-spec-20260918.json. B eligibility text is corrected; this does not certify all sources or satisfy acceptance.artifacts/phase3a-rederived-20260920/pubchem-structure-index.sqlite from the consolidated Phase 3A source. Terminal manifest status is COMPLETE; 444,792 source rows processed, 424,291 valid/unique scorer identities, 0 invalid SMILES, 0 missing identities and 0 identity mismatches. The 131 MB SQLite artifact and progress manifest are present; no active builder process remains. This completes the bounded structure index build, not source provenance/licensing certification or production promotion.manifests/rederived-20260919/pubchem-rederived-consolidated.tsv: 444,792 valid rows, 444,792 unique CIDs, 424,291 scorer identities, 19,801 training-overlap rows, 424,991 structure-only rows and 179,581 novelty-proxy rows. Independently checked schema and SMILES parseability; no duplicate CIDs. Preserved 24 chunk-reported rejected records in pubchem-rederived-exceptions.jsonl. Manifest: reports/phase3a-rederived-20260919/consolidated-manifest.json. This is a bounded consolidated structure pool, not yet a mass/formula index or promoted production source.reports/todo-review-20260918/pubchem-local-audit.json. This did not recompute chemical identities, training-overlap or novelty labels against the rebuilt namespace, and did not physically relocate the file.reports/phase3a-pubchem-identity-reconciliation-20260919.json and .md.reports/review-20260919-1817/AUDIT.md.September 19 audit repair gates (before promotion)
reports/phase3a-pubchem-scaffold-discrepancy-20260919.json and reports/review-20260919-1817/AUDIT.md.artifacts/phase3a/pubchem-structure-index-repaired.sqlite with PASS evidence in artifacts/phase3a/pubchem-structure-index-repaired.repair-verification.json. This closes the repair execution task only; full source provenance/licensing and promotion remain open.identity-cache.sqlite via normalized_smiles, never raw Parquet inchikey14. Cache-aligned P0/P2 combined MRR@25 are 0.3003/0.3024 and Recall@25 81.67%/82.08%. Identity audit still finds 1,559 registered-derived rows across 44 stored→derived mappings and 164 identities overlapping prior fold 4. Do not label the partition untouched.Phase 3B — Formula/mass-gated candidates and non-neural baseline
reports/phase3a-reference-availability-20260919.json and .md. This is a post-freeze stability diagnostic, not prospective validation.reports/phase3a-analogue-fallback-audit-20260919.json and .md.identity14 recomputations match; 1,559 registered-derived rows across 44 stored→derived mappings disagree with Parquet inchikey14; 164 identities overlap prior fold 4. Query extraction is cache-aligned, per-query provenance is persisted, and regression-tested, but the report remains diagnostic because prior exposure is unresolved.reports/phase3a-analogue-transfer-preregistered-corrected-cache-aligned-20260919.json: 500 identities, 480 valid queries, 1,440 policy/query evidence records. P0 combined MRR@25 is 0.3003 / Recall@25 81.67%; P2 is 0.3024 / 82.08%. These remain diagnostic because 164 identities overlap prior fold 4 and the partition is not untouched.reports/phase3a-pubchem-provenance-audit-20260919.json: local SHA-256 6006275e…91152, 240,536 rows and 231,755 declared identities are reproducible, but source_url, upstream version, download date, license and derivation commit are absent. Do not promote or claim complete PubChem coverage/novelty certification until those fields are independently recovered and verified.Phase 3C — Mass-shifted analogue propagation (separate route)
Later representations and forward checking
Gate G3: measurable structure-known/no-spectrum improvement over the non-neural baseline, with leakage-safe spectral inputs, licensed/pinned structure provenance and separate pool coverage/conditional MRR. Database absence and ranking failure must remain distinguishable.
Phase 4 — Novel structures and de novo generation
Deferred: no immediate autoregressive SMILES or large-transformer training. First diagnose residual failures after database/analogue/fingerprint work and approve a bounded experiment budget. A failed de novo gate need not prevent shipping a validated retrieval/database solution, but forbids claiming novelty coverage.
The C-proxy split/leakage audit needed by G1 is validation infrastructure and is not deferred with generation.
Gate G4: de novo candidates add measurable clean-validation MRR or candidate coverage within the offline runtime budget.
Phase 5 — Unified ensemble
Gate G5: robust integrated model beats the retrieval baseline across the selected validation scenarios.
Phase 6 — Kaggle packaging and final submission
submission.csv columns, nulls, duplicate IDs and maximum candidate count.Gate G6: two successful offline reruns, reproducible output and final compliance audit.
Active experiment log
Historical chronology is retained below. Earlier PASS/closure statements and namespace counts were superseded by the 2026-09-17 audit/rebuild; the current gate sections above govern decisions.
| Date | Experiment | Hypothesis | Result | Decision |
|---|---|---|---|---|
| 2026-09-16 | Competition review and workspace setup | Establish a complete execution baseline | Plan and project scaffold created; no model run | Start Phase 0 |
| 2026-09-16 | Phase 0 data acquisition | Verify current Kaggle release and characterize it under low RAM | 3,033,286,496-byte file; SHA-256 recorded; 2,539,608 rows, 21 row groups, 275,810 unique InChIKey14 identities; audit complete | Proceed to grouped split manifests |
| 2026-09-16 | Molecule-grouped split | Prevent spectral/structure leakage in local validation | Deterministic 5-fold manifest validated: 275,810 groups and 2,539,608 rows; fold assignment is group-exclusive | Use fold 0 as initial validation; keep novelty labels separate |
| 2026-09-16 | G0 data/infrastructure gate | Confirm reproducible Phase 0 inputs and assumptions | All eight gate checks passed; config and split hashes recorded in manifests/phase0-gate.json | Close data/infrastructure work; proceed to retrieval baseline |
| 2026-09-16 | Phase 0 audit | Verify that validation grouping uses scorer identity | Failed: 12/100 sampled normalized SMILES differed between stored and scorer-derived InChIKey14; full probe hit pathological tautomer enumeration | Reopen G0; rebuild identity-safe split before retrieval evaluation |
| 2026-09-16 | Phase 0 corrective audit | Verify corrected scorer-identity split and reproducibility gate | Passed: 8 tests; compilation passed; 2,539,608 rows and 274,288 scorer groups validated; cache/checkpoint agree; full G0 gate passed; 100-row identity probe had 0 invalid structures | Proceed to retrieval baseline; retain external novelty regimes as pending |
| 2026-09-16 | Phase 1 retrieval baseline | Build a low-memory, identity-safe spectral retrieval floor | Index built over all 2,539,608 rows; 2,160,035 valid spectra; fold-0 mask excludes all 503,842 held-out rows; molecule-level fixture submission has 400/400 covered molecules and passes schema/count checks | Keep baseline; run grouped MRR/Top-1/Recall evaluation before closing G1 |
| 2026-09-17 | Qualified rebuilt infrastructure | Separate corrected identity namespace from legacy evidence | phase0-qualified-g0-20260917.json records 274,195 identities, 22 passing local tests, A/B guards and source/mask hashes | Qualified local infrastructure only; preserve official parity/acquisition caveats |
| 2026-09-17 | Bounded molecule-level retrieval | Establish a diagnostic floor on the rebuilt namespace | 5,974-identity library/B cohorts plus 250 A identities; one valid query spectrum per identity; saved report and exit-0 timing log | Not all-spectrum G1; audit metric definitions and query selection before promotion |
| 2026-09-18 | All-spectrum G1 execution | Aggregate permitted spectra for each validation molecule | Started 06:26 MYT; still active with no result artifact at the 08:21 checkpoint | Keep G1 open; require semantic, provenance, strict-output and resource audit, not just process completion |
| 2026-09-18 | Community findings + execution-plan review | Improve candidate coverage before heavier models | Prioritized structure-only snapshots/mass-formula candidates, OOF fingerprints and separate analogue propagation; added leakage, adduct, cap and aggregation checks | Adopt revised backlog order; no new model training, submission or G1 closure claimed |
| 2026-09-18 | Evaluator/validator corrections | Repair the morning audit's concrete defects | A row-mask forwarding, uncapped coverage, covered-query conditional MRR and basic invalid/duplicate/over-25 checks implemented | Local fixes verified; not end-to-end G1 acceptance |
| 2026-09-18 | Checkpointed rerun at 12:39 | Expose progress without waiting for the aggregate report | PID 198260: library scan reached 621 batches / 113,376,230 comparisons; B active; final report absent at evening audit | Preserve ongoing work; counters are not resumable state, saved metrics or a reliable ETA |
| 2026-09-18 | Kaggle CLI release verification | Check authoritative training data against a fresh download | Size/hash/schema/rows/row groups all match; verified staged duplicate permanently deleted on request | Retain original data and comparison report; no rebuild justified by release mismatch |
| 2026-09-18 | Non-competing preparation | Improve small contracts and document next steps | Basic formate alias tested; PubChem physical-format/count audit and Phase 3A/optimization drafts recorded | Drafts/source audit have eligibility/provenance limitations; no new index or GPU runner built |
| 2026-09-18 | Evening progress audit | Reconcile source, tests, claims and TODO | 24 lightweight tests pass with 2 integration tests excluded; independent metric fixes verified; wrapper failures, entrypoint failure, metadata loss and strict-output gaps reproduced | Keep G1 open; prioritize A1-A8, narrow completed claims and retain historical evidence |
| 2026-09-19 | Phase 3A fallback/category audit | Verify fallback equivalence and separate structure/reference failure classes | 475 valid audit queries: 28 structure-pool misses; 260 had a fold-eligible local identity in the mass window; 215 used the audit fallback route; fallback pool recall 90.23%, @25 88.84%; 0 mass-order equivalence mismatches; oracle formula gate showed no improvement | Keep formula gate non-deployable; do not call the 260 count a positive spectral-neighbour count; preserve categories; add maintained tests and repeat on a genuinely pre-registered untouched partition |
Blockers and decisions
reports/official-scorer-certification-20260920/evidence.json. C is a proxy, not true hidden novelty. The 1,184-versus-1,151 documentation discrepancy is pinned locally and is not itself a remaining implementation blocker.