← Back

📑 Contents

Phase 3A post-freeze reserved policy readout — 2026-09-19StatusProtocolAggregate resultsInterpretationDecisionProvenance

Phase 3A post-freeze reserved policy readout — 2026-09-19


Status


This is a post-freeze stability readout, not an untouched prospective estimate. The identity partition was created after the earlier full-cohort and fold-4 diagnostics, so it must not be used as an unbiased policy-selection result.


Protocol


  • Cohort: strict-validation structure-only identities.
  • Partition: sha256(identity14) first uint32 modulo 5 == 0.
  • Readout identities: 1,171 of 5,974 cohort identities.
  • Query rows seen: 57,539; valid query rows: 56,150.
  • Structure index: artifacts/phase3a/pubchem-structure-index.sqlite, 231,749 rows.
  • Retrieval fields: query adduct and precursor m/z only. Identity/formula are used only for scoring and diagnostics.
  • Candidate identities are deduplicated before rank/cap metrics.
  • Policies:
  • - P0: all adducts ±10 ppm.

    - P1: [M+Na]+ ±30 ppm; all others ±10 ppm.

    - P2: ±10 ppm first; widen to ±30 ppm only when the strict pool is empty.

  • Caps: @25, @100, @1,000.

  • Aggregate results


    PolicySpectrum recall@25@100@1,000Molecule-any recallCandidate P50/P95/maxFormula-gated P50/P95
    P0 strict1082.180%73.955%82.073%82.180%93.339%19 / 64 / 1559 / 58
    P1 Na±3085.302%75.802%85.195%85.302%93.424%21 / 77 / 1559 / 60
    P2 empty fallback84.037%75.348%83.931%84.037%93.681%20 / 65 / 1559 / 60

    Interpretation


    1. The adduct-specific Na widening effect is stable in direction: P1 improves query recall over P0 by 3.12 percentage points, but expands the candidate P95 from 64 to 77 and does not improve molecule-any recall materially.

    2. P2 retains most of the Na recovery at lower typical pool cost than P1: +1.86 points spectrum recall over P0, +1.39 points @25 recall, and only +1 candidate at P50 / +1 at P95. Its molecule-any recall is the highest of the three.

    3. P2 is therefore the current provisional cost/coverage winner, but not production-frozen: its empty-pool fallback is a simple deployable rule, not a calibrated query-quality detector.

    4. [M+Na]+ remains the principal weakness. Under P0 it has 63.16% spectrum recall and 50.14% @25 recall; P1 raises these to 93.18% and 67.90%; P2 reaches 74.08% and 59.03% because only empty strict pools widen.

    5. The aggregate recall gap is not evidence of PubChem absence. It remains primarily a query precursor-mass compatibility problem, with particularly poor behavior for sodium adducts and lower-mass strata.

    6. No result here is a spectral-ranking score, hidden-test estimate, official CASMI score, or Top-5 evidence.


    Decision


  • Keep P0 as the conservative baseline.
  • Keep P2 as the preferred provisional retrieval policy for the next deployable diagnostic.
  • Do not promote P1/P2 to a final production policy until a genuinely pre-registered untouched identity partition exists.
  • Do not train a ranker until the final retrieval policy and leakage-safe evaluation partition are frozen.

  • Provenance


    Machine-readable companion: reports/phase3a-reserved-policy-readout-20260919.json.

    Script: scripts/readout_phase3a_reserved_policies.py.