← Back

📑 Contents

PubChem identity reconciliation — 2026-09-19ResultInterpretationDecision

PubChem identity reconciliation — 2026-09-19


Result


Read-only comparison of the pinned source's declared identity namespace with the completed Phase 3A index:


  • Source rows: 240,536
  • Source-declared unique identity14: 231,755
  • Index rows/distinct recomputed identities: 231,749 / 231,749
  • Declared-only identities: 578
  • Computed/index-only identities: 572
  • Net declared-minus-indexed difference: 6

  • The exact identity sets are preserved in:


    reports/phase3a-pubchem-identity-reconciliation-20260919.json


    Interpretation


    The six-identity net difference must not be described as six missing catalogue structures. The index intentionally recomputes the authoritative scorer identity from the structure and deduplicates on that recomputed value. The source-declared identity field is therefore an audit comparison field, not the production namespace.


    The existing index remains the correct artifact for scorer-compatible candidate generation, subject to the separate packaging/provenance limitations already recorded:


  • source is physically plain TSV despite the .gz suffix;
  • source novelty is not certified;
  • broader catalogue completeness is not established;
  • the 578/572 set differences remain provenance evidence, not silently discarded data.

  • Decision


  • Keep the production index unchanged.
  • Preserve the full declared-only and computed-only identity lists in the JSON artifact.
  • Do not proceed to learned ranking on the basis of the net count alone.
  • The remaining PubChem source task is the separate P0 re-normalization/packaging audit, not a rebuild triggered by this six-count discrepancy.

  • Script: scripts/reconcile_phase3a_pubchem_identities.py.