The read-only audit independently re-parsed all 231,749 completed index rows from their stored representative SMILES.
Training-overlap labels
Recomputed unique training overlap: 12,331
Stored unique training overlap: 12,331
Training-overlap mismatches: 0
This part of the index label contract agrees.
Scaffold labels
Recomputed scaffold-overlap unique count: 162,495
Stored scaffold-overlap unique count: 163,344
Rows with a scaffold mismatch or scaffold-overlap disagreement: 9,594
20 mismatch examples are retained in the JSON artifact.
The audit deliberately reports both scaffold-string and scaffold-overlap differences together. A follow-up must separate those two causes before any scaffold-based novelty or Regime-C claim is made.
Stored representative structures
Stored rows that failed RDKit parsing: 7
Missing stored identities: 0
This is a material artifact-quality issue. The builder originally counted source SMILES as valid before storing canonical_smiles; the stored representative must itself be parseable for downstream fingerprint/scaffold use.
Decision
Do not mark PubChem source normalization complete.
Keep the current index unchanged and out of any scaffold-based novelty claim.
Do not use stored scaffold labels for C or novelty filtering until the mismatch causes are separated and fixed/rebuilt.
Training-overlap labels are independently consistent, but this does not repair the representative-SMILES/scaffold problem.
The raw source-row identity audit was attempted but terminated after local audit-script failures/timeout; no raw-source recomputation is claimed complete.