Hi @LucaChiesa Just download the NTC release and done some primary data validation, would you please help with the following questions? 1. The dictionary calls SMILES the deduplication key and compound the “BCM compound ID of the first-observed duplicate.” Does each output row represent an original DEL encoding, or should the supplement contain one row per deduplicated canonical SMILES? 2. What exactly does “first-observed duplicate” mean? Were multiple compound IDs or DEL encodings sharing a canonical SMILES reassigned to one representative compound value while their original rows were retained? 3. For repeated compound/SMILES rows, do the rows represent distinct encodings, barcodes, libraries, batches, or replicates? Is a discriminator missing from the published file? 4. How should repeated rows be interpreted when count_NTC or zscore_NTC differs: retain separately, sum counts before normalization, use a representative observation, or apply another BCM aggregation rule? 5. Is the stated count_PGK2 = 0 filter applied to original source rows before representative-ID/SMILES deduplication? Would that explain why 7,403 representative compound values also occur in the original PGK2 export? Also, the description says compounds “bound to PGK2 (count_PGK2 >= 0)”; should that threshold read > 0 or >= 1? 6. Were supplement counts generated from the same processed sequencing output as the original release? If the supplement was reprocessed, were counts changed and were zscore_NTC values recomputed against a different population? 7. The dictionary marks canonical SMILES as required, but 2,300 rows contain empty SMILES. Is this a publication defect, and should those structures be recovered using the representative compound ID or another BCM/library mapping? Cheers, @asfarasimconcerned

Created by Jing Huang asfarasimconcerned
Hello @asfarasimconcerned, thank you for highlighting these issues. I think @vonboss should have the answers you are looking for.
We found that the Aircheck dictionaries clarify that canonical SMILES is the chemical-level deduplication key and that, for the original PGK2 dataset, duplicate counts were summed, z-scores were Stouffer-combined, historic_hits used the maximum, and the first compound ID was retained. The NTC supplement dictionary does not explicitly state that the same aggregation was applied, and the released file still contains 5,713 repeated nonempty-SMILES groups. Applying the original PGK2 count-sum rule gives 7,378 overlapping canonical SMILES, of which 4,438 counts agree and 2,940 disagree. Could you please clarify specifically: 1. Was the supplement’s count_PGK2 = 0 filter applied before or after canonical-SMILES deduplication and aggregation? 2. Are the repeated SMILES rows pre-aggregation source encodings/occurrences, or unintended duplicated output rows? 3. Should the same count-sum and Stouffer-z aggregation be applied to the supplement? 4. Were the supplement counts and z-scores generated in the same processing run and against the same normalization population as the original PGK2 table? We can use canonical-SMILES presence conservatively, but will not use count or z-score magnitude without this clarification.

Questions to the released PGK2_NTC_supplement.parquet page is loading…