Hi organizers, A small inconsistency affects scoring, not just convenience. The Data page states "A library of 400K compounds was screened by Affinity Selective Mass Spectroscopy (ASMS), resulting in 1500 ASMS hits". The Evaluation page states that a chemical series is "a common substructure observed in less than 100 molecules in the full 460K ASMS screening library". The validation and test splits together contain 428,960 compounds, which sits between the two. Since N_clusters is the primary ranking key and a chemical series is defined by an MCS occurring in fewer than 100 molecules of that library, the denominator matters directly for how many series a submission is credited with. Could you confirm which figure is correct, and whether the 428,960 released compounds from the full screening library or a subset of it? Thanks very much,

Created by Hsuan-Kai Wang DoingWell
Hi @LucaChiesa, Thank you — this resolves the scoring-denominator question. We understand that the full screened library contains approximately 460K molecules, and that this full library was used for the “observed in fewer than 100 molecules” substructure-frequency threshold. We also understand that the official chemical-series assignments cannot be reproduced from the released data, so we will not treat any local scaffold or clustering proxy as equivalent to the official N_clusters. One small terminology check, only for dataset provenance: when you refer to compounds being removed from the “test set,” do you mean the complete 428,960-compound released evaluation collection (validation plus blind-test splits), or specifically the 184,632-compound blind-test split? This does not affect our understanding of the official cluster score. Many thanks for the clarification.
Hello @DoingWell, The 400K library and the 460K ASMS library are the same library. While the full library contains 460K molecules, in the test set we removed all compounds that were not properly detected during screening, and ASMS hit compounds not confirmed in the kinase inhibition assay. Frequency of substructures was calculated against the full 460K library. Very important, trying to reproduce the chemical series classification results is not possible with the available data given the nature of the classification algorithm, see [this answer for more information](https://www.synapse.org/Synapse:syn75349604/discussion/threadId=15113?replyid=36732).

ASMS library size - Data page says 400K, Evaluation page says 460K page is loading…