Hi organizers, While exploring the data I noticed that PGK2\_selection.parquet has no rows with count\_PGK2 = 0; every row has at least one read in the standard PGK2 selection. So compounds sequenced in the inhibitor or NTC arms but not sampled in the PGK2 arm appear to be absent from the file entirely. Was that data collected, and could it be released? Two reasons it would help: - **Allosteric binding is hard to model from the provided data even in principle.** The Hit Selection page distinguishes orthosteric from allosteric binders by whether counts are retained in the inhibitor arm. But if the file only contains inhibitor-arm reads landing on compounds also observed in the PGK2 arm, and each arm _independently_ samples only a small fraction of the library, then that overlap is a small and hard-to-quantify subset of the competition data. Summed over the file, count\_PGK2 = 7,965,845 vs count\_PGK2\_with\_inhibitor = 49,707 (and count\_NTC = 5,649). The PGK2 arm covers under 1% of the enumerated library, so if the arms sample independently those 49,707 reads would be roughly 1% of the inhibitor arm, implying several million reads from the inhibitor arm that are not represented in the file. - **Matrix binders are currently almost invisible.** The Hit Selection page also describes non-specific / bead binders as high in NTC and low elsewhere, but a compound with count\_NTC > 0 and count\_PGK2 = 0 wouldn't be in the provided file, so that population can't be modelled - only matrix affinity among compounds that also happened to have a PGK2 read. Thanks! — Alex

Created by Alex Izvorski memes
Alex (@memes), - I am talking with Baylor about generating a file with count_NTC >= 1 regardless of count_PGK2. - I was able to get the count_PGK2_with_inhibitor >= 1 regardless of count_PGK2. I think the DREAM x Cache organizers are still making a decision on if, when, and how to release this data. We will probably wait until we get a firm answer from Baylor on the count_NTC >= 1 regardless of count_PGK2. @LucaChiesa is more in the loop of these conversations. - Unfortunately, we will not be able to get the naive library reads. You might be able to back calculate the naive reads for the compounds with published normalized zscores. Sorry, I have also been very interested to get a hold of this data. Best, Seth
Seth (@vonboss), Thank you for trying to get the additional data! Combining this with the question I had asked in the other thread, there are actually three different datasets of raw read counts that would be very useful: * count\_NTC >= 1 regardless of count\_PGK2 * count\_PGK2\_with\_inhibitor >= 1 regardless of count\_PGK2 * the _naive library reads_ used to measure representation of synthons in the library (I think this is not the same as count\_NTC which only measures matrix binding) I hope it would be possible to get all three from Baylor - each would add a significant new dimension to the competition, and all three together would of course give maximum and complementary information to work with. Thanks again, Alex
Alex (@memes), Another important column that you can use to distinguish matrix/non-specific binders in the `nHH` column (e.i. number of historic hits). The `nHH` column reveals if a compound repeatedly has high counts in multiple selection screening. I am still working on trying to get the full NTC selection results (all compounds with count_NTC >= 1 regardless of count_PGK2). I'll keep you updated. Best, Seth
Alex @memes, I agree this data would be useful. The full count_PGK2_with_inhibitor and count_NTC data should have been collected by Baylor College of Medicine, but it was not provided to us. I will reach out to BCM to see if they can provide the data for the challenge. Seth
Hello Alex ( @memes ), I think @vonboss and @jmchap might have the answers you are looking for. Luca

Are control-arm-only reads available? page is loading…