Hello organizers, Two questions - one on the incentives, one on the data. **1. Co-authorship condition: "and" or "or"?** The Incentives page says: "All participants whose workflow performs better than random selection in the Blind test phase and Active learning phase will be invited as co-authors on an article to be published in a scientific Journal." The main challenge page says: "All participants who advanced to the Prospective phase or whose predictions are better than random in the Blind test step or in the Active learning step will be invited as co-authors on an article to be published in a scientific Journal." Could you confirm which applies? Specifically, is performing better than random in the Blind test step alone sufficient, or is it required in both the Blind test and the Active learning steps? It changes how participants plan for October. **2. Two sub-libraries absent from `PGK2_selection.parquet`** *(Edited only to fix formatting - the underscores in the library names had been swallowed by markdown.)* The Library page lists 12 sub-libraries totalling 898,311,048 enumerated compounds. Parsing the library prefix from the `compound` column of the selection file, only 10 appear: `qDOS18_3` - 2,883,145 rows `qDOS30` - 1,780,061 `qDOS42_2` - 1,025,549 `qDOS11` - 842,700 `qDOS23` - 476,554 `qDOS9` - 362,642 `qDOS19_2` - 101,482 `qDOS14` - 86,768 `qDOS19_1` - 82,564 `qDOS24` - 61,605 `qDOS4` (234,685,621 enumerated compounds) and `qDOS13` (4,287,381) contribute no rows at all. Is this expected - were those two libraries not part of this selection - or is it a gap in the released file? It matters for negative sampling, since a compound from a library that was never assayed is not the same kind of negative as an unobserved compound from a library that was. Separately, the combined building block file `qDOS_all.building_block.parquet` covers 11 libraries and does not include `qDOS19_1`. Thanks very much, Arnav Agarwal

Created by Arnav Agarwal ArnavAgarwal
Hello @ArnavAgarwal, for the first point is an "or", we want to feature in the paper methods leveraging DEL data that perform well on blind hit discovery, and methods that perform well on hit discovery once a few initial hits have been identified. For the missing data @vonboss and @jmchap can probably give you a better answer since they worked on this data. Luca

Co-authorship condition ("and" vs "or"), and two sub-libraries absent from PGK2_selection.parquet page is loading…