Hello, I reviewed the GLASS Data Dictionary, current release schemas, public pipeline/configuration, and relevant discussion threads, but I still have a few unresolved provenance questions. For a sequence-context opportunity analysis, I’m trying to identify the genomic territory underlying the released GLASS WGS mutation calls. I found references in the public pipeline to scattered_wgs_intervals/scatter-50, but not the interval files themselves. Are any of the following available? - the original WGS calling intervals, or a combined BED/interval-list file, including reference build and any exclusions/padding; - sample-specific callable-region masks for the tumor/normal pairs, with aliquot IDs and filtering criteria; - release-specific documentation on whether downstream filtering further restricted the effective territory in which somatic SNVs could be retained. I’m looking only for existing files or provenance information, not raw sequencing or new processing. Separately, I’m exploring a longitudinal timing analysis using treatment as an external ordering constraint. The documentation clarified the basic surgery/sample/aliquot mappings and TMZ field definitions, but I’m still unsure about three points: - Is there an existing source that can clarify prior TMZ exposure before the first recorded surgery on a patient-specific basis? - For longitudinal copy-number analysis, is there an existing release-compatible file or annotation indicating the selected CN/WGD solution and associated QC/provenance? - Is there any processed GLASS resource that can help determine whether a copy-number gain or other event observed before treatment represents the same ancestral event retained at recurrence, rather than a similar event arising independently? If any of these resources were not retained or cannot be shared, that clarification would also be very helpful. Thank you!

Created by Jaden Chun jadenchun5
Thanks for the detailed questions. **Part 1 — Genomic territory underlying the WGS somatic calls** 1. **Original WGS calling intervals: **We do not have the original interval files or a combined BED/interval-list file available to provide, including the associated build, exclusions, and padding details. 2. **Sample-specific callable-region masks:** We do not have callable-region BED masks for the tumor/normal pairs. We do have per-aliquot cumulative coverage summaries in analysis.coverage, keyed by aliquot_barcode. These report base counts at different depth thresholds but do not identify genomic coordinates. 3. **Release-specific filtering documentation:** We do not have a separate document or genomic mask defining how downstream filtering restricted the effective territory for retained somatic SNVs. The released tables contain variant filters and coverage summaries, but these do not reconstruct that territory. The coverage summaries support depth-adjusted mutation-frequency calculations, but they are insufficient to define the genomic territory needed for a sequence-context opportunity analysis. **Part 2 — Longitudinal timing and treatment ordering** 1. **Prior TMZ exposure: **There is no structured field for TMZ exposure before the first recorded surgery. Treatment fields in clinical.surgeries describe treatment between that surgery and the next. Surgery 1 generally represents the treatment-naive primary, but patient-specific confirmation requires review of the clinical records and free-text fields. 2. **Copy-number solutions and WGD:** TITAN and Sequenza solutions, including purity/cellularity, ploidy, and available fit metrics, are stored in variants.titan_params and variants.seqz_params. Corresponding segmentations are in variants.titan_seg and variants.seqz_seg. There is no explicit WGD flag; deriving one would require a defined method using the copy-number data, rather than ploidy alone. 3. **Shared versus independently recurring events:** Joint PyClone-VI results and pairwise mutation summaries support assessment of shared and private SNVs, while separate copy-number comparisons support assessment of shared and private CNVs. However, not all PyClone-VI results are currently included in the release. These resources can support inference of shared ancestry but cannot definitively distinguish an inherited event from an independently recurring event. Best, Erica

Availability of WGS calling intervals or callable-region masks page is loading…