Data status

What the numbers are numbers of.

Onkydra holds two kinds of real data, and they are counted in different units that are easy to confuse and consequential to confuse. A cell is not a donor. An assay record is not a patient. Neither is a simulated draw. Here is every unit, how the counts collapse from files to donors, and what the files still cannot answer.

The single-cell reference

The malignant cells the cell-state work is built on. Read top to bottom: each row is a weaker claim than the one above it. It is called a DMG reference throughout this product and it is a mixed cohort, which the table under it states rather than implies.

Donors4242 distinct donor identifiers, recovered from the P-<n> component of ID_paper. 40 follow that convention and 2 (BT2016062, BT2018022) use a different scheme whose relationship to the P- series cannot be verified from this file, so treat 42 as an upper bound. The 47 sample_id values are GEO records, not donors: four donors were assayed on two platforms and one contributed two samples.
Patient-sample pairs43Distinct ID_paper values, each one donor and one sample. Four donors were assayed on two platforms, so this is fewer than the GEO record count above it.
GEO records47Distinct sample_id values, each a GEO file prefix. This is the number the CellOracle artefact's provenance calls '47 real 10x samples'. It is a file count, not a biological one.
Cells40,000After malignant-cell filtering and subsampling. Cells within one sample are not independent observations.

Source dmg-reference.h5ad, sha256 4fcf27883c2de011b0b6fc11…, derived by backend/services/celloracle/derive_atlas_facts.py.

What that reference is a cohort of

The reference is not a DMG cohort in the sense the name implies. Of 40,000 cells, 19,172 come from a record that is BOTH high-grade glioma and a midline site. The rest are hemispheric high-grade glioma, ependymoma, or a record whose only stated location is the string 'brain'. The ependymoma donors are recorded elsewhere as comparators from the original study; the hemispheric fraction is recorded nowhere and is the new finding here.

Diagnosis as depositedSite groupAssay recordsDonorsShare of cells
high-grade gliomamidline252247.9%
high-grade gliomaunspecified6619.7%
high-grade gliomahemispheric111019.3%
Ependymomamidline5513.1%

H3 K27M status is not a characteristic on any of the 61 GEO records, so no cell in this reference can be confirmed H3 K27M-mutant from the deposit. Diagnosis is an entity label, not a genotype.

The mix is not spread evenly across the cell states, which decides whether it matters to a given number: AC-like is 1.6% non-glioma, OC-like is 0.3% non-glioma, OPC-like is 0.6% non-glioma, Other is 23.5% non-glioma. The comparator carries almost all of it.

Site grouping is Onkydra-derived by literal membership of two stated string sets, with nothing merged. Derived by backend/services/celloracle/validate_states.py.

The molecular anchor

A separate dataset, counted in patients, and the only place in Onkydra where the word patient is correct.

60

real H3 K27M cases, pooled from Gröbner SN et al. The landscape of genomic alterations across childhood cancers. Nature. 2018 (PMID 29489754). DKFZ / Pfister Lab paediatric pan-cancer cohort, cBioPortal study pediatric_dkfz_2017 (n=53) and PMID 33242424 (Petralia, Cell 2020), the CPTAC/CHOP pediatric brain cancer proteogenomic cohort (n=7).

13 binary driver features per case. No expression, no methylation, no copy number, and no treatment-response or survival outcome, so nothing downstream of this can be calibrated against what happened to a patient.

What is known about the 60 cases

Coverage, not values. Each row says how many of the 60 cases carry the field at all, because for most of them the honest answer is a fraction and prose turns a fraction into a description of the whole cohort.

Anatomical location7 of 60

DKFZ records no location field at all, so “midline” is an inference for 53 of the 60. Where it can be checked, one of the seven is cerebellar.

brain_cptac_2020: Midline 6 · Cerebellar 1

Biopsy versus autopsy0 of 60

Recorded by neither study. It matters because autopsy material is post-treatment and post-progression by definition, so pooling it with biopsy material pools two disease states.

pediatric_dkfz_2017: Primary 51 · Relapse 2

brain_cptac_2020: Initial CNS Tumor Surgery 5 · Progressive surgery 1 · Repeat resection 1

Treatment status at collection7 of 60

Seven cases. Within those seven, six are treatment-naive at collection while five record chemotherapy and radiation: different timepoints, and quoting either alone misleads.

brain_cptac_2020: Treatment naive 6 · Post-treatment 1

The study’s own diagnosis60 of 60

Not Onkydra’s label. The pooled entity is not cleanly DMG.

pediatric_dkfz_2017: High-Grade Glioma, NOS 52 · Pilocytic astrocytoma 1

brain_cptac_2020: Pediatric High Grade Gliomas 7

Sex60 of 60

Recorded because an imbalance here is a confounder nobody has controlled for.

pediatric_dkfz_2017: Female 35 · Male 18

brain_cptac_2020: Male 5 · Female 2

Sequencing assay53 of 60

WES and WGS do not share detection sensitivity, so a per-gene frequency pooled across both is pooled across two assays.

pediatric_dkfz_2017: WES 34 · WGS 19

Cell states, and what is in them

Other21,677 cells54.2%A residue, not a biological state. Whatever the source label did not map to OPC, astrocyte or oligodendrocyte.
OPC-like9,955 cells24.9%Scored by OLIG1, OLIG2, PDGFRA: 3 of the 7 intended markers survived the reference's gene selection.
AC-like6,200 cells15.5%Scored by AQP4, APOE: 2 of the 5 intended markers survived the reference's gene selection.
OC-like2,168 cells5.4%Scored by MBP, PLP1, MAG, CLDN11: 4 of the 5 intended markers survived the reference's gene selection.

What treating cells as observations would cost

Standard errors scale with the square root of the number of independent units. Computing one over cells rather than donors would make every interval this much too narrow:

31x

too narrow, if 40,000 cells were treated as 40,000 observations instead of at most 42 donors.

What we asked the source

What we asked of the source data, and what it said. Answered from the GEO submission, which carries per-sample detail the processed reference does not.

How many donors, and are they all glioma?

Answered on 2026-08-10 by reading the GEO submission the reference was built from, which carries per-sample characteristics the processed h5ad does not. 42 distinct donor identifiers across 61 GEO records (43 patient-sample pairs), so 42 is exact as an identifier count rather than an upper bound. 37 are high-grade glioma and 5 are ependymoma (BT2016062, BT2018022, P-2077, P-6292, P-6431), which were comparators in the original study.

Donor counts on this atlas name the glioma subset. See backend/data/registry/atlas-donors.json.

Which GEO series do these records belong to?

Answered on 2026-08-10. GSE210568 is the SuperSeries; GSE210565 is the scATAC SubSeries and GSE210566 the RNA SubSeries, under BioProject PRJNA866178. The notes naming both were describing different levels of the same submission.

Cite GSE210568 for the study and the SubSeries for a modality.

Anatomical location, sex and age per donor.

Present on 61 of 61 samples in the GEO submission, though absent from the processed reference's obs columns. Retrieved 2026-08-10. Biopsy versus autopsy and treatment status are still not stated anywhere in the deposit.

Location, sex and age are available to the atlas. Timing and treatment status are not, and no note may state them.

Files, and their licences

dmg-reference.h5adsingle-cell referenceno licence stated

cell states, marker programme scores, every count in unitChain

sha256 4fcf27883c2de011b0b6fc110ed00118bfea48a64b876b321aa5d82169854d18

No licence is stated on the GEO deposit. Verified 2026-08-07 by reading the accession records and NCBI's own policy, which says NCBI places no restrictions of its own AND that it "cannot provide comment or unrestricted permission concerning the use, copying, or distribution" because depositors transfer no rights to NCBI. So there is no positive grant of rights from anyone. Attribution is requested, not licensed. This is the PROCESSED layer, which is the open one. The raw reads for the same samples are at EGA under controlled access, and their data access agreement forbids disclosing derived material to anyone outside the named project, so they must never be ingested.

dmg-base-grn.parquetbase gene regulatory networkno licence stated

the network every in-silico knockout propagates through

sha256 5c5971bab9cc409ddbd45882f82c941da3eb6e2f35aa3b6c2bee014a5b455a3a

No licence is stated on the GEO deposit. Verified 2026-08-07 by reading the accession records and NCBI's own policy, which says NCBI places no restrictions of its own AND that it "cannot provide comment or unrestricted permission concerning the use, copying, or distribution" because depositors transfer no rights to NCBI. So there is no positive grant of rights from anyone. Attribution is requested, not licensed. Same deposit family as the reference (GSM accessions are contiguous), and the same processed-versus-EGA split applies.

2 of 2 source files carry no grant of rights from anyone. Until that changes, treat commercial use of anything derived from them as unresolved.

The terms have been read: these come from GEO deposits that state no licence, and NCBI's own policy adds no restriction of its own while expressly declining to grant permission on a depositor's behalf. “No restrictions” and “permission granted” are different sentences, and only the first one is on the page.

What we will not do with these

See also current capability and limitations.

Research use only. Onkydra holds no patient-identifiable data.