Limitations

Why a result in a model is not a result in a patient.

This gap is not closed here. This page is what Onkydra knows about the distance between a model system and the patient it stands in for, and what it does not.

What a model match is currently made of

1 of 6 modalities are computed. The rest are not, and a match that ignores them is a match made on partial information.

wiredMutation-set overlapJaccard between the patient's co-mutations and a hand-curated per-line mutation set. Coarse, and the only similarity that runs today.
absentCell-state compositionWould compare the line's state mixture against the patient state. Neither side is ingested: no anchor case and no named line carries a transcriptome.
absentTranscriptome alignmentThe published approach, and it needs expression on both sides. The transcriptomic data this product does hold belongs to donors and to a xenograft that cannot be linked to either side.
absentCopy-number compatibilityIngested on both sides since 2026-08-12 and still not runnable. The anchor carries focal-event calls harmonised from two callers, the model systems carry segment-level calls from a third, and no subject has been measured both ways, so nothing says what a count on one side means against a count on the other.
absentAssay compatibilityNot a similarity, because a patient has no assay compatibility to be compared against. It is a hard filter on which line can produce the readout you need. Recorded per line from retrieved published sources, and re-checked on 2026-08-12: every quote behind it was found in the source it is attributed to, which is a claim this product has been wrong about elsewhere.
absentAvailability, cost, lead timeAlso not a similarity. A route to obtain each line is now recorded from retrieved published sources. No price and no lead time exists for any line, and neither is estimated.

All 6 curated DMG lines now carry a mutation set, so the one wired comparison reaches every model system. That closes a coverage gap and not the limitation above it: mutation-set overlap is still the only similarity that runs, and a line matching well on it has been matched on co-mutations alone.

The hard limits of the method, whatever the indication

These do not change when the disease does. Everything below them is about one indication's evidence and is derived from the registers that hold it.

A model system is not a patient

A cell line, an organoid or a xenograft is an approximation chosen for tractability. It has been through passage, selection and a culture environment with no immune system, no vasculature and no blood-brain barrier. A similarity score is a statement about named features under a named metric, and no score makes it a patient substitute. Onkydra will not call one an avatar or a digital twin.

Nothing here has measured the size of the gap

The useful question is whether a higher similarity predicts a smaller transfer error. Answering it needs paired data: the same perturbation, in a model and in the patient material it was matched to. Onkydra has no such pairs ingested, so it cannot tell you how much to discount a model result, and it says nothing rather than guessing.

The scores are orderings, not probabilities

Every number the product emits ranks candidates under stated assumptions. None is calibrated against an observed outcome, because no outcome dataset exists yet. A number between 0 and 1 here is not a chance of anything.

Cells are not independent observations

Forty thousand cells come from a few dozen sample records. An interval computed over cells would be far narrower than the evidence supports, so the unit of analysis has to be the donor or the sample, and the donor count is currently unresolved.

Simulated profiles are not evidence

Drawing profiles from a joint fitted to sixty real cases creates no new patients, outcomes or statistical power. They are useful for propagating uncertainty in the sixty and for nothing else.

The mechanistic layer covers transcription factors only

The in-silico knockout works over a gene regulatory network, so it applies to factors that appear in that network as regulators. A kinase, a receptor, a scaffold or a drug has no binding motif and is not in it. For those targets the product abstains rather than substituting a heuristic and calling it mechanistic.

Limits per indication

A retraction reaches this page as soon as it lands. An indication nobody has measured has an empty list, and an empty list is not a clean bill of health.

H3 K27M DMGdraft · 54 recorded
Ewing sarcomaplanned · 0 recorded
MPNSTplanned · 0 recorded
ATRTplanned · 0 recorded
High-risk neuroblastoma (MYCN)planned · 0 recorded
Anaplastic thyroidplanned · 0 recorded
Rhabdomyosarcomaplanned · 0 recorded
Paediatric AML (KMT2A)planned · 0 recorded

H3 K27M DMG

capability registry · 7

The library is draft, not general

n=60 anchor via cBioPortal: DKFZ (53), CPTAC (7). 13 driver features. Not expert-reviewed.

lib/capability/registry.ts

Geneformer (independent second opinion) is draft

RUNS, AND ITS ONE EXTERNAL CHECK FOUND NO SUPPORT. Uncalibrated in-silico gene deletion on Cloud Run, Apache 2.0, with no null distribution behind it.

Rest of the entry, 736 more words

Those are different claims and both are load-bearing. Geneformer-V1-10M serves in-silico gene deletion on Cloud Run (CPU, scale to zero, weights baked in and sha256-verified at build time); a served /knockout returns a real number with its provenance. Licence Apache 2.0, confirmed 2026-08-16 from the model card's own YAML frontmatter rather than from CLAUDE.md, with no gating fields and no commercial restriction. Deploying it validated NOTHING, and it did not rescue CellOracle either. THE RESULT, pre-registered at gfscope_6d5882c8b8f5d236489a9f6a before any forward pass and published as run: against the measured pooled CRISPR TF screen in DIPG17 (GSE267415), the identical column and identical 41 genes CellOracle was scored on, Geneformer scores Spearman -0.023, 95% CI [-0.337, 0.295], where CellOracle scored -0.026, 95% CI [-0.339, 0.293] and FAILED. This arm is recorded NOT TESTABLE rather than failed, which is a demotion and not a reprieve: the pre-registered criterion reads the SIGN of the correlation, and the sign of this readout is retracted by the same artefact that computes it, because 33 of 41 targets come out positive for a reason that has nothing to do with the target. A criterion whose input is disowned returns no verdict, so the pre-registered test was not performed rather than failed. Nothing is withheld by saying so: the signed figure does not clear the bar either, its interval contains zero, and it is a pass under no reading. On MAGNITUDE, which is the quantity that may be quoted once the sign is gone, it is 0.092, 95% CI [-0.231, 0.397], also containing zero; that figure is post hoc and carries no verdict. It is absence of external support, NOT refutation, because a CRISPR screen measures what a cell needs to PROLIFERATE and neither model predicts that. Two nulls also license nothing about in-silico perturbation as a class, and that inference was drawn here until 2026-08-16 and is withdrawn: V1-10M was trained on a corpus that excluded malignant cells, so it is OUT OF DISTRIBUTION on DMG tumour cells and its null may be an out-of-distribution failure rather than an absent signal. Head to head the two do not correspond either: raw Spearman 0.086, 95% CI [-0.237, 0.391], and 0.121, 95% CI [-0.207, 0.425] controlling for detection depth, both intervals containing zero. Reported together because the raw figure alone is not reportable. Top-10 overlap against the measured screen is 3 of 10, where two independent 10-subsets of 41 share 2.44 by chance and 3 or more happens 46% of the time; CellOracle's 4 of 10 happens 18% of the time. Neither is distinguishable from chance, and 2 of Geneformer's 3 (MYBL2, NR4A3) are on the screen's OWN common-essentials list, which is the easiest way into the top 10 of a proliferation screen without predicting anything about this tumour. The reason the two orderings differ is concrete and measured: |effect| against detected fraction runs -0.344 for Geneformer and +0.496 for CellOracle, so their magnitude orderings are detection-driven in OPPOSITE directions. Three further limits a reader needs. 33 of 41 targets return a POSITIVE shift, meaning deletion moves cells TOWARD the anchor centroid, because discarding information drifts an embedding toward the population mean; the sign is not a direction of identity change. The largest magnitudes sit on the least-detected genes, and the three largest are POU2F3 over 42 cells, TAL1 over 30 and GATA2 over 71, of 9,955 OPC-like cells against a floor of 25, with all three in the reported top 10. And there is NO null distribution and NO per-target uncertainty anywhere in the sweep: the per-cell readouts are averaged and discarded, nothing is permuted and nothing is bootstrapped, so the ordering has no calibrated floor and no magnitude in it is known to be distinguishable from no effect. The blocker recorded here until 2026-08-16 is CLEARED: the served reference was subset to 3,000 highly variable genes, exposing only 1,798 with a Geneformer token, which is out-of-distribution input for a rank encoder. It was rebuilt full-gene from GSE210568_RAW.tar to 30,092 genes and 17,745 tokenisable, over the SAME 40,000 cells, verified identical cell-for-cell and label-for-label. The service refuses to compute below 12,000 tokenisable genes rather than computing and caveating. Shared with CellOracle: the same cells, the same cell-state labels, the same dissociation confound and the same detection floor. Not shared: no GRN, no fit to this reference, and a corpus that PROVABLY predates the data, since V1-10M was trained 2021-06 and GSE210568 went public 2022-08-04.

lib/capability/registry.ts

DMG dependency rows is draft

11 of 17 rows serve a CCMA CRISPR beta. 0 of 17 cited values were traceable.

lib/capability/registry.ts

Measured TF dependency (GSE267415) is draft

1,430 transcription factors by pooled CRISPR in DIPG17. Measures proliferation, not cell state.

lib/capability/registry.ts

LINCS L1000 reversal is proxy

A literature-derived escape shortlist, cited and unscored. A labelled placeholder, not a model run.

lib/capability/registry.ts

DMG drug-response rows is proxy

NO AUC AND NO IC50 IS SERVED. Each row states what is known instead.

lib/capability/registry.ts

Measured CCMA screens (CC BY 4.0) is draft

CRISPR betas and drug z-scores for six DMG lines. THE TWO ARE NOT EQUALLY SOLID: r=0.95 against 0.58.

lib/capability/registry.ts

failures register · 19

No curated DMG dependency value could be traced to the paper its row cites.

0 of 17 rows traceable, across 3 cited papers read in full. 7 cited rows name a cell line their paper never uses. Clears when: A licensed dependency source that permits commercial use, read per row. Current DepMap releases prohibit it and Cell Model Passports is stricter, so this is a rights problem before it is a data problem.

dependency-source-verification.json

No curated DMG drug-response value could be traced to the paper its row cites.

0 of 15 rows traceable. 0 of the served AUC values have a source that reports an AUC at all, and 4 candidate rows were refuted by reading the paper. Clears when: A licensed per-compound response source. The CCMA screen is licensed and usable and is a z-score rather than an IC50, so it cannot refill these columns without changing what the column means.

drug-response-verification.json

The pre-registered external comparison of the in-silico knockout ordering failed.

Spearman -0.0260 with a 95% interval of [-0.3394, 0.2927] over 41 genes. Verdict fails, against a failure criterion stated before the data was seen. Floor: an interval containing zero, which is the pre-registered failure criterion. Clears when: A measured screen on the endpoint the model predicts. A pooled screen measures proliferation and the knockout predicts a state, so this may never clear and the comparison bounds how far the ordering may be carried rather than condemning it.

crispr-tf-screen-2026.json

Every OLIG1 result on the OPC-like programme was its own expression averaged into a programme it is a marker of.

Leave-target-out effect exactly 0.0000, 0.0000, 0.0000, 0.0000 across the 4 cell states. All zero: true. Floor: exactly zero is the whole result; there is nothing left after the target is removed. Clears when: Nothing. The result is retracted and the retraction is held by smoke-celloracle-claims.ts.

backend/agents/celloracle/artifacts/ko-batch.json

A large minority of in-silico targets return exactly zero on the readout the product serves.

12 of 45 targets return exactly zero on the OPC-like programme in every cell state, and 13 do on the leave-target-out readout. Separately, on the 22 of 41 head-to-head targets that return zero, 2 are measured dependencies, where silence is a worse answer than a lookup that returns a number with an interval.

Rest of the entry, 46 more words

Floor: zero is the modal outcome, so a zero cannot be read as a negative finding. Clears when: A network with more of these targets as regulators, or an honest statement that the target is out of scope. The second is what the product does now.

backend/agents/celloracle/artifacts/ko-batch.json

The product's own signed headline readout orders measured dependencies worse than a coin.

AUROC 0.2973 for the depleted flag against a 0.5 chance line, on 4 positives. Floor: 0.5. Clears when: Nothing about the sign convention, which is documented and is not a defence. The served number invites the reading, so the fix is in how it is presented.

retrieval-vs-inference.json

The copula joint is right more often than its marginal baseline and wrong more confidently.

Brier 0.0704 against 0.0784, which it wins, and log-loss 0.3954 against 0.2725, which it loses by 45%. Leave-one-out over 60 anchor cases. Floor: independent marginals, the simplest thing that could replace the joint. Clears when: Calibration, or fewer genes, or more cases. At n=60 and 13 genes this is the expected shape.

copula-vs-marginal-loo.json

The AC-like programme separates its own cell state worse than one of its two genes does alone.

AUROC 0.6686 for the programme against 0.6864 for APOE alone, a marker coverage of 2 of 5. The programme is a two-gene mean because only two of five markers survived the reference's gene selection. Floor: the programme's own best single marker, plus expression-matched random sets of the same size. Clears when: A reference retaining more of the AC-like markers, which is a gene-selection change and not a modelling one.

programme-reproduction.json

No two evidence layers in this product have ever been asked the same question about the same material.

86 of 133 claim pairs are never comparable, and all 86 cross-layer pairs are among them: 0 cross-layer pairs are comparable. Clears when: One target measured on one endpoint by two layers. Nothing on any surface may be described as one layer confirming another until then.

lib/workspace/perturb-conflicts.ts

A state-selective drug call does not reproduce between two replicate plates of the same experiment.

DiPG6: selectivity reproduces at 0.2488 against a matched-random floor of 0.3476 and a ceiling of 0.5284; SF8628: selectivity reproduces at -0.0174 against a matched-random floor of 0.1438 and a ceiling of 0.4643 Floor: expression-matched random gene sets of the same size, reproduced the same way. Clears when: A cell-state-resolved drug readout, at a timepoint the product's own rule accepts. The held assay reads out at 24 hours and MIN_STATE_CHANGE_HOURS is 48.

state-selective-drugs.json

No compound's viability across six DMG lines tracks any marker programme after multiple testing.

OPC-like: 0 entries survive BH at q<=0.1, and 71 reach nominal p<=0.05 where 102.65 are expected by chance; AC-like: 0 entries survive BH at q<=0.1, and 76 reach nominal p<=0.05 where 102.65 are expected by chance; OC-like: 0 entries survive BH at q<=0.1, and 82 reach nominal p<=0.05 where 102.65 are expected by chance Floor: the exact null over all 720 orderings of six lines, which puts |rho| at 0.829 at the 95th percentile.

Rest of the entry, 20 more words

Clears when: More lines. At six, the exact null is so wide that only a perfect ordering is nominally significant.

state-selective-drugs.json

In the xenograft, the programme score tracks sequencing depth and every drug effect sits inside the untreated spread.

OPC-like: rho 0.8009 against median genes per cell, largest arm effect 1.4295 against an untreated spread of 2.3956; AC-like: rho 0.4426 against median genes per cell, largest arm effect 1.0791 against an untreated spread of 2.6043; OC-like: rho 0.4887 against median genes per cell, largest arm effect 0.4733 against an untreated spread of 1.3274 Floor: the four untreated animals, which received nothing and still differ.

Rest of the entry, 18 more words

Clears when: Depth-matched libraries, more animals per arm, or a composition readout that is not a mixture mean.

state-selective-drugs.json

No dataset here has a combination arm, and neither combination the product names states a reference model.

0 combination arms across 6 perturbation datasets. The product names 2 combinations and 0 state a synergy method. The smallest HSA excess this screen could tell from noise is 2.199 z units, which only 200 of 2053 entries reach as a single agent.

Rest of the entry, 45 more words

Floor: the measured repeatability of the one readout that measures its own. Clears when: A combination arm measured alongside both single agents, with intervals, and a named reference model. Bliss needs a fraction affected and Loewe needs a concentration, and no held readout has either.

combination-evidence.json

The in-silico perturbation layer cannot be scored against the second CRISPR screen, because the libraries barely overlap.

2 genes shared between 45 in-silico targets and a 352-gene library: ERG and JUN. Clears when: A transcription-factor library in more than one DMG line. GSE267415 is one, and it is one line.

heldout-perturbation-ranking.json

Nothing can say whether a model ranking is right, and the two candidate adjudicators contradict each other.

Two independent external adjudicators of patient similarity agree at Spearman -0.0857, exact two-sided p 0.9194 against a null whose |rho| reaches 0.8286 at the 95th percentile.

Rest of the entry, 64 more words

Within the stronger one, the three patient tumours disagree with each other. 3 of 6 models carry identical curated mutation sets and cannot be separated by the ranking at all. Floor: exact, over all 720 orderings of six models. Clears when: An outcome attached to a model choice. Nothing in the prediction ledger has come back, so there is no such pair anywhere yet.

heldout-model-ranking.json

Model-to-model differences in compound response are undemonstrable: the screen disagrees with itself more than the models disagree with each other.

On the 225 compounds screened twice, two DMG lines agree at 0.4857 where one line agrees with ITSELF at 0.4174: models disagree LESS than repeat readings do.

Rest of the entry, 102 more words

The genetic-dependency readout separates cleanly by comparison, at 0.6662 against 0.8616 over 352 genes. In decision units the same inversion holds: the drug cross-model flip rate is 0.0768 against a same-model noise floor of 0.1051. Floor: 10,000 random orderings, cleared by both readouts and not the binding constraint here. Clears when: A drug screen with replicate plates rather than repeated library entries. The repeat figure is a LOWER bound on this screen's repeatability, because the two entries come through different compound libraries, so a fair repeatability could put the ceiling back above the cross-model figure and restore real heterogeneity. Undemonstrable, not absent.

model-heterogeneity.json

The matching engine returns nothing for a quarter of the anchor cases, most of them because the case carries no driver at all.

15 of 60 requests abstain, 9 of them because the case has no driver in the 13-gene panel. Clears when: A wider mutation panel, which is an anchor-curation change rather than an engine change.

method-baseline-inventory.json

More perturbable targets are skipped for no signal than are answered.

49 of 94 perturbable targets are skipped, leaving 45 with a signal. A skipped target and a zero-valued target are different states and both mean the product has nothing to say. Clears when: A denser network or a deeper reference. Both are data problems.

retrieval-vs-inference.json

The only other indication ever built is held rather than offered, and its anchor reproduces.

1 Ewing capability row at maturity planned. Not yet wired. The n=222 anchor is built and reproduces EWSR1::FLI1 prevalence at 86.5% against a literature consensus of about 85%. It is held until DMG is proven out. Clears when: DMG clearing its own evidence gates, which is the stated condition. This is a deliberate hold rather than a defect, and it is in the register because one selectable indication is the denominator behind every coverage claim the product makes.

lib/capability/registry.ts

method gate · 2

base_grn_sensitivity is not served, and would be blocked if it were, missing: abstention

whether the CellOracle knockout ordering depends on the generic human promoter half of the base GRN, or only on the four DMG scATAC samples

backend/data/registry/base-grn-sensitivity.json

cross_study_label_transfer is served under a dated exemption, missing: abstention

transferring the atlas's cell-state vocabulary onto an independent cohort, GSE184357, and scoring it against that cohort's own author annotations Remediation: The served surface is one string carrying four numbers, so the abstention that fits is a condition on the SURFACE rather than on the method: stateCompositionSeries should refuse to render a composition when the transfer agreement for the cohort in question is below its own majority-class floor, rather than rendering it beside a note.

Rest of the entry, 24 more words

Today the numbers are 0.635 against 0.4349, so the condition would not fire; the point is that nothing would happen if they were reversed.

backend/data/registry/atlas-label-transfer-gse184357.json

coupling register · 17

cell state vocabulary: this layer is specific to the indication it was built for

OPC-like is the oligodendrocyte-precursor programme, the malignant stem compartment of this disease, and the product quotes every knockout against it.

Rest of the entry, 79 more words

Which programme is the headline is a biological judgement about the indication, not a default. There is no rule that picks one from an arbitrary reference, and picking the largest cluster would pick `Other`, which is the residue. What would remove it: A second indication's single-cell reference with its own annotated malignant compartments, plus a stated reason for which of them is the readout. The name would then move onto the atlas release rather than into a module constant.

backend/agents/celloracle/axes.ts

cell state vocabulary: this layer is specific to the indication it was built for

The four strata are a 1:1 relabel of one DMG study's own consensus annotation, applied by a substring rule over one author-supplied column.

Rest of the entry, 82 more words

That is a stronger provenance than a clustering this product invented, and it is exactly why it does not transfer: the vocabulary belongs to that deposit's authors, and another indication's deposit will have named different things. What would remove it: An author-annotated reference for the second indication. Absent one, the alternative is clustering its cells here and naming the clusters, which is a scientific claim this product is not in a position to make and would have to be labelled as one.

lib/workspace/evidence-ontology.ts

cell state vocabulary: this layer is specific to the indication it was built for

`Other` is the largest cluster at 21,677 of 40,000 cells and is a residue rather than a state. Which cells fall into it is a property of the DMG reference's annotation, so the rule that excludes it from reporting is a rule about that reference. What would remove it: The second indication's reference and its own residue, with its own decision about what is reportable. The concept transfers; the membership does not.

backend/agents/celloracle/live-response.ts

marker programmes: this layer is specific to the indication it was built for

The three programme scores are means over named marker genes, and the genes are DMG lineage markers. THE LISTS EXIST IN SIX COPIES ACROSS THE PYTHON, AND THEY HAVE ALREADY DIVERGED: fit_program_baseline.py carries CLU and SPARCL1 on AC-like and CNP and SOX10 on OC-like, which the other five do not.

Rest of the entry, 98 more words

That divergence is a correctness risk today and is deliberately NOT fixed here, because every one of those files produced a checked-in artefact and editing a marker list without regenerating would change what a served number means while leaving the number in place. What would remove it: Nothing about a second indication removes this one. It needs the six copies collapsed into a single artefact-carried vocabulary and every dependent artefact regenerated on the same day, which is a scoped piece of work on the DMG side and a prerequisite for the second indication rather than part of it.

backend/services/celloracle/batch_knockout.py

marker programmes: this layer is specific to the indication it was built for

Exported under a generic name from a module called `capabilities`, and its contents are which DMG markers survived the reference's highly-variable-gene filter: OPC-like 3 of 7, AC-like 2 of 5, OC-like 4 of 5.

Rest of the entry, 58 more words

It has to be exported, because `provenance.markers` reported the INTENDED lists until 2026-08-06 and every surface consequently overstated what the numbers were made of. The AC-like programme is two genes. What would remove it: The generated table gaining an indication key, which is downstream of the atlas release gaining one, which is downstream of a second atlas existing.

mcp-server/src/capabilities.generated.ts

mechanistic bridge: this layer is specific to the indication it was built for

One of seven bridge entries is enabled, ACVR1 via SMAD1/SMAD4 to ID1/ID3, and it is enabled because that specific edge has in vivo DMG pharmacodynamic confirmation: an orthotopic DIPG xenograft in which two chemically distinct ALK2 inhibitors ablate ID1.

Rest of the entry, 67 more words

A bridge edge is an explicit modelling assumption and the evidence licensing it is disease-specific by construction. Five further entries abstain on DMG-specific literature, so the abstentions are as coupled as the enablement. What would remove it: Equivalent in vivo pharmacodynamic evidence in the second indication, per edge. Not transferable: an edge is enabled by a study, and there is no study about two diseases at once.

backend/agents/celloracle/bridge.ts

matching engine: this layer is specific to the indication it was built for

`domainShift()` presents itself as the general list of reasons a result may not transfer, and one of its axes is the H3.1 versus H3.3 split with the ACVR1 co-segregation behind it, carried on a `ModelIdentity` field that exists for this disease alone.

Rest of the entry, 66 more words

The axis is real and it is worth having: ACVR1 reads 0.805 in H3.1 against 0.035 in H3.3, so ignoring it would pool two different diseases. What would remove it: The second indication's own molecular axis, whatever it is, plus a transfer-relevant measurement on it. The shape of the mechanism generalises; the axis does not, and inventing one would be a claim rather than a parameterisation.

backend/agents/matching/engine.ts

benchmark scope: this layer is specific to the indication it was built for

The frozen pre-registration excludes OLIG1, OLIG2 and PDGFRA from the head-to-head comparison because they are markers of the programme being scored, so a knockout of one moves the readout by definition.

Rest of the entry, 51 more words

The exclusion is correct and it is specific: the circular genes are whichever genes define that indication's programmes. What would remove it: The second indication's marker sets, from which its own exclusions are derived. The exclusion RULE is general and is already written as one; only the gene list is coupled.

backend/agents/benchmark/scope.ts

benchmark scope: this layer is specific to the indication it was built for

The leave-target-out table is computed for OLIG2, OLIG1 and FOS: two OL-lineage regulators and the immediate-early control that beat them on gradient.

Rest of the entry, 96 more words

The three are chosen because of what they showed on this artefact, and OLIG1's row is the one that retracted a claim (leave-target-out exactly 0.0000 in all four cell states, so its whole effect was its own expression averaged into a programme it is a marker of). What would remove it: The second indication's own lineage regulators and its own handling control. Note the control half is harder than it looks: FOS and JUNB are unusable as comparators here because they are handling-associated on this reference, measured at Cliff's delta 0.63 between single-cell and single-nuclei libraries.

lib/workspace/perturbation-data.ts

benchmark scope: this layer is specific to the indication it was built for

The generic experiment-specification exchange refuses any specification whose state-marker panel this atlas cannot score, and the rule names release dmg-2026-07-29 and its three realised panels. Refusing is correct: a specification asking for a readout the atlas cannot produce is not executable. What is coupled is which panels count as scorable. What would remove it: A second atlas release, at which point the rule reads the release rather than naming one.

backend/tools/fixtures/laboratory-exchange.json

artefact paths: this layer is specific to the indication it was built for

The generic job graph names `dmg-reference.h5ad`, `dmg-base-grn.parquet`, `dmg-oracle.celloracle.oracle` and `dmg-atlas-facts.json` across twelve sites.

Rest of the entry, 77 more words

These are the filenames of files that exist, pinned by sha256 in the checksum registry, and renaming them to a pattern would break every pin for no benefit while there is one indication. What would remove it: A second indication's equivalent artefacts existing on disk, at which point the graph takes an indication and composes the name. Doing it before they exist would produce a lookup that resolves to nothing and a code path nobody can run.

lib/jobs/scientific-jobs.ts

artefact paths: this layer is specific to the indication it was built for

The provenance layer resolves two Cloud Storage objects by literal URI. They are the reference and the base GRN, and their bytes are what the checksums pin, so the URI is part of the identity of a pinned artefact rather than a configuration value. What would remove it: The same thing as the job graph: second-indication objects in the bucket, and an indication key on the source registry that resolves to them.

lib/provenance/object-store.ts

honesty layer: this layer is specific to the indication it was built for

The general method-release gate holds four entries and all four are DMG methods, with their floors and comparators stated in DMG terms: four scATAC samples, 19,634 DMG rows, six DMG lines with no non-DMG comparator.

Rest of the entry, 63 more words

The gate mechanism is general; its contents are a record of what has been served, and a record cannot be about a thing that has not happened. What would remove it: Nothing. This entry exists so the count is honest, not because it is fixable: the register fills as methods land, and a second indication's methods will add rows rather than change these.

lib/methods/gate.ts

honesty layer: this layer is specific to the indication it was built for

The critic is the general honesty layer and its instructions name DMG facts: that the dependency rows come from paediatric HGG papers, and that nothing may be called DMG-selective because there are six DMG lines and no comparator.

Rest of the entry, 56 more words

These are the specific things a writer would otherwise overclaim about this evidence, so a generic version would be a weaker critic. What would remove it: Per-indication critic instructions, which need a second indication's evidence to be written against. Generalising the wording without the evidence would delete a working check and replace it with a placeholder.

backend/agents/honesty-critic/critic.ts

honesty layer: this layer is specific to the indication it was built for

Correctly guarded by an indication check already, and recorded anyway because it is the pattern the other couplings should move toward: a DMG-specific caveat about the anchor being paediatric-weighted while adult DMG is thalamus-dominant, appended only when the library is DMG.

Rest of the entry, 42 more words

The caveat corpus itself is DMG-only, so a second indication gets no caveats rather than the wrong ones. What would remove it: Its own caveat corpus per indication. The mechanism is already right and this is the shape the rest should copy.

backend/agents/cohort-architect/sampler.ts

honesty layer: this layer is specific to the indication it was built for

A generically named validation module groups reference samples by anatomical site and treats midline as the in-scope group, because the composition question it answers is whether a reference named for a midline disease is made of midline cases.

Rest of the entry, 59 more words

The answer, 30 midline against 11 hemispheric with 6 unstated and 5 ependymoma comparators, is why the module exists. What would remove it: The second indication's own site vocabulary and its own composition question. Anatomical site is not even the relevant axis for most cancers, so this is a coupling to the disease's anatomy rather than to its data.

lib/workspace/state-validation.ts

persistence: this layer is specific to the indication it was built for

Eight tables carry `indication` as free text with no enum, no foreign key and no check constraint, so nothing in SQL prevents a row claiming an indication that does not exist.

Rest of the entry, 146 more words

The only defence is application-level: the coordinator overwrites whatever the planner produced with the library's own slug. Recorded rather than fixed, because a check constraint enumerating today's keys is a constraint that has to be migrated every time an indication is added, and it would encode the registry in a place the registry cannot read. What would remove it: An indications table the eight columns reference. Deliberately NOT written as a migration today, and the reason is the same reason it is recorded here rather than fixed: a reference table with one row constrains nothing, and a check constraint enumerating today's keys would have to be migrated every time an indication is added while encoding the registry somewhere the registry cannot read. It becomes worth doing when there is a second row to put in it, which is the same moment the constraint starts catching something.

backend/db/schema.ts

expansion gate · 9

Does the anchor's structure transport to cases it was not fitted on, by more than a correlation between two frequency vectors would give anyway?

The two numbers exist and are leakage-free: Pearson r = 0.922 on 72 held-out same-study cases and r = 0.587 across institutions on n=18, with the anchor sample ids excluded by construction and cross-checked against dmg-anchor-identities.json.

Rest of the entry, 130 more words

Neither carries a floor. The script contains no permutation, no shuffle, no random draw and no baseline of any kind, so r = 0.587 has no zero: both vectors are dominated by the same few high-frequency drivers, and a correlation between two such vectors is high before any biology enters. The earlier r = 0.97 was retracted for circularity, which is the same class of error one level up. What would clear it: Add a permuted-gene-assignment arm to validate_external_cohort.py and report its 95th percentile beside each r, plus a second comparator that is not the anchor at all, such as a generic paediatric high-grade glioma frequency vector. If r = 0.587 does not clear the permuted null on n=18, that is the answer and it goes in the failures register.

backend/scripts/validate_external_cohort.py

Does every method the indication serves state a floor, a comparator and an abstention condition, without an exemption?

cross_study_label_transfer is served at lib/workspace/plot-data.ts and is grandfathered, dated 2026-08-13, because its abstention is `claim_level_only`: four bans on classes of statement, which govern what a person may write, and no condition under which the served string is withheld. base_grn_sensitivity has no abstention at all and passes only because it is not served, so it would block the day anyone renders it.

Rest of the entry, 50 more words

What would clear it: Implement the remediation already written into the grandfathered entry: stateCompositionSeries refuses to render a composition when that cohort's transfer agreement is below its own majority-class floor, rather than rendering it beside a note. Today the numbers are 0.635 against 0.4349, so the condition would not fire.

lib/methods/gate.ts

Does the held-out benchmark beat the cheapest thing that uses no data at all, on the same scale?

The floor half holds comfortably: leave-one-model-out reaches a median Spearman of 0.7806 across six held-out lines, against a permutation floor whose typical 95th percentile is 0.104 over 10,000 draws, and the whole line is held out rather than a random split.

Rest of the entry, 144 more words

The comparator half does not resolve. The cheapest alternative is one bit per gene, is-it-a-core-housekeeping-gene, with no data, no model and no screen, and it reaches a median AUROC of 0.7431 for the held-out line's strongest dependency decile. Spearman over 352 genes and AUROC over a top decile are different quantities, so 0.7806 against 0.7431 is not a margin. The replicate ceiling is 0.8616, which leaves little room for one to exist. What would clear it: Score the housekeeping flag and the five-line mean on the same metric, both ways: Spearman for the flag over all 352 genes, and AUROC for the mean over the same top decile. Two numbers on one scale, reported whichever way they fall. If the one-bit flag matches the five-line mean, the honest conclusion is that the leave-one-out arm is measuring pan-essentiality, which its own leakage note already raises.

backend/data/registry/heldout-perturbation-ranking.json

Has the mechanistic layer been checked against something outside the data it was built from, and did the check clear its stated bar?

Three of the four hold and the fourth does not. The check was pre-registered with its circular exclusions frozen in advance, it ran, and it is published on the Perturbation lab whichever way it fell: Spearman -0.026, 95% CI [-0.339, 0.293], n=41.

Rest of the entry, 125 more words

It did not clear its bar. The result is the absence of external support rather than evidence against, because it compares a predicted cell-state shift against a measured proliferation dependency and 22 of its 41 points are structural zeros. It is not claimed as validation anywhere and must not be. Recording it as met on process while it failed on result is exactly the upgrade the honesty invariants forbid. What would clear it: An external check whose readout is the same quantity the model predicts. A CRISPR screen measures what a cell needs to proliferate; the knockout predicts a cell-state programme shift. Until a perturbation dataset with a state readout is ingested, this gate cannot be cleared by trying harder against the screen already held.

backend/agents/benchmark/scope.ts

Is every source the indication serves licensed for commercial use, with no unresolved obligation?

The anchor's own source is the unresolved one. cBioPortal is ODbL, not ODC-BY, corrected 2026-08-07 against cBioPortal's own documentation after three surfaces had said otherwise.

Rest of the entry, 145 more words

ODbL adds share-alike and a positive §4.6 obligation to offer recipients either the derivative database or the method that made it. Whether backend/cohorts/dmg-matrix.ts is a Derivative Database under §4.4b or only a Produced Work under §4.5b is open, and §4.6 catches a Produced Work derived from a Derivative Database too, so the second route only helps if the first does not apply. The licence table records the obligation as attribution yes, share-alike unresolved. Exposure starts at the first paying customer, so this is currently latent rather than absent. What would clear it: A determination on whether the n=60 anchor matrix is a Derivative Database, from someone qualified to make it, recorded in DATASET-LICENCES.md with its reasoning. If it is, publishing the derivation method alongside the product satisfies §4.6 and the gate clears; the method is already public, so the cost is procedural rather than strategic.

DATASET-LICENCES.md

Does the indication carry any observed treatment response or survival, or only molecular features?

The n=60 anchor is 13 binary driver features and no outcome of any kind: no treatment response, no survival time, no progression record. Everything the product emits is therefore an ordering signal by construction rather than by caution, and no calibration is possible against this anchor at any sample size.

Rest of the entry, 83 more words

This is the deepest of the fourteen and it would hold expansion on its own, because a second indication built the same way inherits the same ceiling. What would clear it: An outcome layer for the indication, from a source that is rights-cleared for commercial use. This is the item on the list most likely to be impossible from open data alone for a rare paediatric cancer, which is a reason to establish it before a second indication is chosen rather than after.

backend/cohorts/dmg-matrix.ts

Can anyone actually run the experiment the product recommends, in a model system for this indication?

In vitro is workable and in vivo is not. Six DMG cultures are characterised and screened, and none is held here. For xenografts the count of H3 K27-altered DMG models in any public model registry read is ZERO: the two records CancerModels.org returns for a DIPG histology search are both reported histone H3 wild-type by the consortium paper that characterises them, read from the CC BY full text rather than an abstract, so a search stopping at the histology label would have reported two.

Rest of the entry, 119 more words

NCI PDMR returns no data for DIPG, midline, pontine, brain or K27. ITCC-P4's 19 are the largest set and are behind an access gate. Direct engraftment of DMG autopsy tissue has also produced histologically convincing pontine tumours composed of MURINE cells, so a record without a species-of-origin check cannot support a drug claim even where one exists. What would clear it: Either an access agreement for a characterised in vivo panel, or a documented decision that the product's recommendations stop at in vitro for this indication and say so on the surface that makes them. The second is cheaper and is a real option, but it has to be stated rather than left as the shape of what exists.

backend/data/registry/model-crosswalk-2026.json

Does a buyer outside the founder's own accounts want the new indication, before it is built?

There are no customers. Measured against the live database on 2026-08-13: two accounts exist, one is the operator's own and one is residue from a verification script, so the external count is zero.

Rest of the entry, 92 more words

Three surfaces asserted a customer base in the present tense anyway and were retracted on that date. With a zero denominator no per-customer quantity is computed anywhere, which is why the productisation economics artefact reports null rather than zero for every rate. What would clear it: One external account that asks for a named second indication. Note the ordering this gate encodes: the buyer comes before the build, not after it. Building an indication and then looking for someone who wants it is the expensive order and it is the default one.

lib/retention/definitions.ts

Would they pay for the workflow, or for the founder writing the scientific answer?

No payment of any kind has been taken, so the distinction has never been tested. Recorded as `not_met` rather than `not_evaluable_yet` because the state of affairs is definite: zero is a measured answer here, not a missing measurement. What would clear it: One paid run whose scientific content came out of the product. Strictly downstream of buyer_named_and_external, and it is the gate that decides whether a second indication multiplies a product or multiplies the founder.

ONKYDRA_CONSENSUS_HANDOFF.md

Ewing sarcoma

Ewing sarcoma is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

MPNST

MPNST is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

ATRT

ATRT is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

High-risk neuroblastoma (MYCN)

High-risk neuroblastoma (MYCN) is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

Anaplastic thyroid

Anaplastic thyroid is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

Rhabdomyosarcoma

Rhabdomyosarcoma is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

Paediatric AML (KMT2A)

Paediatric AML (KMT2A) is planned and nothing has been measured about it, so it has no recorded limitations. That is not a clean bill of health: it is an empty list that has to be filled before the indication could be served, and every entry on the built indication's list is an entry this one will need its own answer to.

What would change these

Paired data. The same perturbation run in a model and in the patient material that model was matched to, enough times to measure whether match quality predicts transfer error. That is a wet-lab programme, not a software release, and until it exists the honest position is the one on this page.

See also current capability and data status.

Research use only. Not a diagnostic, not a medical device, not a treatment recommendation, and not a substitute for wet-lab validation.