The perturbation lab for rare cancer

Ask the workspace:which cells is any of this actually about?what can a cell line actually be matched on?how much of this ranking is artefact?what is connected to what, and what measured it?what would prove this wrong?

Pick a real tumour cell state. Compare the interventions that could move it, find the least-wrong model to test it in, and register the prediction before the experiment runs. Every number is labelled as a model run, a third-party measurement, an assumption or a simulated draw, and the workspace abstains where it has nothing.

State atlas

40,000 measured cells, and nothing simulated

Four cell states over one 40,000-cell malignant DMG reference: about half midline high-grade glioma, the rest hemispheric, site-unstated or ependymoma comparators. Nothing here is simulated, and 46% sits in a scored programme.

  • Five counting levels kept apart and never summed. The 40,000 cells come from at most 42 donors.
  • The AC-like programme is two genes, AQP4 and APOE, after three of five intended markers failed.
  • A second DMG cohort puts OPC-like at 52.0% against 24.9% here. A distance, not a validation.

Private beta · Built on Gemini 3 and a Gaussian-copula cohort engine anchored to 60 real K27M cases

Why now

Rare-cancer drug development just changed shape

For forty years, diffuse midline glioma had no targeted therapy. In August 2025 it did. Three structural shifts converged within twelve months, and together they make a workspace like Onkydra possible for the first time.

01August 2025

DMG got its first targeted therapy

Dordaviprone (Modeyso) won FDA accelerated approval in August 2025 on a pooled n=50 across two trials. The first hit in a cancer that had nothing for forty years. A new generation of rare-cancer biotech founders is being funded around indications like this, and they need in-silico target validation that matches the size of the evidence.

022024 – 2026

Foundation models for biology shipped

CellOracle put in-silico knockouts over an ATAC-derived regulatory network within reach of a single team. MedGemma landed on Vertex AI. The compute is genuinely available now, and the bottleneck moved from infrastructure to methodology. Licensing is the part nobody advertises: the strongest non-coding variant model on the market is non-commercial only, and several dependency screens closed to commercial use between releases.

03June 2026

AI workbenches are the new baseline

Anthropic launched Claude Science on 2026-06-30. Multi-agent deliberation loops with adversarial critics, audit-logged tool calls, reproducible artefacts, structured citations at source. The interaction pattern is now the industry floor. Onkydra takes that quality bar and specialises it to the vertical where curated cohorts and honest CIs decide whether the report gets peer-reviewed.

PUBLIC DATA + AI STACK
cBioPortal

Rare-cancer cohorts (DKFZ + CPTAC)

Children's Brain Tumor Network
OpenPedCan (D3b Center, CHOP)

Paediatric cohort expansion · planned

CCMA
Cellosaurus (SIB cell-line reference)

Cell-line dependency (CCMA CRISPR) · limited

LINCS L1000

Escape candidates (literature) · placeholder

ClinVar

Variant significance · planned

Open Targets

Target-disease atlas · planned

PubMed

Literature retrieval

TCGA / GDC

MedGemma demo tissue (TCGA-LGG)

NCI Thesaurus (National Cancer Institute)
Human Phenotype Ontology

Cancer ontology + phenotype · planned

OpenFDA Drug Labels

Regulatory anchoring · planned

cBioPortal

Rare-cancer cohorts (DKFZ + CPTAC)

Children's Brain Tumor Network
OpenPedCan (D3b Center, CHOP)

Paediatric cohort expansion · planned

CCMA
Cellosaurus (SIB cell-line reference)

Cell-line dependency (CCMA CRISPR) · limited

LINCS L1000

Escape candidates (literature) · placeholder

ClinVar

Variant significance · planned

Open Targets

Target-disease atlas · planned

PubMed

Literature retrieval

TCGA / GDC

MedGemma demo tissue (TCGA-LGG)

NCI Thesaurus (National Cancer Institute)
Human Phenotype Ontology

Cancer ontology + phenotype · planned

OpenFDA Drug Labels

Regulatory anchoring · planned

cBioPortal

Rare-cancer cohorts (DKFZ + CPTAC)

Children's Brain Tumor Network
OpenPedCan (D3b Center, CHOP)

Paediatric cohort expansion · planned

CCMA
Cellosaurus (SIB cell-line reference)

Cell-line dependency (CCMA CRISPR) · limited

LINCS L1000

Escape candidates (literature) · placeholder

ClinVar

Variant significance · planned

Open Targets

Target-disease atlas · planned

PubMed

Literature retrieval

TCGA / GDC

MedGemma demo tissue (TCGA-LGG)

NCI Thesaurus (National Cancer Institute)
Human Phenotype Ontology

Cancer ontology + phenotype · planned

OpenFDA Drug Labels

Regulatory anchoring · planned

Gemini 3

Planner + Honesty Critic

Gemini 3

Workspace Writer

Gemini 3
Google DeepMind

Imaging (MedGemma)

CO

Perturbation model (CellOracle)

Broad Institute

Ensemble second model (Geneformer) · planned

Google Cloud

Vertex AI runtime

Gemini 3

Planner + Honesty Critic

Gemini 3

Workspace Writer

Gemini 3
Google DeepMind

Imaging (MedGemma)

CO

Perturbation model (CellOracle)

Broad Institute

Ensemble second model (Geneformer) · planned

Google Cloud

Vertex AI runtime

Gemini 3

Planner + Honesty Critic

Gemini 3

Workspace Writer

Gemini 3
Google DeepMind

Imaging (MedGemma)

CO

Perturbation model (CellOracle)

Broad Institute

Ensemble second model (Geneformer) · planned

Google Cloud

Vertex AI runtime

THE PROBLEM

The teams curing rare cancers can't afford to model them.

Tumour profiling at scale exists. It is just priced, built and validated for billion-dollar pharma, not for the small teams taking on the cancers pharma won't.

BENCHMARK · ILLUSTRATIVE

Anchor n vs foundation-model requirement

DMG anchor · DKFZ + CPTAC
Ewing anchor · DFCI + Curie
Onkydra draws per query
Foundation-model requirement
DMG anchor · DKFZ + CPTAC
60
Ewing anchor · DFCI + Curie
222
Onkydra draws per query
1000
Foundation-model requirement
1000

The gap is the problem. Onkydra draws simulated profiles over the copula joint fitted to the real anchor. Bootstrap CIs stay on the real n.

How it works

Six surfaces. One world model. A lab that will watch your targets.

You give the workspace a target. The Planner picks the right simulation, the Biology Engine ranks the strata, the Honesty Critic checks the result for itself, and the Workspace Writer hands you the report with citations. The Delta Watcher is the sixth surface: it will watch ClinVar, DepMap and the literature and re-run when something moves the needle. That live monitoring is being wired, not running yet.

Picks the right tools01

Planner

Reads your question and picks the right cohort library, plus the dependency and escape layers, each of which declares what kind of thing it is. Deterministic: no specialised syntax, no nine separate database interfaces.

Builds the world model02

Cohort Architect

Monte Carlo draws from a Gaussian copula joint fitted to your indication's real anchor cohort, with Ledoit-Wolf shrinkage on Σ. n=60 in DMG, pooled from DKFZ (n=53) and CPTAC (n=7). Tested against held-out cases never seen during fitting: single-gene frequencies reproduce at r=0.92 on 72 held-out DKFZ cases and r=0.59 against an independent institution. Every downstream statistic propagates that real n through bootstrap CIs you can see.

Runs the simulation03

Biology Engine

A stated rule set, evaluated rather than fitted, plus cohort statistics plus PubMed retrieval, returning a structured evidence bundle, never prose. CellOracle is real here: in-silico knockouts precomputed over a DMG regulatory network built from four real scATAC-seq samples unioned with CellOracle's published human promoter network and a 40,000-cell reference, served for the 45 transcription-factor regulators that produced signal, plus ACVR1 through an explicit ID1/ID3 bridge. That reference is a mixed cohort, not a pure one: about half its cells come from a record that is both high-grade glioma and a stated midline site, the rest are hemispheric, site-unstated or ependymoma comparators, and no record in the deposit states H3 K27M status. The composition is on the state atlas and in the capability registry. For every other target it abstains rather than substituting a heuristic, the readout is an uncalibrated ordering signal rather than a probability. The dependency rows serve measured CRISPR beta scores where the screen covers the gene and an explicit null where it does not, and the escape layer is a literature-derived shortlist with citations and no score. Nothing is ingested from DepMap: its releases from 25Q1 are portal-only under terms that forbid use in a product. Every report declares which is which.

Checks itself04

Honesty Critic

Flags where the two stated rule sets disagree, and reports the disagreement rather than picking a winner: both are sets of assumptions we state, one drawn from the response model and one from published DMG target-class pathway mechanism, and neither settles the other. Hard-fails on claims without retrieved source. Cross-checks the cohort's co-occurrence structure with a DISCOVER-inspired audit. The disagreement is the product.

Speaks in citations05

Workspace Writer

Consumes the evidence bundle. Cannot emit a sentence without a citation id. Renders the memo into your live workspace, and hands you the whole run (memo, agent trace, cohort manifest, citation ids) as a JSON download. A typeset PDF of the memo is in build, not yet live.

Re-runs on new evidence (in build)06

Delta Watcher

Live monitoring is being wired and is not live yet. It will watch ClinVar, DepMap and PubMed for changes on your targets, re-run the pipeline when something moves the needle, and flag shifts in the ranking. The drift calculation is written and tested; the scheduler that triggers it is not running.

INDICATIONS

Built for the cancers nobody else will model.

One indication runs today: H3 K27M diffuse midline glioma, on a 60-case anchor. Seven more are on the roadmap.

Beachhead · scoped library

Diffuse midline glioma

The hardest cancer in paediatric neuro-oncology. The first targeted approval landed in 2025. New programmes need a defensible way to choose the next experiment today.

Driver

H3 K27M

Incidence

~300 / yr

Method-portable · roadmap

Ewing sarcoma

A fusion-driven bone and soft-tissue cancer of adolescents, with cohorts far too small for conventional modelling.

Driver

EWSR1::FLI1

Incidence

~200 / yr

Method-portable · roadmap

Malignant peripheral nerve sheath tumour

An aggressive nerve-sheath sarcoma, often arising in NF1, with few effective options across small, scattered cohorts.

Driver

NF1 loss

Incidence

~1,000 / yr

Method-portable · roadmap

Atypical teratoid / rhabdoid tumour

An aggressive infant brain tumour with one of the smallest cohorts in oncology: exactly where simulating the co-mutation structure matters most.

Driver

SMARCB1 loss

Incidence

~70 / yr

Method-portable · roadmap

High-risk neuroblastoma

The high-risk, MYCN-amplified subset where survival has barely moved in two decades.

Driver

MYCN amplified

Incidence

~600 / yr

Method-portable · roadmap

Anaplastic thyroid cancer

One of the most aggressive solid tumours known. Rare and fast-moving, where the speed of good trial design genuinely matters.

Driver

BRAF / TP53

Incidence

~700 / yr

Method-portable · roadmap

Rhabdomyosarcoma

The most common paediatric soft-tissue sarcoma. Still rare, and sharply split by fusion status.

Driver

PAX3::FOXO1

Incidence

~350 / yr

Method-portable · roadmap

Paediatric AML, KMT2A-rearranged

The KMT2A-rearranged subtype of paediatric AML, a high-risk group defined by its fusion and in need of better targets.

Driver

KMT2A fusions

Incidence

~120 / yr

PRICING

One ladder, from a free benchmark to a named programme.

An academic lab pays per lab, with no seat count, because a research group loses and gains people by design. A commercial buyer pays per seat, or per team with seats included. A programme is one named indication or target, agreed in a project plan, time-boxed, and expandable only by amendment. Research use only: a hypothesis generator and prioritisation engine, not a diagnostic and not a medical device.

Provisional pricing. These numbers are set from published comparables and are not settled. Expect them to move.

Checkout is not open, so nothing here can be bought today and no card is taken anywhere on this site.

Free

Anyone deciding whether Onkydra can answer their question at all.

€0free forever

One seat.

Academic lab

One named academic lab, where no for-profit owns the intellectual property in the work.

€300per lab

€249 a month billed yearly.

No seat count.

Commercial seat

One named person at a company, or an academic on a programme whose IP a for-profit owns.

€800per seat

€664 a month billed yearly.

One seat.

Commercial team

A rare-cancer R&D team running more than one candidate at a time.

€1,600per team

€1,328 a month billed yearly.

3 seats included. €300 a month per extra seat.

Programme

A team betting the next twelve to eighteen months on one target or one indication.

Quotedfrom €40,000 per year
THE TWO ENGINES

A five-stage pipeline that checks itself, fed by a composite of the real evidence we do have.

Onkydra is two engines working in concert. Hydra runs the five stages that assemble the report. Chimera samples the co-mutation profiles those stages run on. Both are built for cancers where the incumbent stack has too few patients to work with.

Hydra

Five-stage pipeline with a self-check

Five stages run in sequence: Planner, Cohort Architect, Biology Engine, Honesty Critic, Workspace Writer. The Honesty Critic labels every number as a model run, a labelled proxy or an assumption we stated, and where two layers disagree it reports the disagreement rather than picking a winner. It refuses to let a claim ship without a real source. Cut one head off, two grow back.

See the pipeline

Chimera

Simulated co-mutation profiles

Every draw is a simulated co-mutation profile, sampled from a Gaussian copula fitted to n=60 real H3 K27M DMG cases from DKFZ and CPTAC, Ledoit-Wolf shrinkage on Σ. One thousand Monte Carlo draws. Every CI resampled over the real 60, not the simulated 1,000. Same posture as the FDA's 2025 dordaviprone approval on n=50 with 22% ORR: the size of the underlying evidence is the size of the field.

Read the methodology

Founder-product fit

Onkydra is built by Faith Ogundimu, a cancer-genomics researcher, for the teams developing the drugs pharma won't. The workspace answers the questions Faith asks in her own PhD work: which subgroup responds, what breaks first under selection, what does the literature already know that we missed.

Research use onlySee scope

PRIVATE BETA

Get early access to Onkydra.

Onkydra is in private beta and the workspace is open to one address. Join the waitlist and we will reach out with a first run for your programme when it opens to biotech and research teams working on rare cancers.

No spam. We email you once, when your run is ready.

Onkydra is in private betaJoin