Thousands of diseases have no approved therapy, and for most of them nobody has written down a serious mechanistic hypothesis. So GROKTOR runs a two-robot lab. Grok reads the genetics and proposes treatments. Optimus takes each one and asks the bench questions: can this actually be made, can it reach the tissue, has it already failed, and what experiment would settle it.
There are more than 7,000 recognised rare diseases and roughly 300 million people living with one. The overwhelming majority have no approved therapy. Not because the biology is uniformly intractable — because a few thousand patients cannot repay a decade of development. For most of these conditions, nobody has ever sat down and written a serious mechanistic hypothesis.
Language models can write that hypothesis in minutes. The problem is that they can also write a beautiful, confident, completely wrong one, and a plausible falsehood in medicine is worse than silence. Fluency is not correctness, and a single model reviewing its own work grades itself generously.
What is actually running: Grok 4.6 on both sides, in two separate contexts with opposite instructions — one rewarded for proposing, one for finding the flaw. "Optimus" is the name of that second role, not a Tesla robot; no hardware is involved and nothing here is tested in a lab. One model checking itself in a fresh context is a weaker check than an outside expert. It is still far harder to pass than a model marking its own work, and every ruling is published so you can judge the checker yourself.
Is the agent real chemical matter, and is it even aimed at this disease?
The full causal chain, variant to symptom. A missing step is a failed step.
Can it be made, can it reach the tissue, and has it already failed in the clinic?
Name the cheapest experiment that would prove it wrong. No experiment, no publication.
A kill is permanent — there is no appeal and no revision. What reaches the archive has been attacked four separate ways by a model that was rewarded for destroying it. When nothing survives, that is published as a wipeout, because a disease where every plausible idea dies is a real result and worth knowing.
Every claim is tagged at the source: [KG] from the knowledge graph,
[KNOWN] from literature, [INFERRED] reasoned, [SPECULATIVE]
flagged as unsupported. The gauntlet checks the tags too — inference dressed up as established fact
is grounds for a kill.
A model that has read the literature can produce a confident, correct-sounding paragraph about a disease without deriving anything. That failure mode is invisible when every disease in the set is unsolved, because there is no answer to check against.
Spinal muscular atrophy sits in the set as a control. SMA is solved — nusinersen, risdiplam and onasemnogene abeparvovec are approved, and the SMN1/SMN2 mechanism is textbook. When the gauntlet reaches it, the trace can be read against a known answer. Did it reconstruct the SMN2 copy-number logic from the graph, or restate what it already knew?
The control validates nothing on its own. It calibrates how much weight to put on the diseases where there is nothing to check against — which is the entire rest of the set.
Every disease resolves to a MONDO identifier against Monarch before entering the queue. The set grows on its own — the researcher proposes related disorders, and they join the frontier.
Newest first. Click any entry to expand its surviving hypotheses and the ones the critic killed.
Monarch Initiative v3 knowledge graph — a public, keyless API unifying curated rare-disease resources including OMIM, Orphanet and HPO annotations. Per disease we pull the entity record, causal and correlated gene associations, the phenotype spectrum with frequency qualifiers, gene pleiotropy, and phenotype-nearest diseases and mouse model genes by semantic similarity.
Two models in opposition, both reached through OpenRouter. Both see identical evidence — the critic must be able to check every claim against the same source the researcher used.
Nothing here is reviewed by a human. Findings are marked UNREVIEWED because that is
what they are — there is no expert review queue behind this site. Contested stages are flagged.
Failures, refusals and empty results are printed rather than hidden, and the archive publishes
killed hypotheses alongside surviving ones.
/events — SSE stream of the live gauntlet