What we built
Disease Twin is a living computational entity for every rare disease in the world. It's not a page. It's not a document. It's a node in a clinical knowledge graph (genes, phenotypes, drugs, metabolic pathways, clinical trials, SUS coverage), wrapped by an agentic AI layer that does three things continuously:
- Reads the literature — mines new articles, extracts typed relations, decides what deserves to enter the graph.
- Generates hypotheses — proposes mechanistic drug-repurposing candidates from causal genes and known pharmacological targets.
- Self-verifies — every claim is checked against official sources before being published. Nothing is invented. Nothing is kept without a URL.
The result is visible on two open surfaces:
- Each disease has a public research journal that shows exactly what the agents did on that disease over the week, the month, the year. It's the anti-black-box.
- Any AI assistant on Earth (Claude, ChatGPT, Gemini, Copilot) can query the disease directly via the Model Context Protocol [5]. No scraping needed. No private API needed. It's a standard call.
A three-layer architecture
Disease Twin is deliberately simple to describe and deliberately hard to fake. The architecture has three layers, each with one non-negotiable principle.
Layer 1 — The clinical knowledge graph
The foundation is a property graph in Neo4j, populated from canonical open sources: Orphanet (the rare-disease catalog), Human Phenotype Ontology [6], MONDO, OMIM, ClinVar, Open Targets [7], PubTator3 [8], WikiPathways, ClinicalTrials.gov, and Brazil's public health system (DATASUS, CEAF, SUS PCDT). The schema follows the Biolink Model [9], exportable as a KGX dump, aligned with the Monarch Initiative [10]. Individual clinical cases are represented in Phenopackets v2 [11]; federated discovery uses Beacon v2 [12].
Principle: no node is invented. Every entity has a stable identifier (CURIE) and provenance traceable back to the original source. The public disease layer is licensed as CC0; patient data never leaves the patient's wallet. Between the two, a k-anonymous boundary.
Layer 2 — The 24-hour co-scientist
Five specialized agents run as scheduled jobs, each with its own "freshness" window (each disease is re-examined when its specific data goes stale). Collectively, they form what we internally call the co-scientist:
Pulls pre-extracted typed relations from PubTator3 (associate, treat, cause, inhibit, stimulate, prevent, positive and negative correlations) and materializes each as a typed edge in the graph, with the list of PMIDs as evidence.
Inspired by MedKGent [13]: uses an LLM to read abstracts and extract typed triples with self-consistency (a triple survives only if N independent runs agree). Every extraction must anchor to entities that already exist — genes by HGNC symbol, drugs by ChEMBL name, diseases by Orphanet code. Whatever doesn't anchor doesn't enter. This is the anti-hallucination gate.
Applies guilt-by-association: if a drug is effective in a disease that shares causal genes with yours, it's a candidate for investigation in yours. Each hypothesis cites the shared gene, the source disease, and the source of its current use. It's not a clinical recommendation. It's a starting point for the human researcher.
Implements Don Swanson's classic idea of literature-based discovery [14]: A→B→C, where A is the disease, B is a causal gene, and C is a drug that targets B but that the literature hasn't yet connected to A. Implicit connections that no paper stated directly, but that the graph makes visible.
Inspired by KGARevion [15]: before any claim leaves the system, it is re-checked against the graph. If the claim has support, it goes out with the source. If it doesn't, it's rejected. The most visible public component of this layer is the transparency dashboard at raras.org/transparencia.
The design speaks to the line of MedGraphRAG [16] and HippoRAG [17]: graph as the agent's long-term memory, weighted multi-hop retrieval, evidence traceable down to the passage level. Our reasoning layer over the graph is closed (it's our edge), but the output — all output — is open source.
Layer 3 — The public surface
Three interfaces make the agents' work consumable by the world:
- Disease pages at raras.org/doencas. Each shows the recent literature, the genes involved, the treatments covered by SUS, and the agents' activity journal for that disease. It's what a patient, doctor, or caregiver sees.
-
MCP server at raras.org/api/mcp. Full compliance with the Streamable HTTP specification [5]. Exposes ~20 tools (search, disease description, evidence, hypotheses, phenotypic similarity) + resource templates (
disease://{orphaCode}) + a standard prompt. Any MCP client — Claude Desktop, ChatGPT Connectors, custom agents — plugs in and queries. Listed in the official MCP Registry. -
Open dumps on Zenodo (persistent DOI), Hugging Face Datasets, and KGX for integration with Monarch and the NCATS Biomedical Data Translator. Discoverability via Bioregistry (prefix
raras), Croissant, andllms.txt.
How we keep it honest
The natural question about any agentic system is "does it hallucinate?". Our answer is structural, not rhetorical.
(a) Everything comes from canonical databases. No data is invented by an LLM. Every entity in the graph comes from an explicit canonical allowlist declared in a single file of the open repository: Orphanet, HPO, OMIM, ClinVar, MONDO, GenCC, HGNC, Open Targets, ChEMBL, PubTator3, WikiPathways, Reactome, ClinicalTrials.gov, ReBEC, FDA, EMA, ANVISA, CONITEC, DOU (in.gov.br), bvsms.saude.gov.br, DATASUS, WHO. Adding a source is a deliberate, reviewable act. Four sequential gates protect publication: (1) a canonical domain filter before the LLM ever sees any source, (2) a mandatory citation declaration per paragraph, (3) cross-validation of the cited indices against the final allowlist, (4) a coverage decision — below 60% the page is blocked, between 60% and 75% it publishes with a review flag, above 75% it publishes. When the LLM-as-judge extractor proposes a triple from a PubMed abstract, it enters the graph only if both endpoints already exist in these databases (gene by HGNC symbol, drug by ChEMBL name, disease by Orphanet code). A triple without an anchor is discarded. Every LLM action on the graph carries createdBy and can be reverted in a single query.
(b) Self-consistency as confidence. Each triple passes through N independent LLM runs at different temperatures. The final confidence is the fraction of runs in which the triple appeared. Triples below a threshold are rejected. Re-observations reinforce: r' = 1 - (1 - r) · (1 - r_new), an evidence-combination formula used in the MedKGent line [13].
(c) Verification as a public gate. The transparency dashboard shows, per disease, what was verified and what was rejected in the last 48 hours. Anyone can audit it. No SUS coverage claim is published without an official DOU decree URL or one from the Ministry of Health's Virtual Health Library. The principle is non-negotiable: every claim carries its primary source.
(d) Hypotheses are marked as hypotheses. The repurposing engine and Swanson's ABC generate candidates for investigation, not clinical recommendations. Every hypothesis card carries the disclaimer. No API route mixes hypotheses with verified evidence.
Claim-level validation: every sentence is a nanopublication
The most common critique of any health knowledge base is: "and who validated this?". Our answer stopped being a policy and became an artifact that anyone — or any machine — can audit on their own. In Disease Twin, every sentence a patient reads is verifiable in isolation.
Every agent-authored paragraph carries inline [n] citations to canonical sources, and sentences without a source are discarded by the grounding gate before publication. It's verification at the level of the atomic claim, in the line of FActScore [18] and MedRAGChecker [19] — except running in production over 10,468 diseases, not over a benchmark.
The new layer — machine-readable provenance. Every verified claim is exported as a nanopublication [20] compliant with semantic-web standards: W3C PROV-O plus the Nanopublication schema. There are three named graphs per claim — the assertion (what we say), the provenance (the sources and the agent that generated it), and the publication info (license, grounding, and review status). It's not a promise that we cite: it's the citation graph attached to the data, dereferenceable per disease at raras.org/api/provenance/{orphaCode} in TriG, JSON-LD, and N-Quads formats. It's the same mechanism the scientific community uses to make biomedical assertions traceable and composable [21] — now applied to every sentence of every entry. Third parties can ingest, re-query, and cross-reference our claims without depending on our interface. That's the concrete, non-repudiable form of "validated".
Human review in the loop. When a page's grounding falls between 60% and 75%, it's published with the label medical review pending and prioritized in a queue. Geneticists and specialists sign entries as curators — and that co-signature travels into the nanopublication's provenance (prov:wasAttributedTo → the specialist, with their CRM). The seal stops being "trust us" and becomes "see who reviewed it, when, and against which source". It's the Wikipedia–Cochrane strategy: turn the critic into a co-author.
And what about "rare in Portuguese"? Rarity is jurisdictional — a disease is rare in a population, not in the abstract (EU < 1:2,000; Brazil's PNGDR < 1:1,300). Tuberculosis appears in Orphanet but is endemic in Brazil; presenting it as "rare" here would be wrong. That's why we keep an explicit flag (notRareInBR) over ~77 conditions common in the country, each with a primary source (Compulsory Notification decree, Ministry of Health bulletins, INCA estimates), and we replace the "rare disease" pill with an "endemic in Brazil" seal. Far from inflating rarity, the system deflates it where local epidemiology demands. And it takes on the mission no one else took on: informational equity for the Portuguese language, where local epidemiology (dengue, Chagas, tuberculosis) simply doesn't exist in the English-language databases [22]. Closing that gap for ~260 million Portuguese speakers isn't a bug to apologize for — it's the mandate.
What we measured on day 1
The daily loop ran for the first time in production this week. In a run of about three minutes, starting from an already-enriched base, it:
repurposing hypotheses
extracted from the literature via LLM
edges in the graph
via MCP
Concrete examples of discoveries the LLM extracted and the graph absorbed in that run: Glioblastoma → RFC4 associated, YY1 stimulates RFC4, RFC4 causes temozolomide resistance; Cystic fibrosis → elexacaftor/tezacaftor/ivacaftor treats cystic fibrosis; Multiple myeloma → M-protein associated, serum free light chain ratio associated; Systemic lupus erythematosus → hydroxychloroquine treats SLE and prevents retinopathy; Spinal cord injury → dexmedetomidine treats, Liproxstatin-1 inhibits ferroptosis. Each with an evidence PMID. Each reproducible.
The loop is scheduled to run every day at 4 a.m. Brasília time, in a cron on Cloud Run. On each run, new diseases enter the slice. In 120 days the cycle completes a full rotation through the entire catalog.
Why open is the only option
Google's AI Co-Scientist is private. Stanford's Virtual Lab runs in an academic lab. Sakana's AI Scientist is a product. None of the three was made for the rare-disease patient, and there's no structural reason any of them ever will be. The marginal cost of serving a disease with 50 patients in Brazil is the same as serving one with 50,000 patients in the US, but the expected revenue is zero. In any closed system, the ultra-rare disease always loses.
The layer that builds medicine's agentic knowledge needs to be a commons. Open, auditable, governed by the people who live the conditions. The CC0 license on the disease layer isn't cosmetic: it's the open lock that prevents anyone from buying the base and closing it up. The patient-wallet model is symmetric: no individual data leaves the patient's pocket, but the knowledge about the condition belongs to everyone.
That's the political point of the architecture. The technical layer only works if the governance layer protects against capture. That's why the ones who create communities on the platform are patient associations, not just any users; that's why the founder of this platform has to be a patient; that's why the transparency dashboard exists.
How to connect
If you're a developer, a researcher, or building an AI health tool, the shortest path to using Disease Twin is via MCP. In any MCP-compatible client, add the server:
claude mcp add raras --transport http https://raras.org/api/mcp
After that, any agent using that client can call tools like describe_disease, get_evidence, find_phenotypically_similar, get_hypotheses, analyze_clinical_case, search_diseases. All return structured data with provenance. Full technical documentation at raras.org/mcp.
For deeper integration (KGX dump, Beacon v2, Phenopackets v2, SPARQL): raras.org/docs. MCP server discovery at raras.org/.well-known/mcp. Open repository: github.com/rarasAI/raras.
What comes next
Three fronts in active development right now, in parallel with the daily loop in production:
- Verification layer 2.0 — we're closing the verify → revise → republish loop automatically so that, when an official source changes, affected publications are reverted without human intervention. It's the next iteration of the full KGARevion spirit. Today the system verifies and flags; the republication gate is already being written.
- LLM literature coverage — we're expanding the MedKGent-style extraction to the long tail of diseases with little literature indexed by PubTator. That's where the co-scientist's most original contribution lives: relations no automated system had captured yet. The first runs are already writing to the graph (Glioblastoma, CF, Myeloma, Lupus, spinal cord injury, pre-eclampsia).
- Federation — we're preparing the MCP server so that other institutions can publish their own Disease Twin shards and we can federate queries. The Brazilian SUS version is already ready to be queried side by side with European (ERDERA) or North American (NCATS Translator) versions. No one needs to host the entire world.
The point of this work isn't to admire what we built. It's to hand over the method. Karpathy says the next Wikipedia will be an agentic LLM. The question is who owns the layer. We chose rare diseases as proof that communities can build and validate their own, in production, open source, without permission. The recipe is replicable for any knowledge network the market decided to forget.