Eighteen children who had spent years, in some cases more than a decade, without a diagnosis for their rare diseases finally got answers, thanks to an artificial intelligence model that did what teams of human geneticists could not. Researchers at Boston Children's Hospital and OpenAI published their findings Thursday in NEJM AI, showing that a commercially available AI tool identified genomic errors behind rare pediatric conditions in cases that had already been analyzed multiple times by human experts.
The results are modest in percentage terms, roughly 5 percent of the 376 undiagnosed patients in the study received new diagnoses, but each number represents a family that had been waiting, sometimes for fifteen years, for a name to attach to their child's suffering. In a medical system that struggles to recruit enough geneticists to process the growing backlog of genomic data, this is the kind of practical, measurable breakthrough that deserves attention far beyond the research journals.
The research team, led by Catherine Brownstein, scientific director of the genetic investigations arm of the Manton Center for Orphan Disease Research, fed OpenAI's o3 model clinicians' notes, patient symptom descriptions, and filtered gene lists for 376 patients who had never received a diagnosis. The AI processed all 376 genomes. A human research team then reviewed every output before making final diagnostic calls.
The result: 18 new diagnoses across four disease areas. Ten patients had rare neurodevelopmental diseases. Four had neuromuscular disorders. Two had early childhood psychosis illnesses. Two were children who had died suddenly.
Seven of those 18 diagnoses turned out to be what the researchers called "rediscoveries", cases where a diagnosis had been identified elsewhere but never shared globally with the patient's care team. That detail alone points to a systemic communication failure in rare disease medicine that AI may be uniquely positioned to fix.
As NBC News reported, Brownstein framed the significance plainly:
"It got almost 5% new diagnoses, which doesn't sound like a lot, but considering how many times these had already been analyzed, that's a huge number, and each one means an answer for a family."
Kyra Benton's story puts a human face on the numbers. Around age nine, Benton's mother noticed her daughter moved differently from her peers. She walked on her tiptoes and struggled to run with a normal gait. Her mother took her to a neuromuscular disease expert in New York City. The expert couldn't identify the disease. They tried Boston Children's Hospital. Same result.
At thirteen, Benton underwent a tracheotomy. Still no diagnosis.
Then, last summer, about a week before her twentieth birthday, the phone rang. A researcher from the Manton Center lab was on the line. Benton recalled the conversation:
"She said, 'Hi, we know it's been about 15 years, but we have some news for you,' and it kind of just blossomed from there."
The diagnosis: myofibrillar myopathy. After fifteen years of uncertainty, Benton finally had a name for what was wrong. That matters not just emotionally but medically, a diagnosis opens the door to targeted treatment and, as Brownstein noted, ensures patients are "first in line for any new technological development or any new therapy" when treatments do come online.
The story carries echoes of other high-profile cases where rare diagnoses reshaped lives and careers, a reminder that behind every medical statistic sits a real person waiting for answers.
The human genome contains approximately 20,000 protein-coding genes. Analyzing a single patient's genome against the latest medical literature is grueling work. Boston Children's Hospital sees over a thousand children every day. The Manton Center alone works with more than 3,500 individuals globally, across all fifty states, affected by rare diseases.
There simply are not enough trained human analysts to keep up. Brownstein described the scale of the problem in blunt terms:
"There's pages upon pages of these genes that I have to get through for a case, while the LLM doesn't get tired."
Suyash Shringarpure, a technical researcher at OpenAI who focuses on health applications, made a related point about the time dimension. A case might go unsolved on its first review, but a year later new research could clarify the link between a gene and a disease. Human analysts rarely have time to circle back. The AI can.
"A researcher can only spend so much time on a single case. Maybe a case remained unsolved when it came to them first, but a year later a paper was published that clarifies the link between the gene and the disease."
This is not a new problem for Boston Children's. Fox News previously reported on the hospital's earlier partnership with IBM's Watson supercomputer to tackle rare pediatric diseases, noting that doctors were spending an average of 25 hours per case on genomic analysis. That earlier effort focused on kidney disease. The new research with OpenAI's o3 model covers a broader range of conditions and uses a commercially available tool, not a custom supercomputer installation.
The shift from bespoke systems to off-the-shelf AI products is significant. It means the technology could, in principle, be deployed at hospitals and research centers that lack the resources of a world-class institution like Boston Children's.
Outside experts who were not involved in the research offered measured praise. Adam Rodman, a physician at Beth Israel Deaconess Medical Center who studies AI in medicine, called the results meaningful:
"A diagnostic yield of 5% is truly meaningful and could serve as a significant screening tool to help speed up the reanalysis of significant backlogs of cases."
The backlogs he references are real and growing. As genomic sequencing becomes cheaper and more widespread, the bottleneck has shifted from collecting data to interpreting it. AI tools that can serve as a first-pass screening layer, flagging cases for human review rather than replacing human judgment, address that bottleneck directly.
Chunhua Weng, a professor of bioinformatics at Columbia University, called the paper a "wonderful" contribution but added a necessary caution: "The appropriate use of LLMs in diagnosis requires careful attention to trustworthiness." That warning is worth taking seriously. Large language models can produce confident-sounding errors. The Boston Children's team addressed this by requiring human review of every AI output before any diagnosis was finalized.
When medical systems fail to catch serious conditions in time, the consequences can be devastating, as illustrated by recent revelations about missed diagnoses even at the highest levels of government. Any tool that improves diagnostic accuracy deserves rigorous but open-minded evaluation.
Ashley Alexander, OpenAI's head of health, struck a careful tone about the findings, one that conservatives who value honest claims over hype should appreciate:
"We definitely don't want to overhype this. But I also want to make sure that people don't miss what's happening and what's possible with even just the version of ChatGPT that's in their pocket today."
That framing matters. The tool used in this study was not a classified government project or a billion-dollar custom system. It was a commercially available AI model. The research was peer-reviewed and published in a journal affiliated with the New England Journal of Medicine. The human team retained final authority over every diagnosis.
This is how technology should work in medicine: as a force multiplier for human expertise, not a replacement for it. The AI flagged possibilities. The doctors made the calls. Families got answers.
Even Kyra Benton, who described herself as "not all that much in favor of AI," acknowledged the result. Her words carry the weight of someone who lived through fifteen years without a diagnosis:
"Such as in this case, where it can lead to massive breakthroughs that can really change people's lives for the better."
Rare diseases affect roughly one in ten Americans, and the diagnosis journey for many families stretches across years of inconclusive tests and specialist visits. Early, accurate diagnosis can mean the difference between timely treatment and years of avoidable suffering, a reality that resonates well beyond the pediatric context, as other recent high-profile diagnoses have underscored.
The study leaves important questions open. The specific methodology, how clinicians' notes and gene lists were structured for the AI, is not fully described in the public reporting. It remains unclear whether the 5 percent diagnostic yield would hold across different patient populations or disease categories. And while the peer-reviewed publication in NEJM AI lends credibility, the research will need replication at other institutions before anyone can call it a proven standard of care.
The "rediscovery" problem, seven of eighteen diagnoses that had been identified elsewhere but never communicated, also raises uncomfortable questions about information-sharing failures in the medical system. That AI had to find diagnoses that already existed somewhere in the global research literature is not just a technology story. It is an indictment of how poorly the medical establishment shares what it already knows.
None of this diminishes the achievement. Eighteen families have answers they did not have before. The tool that found those answers is available today, not in some distant research pipeline. And the human experts who used it did so responsibly, with peer review and independent verification.
When the private sector builds a tool that works, puts it in the hands of skilled doctors, and delivers real results for real patients, without a government mandate, without a new federal program, without a billion-dollar appropriation, that is worth noticing. And worth building on.