即使有了基因组测序,许多罕见病患者仍无法获得明确的遗传诊断。经过广泛检测和专家评估后,大约一半的患者仍无法确诊。他们的医疗数据中可能包含线索,但找到这些线索需要从数千到数百万个可能的遗传变异、零散的临床记录以及快速变化的科学文献中进行筛选。
随着新的基因-疾病关系、病例报告和分类证据的积累,未解决的病例可能变得可以重新解释。
来自波士顿儿童医院曼顿孤儿病研究中心、哈佛大学和OpenAI的研究人员使用OpenAI o3 Deep Research推理模型,分析了376个此前已分析但仍未解决的病例的去标识化临床和基因组信息。该模型为研究人员和临床医生提供了基于证据的候选解释以供审查。经过专家评审、额外检测和临床确认,医生在18个病例中确立了诊断——在专家先前分析的基础上,诊断率额外提高了4.8%。这项研究于2026年6月18日发表在《NEJM AI》上,展示了AI辅助研究流程如何帮助专家在重新审视一些最困难的病例时找到线索。
其中许多病例多年来一直未被专家分析所破解。在这项研究中,OpenAI o3 Deep Research帮助研究人员识别了线索,这些线索随后通过既定的临床流程进行了评估,表明随着知识的演变,由专家主导的定期重新分析可能变得更加可扩展。该模型并未对任何患者进行诊断或做出任何临床决策。它生成了基于证据的假设,供专家审查,并在适当时通过额外检测进行进一步调查,并在临床实验室中确认。
为什么旧病例可能包含新答案
一个不确定的遗传检测结果并不总是永久性的发现。患者的表型描述、检测结果和家族史可能分散在使用不同标识符、格式和词汇的数据库中。将这些记录关联起来很困难,因此即使是专家也可能漏诊。专家可能在相关基因或其变异与疾病关联之前就对儿童基因组进行了测序。随着科学知识的进步,相同的数据可以揭示以前无法发现的答案。
罕见病的重新分析既是一个科学问题,也是一个维护问题。患者的基因组可能保持不变,但围绕它的证据却在不断变化:研究人员将新的基因和变异与疾病联系起来,实验室重新分类旧变异,病例数据库和论文积累新的观察结果。每次更新都可能使旧的未确诊病例值得重新审视,因此许多机构积累了大量需要与不断变化的知识库保持同步的基因组。
在这项研究中,研究人员设计了工作流程,使模型在现有基因组分析流程之上充当一个以解释为先的推理层。模型不是仅返回一个排序后的基因,而是被要求将临床特征、遗传模式、变异证据和科学文献联系起来,形成一个人类评审员可以质疑的论证。
重新分析是如何进行的
对于每个病例,团队整理了一个去标识化的数据包,包含标准的人类表型本体术语来描述患者的临床表现、偶尔的临床医生笔记和任何描述性临床诊断、年龄和性别等元数据,以及一个过滤后的变异表。该表记录了每个变异的稀有性、其对编码蛋白质的预测影响、ClinVar分类以及可用家庭成员中的信号质量。大多数病例包括来自儿童及其亲生父母的数据。
团队要求模型提出最合理的分子解释,并展示其推理过程。然后,研究人员使用临床实验室用于分类遗传变异的相同ACMG/AMP框架审查输出结果。每个候选结果至少由两名团队成员审查,分歧通过共识解决,模型输出从未被视为诊断。只有在合格专家审查证据、变异被分类为致病性或可能致病性、CLIA认证实验室确认,并且临床团队将结果反馈给家庭后,发现才被计为诊断。
在分析未解决病例之前,团队在已确诊的病例上优化了工作流程。在包含各种罕见疾病的51个病例中,该流程在重复运行中恢复了48个病例的正确基因和变异。在一组57个神经肌肉病例中,该流程在重复运行中为45个病例返回了正确诊断。在一个15个病例的长读长基因组数据集中,该流程在每个病例中都命名了正确的基因,并在12个病例中命名了两个致病等位基因。这些评估有助于提示开发,并显示了专家审查仍然至关重要的地方。
模型自我报告的置信度分数与这些先前已解决病例中的正确诊断相关:一致正确判断的平均最低分数为85.6,而错误或未知判断的平均最低分数为42.1。这些分数并非校准后的概率,团队也未将其用作证据或临床裁决的替代品。但它们有助于指导专家评审员关注最有希望的候选诊断。
研究人员的发现
随后,团队将该流程应用于四组先前未解决的病例:患有神经发育疾病的儿童、患有罕见神经肌肉疾病的人、患有早期精神病的儿童和青少年,以及儿科突发意外死亡病例。这些并非等待首次审查的新病例。许多病例已经经过多个商业或机构流程的检查,并由多学科团队讨论过。
按队列划分的结果
队列病例数发现的诊断数****诊断率 神经发育 100 10 10.0% 神经肌肉疾病 61 4 6.6% 儿科突发意外死亡 200 2 1.0% 早期精神病 15 2 13.3% 总计37618****4.8%
早期精神病队列规模较小,因此其百分比具有较宽的置信区间。诊断率也反映了每个队列具有单基因解释的可能性。
在模型提出候选结果、专家完成审查和临床确认后,医生在4.8%的病例中确立了诊断。这个比率虽然不高,但在这一人群中具有重要意义,因为先前的专家审查未能解决这些病例。类似的重新分析研究报告称,在经过大量审查的病例中,诊断率仅有个位数增长;较高的诊断率通常来自包含新病例或等待遗传确认的已知疾病的研究。
在18项诊断中,有7项属于重新发现:这些诊断是在本地研究流程之外确立的,但团队审查的记录中并未包含。在多个案例中,这些变异已在公共数据库中被列为致病性或可能致病性,凸显了跨数据源整合信息的操作挑战。
识别变异时展现灵活性
在一个早期精神病案例中,模型推断出输入数据中未列出的基因组结构事件。它将22号染色体上一系列低质量测序结果与患儿的心脏、免疫、神经发育及精神特征联系起来,随后假设存在与迪乔治综合征相关的22q11.2缺失。这一假设的变异通过后续基因组测序得到确认。
尽管提示要求寻找单一单基因病因,但模型有时会提出两个基因以更好地解释复杂表现。在一个案例中,_LAMA2_和_FOXP1_的变异共同解释了肌肉和神经发育特征;另一个案例则存在涉及_TTN_和_SRPK3_的、此前未被识别的双基因解释。
生成可验证、生物学上连贯的假说
除诊断外,模型还为一种名为白癜风的疾病提出了可能的新型机制解释。在一个神经发育案例中,模型突显了一名白癜风患者_S1PR1_基因中11个氨基酸的缺失。_S1PR1_编码一种参与信号传导、免疫细胞迁移和组织生物学的细胞表面受体。模型整合证据表明,该缺失可能改变受体结构和信号传导方式,从而减少色素生成,同时帮助免疫细胞在皮肤中持续存在。
所提出的_S1PR1_-白癜风关联需要额外的实验验证,但这展示了人工智能在将结构生物学、免疫学和临床遗传学的零散发现转化为具体、可验证假说方面的强大作用。
团队在神经肌肉队列中也观察到可能的表型扩展。_HSPB8_和_CDK13_的有害变异与这些基因最广为人知的疾病不完全匹配,提示存在更广泛的临床谱系,需要更多病例和实验室工作来验证。
局限性
这项研究表明,通用推理模型能够通过整合表型、遗传模式、变异注释、数据质量模式和科学文献,形成可审查的假说,从而为回顾性基因组再分析做出贡献。这也说明了定期再分析的重要性:有些答案只有在知识进步或碎片化记录被整合后才会浮现。
这项研究并非证明患者、临床医生或客户应使用OpenAI模型来诊断疾病或做出医疗决策。它不描述或认可将OpenAI o3 Deep Research、ChatGPT或任何其他OpenAI产品用于诊断的预期客户用途。模型未对任何参与者进行诊断;所有诊断均由医生和其他合格临床专家通过既定审查、检测和临床确认流程做出。
该研究是回顾性的,队列具有异质性,且审查人员对模型置信度未设盲。研究人员未测量节省的时间、成本、临床医生工作量、假阳性工作负担或护理变化。也未系统评估其他形式的遗传变异,如结构变异、重复扩增、深内含子改变或嵌合体。
大型语言模型可能误读上下文或产生看似合理但经不起推敲的解释。因此,每个结果都经过人工裁决和临床确认。模型拓宽了搜索范围并聚焦了后续的人工主导分析;它并未决定应向家庭返回哪些信息或诊断。
本研究使用了去标识化信息,未在批准环境之外使用或传输受保护的健康信息。更广泛的临床部署将需要对所有医疗护理适用的隐私、安全、可审计性和当地法规给予同等关注。模型访问不能替代测序基础设施、遗传咨询、验证性检测或专家判断。
凯瑟琳·布朗斯坦博士,波士顿儿童医院曼顿孤儿病研究中心
艾伦·贝格斯,曼顿孤儿病研究中心主任
未来展望
前瞻性、多中心研究应将LLM辅助再分析与标准实践在诊断率、候选基因发现时间、临床医生工作量、假阳性负担、成本及对护理的影响方面进行比较。版本化的提示、参考文献核查、审计日志和校准的不确定性对于可重复性和安全性至关重要。此类研究仍需要合格临床医生来评估证据、安排适当检测并做出任何诊断或治疗决策。
本研究使用了OpenAI o3 Deep Research。更新的通用模型可以搜索和综合更多科学材料,而GPT‑Rosalind等专用系统则专为更深入的生命科学工作设计,包括变异对蛋白质结构和功能的影响。这些能力未在此测试,需要各自的评估和访问控制。
虽然OpenAI帮助支持了这项初步研究,但曼顿中心将通过OpenAI基金会的资助领导下一阶段工作。该资助将支持该中心更广泛的努力,开发一个平台无关、低成本的遗传学AI辅助工具,帮助临床团队更快速、一致地分析罕见病案例。
长期的研究机会在于探索专家主导的AI辅助再分析是否有助于科学理解跟上发现的步伐。其前景并非AI取代医生的诊断,而是经过仔细评估的研究工具可能帮助专家识别值得调查的证据。对于成千上万的家庭而言,今天未解答的问题不必永远悬而未决。
Even with genomic sequencing, many people with rare diseases never receive a clear genetic diagnosis. Roughly half remain undiagnosed after extensive testing and specialist review. Their medical data may contain clues but finding them can require sifting through thousands to millions of possible genetic variants, fragmented clinical records, and rapidly changing scientific literature.
As new gene-disease relationships, case reports, and classification evidence accumulate, unsolved cases can become newly interpretable.
Researchers from Boston Children’s Hospital’s Manton Center for Orphan Disease Research, Harvard University, and OpenAI used the OpenAI o3 Deep Research reasoning model to analyze de-identified clinical and genomic information from 376 previously analyzed cases that remained unsolved. The model surfaced evidence-linked candidate explanations for researchers and clinicians to review. Following expert review, additional testing, and clinical confirmation, physicians established diagnoses in 18 cases—an additional diagnostic yield of 4.8% after earlier analysis by specialists. This study was published on June 18, 2026, in NEJM AI and shows how an AI-assisted research workflow can help experts generate leads when revisiting some of the most difficult cases.
Many of these cases had evaded years of expert analysis. In this study, OpenAI o3 Deep Research helped researchers identify leads that were later assessed through established clinical processes, suggesting that expert-led periodic reanalysis could become more scalable as knowledge evolves. The model did not diagnose any patient or make any clinical decision. It produced evidence-linked hypotheses for specialists to review and, where appropriate, investigate through additional testing and confirm in a clinical laboratory.
Why an old case can contain a new answer
An inconclusive genetic test is not always a permanent finding. A patient’s phenotype descriptions, test results, and family history can be split across databases that use different identifiers, formats, and vocabularies. Linking those records is difficult, so even specialists can miss a diagnosis. Experts may also sequence a child’s genome before a relevant gene or its variants have been linked to disease. As scientific knowledge advances, the same data can reveal answers that were previously impossible to uncover.
Rare-disease reanalysis is both a scientific and a maintenance problem. The patient’s genome may stay the same, but the evidence around it keeps changing: researchers link new genes and variants to disease, labs reclassify old variants, and case databases and papers accumulate new observations. Each update can make an old inconclusive case worth revisiting, so many institutions inherit a growing backlog of genomes to keep in sync with a moving knowledge base.
In this study, researchers designed the workflow so that the model acted as an explanation-first reasoning layer on top of existing genomic pipelines. Instead of returning only a ranked gene, it was asked to connect the clinical features, inheritance pattern, variant evidence, and scientific literature into a justification that a human reviewer could interrogate.
How the reanalysis worked
For each case, the team assembled a de-identified packet containing standardized Human Phenotype Ontology terms to describe the patient’s clinical presentation, occasional clinician notes and any descriptive clinical diagnosis, metadata such as age and gender, and a filtered variant table. The table captured each variant’s rarity, its predicted effect on the encoded protein, ClinVar classification, and signal quality across available family members. Most cases included data from the child and both biological parents.
The team asked the model to propose the most plausible molecular explanation and to show its work. Researchers then reviewed the outputs using the same ACMG/AMP framework that clinical labs use to classify genetic variants. At least two team members reviewed each candidate, disagreements were resolved by consensus, and a model output was never treated as a diagnosis. A finding counted as a diagnosis only after qualified experts reviewed the evidence, the variant was classified as pathogenic or likely pathogenic, a CLIA-certified laboratory confirmed it, and the clinical team returned the result to the family.
Before analyzing unsolved cases, the team refined the workflow on cases with established diagnoses. It recovered the correct gene and variant in duplicate runs for 48 of 51 cases that included a variety of rare conditions. In a set of 57 neuromuscular cases, the workflow returned the correct diagnosis in duplicate runs for 45 of the cases. In a 15-case long-read genome set, it named the correct gene in every case and both disease-causing alleles in 12 cases. These evaluations aided in prompt development and showed where expert review remained essential.
The model’s self-reported confidence scores tracked with correct diagnoses in these previously solved cases: the mean minimum score was 85.6 for consistently correct calls and 42.1 for incorrect or unknown calls. The scores were not calibrated probabilities, and the team did not use them as a substitute for evidence or clinical adjudication. But they were helpful in guiding the expert reviewers to focus on the most promising candidate diagnoses.
What the researchers found
The team then applied the workflow to four groups of previously unsolved cases: children with neurodevelopmental conditions, people with rare neuromuscular disease, children and adolescents with early psychosis, and cases of sudden unexpected death in pediatrics. These were not fresh cases awaiting a first review. Many had already been examined by multiple commercial or institutional pipelines and discussed by multidisciplinary teams.
Results by cohort
CohortCasesDiagnoses surfaced****Yield Neurodevelopmental 100 10 10.0% Neuromuscular disease 61 4 6.6% Sudden unexpected death in pediatrics 200 2 1.0% Early psychosis 15 2 13.3% Total37618****4.8%
The early psychosis cohort was small, so its percentage has a wide confidence interval. Yield also reflects how likely each cohort was to have a single-gene explanation.
After the model surfaced candidates and experts completed review and clinical confirmation, physicians established diagnoses in 4.8% of the cases. That rate is modest but meaningful in this population because previous expert reviews had not resolved the cases. Similar reanalysis studies report single-digit gains in heavily reviewed cases; higher yields usually come from studies containing new cases or well-known disorders awaiting genetic confirmation.
Of the 18 diagnoses, 7 were rediscoveries: diagnoses established outside the local research workflow but absent from the record the team reviewed. In several cases, the variants were already listed as pathogenic or likely pathogenic in public databases, highlighting the operational challenge of synthesizing information across data sources.
Demonstrating flexibility when identifying variants
In one early-psychosis case, the model inferred a structural event in the genome that was not listed in the input data. It connected a run of low-quality calls on chromosome 22 with the child’s cardiac, immune, neurodevelopmental, and psychiatric features, then hypothesized a 22q11.2 deletion associated with DiGeorge syndrome. This hypothesized variant was confirmed with follow-up genome sequencing.
Although the prompt asked for one monogenic cause, the model sometimes surfaced two genes that better explained a complex presentation. Variants in LAMA2 and FOXP1 together helped account for muscle and neurodevelopmental features in one case; another had a previously unrecognized digenic explanation involving TTN and SRPK3.
Producing a testable, biologically coherent hypothesis
In addition to diagnoses, the model also identified a possible novel mechanistic explanation for a condition called vitiligo. In one neurodevelopmental case, the model highlighted an 11-amino-acid deletion in S1PR1 in a person with vitiligo. S1PR1 encodes a cell-surface receptor involved in signaling, immune-cell movement, and tissue biology. The model integrated evidence suggesting that the deletion could alter receptor structure and signaling in ways that reduce pigment production while also helping immune cells persist in the skin.
The proposed S1PR1-vitiligo relationship requires additional experimental validation but it illustrates a powerful role for AI in translating scattered findings from structural biology, immunology, and clinical genetics into concrete, testable hypotheses.
The team also saw possible phenotype expansion in the neuromuscular cohort. Damaging variants in HSPB8 and CDK13 did not perfectly match the genes’ best-known disorders, suggesting a broader clinical spectrum that more cases and laboratory work will need to test.
Limitations
This study shows that a general-purpose reasoning model can contribute to retrospective genomic reanalysis by combining phenotype, inheritance, variant annotations, data-quality patterns, and scientific literature into reviewable hypotheses. It also shows why periodic reanalysis matters: some answers surface only after knowledge advances or fragmented records are brought together.
This research is not evidence that patients, clinicians, or customers should use OpenAI models to diagnose disease or make medical decisions. It does not describe or endorse an intended customer use of OpenAI o3 Deep Research, ChatGPT, or any other OpenAI product for diagnosis. The model did not diagnose any participant; physicians and other qualified clinical experts made every diagnosis through established review, testing, and clinical-confirmation processes.
The study was retrospective, the cohorts were heterogeneous, and reviewers were not blinded to model confidence. The researchers did not measure time saved, cost, clinician effort, false-positive workload, or changes in care. Nor did they systematically evaluate other forms of genetic variation such as structural variants, repeat expansions, deep-intronic changes, or mosaicism.
Large language models can misread context or produce plausible explanations that fail upon closer inspection. Therefore, every result passed through human adjudication and clinical confirmation. The model widened the search and focused the subsequent human-led analysis; it did not decide what information or diagnosis should be returned to a family.
This study used de-identified information, with no protected health information utilized or transmitted outside approved environments. Broader clinical deployment will require the same attention to privacy, security, auditability, and local regulation that applies to all medical care. Model access does not replace sequencing infrastructure, genetic counseling, confirmatory testing, or specialist judgment.
Dr. Catherine Brownstein, Boston Children’s Hospital’s Manton Center for Orphan Disease Research
Alan Beggs, director of the Manton Center for Orphan Disease Research
What comes next
Prospective, multi-center studies should compare LLM-assisted reanalysis with standard practice on diagnostic yield, time to a candidate, clinician effort, false-positive burden, cost, and effects on care. Versioned prompts, reference checks, audit logs, and calibrated uncertainty will be important for reproducibility and safety. Such studies would still require qualified clinicians to evaluate evidence, order appropriate tests, and make any diagnosis or treatment decision.
This study used OpenAI o3 Deep Research. Newer general-purpose models can search and synthesize more scientific material, while purpose-built systems such as GPT‑Rosalind are designed for deeper life-sciences work, including variant effects on protein structure and function. Those capabilities were not tested here and will require their own evaluations and access controls.
While OpenAI helped support this initial research study, the Manton Center will lead the next stage of the work through a grant from the OpenAI Foundation. The grant will support the Center's broader effort to develop a platform-agnostic, low-cost genetics AI copilot that helps clinical teams analyze rare disease cases more quickly and consistently.
The longer-term research opportunity is to explore whether expert-led AI-assisted reanalysis can help scientific understanding keep pace with discovery. The promise is not that AI replaces a doctor’s diagnosis, but that carefully evaluated research tools may help specialists identify evidence worth investigating. For thousands of families, today’s unanswered questions do not have to remain unanswered forever.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文