Як Boston Children’s застосував OpenAI для повторного аналізу рідкісних хвороб
Evidence-heavy кейс Boston Children’s + OpenAI: de-identified genomic/clinical data, evidence-linked hypotheses, specialist review, confirmatory testing і clinical authority — без підміни лікаря моделлю.
Картка кейсу
Що тут автоматизовано
Обсяг автоматизації
У дослідженні Boston Children’s Hospital, Harvard University і OpenAI reasoning model використовувався для retrospective reanalysis de-identified clinical/genomic information з раніше нерозв'язаних cases та формування evidence-linked candidate hypotheses. AI-Magister відтворює це як clinician-led research workflow: data eligibility, structured case package, model-assisted hypothesis generation, evidence review, specialist adjudication, additional testing і clinical confirmation. Модель не ставить діагноз і не має clinical decision authority.
Роль людини
Qualified researchers, geneticists і clinicians визначають cohort/data eligibility, оцінюють candidate explanations, відкидають unsupported hypotheses, замовляють додаткові тести й роблять clinical confirmation. Privacy/security owners контролюють de-identification та approved environments. Пацієнтський diagnosis/treatment decision не делегується моделі.
Заявлені результати
- 376 previously unsolved cases were reanalyzed in the published retrospective study — study cohort size
- 18 diagnoses were established after model-surfaced leads, expert review, additional testing and clinical confirmation — study outcome
- 4.8% additional diagnostic yield across the heterogeneous retrospective cohort — study result, not a general clinical accuracy rate
- OpenAI separately reports 40+ rare conditions diagnosed across Boston Children's broader AI-enabled work, plus 60,000 hours saved and $7M+ redeployed labor across operational workflows; these broader hospital/provider-reported metrics are not attributed to this specific 376-case reanalysis study
- The study did not measure time saved, cost, clinician effort, false-positive workload or change in care — explicit limitation from OpenAI's research summary
OpenAI 18 червня 2026 року описала NEJM AI study: OpenAI o3 Deep Research reanalyzed de-identified information from 376 previously unsolved cases and surfaced evidence-linked candidates; після expert review, additional testing і clinical confirmation physicians established 18 diagnoses, тобто 4.8% additional diagnostic yield in this retrospective heterogeneous sample. OpenAI прямо зазначає, що model did not diagnose participants, study не вимірювало time saved, cost, clinician effort або false-positive workload. Boston Children’s Manton Center офіційно описує той самий 376-case/18-diagnosis result. Це research evidence з explicit limitations, а не product claim про автономну діагностику.
Зміст статті
- 01Задача: масштабувати повторний аналіз, не перетворити reasoning model на клінічний oracle
- 02Trigger, input, AI stage, integrations та output
- 03Evidence hierarchy: hypothesis, review, test, confirmation
- 04Human-in-the-loop, privacy і autonomy A2
- 05Error handling і failure modes
- 06Метрики: що вимірювати, а чого не вигадувати
- 07Evaluation contract і staged rollout
- 08Кому підходить і як повторити
Передумови
Задача: масштабувати повторний аналіз, не перетворити reasoning model на клінічний oracle
Rare-disease cases можуть залишатися нерозв'язаними роками: геномні варіанти численні, phenotype неповний, а gene-disease literature змінюється швидше, ніж одна команда може системно перечитувати все вручну. Саме periodic reanalysis є дорогою когнітивною задачею, де reasoning model потенційно корисна як search-and-hypothesis amplifier.
Boston Children’s, Harvard і OpenAI перевірили це ретроспективно на 376 already analyzed unsolved cases. Критична межа кейсу: model surfaced evidence-linked hypotheses; physicians established diagnoses лише після expert adjudication, додаткових tests і clinical confirmation. Тому правильна production abstraction тут — research copilot з autonomy A2, а не autonomous diagnostician.
architecture
Карта системи: Як Boston Children’s застосував OpenAI для повторного аналізу рідкісних хвороб
Trigger, input, AI stage, integrations та output
Trigger — approved retrospective reanalysis event: new literature, updated gene-disease evidence або scheduled review of unresolved cases. Input — de-identified clinical phenotype, genomic information, inheritance/context metadata та дозволена scientific evidence base. До model call діють data-eligibility, de-identification і access controls.
AI stage synthesizes phenotype, variant/context information і literature, формуючи candidate explanations із evidence trail. Integrations у production-like research environment можуть включати case registry, variant annotations, literature retrieval і laboratory workflow, але consequential clinical systems не повинні бути writable за замовчуванням. Output — ranked/reviewable hypotheses з citations, uncertainty і missing-evidence fields.
- Trigger → scheduled/knowledge-driven reanalysis approved by research/clinical owner.
- Input → de-identified, task-eligible clinical/genomic packet + versioned scientific evidence.
- AI → generate evidence-linked candidate hypotheses, не diagnosis.
- Human → specialist adjudication → additional test where indicated → clinical/lab confirmation.
- Output → reviewed research finding; patient-facing decision лишається clinical authority.
timeline
Контрольні точки для практичного застосування
- Trigger → scheduled/knowledge-driven reanalysis approved by research/clinical own…
Контрольна теза з матеріалу статті.
- Input → de-identified, task-eligible clinical/genomic packet + versioned scientif…
Контрольна теза з матеріалу статті.
- AI → generate evidence-linked candidate hypotheses, не diagnosis.
Контрольна теза з матеріалу статті.
- Human → specialist adjudication → additional test where indicated → clinical/lab…
Контрольна теза з матеріалу статті.
- Output → reviewed research finding; patient-facing decision лишається clinical au…
Контрольна теза з матеріалу статті.
- model-routing
Evidence hierarchy: hypothesis, review, test, confirmation
Для такого workflow одна з найнебезпечніших помилок — змішати levels of evidence. Model-generated hypothesis має нижчий статус, ніж specialist interpretation; specialist interpretation — нижчий, ніж validated laboratory/clinical confirmation. UI і data model повинні зберігати цей hierarchy явно, а не показувати всі outputs як однаково 'готові'.
Claim-level provenance має включати source publication, access date/version, genomic annotation version, model/reasoning configuration і reviewer decision. Якщо literature змінилася, попередній hypothesis не стає автоматично false, але його evidence snapshot уже historical. Reanalysis має бути version-aware.
Human-in-the-loop, privacy і autonomy A2
A2 означає, що AI може автоматизувати підготовку й аналіз, але decision loop залишається expert-led. OpenAI прямо зазначає: модель не діагностувала participants і не приймала clinical decisions. Кожен позитивний lead проходив established review/testing/confirmation processes. Це не косметичний disclaimer, а основна архітектура безпеки.
Study використовував de-identified information і не передавав protected health information поза approved environments. Для реального deployment потрібні purpose limitation, least-privilege access, auditability, retention/deletion rules і локальна regulatory review. Навіть дуже сильний model result не дає права розширити data scope постфактум.
Error handling і failure modes
Failure modes: phenotype incomplete; variant annotation stale; literature contradicts candidate; model overweights a plausible paper; citation не підтримує claim; duplicate patient/case identity; data-quality artifact; hypothesis medically coherent but laboratory test negative. System повинен зберігати rejected/uncertain state, а не тиснути reviewer до binary answer.
No-answer є валідним outcome. Якщо evidence недостатньо, модель повертає unresolved gaps і possible next research questions, а не вигадує closure. Negative findings і false leads мають потрапляти в evaluation corpus, інакше система вчитиметься лише на красивих success stories.
Метрики: що вимірювати, а чого не вигадувати
Study result 18/376, або 4.8% additional diagnostic yield, не є '95.2% error rate' і не є general diagnostic accuracy. Cohorts heterogeneous; early psychosis subgroup був малий; reviewers не були blinded to model confidence. OpenAI прямо вказує, що time saved, cost, clinician effort і false-positive workload не вимірювались.
Production KPI мають включати hypothesis recall on adjudicated historical cases, unsupported-citation rate, reviewer workload, time to adjudication, additional-test burden, privacy incidents, clinically confirmed yield і harmful false-reassurance rate. Окремо вимірюється abstention quality. Будь-яка ROI-модель без reviewer/test cost тут буде бухгалтерською фантастикою.
Evaluation contract і staged rollout
Eval set: solved historical cases з hidden outcome, truly unresolved cases, phenotype noise, outdated literature, conflicting papers, ancestry/data-bias slices, missing inheritance info, citation traps і adversarial text inside imported documents. Deterministic checks validate structured fields/citations/data eligibility; specialist graders оцінюють plausibility, evidence support, harmful omission і next-step appropriateness.
Rollout: `retrospective offline study → shadow reanalysis → research-only candidate generation → specialist-reviewed workflow → narrowly governed prospective assistance`. Ніякого automatic patient communication або treatment recommendation. Model upgrade запускає full regression on locked cases; high-severity miss або privacy breach блокує promotion.
Кому підходить і як повторити
Патерн релевантний academic medical centers, genomic research programs і specialist diagnostic teams із curated case registry, scientific-literature access і established confirmation processes. Він не підходить consumer symptom checker, який намагається напряму перенести research model у patient-facing diagnosis.
Pilot починається зі retrospective de-identified dataset і ethics/privacy approval. Зафіксуйте baseline specialist process, lock outcomes, побудуйте provenance-preserving case packet, проганяйте model у shadow і оцінюйте hypotheses blind там, де можливо. Лише після доказу incremental utility та прийнятного reviewer burden переходьте до prospective research assistance.
Практичні приклади
Приклад: unsolved genomic case після появи нового gene-disease evidence
Case registry створює approved reanalysis snapshot без direct identifiers. Model формує evidence-linked candidate і вказує missing confirmation. Geneticist перевіряє source, phenotype fit та inheritance, після чого вирішує, чи потрібен додатковий lab test. Тільки validated clinical process може змінити diagnosis state; model output сам по собі цього не робить.
FAQ
Чи OpenAI model сам поставив 18 діагнозів?
Ні. Model surfaced candidate explanations; diagnoses were established by physicians after expert review, additional testing and clinical confirmation.
Що означає 4.8%?
Additional diagnostic yield у retrospective heterogeneous cohort 376 previously unsolved cases. Це не загальна accuracy моделі й не очікуваний результат для будь-якої клініки.
Чи дослідження довело економію часу або грошей?
Ні. OpenAI прямо зазначає, що study не вимірювало time saved, cost, clinician effort або false-positive workload.
Яка правильна autonomy?
A2: model-assisted research and hypothesis generation з mandatory specialist adjudication та clinical/laboratory confirmation; diagnosis/treatment authority залишається у qualified clinicians.
Пов’язані матеріали
Production-кейс Samsung Electronics + OpenAI: глобальний rollout ChatGPT Enterprise і Codex для R&D, manufacturing, marketing та corporate work — із multi-model governance, permission boundaries і перевіркою фактичних результатів.
Як Cisco зробив Codex частиною enterprise engineeringProduction-кейс Cisco + OpenAI: Codex у великих codebases, defect remediation, framework migrations і product engineering — із plan artifacts, CI/security gates, human merge authority та measured throughput.
Як HSP GRUPPE масштабує ChatGPT Enterprise у податковому консалтингуProduction-кейс HSP GRUPPE: ChatGPT Enterprise працює в tax advisory, legal research, client communication, financial analysis і knowledge sharing, а agentic workflows проходять окремий pilot та governance gate.
Як Australian Payments Plus використовує ChatGPT і Codex у критичній платіжній інфраструктуріProduction-кейс AP+: ChatGPT Enterprise допомагає працювати зі складними rules/specifications, а Codex прискорює technical investigation і functional simulations із human accountability у regulated payments environment.
Як BBVA масштабує ChatGPT Enterprise у банку: 100 000 користувачів, governance і шлях до AI-native bankingProduction-кейс BBVA: ChatGPT Enterprise як керований enterprise layer для knowledge work, custom GPTs і банківських AI-сценаріїв із security, legal, compliance та human-controlled authority.
Red teaming LLM-системПрактичний red teaming перетворює припущення про безпеку LLM-системи на відтворювані атаки, докази та regression-тести. Розглядаємо threat model, ручні й автоматизовані кампанії, triage, безпечну лабораторію та перевірку виправлень.
State machines для агентівState machines для агентів — практичний розбір production-архітектури: відокремлення ймовірнісного рішення моделі від детермінованого життєвого циклу виконання. Матеріал охоплює контракти, межі повноважень, failure modes, оцінювання та контрольований rollout.
Вибір моделей і model routingЯк маршрутизувати запити між моделями та провайдерами за capabilities, якістю, latency, вартістю, ризиком, доступністю і політикою fallback.