Перейти до основного вмісту
Просунутий6 хв1048 слівСкладність 5/5Автоматизація A2

Як Boston Children’s застосував OpenAI для повторного аналізу рідкісних хвороб

Evidence-heavy кейс Boston Children’s + OpenAI: de-identified genomic/clinical data, evidence-linked hypotheses, specialist review, confirmatory testing і clinical authority — без підміни лікаря моделлю.

Картка кейсу

Що тут автоматизовано

Складність 5/5Автоматизація A2

Обсяг автоматизації

У дослідженні Boston Children’s Hospital, Harvard University і OpenAI reasoning model використовувався для retrospective reanalysis de-identified clinical/genomic information з раніше нерозв'язаних cases та формування evidence-linked candidate hypotheses. AI-Magister відтворює це як clinician-led research workflow: data eligibility, structured case package, model-assisted hypothesis generation, evidence review, specialist adjudication, additional testing і clinical confirmation. Модель не ставить діагноз і не має clinical decision authority.

Роль людини

Qualified researchers, geneticists і clinicians визначають cohort/data eligibility, оцінюють candidate explanations, відкидають unsupported hypotheses, замовляють додаткові тести й роблять clinical confirmation. Privacy/security owners контролюють de-identification та approved environments. Пацієнтський diagnosis/treatment decision не делегується моделі.

Заявлені результати

  • 376 previously unsolved cases were reanalyzed in the published retrospective study — study cohort size
  • 18 diagnoses were established after model-surfaced leads, expert review, additional testing and clinical confirmation — study outcome
  • 4.8% additional diagnostic yield across the heterogeneous retrospective cohort — study result, not a general clinical accuracy rate
  • OpenAI separately reports 40+ rare conditions diagnosed across Boston Children's broader AI-enabled work, plus 60,000 hours saved and $7M+ redeployed labor across operational workflows; these broader hospital/provider-reported metrics are not attributed to this specific 376-case reanalysis study
  • The study did not measure time saved, cost, clinician effort, false-positive workload or change in care — explicit limitation from OpenAI's research summary

OpenAI 18 червня 2026 року описала NEJM AI study: OpenAI o3 Deep Research reanalyzed de-identified information from 376 previously unsolved cases and surfaced evidence-linked candidates; після expert review, additional testing і clinical confirmation physicians established 18 diagnoses, тобто 4.8% additional diagnostic yield in this retrospective heterogeneous sample. OpenAI прямо зазначає, що model did not diagnose participants, study не вимірювало time saved, cost, clinician effort або false-positive workload. Boston Children’s Manton Center офіційно описує той самий 376-case/18-diagnosis result. Це research evidence з explicit limitations, а не product claim про автономну діагностику.

Зміст статті
  1. 01Задача: масштабувати повторний аналіз, не перетворити reasoning model на клінічний oracle
  2. 02Trigger, input, AI stage, integrations та output
  3. 03Evidence hierarchy: hypothesis, review, test, confirmation
  4. 04Human-in-the-loop, privacy і autonomy A2
  5. 05Error handling і failure modes
  6. 06Метрики: що вимірювати, а чого не вигадувати
  7. 07Evaluation contract і staged rollout
  8. 08Кому підходить і як повторити

Передумови

Задача: масштабувати повторний аналіз, не перетворити reasoning model на клінічний oracle

Rare-disease cases можуть залишатися нерозв'язаними роками: геномні варіанти численні, phenotype неповний, а gene-disease literature змінюється швидше, ніж одна команда може системно перечитувати все вручну. Саме periodic reanalysis є дорогою когнітивною задачею, де reasoning model потенційно корисна як search-and-hypothesis amplifier.

Boston Children’s, Harvard і OpenAI перевірили це ретроспективно на 376 already analyzed unsolved cases. Критична межа кейсу: model surfaced evidence-linked hypotheses; physicians established diagnoses лише після expert adjudication, додаткових tests і clinical confirmation. Тому правильна production abstraction тут — research copilot з autonomy A2, а не autonomous diagnostician.

architecture

Карта системи: Як Boston Children’s застосував OpenAI для повторного аналізу рідкісних хвороб

Схема побудована з ключових секцій статті та показує послідовність або архітектурні блоки, які потрібно опрацювати.

Trigger, input, AI stage, integrations та output

Trigger — approved retrospective reanalysis event: new literature, updated gene-disease evidence або scheduled review of unresolved cases. Input — de-identified clinical phenotype, genomic information, inheritance/context metadata та дозволена scientific evidence base. До model call діють data-eligibility, de-identification і access controls.

AI stage synthesizes phenotype, variant/context information і literature, формуючи candidate explanations із evidence trail. Integrations у production-like research environment можуть включати case registry, variant annotations, literature retrieval і laboratory workflow, але consequential clinical systems не повинні бути writable за замовчуванням. Output — ranked/reviewable hypotheses з citations, uncertainty і missing-evidence fields.

  • Trigger → scheduled/knowledge-driven reanalysis approved by research/clinical owner.
  • Input → de-identified, task-eligible clinical/genomic packet + versioned scientific evidence.
  • AI → generate evidence-linked candidate hypotheses, не diagnosis.
  • Human → specialist adjudication → additional test where indicated → clinical/lab confirmation.
  • Output → reviewed research finding; patient-facing decision лишається clinical authority.

timeline

Контрольні точки для практичного застосування

Візуалізація використовує тези, приклади та наступні кроки статті як перевірювані контрольні точки, а не декоративні елементи.

Evidence hierarchy: hypothesis, review, test, confirmation

Для такого workflow одна з найнебезпечніших помилок — змішати levels of evidence. Model-generated hypothesis має нижчий статус, ніж specialist interpretation; specialist interpretation — нижчий, ніж validated laboratory/clinical confirmation. UI і data model повинні зберігати цей hierarchy явно, а не показувати всі outputs як однаково 'готові'.

Claim-level provenance має включати source publication, access date/version, genomic annotation version, model/reasoning configuration і reviewer decision. Якщо literature змінилася, попередній hypothesis не стає автоматично false, але його evidence snapshot уже historical. Reanalysis має бути version-aware.

Human-in-the-loop, privacy і autonomy A2

A2 означає, що AI може автоматизувати підготовку й аналіз, але decision loop залишається expert-led. OpenAI прямо зазначає: модель не діагностувала participants і не приймала clinical decisions. Кожен позитивний lead проходив established review/testing/confirmation processes. Це не косметичний disclaimer, а основна архітектура безпеки.

Study використовував de-identified information і не передавав protected health information поза approved environments. Для реального deployment потрібні purpose limitation, least-privilege access, auditability, retention/deletion rules і локальна regulatory review. Навіть дуже сильний model result не дає права розширити data scope постфактум.

Error handling і failure modes

Failure modes: phenotype incomplete; variant annotation stale; literature contradicts candidate; model overweights a plausible paper; citation не підтримує claim; duplicate patient/case identity; data-quality artifact; hypothesis medically coherent but laboratory test negative. System повинен зберігати rejected/uncertain state, а не тиснути reviewer до binary answer.

No-answer є валідним outcome. Якщо evidence недостатньо, модель повертає unresolved gaps і possible next research questions, а не вигадує closure. Negative findings і false leads мають потрапляти в evaluation corpus, інакше система вчитиметься лише на красивих success stories.

Метрики: що вимірювати, а чого не вигадувати

Study result 18/376, або 4.8% additional diagnostic yield, не є '95.2% error rate' і не є general diagnostic accuracy. Cohorts heterogeneous; early psychosis subgroup був малий; reviewers не були blinded to model confidence. OpenAI прямо вказує, що time saved, cost, clinician effort і false-positive workload не вимірювались.

Production KPI мають включати hypothesis recall on adjudicated historical cases, unsupported-citation rate, reviewer workload, time to adjudication, additional-test burden, privacy incidents, clinically confirmed yield і harmful false-reassurance rate. Окремо вимірюється abstention quality. Будь-яка ROI-модель без reviewer/test cost тут буде бухгалтерською фантастикою.

Evaluation contract і staged rollout

Eval set: solved historical cases з hidden outcome, truly unresolved cases, phenotype noise, outdated literature, conflicting papers, ancestry/data-bias slices, missing inheritance info, citation traps і adversarial text inside imported documents. Deterministic checks validate structured fields/citations/data eligibility; specialist graders оцінюють plausibility, evidence support, harmful omission і next-step appropriateness.

Rollout: `retrospective offline study → shadow reanalysis → research-only candidate generation → specialist-reviewed workflow → narrowly governed prospective assistance`. Ніякого automatic patient communication або treatment recommendation. Model upgrade запускає full regression on locked cases; high-severity miss або privacy breach блокує promotion.

Кому підходить і як повторити

Патерн релевантний academic medical centers, genomic research programs і specialist diagnostic teams із curated case registry, scientific-literature access і established confirmation processes. Він не підходить consumer symptom checker, який намагається напряму перенести research model у patient-facing diagnosis.

Pilot починається зі retrospective de-identified dataset і ethics/privacy approval. Зафіксуйте baseline specialist process, lock outcomes, побудуйте provenance-preserving case packet, проганяйте model у shadow і оцінюйте hypotheses blind там, де можливо. Лише після доказу incremental utility та прийнятного reviewer burden переходьте до prospective research assistance.

Практичні приклади

Приклад: unsolved genomic case після появи нового gene-disease evidence

Case registry створює approved reanalysis snapshot без direct identifiers. Model формує evidence-linked candidate і вказує missing confirmation. Geneticist перевіряє source, phenotype fit та inheritance, після чого вирішує, чи потрібен додатковий lab test. Тільки validated clinical process може змінити diagnosis state; model output сам по собі цього не робить.

FAQ

Чи OpenAI model сам поставив 18 діагнозів?

Ні. Model surfaced candidate explanations; diagnoses were established by physicians after expert review, additional testing and clinical confirmation.

Що означає 4.8%?

Additional diagnostic yield у retrospective heterogeneous cohort 376 previously unsolved cases. Це не загальна accuracy моделі й не очікуваний результат для будь-якої клініки.

Чи дослідження довело економію часу або грошей?

Ні. OpenAI прямо зазначає, що study не вимірювало time saved, cost, clinician effort або false-positive workload.

Яка правильна autonomy?

A2: model-assisted research and hypothesis generation з mandatory specialist adjudication та clinical/laboratory confirmation; diagnosis/treatment authority залишається у qualified clinicians.

Пов’язані матеріали

Як Samsung масштабує ChatGPT Enterprise і Codex на глобальну організацію

Production-кейс Samsung Electronics + OpenAI: глобальний rollout ChatGPT Enterprise і Codex для R&D, manufacturing, marketing та corporate work — із multi-model governance, permission boundaries і перевіркою фактичних результатів.

Як Cisco зробив Codex частиною enterprise engineering

Production-кейс Cisco + OpenAI: Codex у великих codebases, defect remediation, framework migrations і product engineering — із plan artifacts, CI/security gates, human merge authority та measured throughput.

Як HSP GRUPPE масштабує ChatGPT Enterprise у податковому консалтингу

Production-кейс HSP GRUPPE: ChatGPT Enterprise працює в tax advisory, legal research, client communication, financial analysis і knowledge sharing, а agentic workflows проходять окремий pilot та governance gate.

Як Australian Payments Plus використовує ChatGPT і Codex у критичній платіжній інфраструктурі

Production-кейс AP+: ChatGPT Enterprise допомагає працювати зі складними rules/specifications, а Codex прискорює technical investigation і functional simulations із human accountability у regulated payments environment.

Як BBVA масштабує ChatGPT Enterprise у банку: 100 000 користувачів, governance і шлях до AI-native banking

Production-кейс BBVA: ChatGPT Enterprise як керований enterprise layer для knowledge work, custom GPTs і банківських AI-сценаріїв із security, legal, compliance та human-controlled authority.

Red teaming LLM-систем

Практичний red teaming перетворює припущення про безпеку LLM-системи на відтворювані атаки, докази та regression-тести. Розглядаємо threat model, ручні й автоматизовані кампанії, triage, безпечну лабораторію та перевірку виправлень.

State machines для агентів

State machines для агентів — практичний розбір production-архітектури: відокремлення ймовірнісного рішення моделі від детермінованого життєвого циклу виконання. Матеріал охоплює контракти, межі повноважень, failure modes, оцінювання та контрольований rollout.

Вибір моделей і model routing

Як маршрутизувати запити між моделями та провайдерами за capabilities, якістю, latency, вартістю, ризиком, доступністю і політикою fallback.

Джерела

  1. Using AI to help physicians diagnose rare genetic diseases affecting childrenофіційне
  2. Manton Center for Orphan Disease Research: Leveraging AI in rare disease researchпервинне
  3. Boston Children’s uses AI to unlock new diagnosesофіційне