Як CyberAgent масштабує ChatGPT Enterprise і Codex
Production-кейс CyberAgent: enterprise AI operating model для research, drafting, design review, code review та Codex execution — з data governance, evals, HITL і cost-per-verified-outcome.
Картка кейсу
Що тут автоматизовано
Обсяг автоматизації
ChatGPT Enterprise є керованим enterprise surface для knowledge work, а Codex — agentic engineering layer для design, review, planning та bounded implementation. Організаційно це A3: окремі engineering subflows можуть доходити до A4 у sandbox/branch, але confidential-data exceptions, merge, release та consequential business decisions залишаються зовнішньою authority.
Роль людини
Employees формують task intent і перевіряють output; architects та reviewers володіють design і merge decisions; security/IT керують data classes, access, logs, connectors і retention; platform owners вимірюють outcome, cost та incidents.
Заявлені результати
- OpenAI reports 93% monthly active usage of ChatGPT Enterprise across CyberAgent; this is provider/customer-reported adoption, not an independent productivity or quality benchmark
- OpenAI reports more than 100 employees attended each of over ten training sessions; this is enablement scale, not evidence of task success
- OpenAI describes a GOODROID game reaching soft launch after about one month with Codex involved; this is a specific team example, not a universal Codex speedup
OpenAI 9 квітня 2026 року описала ChatGPT Enterprise та Codex як центральні компоненти AI operating model CyberAgent і повідомила про 93% monthly active usage. CyberAgent-owned materials окремо описують company-wide IT governance, usage/cost visibility та значний масштаб використання OpenAI Codex поряд з іншими AI tools.
Зміст статті
- 01Бізнес-задача: AI як керований operating layer, а не набір приватних prompt-ів
- 02Trigger, input, AI stage, integrations та output
- 03Automation workflow і human-in-the-loop
- 04Governance: централізуємо policy та observability, а не кожен prompt
- 05Error handling і controls
- 06Evaluation contract: adoption не дорівнює verified outcome
- 07Frequency, scalability та cost model
- 08Як повторити: 7-кроковий rollout
Бізнес-задача: AI як керований operating layer, а не набір приватних prompt-ів
CyberAgent працює одночасно в рекламі, медіа, IP та game development, тому задача не зводиться до видачі coding assistant розробникам. Потрібне середовище, де різні ролі можуть безпечно використовувати ChatGPT Enterprise для research, drafting і структурування роботи, а engineering-команди — Codex для design review, implementation planning, code review та документації.
Production-відтворення починається не з seats. Спочатку фіксуються task classes, data classes, authority та acceptance: які дані можуть потрапляти в model context, які outputs залишаються draft, які дії вимагають deterministic validation і хто має право затвердити consequential outcome. Інакше 93% adoption перетворюється на 93% масштабованої невизначеності.
architecture
Карта системи: Як CyberAgent масштабує ChatGPT Enterprise і Codex
Trigger, input, AI stage, integrations та output
Trigger-и залежать від ролі: business user запускає research або drafting task; product/engineering team — design review, specification чи implementation; code review — diff або pull request; enablement team — оновлення knowledge artifact чи internal usage tooling. Input включає approved enterprise context, design docs, repository state, requirements, review comments та актуальні guidelines.
AI stage розділяється на reasoning та execution. ChatGPT Enterprise допомагає аналізувати, порівнювати варіанти та готувати artifacts; Codex може дослідити codebase, створити plan, змінити код і запустити перевірки в дозволеному середовищі. Output — versioned draft, design packet, review evidence або branch/PR, а не прихований direct write у production.
- Trigger → research, design review, specification, code review або bounded implementation task.
- Input → approved context + repo/design artifacts + policy revision + acceptance criteria.
- AI → analyze, pressure-test, draft, implement, verify.
- Output → reviewable artifact із provenance та explicit terminal state.
decision-tree
Контрольні точки для практичного застосування
Контрольна теза з матеріалу статті.
Контрольна теза з матеріалу статті.
Контрольна теза з матеріалу статті.
Контрольна теза з матеріалу статті.
Automation workflow і human-in-the-loop
Engineering flow: `task contract → context collection → Codex analysis → options/design rationale → implementation plan → isolated execution → tests/scans → reviewer decision → merge`. Knowledge-work flow: `question → permitted sources → draft → factual/policy checks → owner review → publish/use`. Для окремого coding task це може бути A4, але весь organizational case лишається A3, бо значна частина роботи advisory або review-gated.
HITL не повинен бути декоративним. Reviewer бачить intent, touched systems, assumptions, test results, policy verdict і unresolved risks. Merge, production release, legal/commercial commitments, security exceptions та використання restricted data не успадковують authority від моделі, навіть якщо tool technically доступний.
Governance: централізуємо policy та observability, а не кожен prompt
OpenAI описує enterprise account management, usage visibility та internal guidelines для confidential information. CyberAgent не робить blanket mandate для всіх teams; adoption підтримується training, workshops, knowledge sharing і visibility usage. Це сильніший operating pattern, ніж примусова метрика кількості prompt-ів.
Для повторення потрібні SSO/SCIM, role-based access, approved connectors, retention contract, logging policy, data-classification gate, incident path і owner для кожного workflow. Usage dashboard не є business-value dashboard: окремо вимірюються accepted outcomes, correction rate, quality deltas та cost per verified task.
Error handling і controls
Критичні failure modes: restricted context пішов у недозволений workflow; design proposal суперечить architecture constraint; agent працює на stale base; code-review suggestion послаблює security; knowledge file містить outdated instruction; implementation змінює tests так, щоб приховати regression.
Контролі: data-class gate до model context, instruction/release fingerprint, stale-base check, protected tests і CI config, dependency/secret scanning, deterministic acceptance commands, rollback і explicit clarification state. Якщо external action має unknown result після timeout, спочатку authoritative reconciliation, а не blind retry.
Evaluation contract: adoption не дорівнює verified outcome
93% monthly active usage — корисний сигнал adoption, але не quality benchmark. Eval corpus розділяється на research/drafting, design review, code review, implementation та non-developer specification tasks. Для кожного slice потрібні baseline, expected evidence, hard blockers і human correction rubric.
Engineering graders оцінюють correctness, architecture compliance, test quality, security, unnecessary churn і final system state. Knowledge-work graders — factual support, source coverage, policy compliance та usefulness. Unauthorized write, secret exposure, disabled control або fabricated evidence — hard blocker незалежно від aggregate score.
Frequency, scalability та cost model
Коли AI використовується майже в усіх departments, scaling bottleneck — не лише tokens. Це identity lifecycle, connector permissions, support, training, review capacity, CI compute, telemetry, model-routing policy та platform operations. Engineering concurrency лімітується за repo/service і blast radius, щоб паралельні agents не створювали взаємно конфліктні зміни.
Full cost = enterprise seats + model/agent usage + sandbox/CI + connectors + observability + enablement + human review + failed runs + incident handling. Практичні denominator-и — cost per accepted artifact, per verified merged change або per completed business task, а не ціна одного запиту.
Як повторити: 7-кроковий rollout
1) Визначити 3–5 task classes із різним risk. 2) Зафіксувати data/authority matrix. 3) Запустити read/draft workflows. 4) Для Codex додати isolated workspace і deterministic verification. 5) Побудувати role-specific evals. 6) Додати usage + outcome telemetry. 7) Розширювати autonomy лише після stable regression history і rollback drill.
Кейс підходить multi-business organizations, де AI вже використовується неформально і потрібен керований operating layer. Не підходить компанії, яка хоче копіювати 93% adoption як KPI без quality, security та outcome gates.
Практичні приклади
Design proposal → Codex implementation → verified PR
Product engineer дає design context і acceptance criteria; Codex pressure-tests proposal, готує plan та implementation у branch; deterministic tests і security checks формують evidence; reviewer має merge authority. Usage telemetry зберігається окремо від outcome telemetry.
FAQ
Чи 93% usage означає 93% productivity uplift?
Ні. Це OpenAI-reported monthly active usage. Business outcome, quality та safety треба вимірювати task-level evals.
Який рівень автономності доречний?
Організаційно A3; bounded coding tasks можуть бути A4 у sandbox/branch, але merge, release та high-impact decisions лишаються зовнішньою authority.
З чого почати повторення?
З data/authority matrix і 3–5 task classes, потім read/draft workflows, isolated Codex execution, role-specific evals і лише після цього ширший rollout.
Пов’язані матеріали
Production-кейс TRUSTBANK + Recursive: multi-agent recommendation system для каталогу приблизно 760k gifts — routing, RAG, personalization, model routing, evals і bounded transaction authority.
Як Notion використовує Codex для one-shot engineeringProduction-кейс Notion: spec + reference implementation + verification harness → autonomous Codex run → tested PR — із A4 autonomy, parallel work, failure gates і cost per accepted change.
Як Endava перебудовує software delivery навколо ChatGPT і CodexProduction-кейс Endava: Codex і ChatGPT Enterprise від requirements та architecture до build, client collaboration і operations — з encoded senior expertise, bounded autonomy, deterministic verification та human release authority.
Як Wayfair масштабує catalog quality і supplier support з OpenAIProduction-кейс Wayfair + OpenAI: reusable catalog classification, Wilma agentic support, confidence-based autonomy, tool use, human validation, reconciliation, evals і cost controls.
Observability для LLM-системЯкі traces, metrics, logs і evaluation signals потрібні для LLM: prompts, retrieval, tool calls, usage, quality, privacy, cardinality і розслідування інцидентів.
Оцінювання AI-агентівОцінювання AI-агентів — практичний розбір production-архітектури: вимірювання не лише фінальної відповіді, а всієї траєкторії рішень, дій, витрат і безпечного завершення. Матеріал охоплює контракти, межі повноважень, failure modes, оцінювання та контрольований rollout.