AI Product Operations:运营节奏与结果复盘架构
AI Product Operations 是上线后的证据系统: 它把运行时行为、采用、结果、风险、成本、事故和变更发布转成可重复的产品与组合决策。
AI Product Operations / Operating Cadence / Outcome Review Architecture 解读
配对阅读:本篇的操作手册版(模板/RACI/门禁/runbook)是
docs/AI_PRODUCT_OPERATIONS_OPERATING_CADENCE_OUTCOME_REVIEW_PLAYBOOK.md。第一遍读本篇建立原理与架构判断;第二遍做案例时再用 playbook 查表落地,两者不需要重复精读。
核心问题: AI 产品上线后, 如何用 weekly ops review、monthly value review、quarterly portfolio review 和 release / experiment / incident loops, 把真实运营证据转化为 scale、restrict、redesign、retire 和投资决策?
重要说明: 本文是学习和内部架构训练材料, 不构成法律意见、监管解释、合规确认、审计意见、模型验证结论、风险接受决定、财务投资建议或生产上线批准。正式项目中的审批权、残余风险接受、监管沟通、审计依赖、客户影响判断和发布授权必须由机构授权角色结合司法辖区、产品、客户群、风险偏好、内部政策、模型风险、信息安全、隐私、供应商合同和运营能力确认。访问日期按 2026-06-30 记录。
Source Anchors
以下来源用于组织 AI 风险管理、AI 管理体系、需求工程、工程绩效、可观测性和服务可靠性语言。本文只将这些来源作为产品、架构和内部 product operations 的设计锚点, 不声称任何运营节奏、dashboard 或 review pack 自动构成法律、监管、审计或模型验证批准。
| Source | Link | 本文采用的思想 |
|---|---|---|
| NIST AI Risk Management Framework | https://www.nist.gov/itl/ai-risk-management-framework | 用 Govern / Map / Measure / Manage 组织 post-launch risk review、monitoring、incident learning 和 action closure。(AI RMF 1.0 发布 2023-01;访问日期: 2026-07-01) |
| ISO/IEC 42001 AI management system | https://www.iso.org/standard/81230.html | 用 management system 的 policy、objectives、operation、performance evaluation、internal audit、improvement 语言定义 AI Product Ops。(访问日期: 2026-07-01) |
| ISO/IEC/IEEE 29148 Requirements Engineering | https://www.iso.org/standard/72089.html | 用 requirements quality、stakeholder needs、verification、traceability 思路设计 metric contract、assumption ledger 和 decision log。(访问日期: 2026-07-01) |
| DORA | https://dora.dev/ | 用 software delivery performance 和 reliability mindset 连接 release cadence、change fail rate、restore time 和 learning loop。(最新年度报告为《State of AI-assisted Software Development 2025》,见文末 SOTA 检查;访问日期: 2026-07-01) |
| OpenTelemetry Documentation | https://opentelemetry.io/docs/ | 用 traces、metrics、logs 和 semantic conventions 的思想设计 AI product operations telemetry。(GenAI semconv 状态见文末 SOTA 检查;访问日期: 2026-07-01) |
| Google SRE: Service Level Objectives | https://sre.google/sre-book/service-level-objectives/ | 用 SLO、error budget 和 service reliability 语言定义 AI product operational thresholds。(访问日期: 2026-07-01) |
核心导读
AI Product Operations 是上线后的证据系统: 它把运行时行为、采用、结果、风险、成本、事故和变更发布转成可重复的产品与组合决策。
post-launch telemetry
-> evidence review pack
-> cadence-specific decisions
-> backlog and release calendar
-> action closure
-> outcome and risk learning
-> portfolio allocation
上线前的核心问题是“能不能做”。上线后的核心问题变成“是否仍然值得运行、如何运行得更好、何时扩大、何时限制、何时重构、何时退役”。AI Product Ops 的价值就在于让这个判断不依赖主观热情, 而依赖 metric contract、evidence pack、decision log、assumption ledger、release calendar、incident learning 和 action closure。
1. 问题定义: AI 上线后的运营断裂
很多 AI 产品的失败发生在上线之后:
- Pilot 证明了模型能回答问题, 但上线后 adoption 只停留在少数 champion。
- Usage 很高, 但流程周期、质量、投诉或风险控制没有改善。
- Prompt、知识库、模型、tool permission 和 policy pack 持续变更, 但没有统一 release calendar。
- 事故复盘只产生修复 ticket, 没有进入 roadmap、metric contract、training、policy 或 control design。
- 成本增长被解释为“用户增长”, 但没有 case-level unit economics 和 capacity review。
- 管理层每月看到 dashboard, 但没有 decision log、assumption ledger 和 action closure。
AI Product Operations 的目标不是多开几个会议。它是一套运营架构, 把七类证据放进同一个节奏:
| Evidence lane | 核心问题 | 常见证据 |
|---|---|---|
| Outcome | 是否改善目标业务结果 | cycle time、first-pass yield、AHT、loss avoided、conversion、complaint rate |
| Adoption | 目标角色是否在正确工作步骤采用 | qualified use、accept/edit/reject、cohort durability、manager reinforcement |
| Quality | 输出质量是否稳定并适配 case mix | eval pass rate、QA defects、hallucination class、retrieval freshness |
| Risk / Control | 风险是否仍在 appetite 内 | override、escalation、policy breach、customer harm、audit finding |
| Cost / Capacity | 单位经济是否成立 | cost per case、token/tool cost、review load、queue aging、support effort |
| Incident Learning | 失败是否被转成系统改进 | incident taxonomy、root cause、corrective action、recurrence signal |
| Roadmap | 证据是否改变投资和优先级 | decision log、assumption ledger、experiment result、release calendar |
本文聚焦 post-launch product operations cadence and outcome evidence。它不重复团队授权、product trio 或 decision rights 的基础模型, 而是假设 AI capability 已经完成 controlled pilot 或初步上线, 接下来要建立持续运营节奏。
| Dimension | AI Product Operating Model | AI Product Operations Cadence |
|---|---|---|
| 关注点 | 团队授权、decision rights、guardrails | post-launch review、evidence、action closure、roadmap decision |
| 时间位置 | discovery 到 launch 前后 | controlled pilot、production、scale、refresh、retire |
| 主要问题 | 团队能否在 guardrails 内解决问题 | 产品是否仍创造价值且风险可控 |
| 核心对象 | team、decision rights、gates | metric contract、evidence pack、operating calendar |
| 成功标志 | 能发现、交付、治理 AI capability | 能持续证明、调整、扩展、限制或退役 capability |
2. 架构模型: AI Product Ops 作为结果证据运行系统
AI Product Ops 的最小闭环:
Observe
-> interpret
-> decide
-> act
-> verify closure
-> update assumptions and roadmap
如果缺少任一环, cadence 就会退化。
| Missing piece | 退化表现 | 结果 |
|---|---|---|
| Observe | 只有主观反馈, 没有 trace / metric / sample | 无法区分真实风险和噪声 |
| Interpret | dashboard 多, 但没有 root cause language | 数字变化不产生决策 |
| Decide | 会议讨论多, 没有 decision log | 同一争议反复出现 |
| Act | action 没 owner / due date / evidence | 会议变成汇报仪式 |
| Verify closure | ticket closed, 但 outcome 未复核 | 修复不等于问题解决 |
| Update roadmap | 事故和学习不改变优先级 | 产品继续按旧假设投资 |
2.1 Operating Model Components
| Component | Purpose | Accountability question |
|---|---|---|
| Operating calendar | 定义 weekly、monthly、quarterly、release、incident、experiment review 节奏 | 哪些节奏负责哪些决策, 哪些信号必须升级 |
| Metric contract | 定义指标口径、owner、阈值、数据源、行动规则 | 这个指标是否能支撑明确行动, 而不是只做展示 |
| Evidence review pack | 把 adoption、outcome、quality、risk、cost、incident、roadmap 整合为决策材料 | 一页证据是否足以解释建议的 scale / restrict / redesign / retire |
| Decision log | 记录 scale、restrict、release、rollback、policy、roadmap 决策及依据 | 未来是否能重建当时为什么这样决策 |
| Assumption ledger | 记录价值、行为、风险、成本和 capacity 假设是否仍成立 | 哪些旧假设已经过期, 哪些仍可用于投资判断 |
| Experiment registry | 记录实验目的、population、hypothesis、metric、risk guardrail、结果 | 实验是否留下可复用学习, 而不是只留下局部胜负 |
| Release calendar | 管理 model、prompt、data、knowledge、tool、policy、UX、workflow 变更 | AI 行为变化是否可追踪、可回滚、可解释 |
| Incident learning loop | 将 incident / complaint / near miss 转为 corrective action 和 roadmap item | 失败是否买到了系统学习 |
| Action closure register | 跟踪 action owner、due date、closure evidence、reopen trigger | 行动是否真的改变指标、流程或控制状态 |
| Portfolio review pack | 支撑 fund / scale / pause / retire / consolidate 决策 | 资源是否流向净价值最高且风险可控的能力 |
2.2 Control Planes
AI Product Ops 至少覆盖九个控制面:
| Plane | Review question |
|---|---|
| Value | 业务结果是否移动, benefit 是否净实现 |
| Adoption | 目标用户是否持续正确采用 |
| Quality | 输出质量和 workflow fit 是否稳定 |
| Reliability | latency、availability、restore、fallback 是否达标 |
| Risk | customer harm、model risk、policy breach、over-reliance 是否受控 |
| Cost | unit cost、support load、review load、capacity 是否可承受 |
| Change | model/prompt/data/tool/policy 变更是否可追溯 |
| Incident | 失败是否被学习、关闭并防止复发 |
| Roadmap | 新证据是否改变投资方向 |
2.3 Product Ops Data Objects
| Object | Key fields | Why it matters |
|---|---|---|
| Metric contract | metric_id、definition、owner、source、threshold、action | 防止每次 review 重新争论指标口径 |
| Review pack | period、population、evidence、decision request、actions | 让会议从汇报转成决策 |
| Release item | object_type、version、change reason、risk tier、rollback | 把 AI 变更纳入可追溯 calendar |
| Experiment record | hypothesis、cohort、guardrails、duration、result、decision | 防止实验结果丢失或被选择性引用 |
| Incident record | severity、impact、root cause、affected versions、corrective action | 让事故进入 learning loop |
| Assumption | statement、evidence、confidence、expiration、owner | 管理价值和风险叙事的有效期 |
| Action closure | action、owner、due date、evidence、reviewer、reopen trigger | 防止会议行动消失 |
3. 关键机制与生命周期: Cadence Architecture
AI 产品上线后的 cadence 不是单一会议, 而是不同时间尺度的证据处理栈。
Daily signal triage
-> weekly ops review
-> monthly value review
-> quarterly portfolio review
-> annual / semiannual management system review
| Cadence | Primary lens | Typical decisions |
|---|---|---|
| Daily signal triage | incidents、latency、availability、complaint spikes、cost anomaly | mitigate、rollback、escalate、sample、hotfix |
| Weekly ops review | adoption、quality、reliability、capacity、open actions | prioritize fixes、adjust release、assign owners |
| Monthly value review | outcome、unit economics、benefit leakage、risk trend | scale、restrict、redesign、update business case |
| Quarterly portfolio review | use-case portfolio、platform reuse、risk concentration、funding | fund、pause、consolidate、retire、reallocate capacity |
| Management system review | policy effectiveness、audit findings、objectives、continual improvement | update operating policy、control library、governance model |
3.1 Weekly Ops Review
Weekly ops review 是 tactical learning forum, 不应退化为 status meeting。
| Input | Review question | Output |
|---|---|---|
| Adoption funnel by cohort | 哪些用户、case type、manager group 掉队 | enablement action、product fix、workflow change |
| Quality sample and eval result | 哪类 failure 正在上升 | prompt/index/model/tool fix、sampling change |
| Reliability and SLO | latency、availability、restore 是否影响工作 | platform action、fallback adjustment |
| Cost / capacity | review queue、token/tool cost、support load 是否异常 | capacity rebalance、cost guardrail |
| Incident / complaint signals | 是否存在 customer harm 或 policy drift | incident triage、risk escalation |
| Open action register | 上周行动是否关闭, closure evidence 是否充分 | close、reopen、escalate |
Weekly outputs 必须可执行:
- action owner。
- due date。
- closure evidence。
- decision log entry。
- backlog item 或 release calendar update。
- escalation path。
3.2 Monthly Value Review
Monthly value review 是 outcome and investment forum。它回答“这个 AI capability 是否仍然值得继续投资”。
| Review block | Evidence |
|---|---|
| Outcome movement | baseline vs current、cohort trend、seasonality adjustment |
| Adoption durability | returning qualified use、manager reinforcement、work-as-done evidence |
| Value leakage | human review load、rework、support cost、exception queue、customer redress |
| Risk trend | complaints、overrides、policy breaches、fairness / conduct signals |
| Cost-to-serve | unit cost per case、marginal cost、platform capacity |
| Release impact | recent model/prompt/data/tool releases and outcome changes |
| Decision request | scale、hold、restrict、redesign、retire、continue experiment |
Monthly review 的关键是把数据转成明确决策:
Continue because evidence is improving and risk is stable.
Scale because outcome lift is durable and marginal cost is acceptable.
Restrict because specific cohorts or case types show harm or poor reliability.
Redesign because usage is high but value leakage removes benefit.
Retire because assumptions failed and no credible path remains.
3.3 Quarterly Portfolio Review
Quarterly portfolio review 把单个 use case 上升到 enterprise AI allocation。
| Portfolio lens | Questions |
|---|---|
| Value concentration | 哪些 use cases 贡献主要净收益, 哪些只有 activity |
| Risk concentration | 是否在同一 customer segment、model provider、data source 或 control weakness 上集中 |
| Platform leverage | 哪些 capabilities 应产品化复用, 哪些 bespoke build 应合并 |
| Talent / capacity | SME review、risk review、data engineering、platform support 是否成为瓶颈 |
| Policy drift | 业务规则、监管解释、模型能力、供应商条款是否改变 |
| Roadmap reallocation | 哪些主题应加速, 哪些应暂停、合并或退役 |
Quarterly review 的输出不是“下季度计划”, 而是 funding change、platform investment decision、risk appetite adjustment、capability retirement decision、governance process improvement 和 portfolio-level assumption update。
4. Outcome Review and Metric Contract
Outcome review 不是看一个 North Star metric, 而是看 outcome chain。
AI release / experiment
-> target exposure
-> qualified adoption
-> workflow behavior change
-> quality and control movement
-> business outcome
-> net value after risk and cost
| Layer | Evidence | Failure interpretation |
|---|---|---|
| Release | version changed、target cohort exposed | release did not reach workflow |
| Adoption | accepted / edited / rejected / escalated | users do not trust、do not need、or cannot use |
| Behavior | artifact、decision、handoff changed | AI used as side tool but not embedded |
| Quality | first-pass yield、QA、eval、defect class | output not fit for production work |
| Control | override、escalation、policy breach、complaint | value is creating hidden risk |
| Outcome | cycle、conversion、AHT、loss、STP | business result did not move |
| Net value | benefit minus operating、review、support、risk cost | gross benefit does not survive operations |
Outcome review 要避免两个陷阱: 把 release 当成结果, 以及把单点结果改善当成可持续价值。
| Example | Outcome claim | Required counter-evidence |
|---|---|---|
| Contact-center agent assist | AHT 下降 | repeat contact、complaints、script compliance、hold transfer |
| Complaint intelligence | root cause identification 更快 | misclassification、regulatory breach、remediation delay |
| KYC onboarding | cycle time 下降 | false pass、rework、document chase、vulnerable customer impact |
| Collections hardship | arrangement completion 上升 | unfair pressure、complaints、broken promises、agent override |
| AML triage | alert closure 更快 | suspicious activity miss、escalation quality、audit sampling |
| Personalized pricing | margin / conversion uplift | unfair treatment、explainability、opt-out、complaint trend |
4.1 Metric Contract
AI product review 经常争论指标为什么变了、dashboard 和 finance number 为什么不一致、贡献来自 AI 还是 seasonality、usage 是否等于 adoption、成本上升是 scale 信号还是经济性恶化。Metric contract 是对指标的产品需求说明和治理文件。
| Field | Description |
|---|---|
| metric_id | Stable identifier, such as kyc_ai.first_pass_yield |
| business question | 这个指标要回答什么决策问题 |
| definition | 精确定义, 包括 numerator / denominator |
| population | 用户、case type、channel、risk tier、time window |
| source system | telemetry、workflow system、finance ledger、QA、complaint platform |
| owner | 对口径和解释负责的人 |
| review cadence | daily、weekly、monthly、quarterly |
| threshold | target、warning、breach、stop rule |
| guardrail | 防止局部优化伤害其他目标 |
| segmentation | 必须按哪些 cohort 拆解 |
| action rule | 指标越界时触发什么行动 |
| evidence quality | observed、sampled、inferred、survey、finance-certified |
| expiry / review date | 何时重新审查口径是否仍适用 |
4.2 Metric Taxonomy and Governance
| Metric type | Example | Cadence |
|---|---|---|
| Outcome | complaint cycle time、AHT、KYC approval cycle、AML aging | monthly |
| Adoption | qualified use、acceptance、edit rate、rejection reason | weekly |
| Quality | eval pass、QA defect、hallucination class、retrieval hit quality | weekly |
| Reliability | SLO、latency、availability、fallback success、restore time | daily / weekly |
| Risk | policy breach、override、escalation、customer harm、fairness signal | weekly / monthly |
| Cost | cost per case、token/tool cost、support effort、review load | weekly / monthly |
| Learning | experiment velocity、action closure、incident recurrence | weekly / monthly |
| Portfolio | net value、risk exposure、platform reuse、retirement rate | quarterly |
Metric governance is product governance:
- 每个 metric 有 owner, 没有 owner 的 metric 不进入 executive review。
- 每个 metric 有 action rule, 没有 action rule 的 metric 只是观察值。
- 每个 metric 有 segmentation, 否则会隐藏 vulnerable cohort。
- 每个 metric 有 validity period, 因为流程、模型、政策和用户行为会漂移。
- 每个 metric 有 evidence quality rating, 区分 telemetry、sampling、survey 和 finance-certified value。
5. 证据与控制: Review Pack、Traceability、Release Calendar
Evidence review pack 是每次 review 的共同材料。它不追求信息多, 追求能产生决策。
5.1 Review Pack Structure
| Section | Content |
|---|---|
| Decision requested | continue、scale、restrict、redesign、retire、release、rollback |
| Scope and version | product area、population、model/prompt/data/tool version |
| Outcome summary | baseline、current、movement、confidence |
| Adoption summary | cohort funnel、qualified use、durability |
| Quality summary | eval、QA sample、failure taxonomy |
| Risk/control summary | incidents、complaints、overrides、policy drift |
| Cost/capacity summary | unit cost、review load、support load |
| Release and experiment summary | recent changes、experiments、observed effects |
| Open assumptions | assumptions confirmed、weakened、invalidated |
| Action closure | last actions、evidence、unresolved blockers |
| Recommendation | specific decision and next review trigger |
5.2 Evidence Quality and Traceability
| Level | Description | Review use |
|---|---|---|
| E1 Anecdotal | isolated feedback or demo observation | signal only |
| E2 Sampled | QA sample、complaint sample、interview sample | weekly interpretation |
| E3 Instrumented | production telemetry joined to workflow context | weekly / monthly decision |
| E4 Causal or quasi-causal | controlled experiment、matched cohort、difference analysis | scale / restrict decision |
| E5 Finance / risk certified | reconciled benefit、validated risk and audit-ready evidence | portfolio investment decision |
Post-launch evidence should trace:
metric -> source event -> workflow context -> version -> decision -> action -> closure evidence -> next metric movement
OpenTelemetry-inspired traces 与 ISO/IEC/IEEE 29148-inspired traceability 在这里相遇。重点不是技术优雅, 而是能解释 roadmap 为什么改变。
5.3 Experiment and Release Calendar
AI products change through more than code deploys.
| Change object | Example | Risk |
|---|---|---|
| Model | provider upgrade、model class change、fallback model | quality shift、cost shift、latency、data boundary |
| Prompt | system prompt、tool instruction、refusal wording | behavior shift、policy drift、regression |
| Data | feature change、label change、training data refresh | bias、leakage、stale assumptions |
| Knowledge | RAG corpus、policy document、product catalog | outdated guidance、retrieval mismatch |
| Tool | CRM write action、fee waiver API、case closure action | side effect、authorization、audit |
| Policy | hardship treatment rule、complaint taxonomy、KYC requirement | compliance breach、inconsistent handling |
| Workflow | UI step、queue routing、human review threshold | adoption change、capacity shift |
| Experiment | cohort change、A/B treatment、canary | interpretation error、customer impact |
Release calendar fields:
| Field | Description |
|---|---|
| release_id | Stable release identifier |
| object_type | model、prompt、data、knowledge、tool、policy、workflow |
| affected population | cohort、channel、case type、geography |
| evidence required | eval、QA、risk、cost、regression、rollout plan |
| canary plan | first users、duration、guardrails |
| rollback path | technical and operational rollback |
| communication | frontline、risk、support、manager notes |
| review date | when impact is reviewed |
| decision log link | why release was approved |
A release without impact review is an uncontrolled change. An experiment without release trace is an unrepeatable learning.
6. Incident-to-Roadmap and Backlog Governance
AI incident management becomes product strategy when failures reveal weak assumptions.
6.1 Incident Sources
| Source | Example |
|---|---|
| Customer complaint | customer claims AI-generated explanation was misleading |
| Frontline override | agent repeatedly rejects a suggested hardship script |
| QA defect | complaint classifier misses regulatory complaint |
| Model drift | AML triage quality drops for a new fraud typology |
| Cost anomaly | tool calls spike after prompt change |
| Policy drift | knowledge base uses outdated pricing exception rule |
| Near miss | human reviewer catches a high-impact hallucination |
| External change | regulation、product terms、vendor model behavior changes |
6.2 Learning Loop
detect signal
-> classify severity and affected population
-> contain or rollback
-> root cause across model / prompt / data / tool / workflow / policy / training
-> corrective action
-> metric contract update
-> backlog / roadmap update
-> action closure evidence
-> recurrence review
6.3 Root Cause Taxonomy
| Cause class | Example action |
|---|---|
| Model behavior | change model、add eval、adjust fallback |
| Prompt instruction | revise prompt、add regression case、review release path |
| Knowledge freshness | update corpus、add freshness SLO、assign knowledge owner |
| Tool permission | restrict tool、add approval、update authorization |
| Workflow design | change handoff、add human review、revise UI |
| Training / adoption | manager coaching、SOP update、new refusal guidance |
| Metric design | add missing guardrail、segment by cohort、revise threshold |
| Policy interpretation | update policy pack、legal review、communication note |
Incident learning must enter the roadmap. Otherwise the organization pays for failure without buying learning.
6.4 Backlog Governance
AI Product Ops backlog is not just feature backlog. It is an evidence-driven decision queue.
| Backlog class | Examples | Priority logic |
|---|---|---|
| Outcome gap | no movement in target KPI | high if adoption is strong and value thesis remains |
| Adoption gap | low qualified use in target cohort | high if workflow value depends on broad behavior change |
| Quality gap | recurring failure class | high if blocks trust or control |
| Risk gap | policy breach、over-reliance、customer harm | high by severity and regulatory impact |
| Cost gap | unit cost or review load exceeds threshold | high if scale economics fail |
| Reliability gap | SLO breach、fallback failure | high if workflow depends on real-time AI |
| Evidence gap | weak measurement、missing join、poor traceability | high before scale decision |
| Platform gap | repeated bespoke fixes across products | high if unlocks multiple teams |
| Retirement candidate | weak value、high risk、poor fit | high if capacity should be released |
Backlog governance rules:
- Every high-priority backlog item references a metric、incident、assumption or decision。
- Every roadmap item names the expected evidence movement。
- Risk and reliability items can preempt value features when thresholds are breached。
- Cost and capacity items are first-class roadmap work, not operational noise。
- Retirement is a valid backlog outcome。
7. Dashboard and Decision Protocols
Dashboard 不是越多越好。AI Product Ops dashboard 要支持对应 cadence, 并把指标转成行动。
| Dashboard | Users | Cadence | Decision |
|---|---|---|---|
| Runtime signal board | product、architecture、operations、platform | daily | triage、rollback、escalate |
| Weekly ops board | product、workflow、operations、risk、analytics | weekly | fix、assign、close、release adjustment |
| Monthly value board | sponsor、product、finance、risk | monthly | scale、restrict、redesign、retire |
| Portfolio board | executive、value office、platform | quarterly | allocate funding and capacity |
| Evidence binder | risk、audit、product governance | as needed | explain decision and traceability |
7.1 Weekly Ops Board Sections
| Section | Required segmentation |
|---|---|
| Adoption funnel | role、team、manager、case type、risk tier |
| Quality defects | failure class、version、knowledge source、cohort |
| Reliability / SLO | channel、workflow step、provider、fallback |
| Cost / capacity | case type、tool call、review queue、support category |
| Risk signals | severity、affected population、control、customer impact |
| Open actions | owner、age、due date、closure evidence |
7.2 Design Principles
- Use stable metric names and definitions from metric contract。
- Show version overlays for model / prompt / data / tool releases。
- Show thresholds and action rules, not only trend lines。
- Separate leading indicators from outcome indicators。
- Include small sample narratives for complaints and incidents。
- Make action closure visible in the dashboard。
- Do not mix portfolio metrics and operational triage metrics on the same visual。
7.3 Decision Protocol
每个 review 都应输出可以追踪的 decision record:
Decision requested:
Evidence reviewed:
Interpretation:
Decision:
Conditions:
Action owner:
Closure evidence:
Reopen / stop trigger:
Next review:
这个 protocol 把 dashboard 从“观察界面”变成“运营控制界面”。没有 decision requested 的 dashboard 只是报告; 没有 closure evidence 的 action 只是愿望。
8. 金融零售场景
8.1 Contact-Center Agent Assist
| Ops question | Evidence |
|---|---|
| Are agents using suggestions in eligible calls? | suggestion exposure、accept/edit/reject、call reason |
| Is AHT improvement real? | AHT by call type、repeat contact、transfer、hold time |
| Is compliance stable? | QA script defects、complaint mentions、supervisor overrides |
| Is cost justified? | cost per assisted call、human review、support tickets |
| What enters roadmap? | knowledge gaps、high-edit intents、low-trust product areas |
Weekly review catches issue classes. Monthly review decides whether to expand to new call intents or restrict to low-risk intents.
8.2 Complaint Intelligence
| Ops question | Evidence |
|---|---|
| Is complaint classification improving speed and accuracy? | classification precision sample、cycle time、re-open rate |
| Are regulatory complaints missed? | false negative sampling、QA escalation、regulator response |
| Are root causes actionable? | root cause cluster adoption、remediation closure |
| Is policy drift visible? | taxonomy change log、product policy updates |
Incident-to-roadmap loop is critical: a misclassified regulatory complaint should update taxonomy、eval set、workflow routing and training.
8.3 KYC Onboarding
| Ops question | Evidence |
|---|---|
| Is onboarding cycle time reduced without weaker controls? | document completeness、rework、EDD escalation、false pass sample |
| Which segments suffer value leakage? | entity type、geography、channel、document type |
| Does AI create customer friction? | document chase frequency、complaint text、abandonment |
| What changes in release calendar? | policy rules、document parser、knowledge guidance、threshold |
Monthly value review should not scale if cycle time improves by pushing work into downstream remediation.
8.4 Collections Hardship
| Ops question | Evidence |
|---|---|
| Does AI improve appropriate hardship treatment? | arrangement suitability、kept promises、broken arrangement rate |
| Are vulnerable customers protected? | vulnerability flags、agent override、complaint、QA sample |
| Are agents over-relying? | copy rate、edit rate、supervisor escalation、script deviations |
| What roadmap changes? | policy clarification、conversation guidance、escalation UI |
Here the guardrail metrics may matter more than conversion metrics.
8.5 AML Triage
| Ops question | Evidence |
|---|---|
| Does AI reduce triage aging without missed suspicious activity? | alert aging、escalation quality、audit sampling |
| Does case narrative quality improve? | evidence completeness、reviewer edit distance、SAR prep defects |
| Are new typologies captured? | drift signal、investigator feedback、typology update calendar |
| What enters backlog? | retrieval source、scenario-specific evals、explanation format |
Quarterly portfolio review should examine whether AML AI creates platform capabilities reusable for fraud、sanctions or complaints.
8.6 Personalized Pricing Governance
| Ops question | Evidence |
|---|---|
| Is pricing optimization improving outcome without unfair treatment? | margin、conversion、segment-level impact、complaint |
| Are explanations and overrides adequate? | reason code quality、branch override、audit sample |
| Is policy drift controlled? | pricing policy version、eligibility criteria、exception log |
| What decisions are needed? | restrict segment、add fairness guardrail、update risk appetite |
Personalized pricing needs strong metric governance because local conversion lift can hide conduct risk.
9. 反模式
| Anti-pattern | Symptom | Correction |
|---|---|---|
| Launch theater | 上线后只汇报 usage 和 demo feedback | evidence review pack with outcome、risk、cost and action closure |
| Dashboard without decisions | 指标很多, 没有 decision request | every review starts with decision requested |
| Meeting as memory | 决策靠口头共识 | decision log and assumption ledger |
| Action without closure evidence | ticket closed but metric unchanged | closure requires evidence and reopen trigger |
| Release calendar only for code | prompt / knowledge / tool changes invisible | unified AI release calendar |
| Incident as one-off | 事故修复后不改变 roadmap | incident-to-roadmap loop |
| Value review without risk | 只看 efficiency lift | include complaint、override、policy breach、customer harm |
| Risk review without value | 只看 control checklist | connect controls to outcome and adoption |
| Cost treated as platform problem | token/tool spend not tied to product decisions | cost per case and capacity review |
| Portfolio review as show-and-tell | 每个团队展示进展 | fund / scale / pause / retire decisions |
10. 最终心智模型
AI Product Ops is not governance overhead. It is the operating rhythm that keeps an AI product honest after launch.
No metric contract -> no trusted review.
No evidence pack -> no decision quality.
No release calendar -> no controlled change.
No incident-to-roadmap loop -> no learning.
No action closure -> no operational integrity.
No portfolio review -> no disciplined investment.
最终模型可以压缩为:
Runtime telemetry
-> metric contract
-> evidence review pack
-> cadence-specific decision
-> action closure
-> release / experiment / incident learning
-> backlog and roadmap change
-> scale / restrict / redesign / retire
-> portfolio allocation
成熟问题不是 “Did we launch AI?” 而是:
Are we continuously proving that this AI capability improves outcomes,
stays within risk appetite,
earns its cost,
teaches us from failure,
and deserves its next roadmap decision?
SOTA 检查 (2026-07-01)
- DORA 锚点已换代:本篇引用的 dora.dev 已于 2025 年发布首份《State of AI-assisted Software Development 2025》年度报告(约 5,000 名从业者调研,2025 年发布):90% 受访者在工作中使用 AI、80%+ 报告生产力提升、但 30% 对 AI 生成代码"几乎不信任";核心结论是 "AI 是放大器,不是修复器"——没有自动化测试、快反馈回路和松耦合架构的团队,变更量上升直接转化为不稳定。这与本篇 §2 的 Observe→Verify closure 闭环论证一致,且 DORA 同期发布了 AI Capabilities Model(7 项放大 AI 收益的组织能力,2025-12),可作为本篇 operating calendar 的组织能力前置检查表。
- 可观测性锚点已具体化:OpenTelemetry GenAI semantic conventions 截至 2026-05 整体仍处 Development 状态,但
gen_aiclient spans 与 metrics 已于 2025 年末稳定,agent/framework spans 仍为 experimental(2026 Q1 实践中已趋稳),主流 agent 框架与 Datadog/Google Cloud/AWS/Azure 等平台已跟进 emitter。本篇 §5.2 的 "metric → source event → version → decision" 追溯链落地时应直接采用该 semconv,而非自造字段;本库配套实操见docs/aipa/day22-otel-genai-semconv.md、docs/aipa/day24-instrumentation-full-chain.md、docs/aipa/day25-langfuse-self-host.md。 - 风险框架锚点仍现役但已更新:NIST AI RMF 1.0(2023-01)仍是主流组织语言,2025-03 的更新方向覆盖生成式 AI 风险、供应链与第三方模型评估,并有 Generative AI Profile(NIST AI 600-1)与官方 NIST AI RMF ↔ ISO/IEC 42001 crosswalk——即本篇同时引 NIST 与 ISO 42001 的做法在 2026 年有官方映射支撑,两套语言可互认,不需二选一。
- 合规时间线注意:EU AI Act Annex III 高风险义务已推迟至 2027-12-02(Digital Omnibus,2026-05-07 确认);本篇 §3 的 management system review 与 §6.1 "External change" 事故源在做合规倒排时应按新时间线更新假设(这正是 assumption ledger 的 expiration 字段要捕捉的典型漂移)。
- 不随版本过时的框架性结论:metric contract、evidence review pack、decision log、assumption ledger、release calendar(覆盖 model/prompt/data/knowledge/tool/policy 八类变更对象)、incident-to-roadmap loop、action closure register 这套"证据运行系统"与具体模型/平台无关,是本篇的耐久内核;随版本演化的只是 telemetry 层(OTel GenAI semconv、Langfuse 类 tracing)与 eval 门禁工具层(本库的阻塞式 CI eval gate 实作见
docs/aipa/day19-blocking-ci-eval-gate.md)。 - 本篇主线仍成立:weekly ops / monthly value / quarterly portfolio 的分层 cadence 与 2026 年企业实践一致(板级 AI 治理普遍要求 ownership、风险分级与 review cadence、季度模型绩效复盘留痕);行业普遍短板恰是本篇强调的 cost 控制面——多数企业尚未把 LLM 成本控制纳入运营节奏,印证 §4.2 "cost per case 为一等指标" 的取向。