返回 Papers
AI 底层逻辑 / 经典论文

AI Product Operations:运营节奏与结果复盘架构

AI Product Operations 是上线后的证据系统: 它把运行时行为、采用、结果、风险、成本、事故和变更发布转成可重复的产品与组合决策。

638ai-foundations/papers/156-ai-product-operations-operating-cadence-outcome-review-architecture.md

AI Product Operations / Operating Cadence / Outcome Review Architecture 解读

配对阅读:本篇的操作手册版(模板/RACI/门禁/runbook)是 docs/AI_PRODUCT_OPERATIONS_OPERATING_CADENCE_OUTCOME_REVIEW_PLAYBOOK.md。第一遍读本篇建立原理与架构判断;第二遍做案例时再用 playbook 查表落地,两者不需要重复精读。

核心问题: AI 产品上线后, 如何用 weekly ops review、monthly value review、quarterly portfolio review 和 release / experiment / incident loops, 把真实运营证据转化为 scale、restrict、redesign、retire 和投资决策?

重要说明: 本文是学习和内部架构训练材料, 不构成法律意见、监管解释、合规确认、审计意见、模型验证结论、风险接受决定、财务投资建议或生产上线批准。正式项目中的审批权、残余风险接受、监管沟通、审计依赖、客户影响判断和发布授权必须由机构授权角色结合司法辖区、产品、客户群、风险偏好、内部政策、模型风险、信息安全、隐私、供应商合同和运营能力确认。访问日期按 2026-06-30 记录。


Source Anchors

以下来源用于组织 AI 风险管理、AI 管理体系、需求工程、工程绩效、可观测性和服务可靠性语言。本文只将这些来源作为产品、架构和内部 product operations 的设计锚点, 不声称任何运营节奏、dashboard 或 review pack 自动构成法律、监管、审计或模型验证批准。

SourceLink本文采用的思想
NIST AI Risk Management Frameworkhttps://www.nist.gov/itl/ai-risk-management-framework用 Govern / Map / Measure / Manage 组织 post-launch risk review、monitoring、incident learning 和 action closure。(AI RMF 1.0 发布 2023-01;访问日期: 2026-07-01)
ISO/IEC 42001 AI management systemhttps://www.iso.org/standard/81230.html用 management system 的 policy、objectives、operation、performance evaluation、internal audit、improvement 语言定义 AI Product Ops。(访问日期: 2026-07-01)
ISO/IEC/IEEE 29148 Requirements Engineeringhttps://www.iso.org/standard/72089.html用 requirements quality、stakeholder needs、verification、traceability 思路设计 metric contract、assumption ledger 和 decision log。(访问日期: 2026-07-01)
DORAhttps://dora.dev/用 software delivery performance 和 reliability mindset 连接 release cadence、change fail rate、restore time 和 learning loop。(最新年度报告为《State of AI-assisted Software Development 2025》,见文末 SOTA 检查;访问日期: 2026-07-01)
OpenTelemetry Documentationhttps://opentelemetry.io/docs/用 traces、metrics、logs 和 semantic conventions 的思想设计 AI product operations telemetry。(GenAI semconv 状态见文末 SOTA 检查;访问日期: 2026-07-01)
Google SRE: Service Level Objectiveshttps://sre.google/sre-book/service-level-objectives/用 SLO、error budget 和 service reliability 语言定义 AI product operational thresholds。(访问日期: 2026-07-01)

核心导读

AI Product Operations 是上线后的证据系统: 它把运行时行为、采用、结果、风险、成本、事故和变更发布转成可重复的产品与组合决策。

post-launch telemetry
  -> evidence review pack
  -> cadence-specific decisions
  -> backlog and release calendar
  -> action closure
  -> outcome and risk learning
  -> portfolio allocation

上线前的核心问题是“能不能做”。上线后的核心问题变成“是否仍然值得运行、如何运行得更好、何时扩大、何时限制、何时重构、何时退役”。AI Product Ops 的价值就在于让这个判断不依赖主观热情, 而依赖 metric contract、evidence pack、decision log、assumption ledger、release calendar、incident learning 和 action closure。


1. 问题定义: AI 上线后的运营断裂

很多 AI 产品的失败发生在上线之后:

  • Pilot 证明了模型能回答问题, 但上线后 adoption 只停留在少数 champion。
  • Usage 很高, 但流程周期、质量、投诉或风险控制没有改善。
  • Prompt、知识库、模型、tool permission 和 policy pack 持续变更, 但没有统一 release calendar。
  • 事故复盘只产生修复 ticket, 没有进入 roadmap、metric contract、training、policy 或 control design。
  • 成本增长被解释为“用户增长”, 但没有 case-level unit economics 和 capacity review。
  • 管理层每月看到 dashboard, 但没有 decision log、assumption ledger 和 action closure。

AI Product Operations 的目标不是多开几个会议。它是一套运营架构, 把七类证据放进同一个节奏:

Evidence lane核心问题常见证据
Outcome是否改善目标业务结果cycle time、first-pass yield、AHT、loss avoided、conversion、complaint rate
Adoption目标角色是否在正确工作步骤采用qualified use、accept/edit/reject、cohort durability、manager reinforcement
Quality输出质量是否稳定并适配 case mixeval pass rate、QA defects、hallucination class、retrieval freshness
Risk / Control风险是否仍在 appetite 内override、escalation、policy breach、customer harm、audit finding
Cost / Capacity单位经济是否成立cost per case、token/tool cost、review load、queue aging、support effort
Incident Learning失败是否被转成系统改进incident taxonomy、root cause、corrective action、recurrence signal
Roadmap证据是否改变投资和优先级decision log、assumption ledger、experiment result、release calendar

本文聚焦 post-launch product operations cadence and outcome evidence。它不重复团队授权、product trio 或 decision rights 的基础模型, 而是假设 AI capability 已经完成 controlled pilot 或初步上线, 接下来要建立持续运营节奏。

DimensionAI Product Operating ModelAI Product Operations Cadence
关注点团队授权、decision rights、guardrailspost-launch review、evidence、action closure、roadmap decision
时间位置discovery 到 launch 前后controlled pilot、production、scale、refresh、retire
主要问题团队能否在 guardrails 内解决问题产品是否仍创造价值且风险可控
核心对象team、decision rights、gatesmetric contract、evidence pack、operating calendar
成功标志能发现、交付、治理 AI capability能持续证明、调整、扩展、限制或退役 capability

2. 架构模型: AI Product Ops 作为结果证据运行系统

AI Product Ops 的最小闭环:

Observe
  -> interpret
  -> decide
  -> act
  -> verify closure
  -> update assumptions and roadmap

如果缺少任一环, cadence 就会退化。

Missing piece退化表现结果
Observe只有主观反馈, 没有 trace / metric / sample无法区分真实风险和噪声
Interpretdashboard 多, 但没有 root cause language数字变化不产生决策
Decide会议讨论多, 没有 decision log同一争议反复出现
Actaction 没 owner / due date / evidence会议变成汇报仪式
Verify closureticket closed, 但 outcome 未复核修复不等于问题解决
Update roadmap事故和学习不改变优先级产品继续按旧假设投资

2.1 Operating Model Components

ComponentPurposeAccountability question
Operating calendar定义 weekly、monthly、quarterly、release、incident、experiment review 节奏哪些节奏负责哪些决策, 哪些信号必须升级
Metric contract定义指标口径、owner、阈值、数据源、行动规则这个指标是否能支撑明确行动, 而不是只做展示
Evidence review pack把 adoption、outcome、quality、risk、cost、incident、roadmap 整合为决策材料一页证据是否足以解释建议的 scale / restrict / redesign / retire
Decision log记录 scale、restrict、release、rollback、policy、roadmap 决策及依据未来是否能重建当时为什么这样决策
Assumption ledger记录价值、行为、风险、成本和 capacity 假设是否仍成立哪些旧假设已经过期, 哪些仍可用于投资判断
Experiment registry记录实验目的、population、hypothesis、metric、risk guardrail、结果实验是否留下可复用学习, 而不是只留下局部胜负
Release calendar管理 model、prompt、data、knowledge、tool、policy、UX、workflow 变更AI 行为变化是否可追踪、可回滚、可解释
Incident learning loop将 incident / complaint / near miss 转为 corrective action 和 roadmap item失败是否买到了系统学习
Action closure register跟踪 action owner、due date、closure evidence、reopen trigger行动是否真的改变指标、流程或控制状态
Portfolio review pack支撑 fund / scale / pause / retire / consolidate 决策资源是否流向净价值最高且风险可控的能力

2.2 Control Planes

AI Product Ops 至少覆盖九个控制面:

PlaneReview question
Value业务结果是否移动, benefit 是否净实现
Adoption目标用户是否持续正确采用
Quality输出质量和 workflow fit 是否稳定
Reliabilitylatency、availability、restore、fallback 是否达标
Riskcustomer harm、model risk、policy breach、over-reliance 是否受控
Costunit cost、support load、review load、capacity 是否可承受
Changemodel/prompt/data/tool/policy 变更是否可追溯
Incident失败是否被学习、关闭并防止复发
Roadmap新证据是否改变投资方向

2.3 Product Ops Data Objects

ObjectKey fieldsWhy it matters
Metric contractmetric_id、definition、owner、source、threshold、action防止每次 review 重新争论指标口径
Review packperiod、population、evidence、decision request、actions让会议从汇报转成决策
Release itemobject_type、version、change reason、risk tier、rollback把 AI 变更纳入可追溯 calendar
Experiment recordhypothesis、cohort、guardrails、duration、result、decision防止实验结果丢失或被选择性引用
Incident recordseverity、impact、root cause、affected versions、corrective action让事故进入 learning loop
Assumptionstatement、evidence、confidence、expiration、owner管理价值和风险叙事的有效期
Action closureaction、owner、due date、evidence、reviewer、reopen trigger防止会议行动消失

3. 关键机制与生命周期: Cadence Architecture

AI 产品上线后的 cadence 不是单一会议, 而是不同时间尺度的证据处理栈。

Daily signal triage
  -> weekly ops review
  -> monthly value review
  -> quarterly portfolio review
  -> annual / semiannual management system review
CadencePrimary lensTypical decisions
Daily signal triageincidents、latency、availability、complaint spikes、cost anomalymitigate、rollback、escalate、sample、hotfix
Weekly ops reviewadoption、quality、reliability、capacity、open actionsprioritize fixes、adjust release、assign owners
Monthly value reviewoutcome、unit economics、benefit leakage、risk trendscale、restrict、redesign、update business case
Quarterly portfolio reviewuse-case portfolio、platform reuse、risk concentration、fundingfund、pause、consolidate、retire、reallocate capacity
Management system reviewpolicy effectiveness、audit findings、objectives、continual improvementupdate operating policy、control library、governance model

3.1 Weekly Ops Review

Weekly ops review 是 tactical learning forum, 不应退化为 status meeting。

InputReview questionOutput
Adoption funnel by cohort哪些用户、case type、manager group 掉队enablement action、product fix、workflow change
Quality sample and eval result哪类 failure 正在上升prompt/index/model/tool fix、sampling change
Reliability and SLOlatency、availability、restore 是否影响工作platform action、fallback adjustment
Cost / capacityreview queue、token/tool cost、support load 是否异常capacity rebalance、cost guardrail
Incident / complaint signals是否存在 customer harm 或 policy driftincident triage、risk escalation
Open action register上周行动是否关闭, closure evidence 是否充分close、reopen、escalate

Weekly outputs 必须可执行:

  • action owner。
  • due date。
  • closure evidence。
  • decision log entry。
  • backlog item 或 release calendar update。
  • escalation path。

3.2 Monthly Value Review

Monthly value review 是 outcome and investment forum。它回答“这个 AI capability 是否仍然值得继续投资”。

Review blockEvidence
Outcome movementbaseline vs current、cohort trend、seasonality adjustment
Adoption durabilityreturning qualified use、manager reinforcement、work-as-done evidence
Value leakagehuman review load、rework、support cost、exception queue、customer redress
Risk trendcomplaints、overrides、policy breaches、fairness / conduct signals
Cost-to-serveunit cost per case、marginal cost、platform capacity
Release impactrecent model/prompt/data/tool releases and outcome changes
Decision requestscale、hold、restrict、redesign、retire、continue experiment

Monthly review 的关键是把数据转成明确决策:

Continue because evidence is improving and risk is stable.
Scale because outcome lift is durable and marginal cost is acceptable.
Restrict because specific cohorts or case types show harm or poor reliability.
Redesign because usage is high but value leakage removes benefit.
Retire because assumptions failed and no credible path remains.

3.3 Quarterly Portfolio Review

Quarterly portfolio review 把单个 use case 上升到 enterprise AI allocation。

Portfolio lensQuestions
Value concentration哪些 use cases 贡献主要净收益, 哪些只有 activity
Risk concentration是否在同一 customer segment、model provider、data source 或 control weakness 上集中
Platform leverage哪些 capabilities 应产品化复用, 哪些 bespoke build 应合并
Talent / capacitySME review、risk review、data engineering、platform support 是否成为瓶颈
Policy drift业务规则、监管解释、模型能力、供应商条款是否改变
Roadmap reallocation哪些主题应加速, 哪些应暂停、合并或退役

Quarterly review 的输出不是“下季度计划”, 而是 funding change、platform investment decision、risk appetite adjustment、capability retirement decision、governance process improvement 和 portfolio-level assumption update。


4. Outcome Review and Metric Contract

Outcome review 不是看一个 North Star metric, 而是看 outcome chain。

AI release / experiment
  -> target exposure
  -> qualified adoption
  -> workflow behavior change
  -> quality and control movement
  -> business outcome
  -> net value after risk and cost
LayerEvidenceFailure interpretation
Releaseversion changed、target cohort exposedrelease did not reach workflow
Adoptionaccepted / edited / rejected / escalatedusers do not trust、do not need、or cannot use
Behaviorartifact、decision、handoff changedAI used as side tool but not embedded
Qualityfirst-pass yield、QA、eval、defect classoutput not fit for production work
Controloverride、escalation、policy breach、complaintvalue is creating hidden risk
Outcomecycle、conversion、AHT、loss、STPbusiness result did not move
Net valuebenefit minus operating、review、support、risk costgross benefit does not survive operations

Outcome review 要避免两个陷阱: 把 release 当成结果, 以及把单点结果改善当成可持续价值。

ExampleOutcome claimRequired counter-evidence
Contact-center agent assistAHT 下降repeat contact、complaints、script compliance、hold transfer
Complaint intelligenceroot cause identification 更快misclassification、regulatory breach、remediation delay
KYC onboardingcycle time 下降false pass、rework、document chase、vulnerable customer impact
Collections hardshiparrangement completion 上升unfair pressure、complaints、broken promises、agent override
AML triagealert closure 更快suspicious activity miss、escalation quality、audit sampling
Personalized pricingmargin / conversion upliftunfair treatment、explainability、opt-out、complaint trend

4.1 Metric Contract

AI product review 经常争论指标为什么变了、dashboard 和 finance number 为什么不一致、贡献来自 AI 还是 seasonality、usage 是否等于 adoption、成本上升是 scale 信号还是经济性恶化。Metric contract 是对指标的产品需求说明和治理文件。

FieldDescription
metric_idStable identifier, such as kyc_ai.first_pass_yield
business question这个指标要回答什么决策问题
definition精确定义, 包括 numerator / denominator
population用户、case type、channel、risk tier、time window
source systemtelemetry、workflow system、finance ledger、QA、complaint platform
owner对口径和解释负责的人
review cadencedaily、weekly、monthly、quarterly
thresholdtarget、warning、breach、stop rule
guardrail防止局部优化伤害其他目标
segmentation必须按哪些 cohort 拆解
action rule指标越界时触发什么行动
evidence qualityobserved、sampled、inferred、survey、finance-certified
expiry / review date何时重新审查口径是否仍适用

4.2 Metric Taxonomy and Governance

Metric typeExampleCadence
Outcomecomplaint cycle time、AHT、KYC approval cycle、AML agingmonthly
Adoptionqualified use、acceptance、edit rate、rejection reasonweekly
Qualityeval pass、QA defect、hallucination class、retrieval hit qualityweekly
ReliabilitySLO、latency、availability、fallback success、restore timedaily / weekly
Riskpolicy breach、override、escalation、customer harm、fairness signalweekly / monthly
Costcost per case、token/tool cost、support effort、review loadweekly / monthly
Learningexperiment velocity、action closure、incident recurrenceweekly / monthly
Portfolionet value、risk exposure、platform reuse、retirement ratequarterly

Metric governance is product governance:

  • 每个 metric 有 owner, 没有 owner 的 metric 不进入 executive review。
  • 每个 metric 有 action rule, 没有 action rule 的 metric 只是观察值。
  • 每个 metric 有 segmentation, 否则会隐藏 vulnerable cohort。
  • 每个 metric 有 validity period, 因为流程、模型、政策和用户行为会漂移。
  • 每个 metric 有 evidence quality rating, 区分 telemetry、sampling、survey 和 finance-certified value。

5. 证据与控制: Review Pack、Traceability、Release Calendar

Evidence review pack 是每次 review 的共同材料。它不追求信息多, 追求能产生决策。

5.1 Review Pack Structure

SectionContent
Decision requestedcontinue、scale、restrict、redesign、retire、release、rollback
Scope and versionproduct area、population、model/prompt/data/tool version
Outcome summarybaseline、current、movement、confidence
Adoption summarycohort funnel、qualified use、durability
Quality summaryeval、QA sample、failure taxonomy
Risk/control summaryincidents、complaints、overrides、policy drift
Cost/capacity summaryunit cost、review load、support load
Release and experiment summaryrecent changes、experiments、observed effects
Open assumptionsassumptions confirmed、weakened、invalidated
Action closurelast actions、evidence、unresolved blockers
Recommendationspecific decision and next review trigger

5.2 Evidence Quality and Traceability

LevelDescriptionReview use
E1 Anecdotalisolated feedback or demo observationsignal only
E2 SampledQA sample、complaint sample、interview sampleweekly interpretation
E3 Instrumentedproduction telemetry joined to workflow contextweekly / monthly decision
E4 Causal or quasi-causalcontrolled experiment、matched cohort、difference analysisscale / restrict decision
E5 Finance / risk certifiedreconciled benefit、validated risk and audit-ready evidenceportfolio investment decision

Post-launch evidence should trace:

metric -> source event -> workflow context -> version -> decision -> action -> closure evidence -> next metric movement

OpenTelemetry-inspired traces 与 ISO/IEC/IEEE 29148-inspired traceability 在这里相遇。重点不是技术优雅, 而是能解释 roadmap 为什么改变。

5.3 Experiment and Release Calendar

AI products change through more than code deploys.

Change objectExampleRisk
Modelprovider upgrade、model class change、fallback modelquality shift、cost shift、latency、data boundary
Promptsystem prompt、tool instruction、refusal wordingbehavior shift、policy drift、regression
Datafeature change、label change、training data refreshbias、leakage、stale assumptions
KnowledgeRAG corpus、policy document、product catalogoutdated guidance、retrieval mismatch
ToolCRM write action、fee waiver API、case closure actionside effect、authorization、audit
Policyhardship treatment rule、complaint taxonomy、KYC requirementcompliance breach、inconsistent handling
WorkflowUI step、queue routing、human review thresholdadoption change、capacity shift
Experimentcohort change、A/B treatment、canaryinterpretation error、customer impact

Release calendar fields:

FieldDescription
release_idStable release identifier
object_typemodel、prompt、data、knowledge、tool、policy、workflow
affected populationcohort、channel、case type、geography
evidence requiredeval、QA、risk、cost、regression、rollout plan
canary planfirst users、duration、guardrails
rollback pathtechnical and operational rollback
communicationfrontline、risk、support、manager notes
review datewhen impact is reviewed
decision log linkwhy release was approved

A release without impact review is an uncontrolled change. An experiment without release trace is an unrepeatable learning.


6. Incident-to-Roadmap and Backlog Governance

AI incident management becomes product strategy when failures reveal weak assumptions.

6.1 Incident Sources

SourceExample
Customer complaintcustomer claims AI-generated explanation was misleading
Frontline overrideagent repeatedly rejects a suggested hardship script
QA defectcomplaint classifier misses regulatory complaint
Model driftAML triage quality drops for a new fraud typology
Cost anomalytool calls spike after prompt change
Policy driftknowledge base uses outdated pricing exception rule
Near misshuman reviewer catches a high-impact hallucination
External changeregulation、product terms、vendor model behavior changes

6.2 Learning Loop

detect signal
  -> classify severity and affected population
  -> contain or rollback
  -> root cause across model / prompt / data / tool / workflow / policy / training
  -> corrective action
  -> metric contract update
  -> backlog / roadmap update
  -> action closure evidence
  -> recurrence review

6.3 Root Cause Taxonomy

Cause classExample action
Model behaviorchange model、add eval、adjust fallback
Prompt instructionrevise prompt、add regression case、review release path
Knowledge freshnessupdate corpus、add freshness SLO、assign knowledge owner
Tool permissionrestrict tool、add approval、update authorization
Workflow designchange handoff、add human review、revise UI
Training / adoptionmanager coaching、SOP update、new refusal guidance
Metric designadd missing guardrail、segment by cohort、revise threshold
Policy interpretationupdate policy pack、legal review、communication note

Incident learning must enter the roadmap. Otherwise the organization pays for failure without buying learning.

6.4 Backlog Governance

AI Product Ops backlog is not just feature backlog. It is an evidence-driven decision queue.

Backlog classExamplesPriority logic
Outcome gapno movement in target KPIhigh if adoption is strong and value thesis remains
Adoption gaplow qualified use in target cohorthigh if workflow value depends on broad behavior change
Quality gaprecurring failure classhigh if blocks trust or control
Risk gappolicy breach、over-reliance、customer harmhigh by severity and regulatory impact
Cost gapunit cost or review load exceeds thresholdhigh if scale economics fail
Reliability gapSLO breach、fallback failurehigh if workflow depends on real-time AI
Evidence gapweak measurement、missing join、poor traceabilityhigh before scale decision
Platform gaprepeated bespoke fixes across productshigh if unlocks multiple teams
Retirement candidateweak value、high risk、poor fithigh if capacity should be released

Backlog governance rules:

  • Every high-priority backlog item references a metric、incident、assumption or decision。
  • Every roadmap item names the expected evidence movement。
  • Risk and reliability items can preempt value features when thresholds are breached。
  • Cost and capacity items are first-class roadmap work, not operational noise。
  • Retirement is a valid backlog outcome。

7. Dashboard and Decision Protocols

Dashboard 不是越多越好。AI Product Ops dashboard 要支持对应 cadence, 并把指标转成行动。

DashboardUsersCadenceDecision
Runtime signal boardproduct、architecture、operations、platformdailytriage、rollback、escalate
Weekly ops boardproduct、workflow、operations、risk、analyticsweeklyfix、assign、close、release adjustment
Monthly value boardsponsor、product、finance、riskmonthlyscale、restrict、redesign、retire
Portfolio boardexecutive、value office、platformquarterlyallocate funding and capacity
Evidence binderrisk、audit、product governanceas neededexplain decision and traceability

7.1 Weekly Ops Board Sections

SectionRequired segmentation
Adoption funnelrole、team、manager、case type、risk tier
Quality defectsfailure class、version、knowledge source、cohort
Reliability / SLOchannel、workflow step、provider、fallback
Cost / capacitycase type、tool call、review queue、support category
Risk signalsseverity、affected population、control、customer impact
Open actionsowner、age、due date、closure evidence

7.2 Design Principles

  • Use stable metric names and definitions from metric contract。
  • Show version overlays for model / prompt / data / tool releases。
  • Show thresholds and action rules, not only trend lines。
  • Separate leading indicators from outcome indicators。
  • Include small sample narratives for complaints and incidents。
  • Make action closure visible in the dashboard。
  • Do not mix portfolio metrics and operational triage metrics on the same visual。

7.3 Decision Protocol

每个 review 都应输出可以追踪的 decision record:

Decision requested:
Evidence reviewed:
Interpretation:
Decision:
Conditions:
Action owner:
Closure evidence:
Reopen / stop trigger:
Next review:

这个 protocol 把 dashboard 从“观察界面”变成“运营控制界面”。没有 decision requested 的 dashboard 只是报告; 没有 closure evidence 的 action 只是愿望。


8. 金融零售场景

8.1 Contact-Center Agent Assist

Ops questionEvidence
Are agents using suggestions in eligible calls?suggestion exposure、accept/edit/reject、call reason
Is AHT improvement real?AHT by call type、repeat contact、transfer、hold time
Is compliance stable?QA script defects、complaint mentions、supervisor overrides
Is cost justified?cost per assisted call、human review、support tickets
What enters roadmap?knowledge gaps、high-edit intents、low-trust product areas

Weekly review catches issue classes. Monthly review decides whether to expand to new call intents or restrict to low-risk intents.

8.2 Complaint Intelligence

Ops questionEvidence
Is complaint classification improving speed and accuracy?classification precision sample、cycle time、re-open rate
Are regulatory complaints missed?false negative sampling、QA escalation、regulator response
Are root causes actionable?root cause cluster adoption、remediation closure
Is policy drift visible?taxonomy change log、product policy updates

Incident-to-roadmap loop is critical: a misclassified regulatory complaint should update taxonomy、eval set、workflow routing and training.

8.3 KYC Onboarding

Ops questionEvidence
Is onboarding cycle time reduced without weaker controls?document completeness、rework、EDD escalation、false pass sample
Which segments suffer value leakage?entity type、geography、channel、document type
Does AI create customer friction?document chase frequency、complaint text、abandonment
What changes in release calendar?policy rules、document parser、knowledge guidance、threshold

Monthly value review should not scale if cycle time improves by pushing work into downstream remediation.

8.4 Collections Hardship

Ops questionEvidence
Does AI improve appropriate hardship treatment?arrangement suitability、kept promises、broken arrangement rate
Are vulnerable customers protected?vulnerability flags、agent override、complaint、QA sample
Are agents over-relying?copy rate、edit rate、supervisor escalation、script deviations
What roadmap changes?policy clarification、conversation guidance、escalation UI

Here the guardrail metrics may matter more than conversion metrics.

8.5 AML Triage

Ops questionEvidence
Does AI reduce triage aging without missed suspicious activity?alert aging、escalation quality、audit sampling
Does case narrative quality improve?evidence completeness、reviewer edit distance、SAR prep defects
Are new typologies captured?drift signal、investigator feedback、typology update calendar
What enters backlog?retrieval source、scenario-specific evals、explanation format

Quarterly portfolio review should examine whether AML AI creates platform capabilities reusable for fraud、sanctions or complaints.

8.6 Personalized Pricing Governance

Ops questionEvidence
Is pricing optimization improving outcome without unfair treatment?margin、conversion、segment-level impact、complaint
Are explanations and overrides adequate?reason code quality、branch override、audit sample
Is policy drift controlled?pricing policy version、eligibility criteria、exception log
What decisions are needed?restrict segment、add fairness guardrail、update risk appetite

Personalized pricing needs strong metric governance because local conversion lift can hide conduct risk.


9. 反模式

Anti-patternSymptomCorrection
Launch theater上线后只汇报 usage 和 demo feedbackevidence review pack with outcome、risk、cost and action closure
Dashboard without decisions指标很多, 没有 decision requestevery review starts with decision requested
Meeting as memory决策靠口头共识decision log and assumption ledger
Action without closure evidenceticket closed but metric unchangedclosure requires evidence and reopen trigger
Release calendar only for codeprompt / knowledge / tool changes invisibleunified AI release calendar
Incident as one-off事故修复后不改变 roadmapincident-to-roadmap loop
Value review without risk只看 efficiency liftinclude complaint、override、policy breach、customer harm
Risk review without value只看 control checklistconnect controls to outcome and adoption
Cost treated as platform problemtoken/tool spend not tied to product decisionscost per case and capacity review
Portfolio review as show-and-tell每个团队展示进展fund / scale / pause / retire decisions

10. 最终心智模型

AI Product Ops is not governance overhead. It is the operating rhythm that keeps an AI product honest after launch.

No metric contract -> no trusted review.
No evidence pack -> no decision quality.
No release calendar -> no controlled change.
No incident-to-roadmap loop -> no learning.
No action closure -> no operational integrity.
No portfolio review -> no disciplined investment.

最终模型可以压缩为:

Runtime telemetry
  -> metric contract
  -> evidence review pack
  -> cadence-specific decision
  -> action closure
  -> release / experiment / incident learning
  -> backlog and roadmap change
  -> scale / restrict / redesign / retire
  -> portfolio allocation

成熟问题不是 “Did we launch AI?” 而是:

Are we continuously proving that this AI capability improves outcomes,
stays within risk appetite,
earns its cost,
teaches us from failure,
and deserves its next roadmap decision?

SOTA 检查 (2026-07-01)

  • DORA 锚点已换代:本篇引用的 dora.dev 已于 2025 年发布首份《State of AI-assisted Software Development 2025》年度报告(约 5,000 名从业者调研,2025 年发布):90% 受访者在工作中使用 AI、80%+ 报告生产力提升、但 30% 对 AI 生成代码"几乎不信任";核心结论是 "AI 是放大器,不是修复器"——没有自动化测试、快反馈回路和松耦合架构的团队,变更量上升直接转化为不稳定。这与本篇 §2 的 Observe→Verify closure 闭环论证一致,且 DORA 同期发布了 AI Capabilities Model(7 项放大 AI 收益的组织能力,2025-12),可作为本篇 operating calendar 的组织能力前置检查表。
  • 可观测性锚点已具体化:OpenTelemetry GenAI semantic conventions 截至 2026-05 整体仍处 Development 状态,但 gen_ai client spans 与 metrics 已于 2025 年末稳定,agent/framework spans 仍为 experimental(2026 Q1 实践中已趋稳),主流 agent 框架与 Datadog/Google Cloud/AWS/Azure 等平台已跟进 emitter。本篇 §5.2 的 "metric → source event → version → decision" 追溯链落地时应直接采用该 semconv,而非自造字段;本库配套实操见 docs/aipa/day22-otel-genai-semconv.mddocs/aipa/day24-instrumentation-full-chain.mddocs/aipa/day25-langfuse-self-host.md
  • 风险框架锚点仍现役但已更新:NIST AI RMF 1.0(2023-01)仍是主流组织语言,2025-03 的更新方向覆盖生成式 AI 风险、供应链与第三方模型评估,并有 Generative AI Profile(NIST AI 600-1)与官方 NIST AI RMF ↔ ISO/IEC 42001 crosswalk——即本篇同时引 NIST 与 ISO 42001 的做法在 2026 年有官方映射支撑,两套语言可互认,不需二选一。
  • 合规时间线注意:EU AI Act Annex III 高风险义务已推迟至 2027-12-02(Digital Omnibus,2026-05-07 确认);本篇 §3 的 management system review 与 §6.1 "External change" 事故源在做合规倒排时应按新时间线更新假设(这正是 assumption ledger 的 expiration 字段要捕捉的典型漂移)。
  • 不随版本过时的框架性结论:metric contract、evidence review pack、decision log、assumption ledger、release calendar(覆盖 model/prompt/data/knowledge/tool/policy 八类变更对象)、incident-to-roadmap loop、action closure register 这套"证据运行系统"与具体模型/平台无关,是本篇的耐久内核;随版本演化的只是 telemetry 层(OTel GenAI semconv、Langfuse 类 tracing)与 eval 门禁工具层(本库的阻塞式 CI eval gate 实作见 docs/aipa/day19-blocking-ci-eval-gate.md)。
  • 本篇主线仍成立:weekly ops / monthly value / quarterly portfolio 的分层 cadence 与 2026 年企业实践一致(板级 AI 治理普遍要求 ownership、风险分级与 review cadence、季度模型绩效复盘留痕);行业普遍短板恰是本篇强调的 cost 控制面——多数企业尚未把 LLM 成本控制纳入运营节奏,印证 §4.2 "cost per case 为一等指标" 的取向。