返回 Papers
AI 扩展计划 / Playbooks

AI 角色能力矩阵 2026

2026+ 企业 AI 能力正在从单点 prompt / chatbot 能力, 转向跨职能的 AI operating capability.

671AI_ROLE_COMPETENCY_MATRIX_2026.md

AI Capability Boundary Matrix 2026+

目标: 为企业 AI 的业务分析、产品经营、方案架构、企业架构、EvalOps 和现场交付建立一张可执行的能力边界矩阵. 背景: 面向 2026+ 企业 AI 需求, 重点从 "会做 AI demo" 升级到 "能发现真问题, 设计可评估系统, 推动治理和 adoption, 讲清业务价值". 使用方式: 每月用本文做一次能力自评, 每季度产出一组 evidence artifacts, 每半年整理成可被设计审查、治理审查和运行复盘追问的系统证据.


1. 总体定位

2026+ 企业 AI 能力正在从单点 prompt / chatbot 能力, 转向跨职能的 AI operating capability.

真正有复用价值的能力体系, 需要同时回答六类问题:

  1. Business: 这个问题是否值得用 AI, 不做或不用 AI 的机会成本是什么?
  2. Product: 用户是谁, workflow 如何变化, 价值指标和 adoption 指标是什么?
  3. 业务分析: 现有流程, 决策, 数据, 规则, 异常, 干系人冲突在哪里?
  4. Architecture: 应该用 RAG, workflow, agent, fine-tuning, rule engine, vendor product, or hybrid?
  5. EvalOps: 需求如何变成 eval, release gate, drift monitor, incident review?
  6. Governance: 如何满足风险, 合规, 审计, 人工监督, model lifecycle, vendor accountability?

本文的核心判断:

  • 业务分析能力不是会议纪要自动化, 而是 AI 项目的 problem evidence owner.
  • 产品经营能力不是功能排期, 而是 AI 产品价值, 用户行为改变, adoption 的 owner.
  • AI Solutions Architect 不是画大框图的人, 而是需求到架构控制, 集成, 成本, 安全, 可观测性的 owner.
  • AI Enterprise Architect 不是只管标准的人, 而是企业 AI capability, portfolio, governance, target architecture 的 owner.
  • AI Product Operations / EvalOps 不是 QA 执行者, 而是 AI system quality loop 和 production evidence 的 owner.
  • Field AI Engineer / Forward Deployed Engineer 不是纯交付工程师, 而是在客户现场把 ambiguous problem 变成 working AI system 的 owner.

2. 能力边界与协作面

能力域核心问题主要责任典型 deliverables成功信号
业务分析能力What is the real business problem?问题定义, 干系人证据, 流程, 需求, 决策规则AI Opportunity Canvas, Stakeholder Evidence Map, BPMN, Requirements-to-Eval Matrix能把模糊需求变成可验证的业务决策
产品经营能力What product should we build and why now?用户价值, scope, roadmap, metric tree, adoptionProduct Brief, PRD, Prototype Report, Adoption Dashboard, Business Case能证明用户工作方式和业务指标发生变化
AI Solutions ArchitectHow should the solution be designed safely?系统架构, 集成, RAG/agent 选型, security, cost, observabilityC4, ADR, Data/Control Pack, Vendor Assessment, NFR能把需求转成可落地, 可控, 可运维的方案
AI Enterprise ArchitectHow does this fit enterprise strategy?capability map, target architecture, governance, portfolio, standardsAI Capability Map, Target Architecture, Roadmap, Architecture Principles, Review Pack能把多个 AI 项目纳入企业级能力演进
AI Product Operations / EvalOpsHow do we know it keeps working?eval design, release gates, monitoring, feedback loop, incident reviewEval Suite, Golden Dataset, Release Gate, Drift Dashboard, Incident Review能让 AI 质量从一次性验收变成持续运营
Field AI Engineer / FDEHow do we make it work in the real client context?discovery-to-build, integration, prototype, deployment, user feedbackWorking Prototype, Integration Adapter, Field Notes, Pilot Report, Handoff Pack能在不完整信息下快速交付可验证系统

2.1 主要差异

  • 业务分析能力从 "what users ask for" 追溯到 "what decision or workflow must improve".
  • 产品经营能力从 "feature backlog" 上升到 "behavior change, value realization, and market positioning".
  • AI Solutions Architect 从 "can it call a model" 深入到 "knowledge, permission, latency, cost, fallback, audit, and resilience".
  • AI Enterprise Architect 从 "one system design" 上升到 "enterprise capability portfolio and governance operating model".
  • EvalOps 从 "testing after build" 前移到 "requirements as eval contracts".
  • FDE 从 "implement tickets" 升级到 "discover, design, integrate, iterate, and transfer ownership in the field".

2.2 主要重叠

  • 业务分析能力和产品经营能力重叠在 problem framing, user research, success metrics.
  • 业务分析能力和 Solutions Architect 重叠在 requirements-to-eval, data readiness, domain modeling.
  • 产品经营能力和 EvalOps 重叠在 product quality, launch gates, adoption metrics.
  • Solutions Architect 和 Enterprise Architect 重叠在 target architecture, standards, reusable platform patterns.
  • Enterprise Architect 和 Governance 重叠在 systemic risk, policy, operating model, vendor accountability.
  • FDE 和所有角色都有重叠, 但更强调现场约束, speed, integration, feedback.

2.3 组合定位

组合能力的主线:

将金融零售业务问题、流程证据、产品价值、AI 架构、EvalOps、风险治理和组织 adoption 放进同一条可验证路线图。


3. Source Anchors and Standards References

这些标准不是岗位说明书, 而是能力矩阵的 anchor. 使用时要转成 artifact, control, eval, or review evidence.

SourceOfficial / Primary Link在本文中的用法
NIST AI RMFhttps://www.nist.gov/itl/ai-risk-management-framework用 Govern, Map, Measure, Manage 思路把 AI 风险转成需求, eval, release gate, owner
EU AI Acthttps://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng and https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai用 risk-based obligations, transparency, human oversight, high-risk system logic 设计治理证据
IIBA / BABOKhttps://www.iiba.org/career-resources/a-business-analysis-professionals-foundation-for-success/babok/用 elicitation, requirements analysis, strategy analysis, solution evaluation 规范 业务分析能力证据
TOGAF Standardhttps://www.opengroup.org/togaf and https://www.opengroup.org/togaf-standard-10th-edition-downloads用 ADM, architecture governance, capability-based planning, roadmap 管理 AI transformation
BPMN 2.0.2https://www.omg.org/spec/BPMN/2.0.2/About-BPMN用标准流程语言表达 human + AI workflow, exception flow, control point
BIAN Service Landscapehttps://bian.org/deliverables/service-landscape/用银行服务域和语义服务视角做 financial domain modeling
OWASP LLM Top 10https://owasp.org/www-project-top-10-for-large-language-model-applications/用 prompt injection, sensitive information disclosure, excessive agency 等风险指导架构控制
ISO/IEC 42001:2023https://www.iso.org/standard/81230.html用 AI management system 思路设计 policy, accountability, lifecycle, continual improvement

Standards-to-artifacts translation:

  • NIST AI RMF -> AI Control Pack, risk register, eval gate, monitoring plan.
  • EU AI Act -> risk classification memo, transparency checklist, human oversight design, audit evidence.
  • IIBA -> stakeholder map, elicitation plan, requirements model, solution evaluation report.
  • TOGAF -> capability map, target architecture, transition roadmap, architecture governance pack.
  • BPMN -> current-state and future-state workflow with SLA, exception, handoff, control point.
  • BIAN -> banking capability map, service domain boundary, semantic API and integration scope.
  • OWASP LLM Top 10 -> threat model, prompt/tool/data controls, red team backlog.
  • ISO/IEC 42001 -> AI management operating model, responsibility matrix, lifecycle review cadence.

4. Cross-Capability Matrix

Legend:

  • A = Accountable, owns final quality.
  • C = Contributor, provides inputs.
  • R = Reviewer, sets guardrails or acceptance.
  • P = Practitioner, implements or operates.
Capability业务分析产品经营AI Solutions ArchitectAI Enterprise ArchitectEvalOpsFDE
Business problem framingAACRCC
Stakeholder discoveryACCRCA
Process modelingACCRCC
Domain modelingACAACC
Data readinessACARAP
RAG / agent architectureCRARCP
Evaluation designAACRAP
Governance / riskCRAAAC
ROI / business caseCACACC
Change adoptionCACACC
Vendor assessmentCAAACC
System evidence synthesisCACACC

4.1 Capability output standards

CapabilityMinimum artifactStrong artifact
Business problem framingProblem statement with measurable baselineOpportunity Canvas with alternatives, no-AI option, risk-adjusted value
Stakeholder discoveryStakeholder listEvidence map with influence, pain, decision power, objections, review evidence
Process modelingHappy-path flowBPMN with exception flows, waits, rework, SLA, control points, handoffs
Domain modelingEntity listDomain model with bounded contexts, BIAN/service domain mapping, policy concepts
Data readinessData source inventoryData readiness pack with quality, lineage, labels, PII, access, retention, freshness
RAG / agent architectureHigh-level diagramADR set with option analysis, reversal triggers, cost, latency, security, fallback
Evaluation designManual test casesRequirements-to-Eval Matrix with gold data, graders, thresholds, owners, sampling
Governance / riskChecklistControl pack mapped to NIST/EU/ISO/OWASP with release gates and audit evidence
ROI / business caseBenefit guessBaseline, unit economics, sensitivity, funding stages, risk-adjusted ROI
Change adoptionTraining planAdoption dashboard, SOP, RACI, champion model, feedback and escalation loop
Vendor assessmentFeature comparisonDue diligence pack with data, security, legal, model, cost, lock-in, exit plan
System evidence synthesisFolder of docs10-minute narrative: problem, evidence, decisions, architecture, evals, results

5. L1-L5 Capability Growth Levels

5.1 业务分析能力

L1 - AI-aware problem framing

  • 行为边界: 能描述 AI use case, 区分 automation, copilot, workflow, agent.
  • 产出: 简单 problem statement, stakeholder list, user story draft.
  • 系统证据: 1 个 AI Opportunity Canvas v0.1, 1 份访谈问题清单.
  • 能力证据: 能说明 "为什么不一定需要 AI", 能把需求改写成可测指标.
  • 常见失误: 把业务要求直接变成 prompt; 只写功能, 不写流程和证据.

L2 - Evidence-driven workflow analysis

  • 行为边界: 主动追问 baseline, volume, error rate, SLA, rework, risk, exception.
  • 产出: Stakeholder Evidence Map, BPMN current-state, pain metrics baseline.
  • 系统证据: 3 个业务场景的流程图和痛点量化表.
  • 能力证据: 能解释 stakeholder conflict, human-in-the-loop, no-AI boundary.
  • 常见失误: 只访谈 manager, 不访谈一线; 流程图没有异常流和控制点.

L3 - Requirements-to-Eval analysis

  • 行为边界: 把每条需求绑定 eval data, expected behavior, threshold, owner, fallback.
  • 产出: Requirements-to-Eval Matrix, decision model, data readiness inputs.
  • 系统证据: 1 套 golden dataset, 20-50 条测试样本, 需求追溯表.
  • 能力证据: 能说明 "需求如何验收", "模型失败时谁处理", "哪些样本最难".
  • 常见失误: 只写 acceptance criteria, 不写 eval design; 把模型准确率当唯一指标.

L4 - AI transformation analysis

  • 行为边界: 能跨业务, 风控, 法务, 数据, IT 推动 workflow redesign.
  • 产出: future-state BPMN, operating model, RACI, adoption risk register.
  • 系统证据: 端到端 case pack: problem -> workflow -> requirements -> eval -> adoption.
  • 能力证据: 能讲清某个 AI 项目上线后岗位, SOP, control, escalation 如何变化.
  • 常见失误: 忽略组织阻力; 只做系统需求, 不设计工作方式变化.

L5 - AI business architecture

  • 行为边界: 能用 capability, value stream, domain model 设计 AI transformation portfolio.
  • 产出: AI capability map, decision inventory, domain service map, governance roadmap.
  • 系统证据: 3-5 个金融零售场景组成的 evidence system, 每个都有 evidence artifact.
  • 能力证据: 能从战略, 业务能力, 风险, 架构, adoption 讲一个完整 transformation story.
  • 常见失误: 叙事过大但证据不足; 只讲愿景, 不能落到 eval 和 operating model.

5.2 产品经营能力

L1 - AI feature framing

  • 行为边界: 能定义用户, use case, basic PRD, prompt/prototype demo.
  • 产出: AI Product Brief, user journey, simple success metrics.
  • 系统证据: 一个可点击原型或 demo recording, 一页 PRD.
  • 能力证据: 能说明用户任务, 输入输出, basic constraints.
  • 常见失误: 把 "加 AI" 当价值; demo 有趣但没有 adoption path.

L2 - Workflow product operating model

  • 行为边界: 从用户 job 和 workflow friction 定义产品机会, 而不是从模型能力出发.
  • 产出: JTBD, metric tree, MVP scope, prototype test plan.
  • 系统证据: 5 人 usability test notes, before/after workflow comparison.
  • 能力证据: 能讲清用户何时信任 AI, 何时需要人工复核, 何时拒绝使用.
  • 常见失误: 只优化 UI, 不改变 workflow; 只看 usage, 不看质量和风险.

L3 - Eval-native product operations

  • 行为边界: 将产品需求写成 eval-backed requirements, 定义 launch gate.
  • 产出: PRD + Requirements-to-Eval Matrix, release criteria, adoption dashboard.
  • 系统证据: 1 个产品的 metric tree, eval suite, release decision memo.
  • 能力证据: 能平衡 precision/recall, time saved, trust, error cost, compliance risk.
  • 常见失误: 只追模型 benchmark; 不定义 bad answer 的业务损失.

L4 - AI growth and operations

  • 行为边界: 管理从 pilot 到 rollout 的 adoption funnel, cohort, training, feedback loop.
  • 产出: rollout plan, SOP, training, feedback taxonomy, ROI model.
  • 系统证据: 30/60/90 launch dashboard, user cohort analysis, business case.
  • 能力证据: 能解释 why pilot succeeds but enterprise rollout fails.
  • 常见失误: pilot 后不跟踪行为改变; 忽略 manager incentives 和 process ownership.

L5 - AI Product Strategy Lead

  • 行为边界: 能定义 AI product portfolio, build-vs-buy, platform leverage, moat.
  • 产出: product strategy, portfolio roadmap, vendor strategy, executive narrative.
  • 系统证据: 3 个产品线的 prioritization model, investment memo, evidence map.
  • 能力证据: 能连接市场机会, enterprise architecture, governance cost, product velocity.
  • 常见失误: 只讲愿景和路线图; 无法证明优先级和 ROI 假设.

5.3 AI Solutions Architect

L1 - AI application architect

  • 行为边界: 能画基本 LLM app 架构, 了解 model API, vector DB, RAG, prompt layer.
  • 产出: high-level C4, component list, integration assumptions.
  • 系统证据: 一个 RAG 或 copilot demo 的 architecture note.
  • 能力证据: 能解释 RAG 和 fine-tuning 的基本取舍.
  • 常见失误: 架构图只有 LLM box; 忽略权限, 审计, fallback, latency.

L2 - Enterprise RAG / workflow architect

  • 行为边界: 能设计 ingestion, chunking, metadata, retrieval, rerank, citation, RBAC.
  • 产出: RAG ADR, data flow, security design, observability events.
  • 系统证据: Enterprise RAG ADR, eval sample, source citation policy.
  • 能力证据: 能说明 freshness, source-of-truth, access control, grounding 失败处理.
  • 常见失误: 把 vector search 当知识治理; 没有 data lifecycle 和 content ownership.

L3 - Agentic architecture architect

  • 行为边界: 能设计 tool gateway, policy check, state, memory, approval, idempotency, audit.
  • 产出: agent workflow C4, tool risk classification, control pack, cost model.
  • 系统证据: agent ADR set, OWASP LLM threat model, eval and red team plan.
  • 能力证据: 能解释 excessive agency, prompt injection, tool abuse, human approval boundary.
  • 常见失误: 过早多 agent; agent 可直接执行高风险动作; 无停止条件.

L4 - Production AI architect

  • 行为边界: 能管理 model gateway, prompt/versioning, telemetry, eval gate, incident response.
  • 产出: production readiness review, SLO/SLI, runbook, governance integration.
  • 系统证据: release gate checklist, trace design, incident review template.
  • 能力证据: 能从 reliability, security, compliance, cost, data quality 讲上线条件.
  • 常见失误: 只做 PoC; 没有 rollback, audit trail, budget guardrail, vendor exit.

L5 - AI platform and transformation architect

  • 行为边界: 能抽象 reusable AI platform capabilities, 支持多个业务域的受控复用.
  • 产出: reference architecture, platform capability map, standards, migration roadmap.
  • 系统证据: AI platform blueprint, build-vs-buy matrix, multi-use-case control model.
  • 能力证据: 能说明何时平台化, 何时场景定制, 如何降低 marginal delivery cost.
  • 常见失误: 平台先行但无业务拉动; 标准过重导致交付停滞.

5.4 AI Enterprise Architect

L1 - AI-aware EA

  • 行为边界: 能把 AI 项目放入业务能力, 应用, 数据, 技术视图.
  • 产出: simple capability map, context diagram, architecture principles draft.
  • 系统证据: 一个 AI use case 的 business/data/application/technology view.
  • 能力证据: 能解释 AI capability 和普通 application capability 的差异.
  • 常见失误: 把 AI 当单独应用; 不关心业务能力和 operating model.

L2 - Capability and portfolio EA

  • 行为边界: 用 capability-based planning 对 AI use cases 分组和排优先级.
  • 产出: AI capability heatmap, use case evidence map, dependency map.
  • 系统证据: 10 个金融零售 AI use cases 的 scorecard 和 roadmap.
  • 能力证据: 能说明 value, risk, readiness, reuse, regulatory exposure 的权衡.
  • 常见失误: 按部门需求排队; 不评估数据和治理 readiness.

L3 - Target architecture EA

  • 行为边界: 设计 enterprise target architecture: model gateway, knowledge, eval, governance, observability.
  • 产出: target architecture, transition states, standards, architecture review checklist.
  • 系统证据: TOGAF-style roadmap, architecture decision log, transition architecture.
  • 能力证据: 能连接 business strategy, data architecture, security, product delivery.
  • 常见失误: 架构原则抽象; 无迁移路径和 funding logic.

L4 - AI governance EA

  • 行为边界: 将 NIST/EU/ISO/OWASP 映射为企业级政策, 控制, review cadence.
  • 产出: AI governance model, risk tiering, architecture review board pack.
  • 系统证据: risk classification matrix, control library, compliance evidence map.
  • 能力证据: 能解释 high-risk AI, human oversight, transparency, model lifecycle.
  • 常见失误: 合规 checklist 化; 没有 owner, evidence, exception handling.

L5 - Enterprise AI transformation EA

  • 行为边界: 能设计从 local pilots 到 enterprise AI operating model 的演进.
  • 产出: 18-month roadmap, platform operating model, investment portfolio, capability maturity model.
  • 系统证据: multi-domain transformation deck with financial retail examples and measurable gates.
  • 能力证据: 能和 CIO/CTO/CRO/COO 讨论组织能力, 技术架构, 风险和价值.
  • 常见失误: 太战略, 不可执行; 不知道一线 adoption 和 eval data 如何生产.

5.5 AI Product Operations / EvalOps

L1 - AI QA operator

  • 行为边界: 能执行人工测试, 记录 bad outputs, 跟踪问题.
  • 产出: test checklist, bug report, sample failure log.
  • 系统证据: 20 条 AI 输出测试记录和分类.
  • 能力证据: 能描述 hallucination, refusal, grounding failure, unsafe answer.
  • 常见失误: 只看 pass/fail, 不建 failure taxonomy.

L2 - Eval analyst

  • 行为边界: 能设计 gold questions, labels, rubrics, human review protocol.
  • 产出: golden dataset, evaluation rubric, sampling plan.
  • 系统证据: 50-100 条样本, grader rubric, reviewer agreement notes.
  • 能力证据: 能解释 exact match, semantic grading, human review, adversarial cases.
  • 常见失误: 样本太容易; 没有 hard cases, negative cases, edge cases.

L3 - Release gate owner

  • 行为边界: 将 eval 连接 CI/release, product metrics, monitoring, incident workflow.
  • 产出: release gate, daily eval dashboard, regression test suite.
  • 系统证据: evaluation report with threshold, trend, release decision.
  • 能力证据: 能说明何时不发布, 何时回滚, 何时增加人工审核.
  • 常见失误: 阈值随意; release gate 与业务风险不匹配.

L4 - AI quality operations lead

  • 行为边界: 建立 production feedback loop: trace, sampling, drift, complaint, incident, retraining input.
  • 产出: EvalOps runbook, telemetry schema, incident review, change control.
  • 系统证据: 30-day production quality review, failure taxonomy trend, corrective actions.
  • 能力证据: 能讲清 AI system quality 如何持续维护, 而不是上线即结束.
  • 常见失误: 只做离线 eval; 不监控 real workflow impact 和 user trust.

L5 - Enterprise EvalOps architect

  • 行为边界: 定义跨产品 eval platform, rubric standards, governance reporting, quality portfolio.
  • 产出: enterprise eval framework, model/product scorecard, risk-tiered eval policy.
  • 系统证据: 多场景 eval library, evaluator calibration protocol, executive quality dashboard.
  • 能力证据: 能将 eval 变成企业 AI 管理体系的一部分.
  • 常见失误: eval 平台化过早; 忽略每个业务域的不同 error cost.

5.6 Field AI Engineer / Forward Deployed Engineer

L1 - AI implementation engineer

  • 行为边界: 能根据明确需求接入 model API, RAG, simple workflow.
  • 产出: prototype, integration script, demo notes.
  • 系统证据: 一个 end-to-end demo with README and limitations.
  • 能力证据: 能解释实现边界和已知风险.
  • 常见失误: 只追 demo speed; 不记录假设和现场约束.

L2 - Discovery-to-prototype FDE

  • 行为边界: 在客户现场快速访谈, 抽取 workflow, 2-3 天做可试用原型.
  • 产出: field discovery notes, prototype, pilot test plan.
  • 系统证据: before/after workflow video or annotated screenshots, user feedback.
  • 能力证据: 能讲清如何从 ambiguous requirement 找到 first wedge.
  • 常见失误: 用户说什么做什么; 没有 problem framing 和 success metric.

L3 - Integration and deployment FDE

  • 行为边界: 能处理 SSO, RBAC, data connector, audit log, environment, monitoring.
  • 产出: integration adapter, deployment checklist, observability plan, handoff doc.
  • 系统证据: pilot deployment pack with risks, rollback, access control.
  • 能力证据: 能说明 enterprise environment 中最容易卡住的技术和组织问题.
  • 常见失误: 本地可跑但企业不可部署; 忽略 security review 和 data access.

L4 - Field product architect

  • 行为边界: 将多个客户现场反馈抽象为 reusable product / platform capability.
  • 产出: field pattern catalog, product gap analysis, reference implementation.
  • 系统证据: 3 个现场案例的共性需求和产品化建议.
  • 能力证据: 能在 customization 和 productization 之间做取舍.
  • 常见失误: 永远定制; 不能沉淀 reusable architecture.

L5 - Strategic FDE / Forward Deployed Architect

  • 行为边界: 能与客户高层定义 AI transformation case, 同时带队完成 pilot 到 production.
  • 产出: executive roadmap, joint success plan, production scale plan, value report.
  • 系统证据: end-to-end transformation case: discovery -> architecture -> pilot -> launch -> ROI.
  • 能力证据: 能跨 business, engineering, legal, security, operations 形成共同决策.
  • 常见失误: 现场英雄主义; 缺少可移交的 operating model 和 internal owner.

6. Financial Retail Examples

6.1 AML / KYC Investigation Copilot

  • Business problem: alert backlog 高, false positive 多, case narrative 质量不稳定.
  • 业务分析视角: alert triage workflow, investigator pain, SAR / case narrative decision rules.
  • 产品视角: investigator time saved, first-pass quality, reviewer acceptance, adoption.
  • Architect focus: evidence retrieval, case graph, policy RAG, audit trail, HITL approval.
  • EvalOps focus: groundedness, citation correctness, suspicious typology coverage, unsafe recommendation.
  • Governance focus: human oversight, explainability, audit evidence, model risk management.
  • 系统证据: AML Opportunity Canvas, BPMN, Requirements-to-Eval Matrix, AI Control Pack, SAR quality eval.

6.2 Lending Underwriting Assistant

  • Business problem: loan review cycle time 长, policy interpretation 不一致, manual exception 多.
  • 业务分析视角: underwriting decision inventory, policy exception flow, credit officer workflow.
  • 产品视角: approval cycle time, applicant experience, override rate, adverse action compliance.
  • Architect focus: policy RAG, scoring system integration, decision support only, audit log.
  • EvalOps focus: policy-grounded answer, missing-data detection, bias monitoring, manual review trigger.
  • Governance focus: high-risk AI exposure, fair lending, explainability, appeal process.
  • 系统证据: lending decision model, data readiness pack, human oversight memo, release gate.

6.3 Fraud Operations Copilot

  • Business problem: fraud queue volatility, chargeback response time, analyst overload.
  • 业务分析视角: case intake, evidence gathering, escalation, exception paths.
  • 产品视角: analyst throughput, false positive handling, case closure quality.
  • Architect focus: event stream, feature store, rule/model explainability, tool action limits.
  • EvalOps focus: hard-negative cases, emerging typology sampling, drift and incident review.
  • Governance focus: customer harm, account freeze controls, dual approval for high-impact actions.
  • 系统证据: fraud workflow BPMN, tool risk matrix, incident playbook, eval dashboard.

6.4 Service Copilot

  • Business problem: customer service AHT 高, policy answers inconsistent, escalation unclear.
  • 业务分析视角: intent taxonomy, knowledge source ownership, escalation criteria.
  • 产品视角: containment, customer satisfaction, agent trust, coaching loops.
  • Architect focus: contact center integration, retrieval permissions, response guardrails, citation.
  • EvalOps focus: answer correctness, tone, refusal, escalation, compliance phrase checks.
  • Governance focus: transparent AI use, privacy, complaint handling, human fallback.
  • 系统证据: service journey, RAG ADR, gold Q&A set, adoption dashboard.

6.5 Payments Operations

  • Business problem: payment exceptions, reconciliation breaks, refund disputes, settlement delays.
  • 业务分析视角: exception taxonomy, reconciliation decision flow, SLA and ownership.
  • 产品视角: ops productivity, break resolution time, error reduction, merchant experience.
  • Architect focus: payment gateway data, ledger constraints, idempotency, tool permissions.
  • EvalOps focus: numeric consistency, source traceability, action safety, regression samples.
  • Governance focus: financial loss, auditability, segregation of duties, change control.
  • 系统证据: payment exception BPMN, domain model, ADR for read-only vs write tools.

6.6 Wealth Advisory Compliance

  • Business problem: advisors need compliant research support, suitability review, disclosure consistency.
  • 业务分析视角: suitability constraints, advice boundary, disclosure workflow, review evidence.
  • 产品视角: advisor productivity, compliance confidence, client communication quality.
  • Architect focus: policy RAG, portfolio data access, capability-based guardrails, no personalized advice without controls.
  • EvalOps focus: suitability red flags, hallucinated product facts, prohibited recommendation language.
  • Governance focus: advisor oversight, regulatory disclosure, audit trail, model output retention.
  • 系统证据: wealth advisory control pack, eval rubric, human approval workflow, compliance memo.

7. Practical Evidence Artifacts from Existing ABPA Templates

Use the existing structures under docs/abpa/templates/ instead of inventing new formats.

TemplateBest capability evidenceHow to use it
01-ai-opportunity-canvas.md业务分析能力, 产品经营能力Prove problem framing, no-AI alternatives, measurable value
02-stakeholder-evidence-map.md业务分析能力, FDEMap business, risk, legal, data, IT, frontline objections
03-bpmn-pain-metrics.md业务分析能力, Enterprise ArchitectCapture current-state workflow, exception flow, pain metrics
04-requirements-to-eval-matrix.md业务分析能力, 产品经营能力, EvalOpsConvert requirements into eval data, rubric, threshold, owner
05-ai-control-pack.mdArchitect, EvalOps, Enterprise ArchitectMap NIST/EU/ISO/OWASP risks to controls and release gates
06-executive-decision-memo.md产品经营能力, Enterprise ArchitectSummarize decision, options, risks, next 30 days
07-data-readiness-pack.md业务分析能力, Architect, EvalOpsProve data quality, labels, access, PII, lineage, freshness
08-ai-architecture-adr-set.mdSolutions Architect, FDEDocument RAG, agent, model, gateway, HITL, audit choices
09-operating-model-raci.mdEnterprise Architect, 产品经营能力Assign owners for product, data, evals, risk, operations
10-adoption-dashboard.md产品经营能力, EvalOpsTrack usage, trust, quality, time saved, fallback, value
11-business-case.md产品经营能力, Enterprise ArchitectConvert workflow impact into cost, benefit, sensitivity, funding gate
12-portfolio-evidence-map.mdAll capabilitiesTurn notes, code, PRD, evals, diagrams into review evidence

Recommended evidence packages:

  • 业务分析证据包: Opportunity Canvas + Stakeholder Evidence Map + BPMN + Requirements-to-Eval Matrix.
  • 产品经营证据包: Product Brief + Metric Tree + Prototype Report + Adoption Dashboard + Business Case.
  • Solutions Architect package: Data Readiness Pack + Architecture ADR Set + Control Pack + C4 diagrams.
  • Enterprise Architect package: Capability Map + Target Architecture + Operating Model + Roadmap + Review Pack.
  • EvalOps package: Golden Dataset + Rubric + Release Gate + Failure Taxonomy + Production Quality Review.
  • FDE package: Field Notes + Prototype + Integration Adapter + Pilot Report + Handoff Pack.

Evidence conversion rule:

  • Every artifact should answer: what decision does this support?
  • Every artifact should show: what evidence backs it?
  • Every artifact should state: what uncertainty remains?
  • Every artifact should propose: what should be tested next?

8. 30 / 60 / 90 / 180-Day Growth Roadmap

First 30 days - foundation and capability clarity

  • Week 1: Read this matrix, ABPA README, source anchors, and existing AML Copilot docs.
  • Week 1 output: personal capability positioning memo: 业务分析能力 vs 产品经营能力 vs Architect vs EvalOps.
  • Week 2: Build one AML/KYC AI Opportunity Canvas and Stakeholder Evidence Map.
  • Week 2 output: 8 stakeholder groups, 10 evidence questions, 5 explicit assumptions.
  • Week 3: Create current-state BPMN for AML or payments operations.
  • Week 3 output: BPMN with exception flow, SLA, handoff, rework, control points.
  • Week 4: Convert 10 requirements into eval-ready form.
  • Week 4 output: Requirements-to-Eval Matrix with gold samples, thresholds, owners, failure modes.
  • 30-day review signal: 能从业务问题讲到 eval, 而不是只讲 AI 技术.

First 60 days - architecture and governance

  • Week 5: Produce Data Readiness Pack for AML, lending, or service copilot.
  • Week 6: Write Architecture ADR Set: RAG vs long context, workflow vs agent, model gateway, HITL.
  • Week 7: Map OWASP LLM Top 10 and NIST AI RMF risks into an AI Control Pack.
  • Week 8: Create one executive decision memo with go/no-go recommendation.
  • 60-day review signal: 能讲清 "why this architecture", "why now", "why safe enough", "what to monitor".

First 90 days - pilot and operations

  • Week 9: Build or document a small prototype, even if mock-based, tied to a real workflow.
  • Week 10: Design EvalOps release gate and failure taxonomy.
  • Week 11: Build Adoption Dashboard: usage, trust, quality, time saved, fallback, escalation.
  • Week 12: Write Business Case with baseline, cost, benefit, sensitivity, risk adjustment.
  • 90-day review signal: 能展示 one end-to-end case: problem -> workflow -> architecture -> eval -> launch gate -> business case.

First 180 days - portfolio and senior narrative

  • Days 91-120: Add two more scenarios: lending assistant and service copilot.
  • Days 121-140: Build enterprise AI capability map and target architecture across all scenarios.
  • Days 141-160: Create vendor assessment and build-vs-buy matrix for one scenario.
  • Days 161-180: Assemble evidence map and 10-minute executive story.
  • 180-day review signal: 能同时胜任 业务分析能力 / 产品经营能力 / Solutions Architect 讨论, 并能升级到 Enterprise Architect 视角.

9. Weekly Training Loop

Use one business scenario per week. Do not only read. Every week must produce artifact evidence.

Monday - Problem and baseline

  • Pick one scenario: AML/KYC, lending, fraud, service copilot, payments ops, wealth compliance.
  • Write one measurable problem statement.
  • Identify baseline metric: cycle time, false positive rate, AHT, error rate, rework, backlog, cost per case.
  • Write one no-AI alternative.

Tuesday - Stakeholders and workflow

  • Map users, approvers, risk owners, data owners, IT, audit, legal, operations.
  • Draft 8-12 evidence questions.
  • Draw current-state flow with exception path.
  • Mark pain, wait, handoff, control, evidence point.

Wednesday - Requirements and evals

  • Write 5-10 requirements.
  • For each requirement, define gold sample, expected behavior, grader, threshold, owner.
  • Add at least 3 negative or adversarial cases.
  • Define when the system must ask for human help.

Thursday - Architecture and data

  • Draft C4 context/container diagram in text or Mermaid.
  • Write 2 ADRs: one about RAG/knowledge, one about workflow/agent/action boundary.
  • Check data readiness: source, owner, access, freshness, labels, PII, lineage.
  • Identify one vendor or platform option and one build option.

Friday - Risk, governance, and operations

  • Map top risks to NIST AI RMF, EU AI Act logic, OWASP LLM Top 10, ISO/IEC 42001.
  • Define controls: RBAC, citation, HITL, approval, logging, rate limits, tool permissions, red team.
  • Define release gate and monitoring metrics.
  • Write incident scenario and rollback trigger.

Saturday - Business case and adoption

  • Estimate value: time saved, quality improvement, risk reduction, revenue, cost avoidance.
  • Estimate cost: model, engineering, data prep, review, vendor, governance, training.
  • Define adoption funnel: eligible users, active users, repeated users, trusted outputs, abandoned flows.
  • Create rollout plan: pilot group, training, champion, support, feedback cadence.

Sunday - System evidence synthesis

  • Write a 10-minute narrative.
  • Structure: situation, baseline, root cause, options, architecture, eval, risk, value, adoption, next step.
  • Update 12-portfolio-evidence-map.md with artifact links.
  • Write one design-review note: decision, evidence, risk, and next validation step.

10. Common Cross-Capability Failure Modes

  • Demo trap: building a slick prototype without business baseline or adoption plan.
  • Model trap: optimizing model score while ignoring workflow, data ownership, and human review.
  • Governance trap: writing compliance checklists with no owner, evidence, or release gate.
  • Architecture trap: drawing target architecture without transition states and funding logic.
  • Problem-framing trap: accepting stakeholder requests as truth without evidence and conflict analysis.
  • Product-operating trap: measuring usage while ignoring quality, trust, cost, and risk.
  • Eval trap: testing easy examples and missing hard negatives, adversarial inputs, policy boundaries.
  • FDE trap: over-customizing one client solution without productizing reusable patterns.
  • Enterprise trap: creating standards so heavy that teams bypass them.
  • Evidence trap: listing many artifacts without a crisp story of decisions and proof.

11. Self-Assessment Rubric

Score each capability from 1 to 5 every month.

ScoreMeaningEvidence required
1Can explain conceptNotes, glossary, simple example
2Can produce basic artifactOne structure filled with assumptions marked
3Can apply to realistic caseEvidence, metrics, trade-offs, eval, risk controls
4Can lead cross-functional decisionStakeholder conflicts, governance, adoption, business case
5Can generalize across portfolioReusable patterns, target architecture, standards, executive narrative

Monthly review: identify the strongest capability, the bottleneck capability, the weakest artifact evidence, the best financial retail proof scenario, the source anchor to refresh, and the review story to rehearse.


Use an evidence index with six tabs:

  1. Problem and Business Case.
  2. Workflow and Requirements.
  3. Data and Domain Model.
  4. AI Architecture and Governance.
  5. EvalOps and Production Quality.
  6. Adoption and Executive Story.

Each case should have:

  • One-page executive decision memo.
  • Current-state and future-state workflow.
  • Requirements-to-Eval Matrix.
  • Architecture ADR Set.
  • AI Control Pack.
  • Adoption Dashboard.
  • Business Case.
  • Design review script.

Minimum evidence system by 180 days:

  • AML/KYC Copilot: strongest governance and eval case.
  • Lending Assistant: strongest high-risk AI and policy reasoning case.
  • Payments Operations Copilot: strongest architecture and operational reliability case.
  • Service Copilot: strongest adoption and customer workflow case.
  • Wealth Advisory Compliance: strongest compliance, human oversight, and product boundary case.

13. Final Operating Principle

The target is not to become six separate people.

The target is to become the person who can translate across the six capabilities:

  • from business pain to measurable baseline,
  • from stakeholder conflict to decision evidence,
  • from process model to AI workflow,
  • from requirement to eval,
  • from architecture to control,
  • from pilot to adoption,
  • from artifact folder to executive story.

That translation ability is the durable advantage for 2026+ enterprise AI capabilities.