AI 面试作品集叙事 Playbook
把已有 AI、金融零售、架构、EvalOps、治理和 Web3 学习资产整理成可复盘、可审查、可持续升级的复杂系统经验。
复杂 AI 系统经验叙事与证据组织手册
核心导读: 把已有 AI、金融零售、架构、EvalOps、治理和 Web3 学习资产整理成可复盘、可审查、可持续升级的复杂系统经验。 定位: 本文件是系统叙事与证据组织层, 不替代旧学习计划, 不删除旧笔记, 不重写已有 case。它关注如何用证据说明系统边界、设计取舍、评测门槛、风险控制和组织采用。 使用范围: 学习和系统复盘材料, 不是法律意见、合规意见或监管解释。标准和法规只作为方法锚点。
0. 与既有资产的连接方式
复用资产:
docs/abpa/README.md: 业务问题、流程、评测、治理和运营模板入口。docs/AI_ROLE_COMPETENCY_MATRIX_2026.md: 需求流程视角、产品架构、解决方案架构、EvalOps、FDE 能力边界。- 金融零售 AI 深度案例库: 金融场景 case source。
docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md: Capability Map, BPMN, C4, RAG, Agent, Eval, Risk/Control 图谱.
本文件新增价值:
- 把 "学过什么" 转成 "能证明什么".
- 把每个 case 整理成系统摘要、复盘主线、深度审查三层材料.
- 把同一个 case 按能力视角切换表达.
- 把标准, 架构, eval, control, adoption, ROI 统一到同一条 evidence narrative spine.
核心原则:
Preserve old learning assets -> Select evidence -> Package evidence narrative -> Rehearse proof -> Refresh gaps
1. Source And Standard Anchors
这些 anchor 不需要背全文. 系统叙事中要把它们转成 artifact, control, eval, architecture decision, operating model.
| Anchor | 系统叙事中如何使用 | 可展示 artifact |
|---|---|---|
| NIST AI RMF | 用 Govern, Map, Measure, Manage 说明 AI risk lifecycle | AI Control Pack, release gate, monitoring plan |
| EU AI Act | 用 risk-based lens 说明 high-risk context, transparency, human oversight | Risk classification memo, audit evidence |
| ISO/IEC 42001 | 用 AI management system 语言讲 accountability 和 continual improvement | AI operating model, RACI, governance cadence |
| OWASP LLM Top 10 | 用 prompt injection, sensitive information disclosure, excessive agency 设计 controls | Threat model, red-team backlog, tool permission matrix |
| TOGAF | 用 capability, target architecture, roadmap, architecture governance 讲转型 | Capability map, architecture review pack |
| BPMN | 用 AS-IS, TO-BE, exception flow, handoff, human task 表达流程 | BPMN process map, pain metrics |
| BIAN | 用 banking service domain 组织金融零售 capability 和 integration boundary | Domain map, service boundary |
System Narrative translation:
- 不说 "熟悉 NIST AI RMF". 要说 "把 Govern, Map, Measure, Manage 转成 control register, eval gate, incident review 和 owner cadence".
- 不说 "了解 EU AI Act". 要说 "在 credit 或 wealth compliance 场景里, 先界定 decision boundary, human oversight, documentation 和 audit trail".
2. 能力视角叙事
| 能力视角 | 叙事重点 | 审查要点 | 主要证据 | 常见误区 |
|---|---|---|---|---|
| Solution Architecture | 将金融零售业务流程, 数据边界, RAG/Agent 架构, eval gate, risk control, observability 和 implementation constraints 连成可落地的 AI solution. | 能否从业务流程推导系统边界, 解释 RAG/agent/rules/vendor/custom build 取舍, 处理权限, 审计, latency, cost, fallback. | C4, ADR, data readiness, eval architecture, risk/control architecture, runbook. | "接模型 API, 用向量数据库, 做 chatbot." |
| 业务架构视角 | 将企业 AI 机会从单点 use case 提升到 capability map, value stream, operating model, governance 和 transformation roadmap. | 能否把高层 AI 战略落到业务能力, 处理效率/合规/数据/技术冲突, 组合 evidence. | AI capability map, case evidence, cross-case pattern, business case, RACI. | "推动公司用 AI, 大家都可以提需求." |
| 产品价值视角 | 将 AI 机会变成可验证的用户价值, MVP scope, success metrics, eval gate, launch plan 和 adoption loop. | 能否选择 worth-building opportunity, 定义 trust/quality/outcome/adoption/ROI, 平衡速度和风险. | Opportunity canvas, PRD, metric tree, requirements-to-eval, adoption dashboard, executive memo. | "设计 AI 功能, 用户可以聊天, 效率会提升." |
| 需求流程视角 | 将模糊 AI 需求转成业务问题证据, stakeholder map, BPMN workflow, decision rules, requirements-to-eval 和 acceptance criteria. | 能否区分 pain/want/constraint, 建模 exception path 和 handoff, 把需求变成 eval. | Stakeholder map, BPMN, requirements-to-eval, decision inventory, data readiness inputs. | "收集需求, 然后让技术团队实现 AI." |
| EvalOps/Product Ops | 将 AI quality 从一次性测试变成持续运营, 包括 golden dataset, eval suite, release gate, monitoring, feedback loop 和 incident review. | 能否定义上线门槛, 监控 drift/retrieval/unsupported claims/override/cost, 把反馈接回 roadmap. | Eval design, release gate, quality dashboard, feedback taxonomy, incident review. | "上线后看用户反馈, 有问题再调 prompt." |
| FDE/Forward Deployed AI Engineer | 能在客户现场从 ambiguous problem 出发, 快速做 discovery, prototype, integration, eval, pilot, handoff. | 能否处理 messy data, legacy systems, changing requirements, 并把 demo 变成可运营 pilot. | Field notes, prototype README, integration assumptions, pilot report, handoff checklist. | "现场快速帮客户做一个 AI 工具." |
能力视角的强表达:
- Solution Architecture: "LLM 不是系统中心, workflow, evidence, policy, eval, control 和 audit log 才是系统中心."
- 业务架构视角: "AI transformation 应讲成 capability change, 不是 tool rollout."
- 产品价值视角: "不能用 usage 替代 value, 要同时看 task success, quality, trust, control, and business impact."
- 需求流程视角: "需求流程视角 的核心不是写 prompt, 是把不确定需求变成可验证的业务和质量契约."
- EvalOps: "对 AI 产品来说, eval 是产品能力, 不是 QA 后置工作."
- FDE: "demo 必须推进到客户 workflow, eval gate, operating owner 和 handoff pack."
3. Reusable Evidence Narrative Framework
所有 evidence narrative 都用同一条主线:
Claim -> Evidence -> Architecture -> Eval -> Control -> Business Value -> Adoption -> Reflection
| Layer | 系统叙事中要说清楚什么 | Strong proof | Weak pattern |
|---|---|---|---|
| Claim | 你要证明的能力, 不是身份标签 | "我能把高风险金融运营设计成 HITL copilot." | "我熟悉 AI." |
| Evidence | 可展示, 可追问, 可复述的资产 | diagram, matrix, PRD, ADR, dashboard, prototype, eval set | 只有学习笔记 |
| Architecture | workflow, data, AI pattern, integration, control, operating boundary | C4, BPMN, RAG/Agent, risk architecture | 只有 LLM box |
| Eval | 需求如何变成上线门槛 | requirement, eval data, grader, threshold, owner, cadence | 只说准确率 |
| Control | 为什么能在 regulated setting 使用 | preventive, detective, corrective, governance controls | prompt 当控制 |
| Business Value | baseline 到 outcome 的变化 | cycle time, QA defect, backlog, loss avoided, ROI | "效率提升" 无基线 |
| Adoption | 用户如何安全改变工作方式 | activation, repeat usage, override, trust, manager cadence | 培训完成即 adoption |
| Reflection | 证据如何改变判断 | assumption -> evidence -> decision -> residual risk -> next step | 只讲成功 |
Minimum proof pack for each narrative:
- One artifact link.
- One metric or eval.
- One architecture decision.
- One risk control.
- One adoption or business value signal.
- One reflection sentence.
Architecture rehearsal question:
If this AI system fails, where would you notice it, who owns it, and how do you stop damage?
Eval categories to reuse:
- Groundedness and citation accuracy.
- Retrieval hit rate and source freshness.
- Red flag recall.
- Unsupported claim rate.
- Policy violation rate.
- Human override rate.
- Latency and cost per task.
- Adoption and trust.
Control categories to reuse:
- Preventive: RBAC, source allowlist, policy guardrail, tool permission, PII minimization.
- Detective: citation validation, output monitoring, audit sampling, anomaly alerts.
- Corrective: fallback, rollback, human escalation, incident review, eval set refresh.
- Governance: release gate, model change approval, owner cadence, risk acceptance.
System Narrative rule:
Never describe high-risk financial AI as autonomous final decisioning unless the case explicitly supports it.
Reflection structure:
Initial assumption -> Evidence change -> Decision made -> Residual risk -> Next improvement
4. Flagship Case Selection Matrix
| Case | Best role fit | Risk level | Best proof angle | First artifact to show |
|---|---|---|---|---|
| AML Copilot | Solution Architecture, 需求流程视角, EvalOps | High | Investigation workflow, red-flag eval, control pack | BPMN + requirements-to-eval |
| KYC Remediation | 需求流程视角, 产品价值视角, 业务架构 | High | Data quality, remediation workflow, customer outreach | Opportunity canvas + data readiness |
| Customer Service RAG | 产品价值视角, Solutions Architect | Medium | RAG governance, agent assist, adoption metrics | RAG architecture + adoption dashboard |
| Payments Exception Agent | FDE, Solutions Architect, 产品架构 | Medium | Exception workflow, tool-use boundary, ROI | Sequence diagram + control matrix |
| Lending Assistant | Solutions Architect, EvalOps, 业务架构 | High | Decision support, fair lending, reason codes | Control pack + ADR |
| Fraud Operations | FDE, 产品架构, Solutions Architect | High | Speed-control tradeoff, action agency levels | Agent workflow + permission matrix |
| Regulatory Change Impact | 业务架构, 业务分析, Solutions Architect | High | Capability impact, obligation mapping, governance | Capability heatmap + executive memo |
| Wealth Compliance Guardrail | 产品架构, EvalOps, Solutions Architect | High | Policy guardrails, advisor workflow, supervision | Risk/control architecture |
| Product Knowledge RAG | 产品架构, 业务分析, FDE | Low/Medium | Low-risk MVP, eval set, content governance | Data readiness + eval set |
| AI Governance/EvalOps Platform | EvalOps, 业务架构, Solutions Architect | Enterprise | Platform operating model, release gates, monitoring | Eval architecture + governance dashboard |
5. Flagship Evidence Cases
Case 01: AML Copilot
- 能力观察点: Solution Architecture, 需求流程视角, EvalOps/Product Ops, 业务架构视角.
- 系统经验摘要: AML Copilot 的核心不是让模型决定是否提交 SAR, 而是帮助 investigator 更快聚合证据, 检查 red flags, 生成有 citation 的 case narrative, 并保留 human approval 和 audit trail. 系统应设计成 RAG + workflow copilot + control pack, 用 evidence recall, citation precision, unsupported claim rate, QA rework 和 cycle time 衡量是否值得扩展.
- 复盘主线: Context: analyst 在 alerts, KYC, adverse media, sanctions, historical cases, SOP 之间切换. Claim: 高风险合规工作要做 decision support, 不是 autonomous decision engine. Architecture: Case UI -> orchestration -> rules/policy -> retrieval -> transaction/entity features -> summarizer -> reviewer workflow -> audit log. Eval: evidence coverage, red flag recall, citation precision, unsupported claim rate. Control: no autonomous SAR decision, maker-checker, RBAC, source citation, stop rule. Value: time-to-first-draft, QA defect, case completeness. Adoption: evidence summary pilot -> narrative draft.
- 深度复盘结构: Claim: 可审计的人机协同调查系统. Evidence: BPMN, requirements-to-eval, control pack, redacted before/after narrative. Architecture: evidence retrieval, checklist, draft, review, decision 分层. Eval: critical miss 不能被 overall accuracy 掩盖. Control: NIST AI RMF + OWASP LLM controls. Business Value: backlog, touch time, QA rework, audit issue. Adoption: trust 来自 citation, editable draft, override capture. Reflection: first feature should be evidence pack and missing evidence prompt.
- 系统证据材料:
docs/abpa/templates/03-bpmn-pain-metrics.md,docs/abpa/templates/04-requirements-to-eval-matrix.md,docs/abpa/templates/05-ai-control-pack.md,docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md. - 设计审查问题: How do you prevent hallucinated AML conclusions? What is the minimum release gate? What should the model never do? How do you handle analyst over-reliance?
- 常见误区: "AML 很适合 AI, 因为模型可以自动分析交易并判断可疑行为."
Case 02: KYC Remediation
- 能力观察点: 需求流程视角, 产品价值视角, 业务架构视角, FDE.
- 系统经验摘要: KYC remediation 的价值不在于生成客户邮件, 而在于把 data quality gaps, policy validation, document collection, customer outreach, reviewer approval 和 customer master update 串成可审计闭环. AI 做 gap classification, outreach draft, document extraction 和 prioritization, 高风险客户和 UBO 变更进入人工审批.
- 复盘主线: Context: 周期 review 或 regulatory remediation 中, UBO, tax, source of funds 等字段缺失. Claim: 数据质量和客户体验可设计成 AI-assisted remediation workflow. Architecture: data quality engine -> gap classifier -> risk queue -> outreach generator -> OCR/extraction -> policy RAG -> reviewer workbench -> golden source update. Eval: missing-field recall, extraction accuracy, invalid-document false accept, cycle time. Control: PII minimization, jurisdiction policy, approved messaging, source lineage. Value: backlog burn-down, days to complete, manual touches, deadline hit rate. Adoption: expired ID reminder 起步.
- 深度复盘结构: Claim: data + workflow + compliance + customer communication. Evidence: opportunity canvas, data readiness, BPMN, control register, dashboard. Architecture: classifier, RAG, OCR, CRM task, golden source update 分层. Eval: false accept 比 false reject 更危险. Control: high-risk customers, UBO changes, sanctions/PEP mandatory review. Business Value: SLA, backlog, contact success. Adoption: RM/Ops 需要 clear priority 和 status. Reflection: bottleneck 可能是 source-of-truth ownership.
- 系统证据材料:
docs/abpa/templates/01-ai-opportunity-canvas.md,docs/abpa/templates/07-data-readiness-pack.md,docs/abpa/templates/05-ai-control-pack.md,docs/abpa/templates/10-adoption-dashboard.md. - 设计审查问题: How do you handle jurisdiction policies? What data should not enter model context? How do you prevent invalid documents? What is the no-AI boundary?
- 常见误区: "AI 可以自动补齐客户 KYC 信息."
Case 03: Customer Service RAG
- 能力观察点: 产品价值视角, Solution Architecture, 需求流程视角, EvalOps.
- 系统经验摘要: Customer Service RAG 的目标不是建 generic chatbot, 而是让客服在已认证客户上下文和 approved knowledge base 之间更快给出准确, 有引用, 可审计的答复. 知识治理, metadata, effective date, citation, no-answer behavior, policy guardrail 和 QA feedback 必须放在产品核心.
- 复盘主线: Context: agents 在 CRM, core system, KB, SOP, complaint rules 之间切换. Claim: RAG 要从 demo 变成 governed agent-assist product. Architecture: desktop -> authenticated context -> intent detector -> knowledge RAG -> policy guardrail -> response composer -> QA analytics. Eval: correctness, citation coverage, retrieval recall, policy violation, no-answer correctness, AHT, CSAT. Control: no disclosure before authentication, RBAC retrieval, source versioning, stale content block. Value: AHT, FCR, transfer rate, QA defect, ramp time. Adoption: internal assist first.
- 深度复盘结构: Claim: service AI 是受控知识和客户上下文的工作流助手. Evidence: RAG spec, eval matrix, dashboard, memo. Architecture: RAG 不替代 authorized account tools. Eval: answer must satisfy correctness, citation, compliance, escalation. Control: no-answer path 比编造答案重要. Business Value: AHT and QA score together. Adoption: source link, editability, feedback. Reflection: key dependency is knowledge management maturity.
- 系统证据材料:
docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md,docs/abpa/templates/04-requirements-to-eval-matrix.md,docs/abpa/templates/08-ai-architecture-adr-set.md,docs/abpa/templates/10-adoption-dashboard.md. - 设计审查问题: How do you handle outdated articles? What if sources conflict? Agent-facing or customer-facing first? How do you prioritize MVP intents?
- 常见误区: "把知识库放进向量数据库, 客服就可以问问题."
Case 04: Payments Exception Agent
- 能力观察点: FDE, Solution Architecture, 产品价值视角, 需求流程视角.
- 系统经验摘要: Payments exception handling 适合展示 bounded agent thinking. AI 可以查询 payment status, 识别 return reason, 生成 investigation summary, 推荐 next action, pre-fill case tasks, 但 refund, reversal, manual repair, customer communication 必须受权限, rule engine 和 human approval 控制.
- 复盘主线: Context: ACH, card, wire, instant payment 或 merchant settlement exception 跨 gateway, ledger, network code, dispute system, CRM. Claim: agent 是 bounded tool-use workflow, 不是 unrestricted automation. Architecture: queue -> status tool -> return code RAG -> ledger check -> summarizer -> recommender -> human approval -> action API -> audit log. Eval: taxonomy, source accuracy, recommendation precision, wrong-action prevention, tool-call success. Control: tool allowlist, least privilege, idempotency, action risk tiers, reconciliation. Value: aging, handoffs, investigation time, customer resolution. Adoption: assistant -> pre-fill -> low-risk task creation.
- 深度复盘结构: Claim: agentic workflow 必须可控, 可回滚, 可审计. Evidence: sequence diagram, taxonomy, permission matrix, ROI model, runbook. Architecture: read-only tools 和 write/action tools 分层. Eval: final answer + tool path correctness. Control: excessive agency, injection, sensitive information disclosure 在 gateway 处理. Business Value: SLA and reconciliation breaks. Adoption: users need evidence, reason code, next step. Reflection: convenience 不能绕过 ledger controls.
- 系统证据材料:
docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md,docs/abpa/templates/03-bpmn-pain-metrics.md,docs/abpa/templates/08-ai-architecture-adr-set.md,docs/abpa/templates/11-business-case.md. - 设计审查问题: Which actions can be automated? How do you make tool calls idempotent? What if status sources disagree? How calculate ROI?
- 常见误区: "Agent 可以自动处理支付异常并通知客户."
Case 05: Lending Assistant
- 能力观察点: Solution Architecture, EvalOps, 业务架构视角, 产品价值视角.
- 系统经验摘要: Lending Assistant 不能被包装成黑盒 credit decision model. 系统应把 deterministic calculations, policy eligibility, document summarization, reason-code suggestion, memo drafting 和 fair lending controls 分层. AI 辅助 underwriter, 但 regulated credit decision, adverse action, exception approval 保留人类责任和审计证据.
- 复盘主线: Context: 贷款审批整合 application, income, debt, collateral, bureau, policy, exception memo. Claim: regulated decision support, not LLM final decision. Architecture: LOS -> document processing -> financial extraction -> deterministic rules -> policy RAG -> underwriting assistant -> reason-code service -> review -> decision record. Eval: extraction, policy citation, missing-risk detection, memo completeness, reason-code consistency, subgroup performance. Control: fair lending review, no protected-class proxy, deterministic calculations separated from prose. Value: faster review, fewer rework cycles. Adoption: already decisioned files -> shadow mode.
- 深度复盘结构: Claim: 区分 evidence, calculation, recommendation, decision. Evidence: requirements-to-eval, fair lending control pack, ADR, ROI. Architecture: rules/calculations deterministic, LLM handles explanation. Eval: reason-code consistency and subgroup error rates visible. Control: EU AI Act lens where applicable. Business Value: cycle time, committee prep, exception leakage. Adoption: citation, checklist, editable memo. Reflection: safer MVP is post-decision memo quality.
- 系统证据材料:
docs/abpa/templates/04-requirements-to-eval-matrix.md,docs/abpa/templates/05-ai-control-pack.md,docs/abpa/templates/07-data-readiness-pack.md,docs/abpa/templates/11-business-case.md. - 设计审查问题: How prevent unfair treatment? How separate rules from LLM output? What data excluded? What stops rollout?
- 常见误区: "模型可以根据客户资料判断是否批贷."
Case 06: Fraud Operations
- 能力观察点: FDE, 产品价值视角, Solution Architecture, EvalOps.
- 系统经验摘要: Fraud Operations 的难点是速度和控制并存. AI 可以解释 alert, 聚合 device/session/transaction evidence, 生成 customer contact script, 推荐 next step, 但冻结账户, 拒赔, law enforcement referral 等动作需要按风险等级进入 analyst confirm 或 supervisor approval.
- 复盘主线: Context: fraud ops 面对 CNP, ATO, mule, scam reimbursement, false positives. Claim: action agency 分成 suggest, pre-fill, human approve, execute-with-control. Architecture: console -> entity resolution -> feature store -> rules/model score -> explainer -> recommendation -> approval -> action API. Eval: missed fraud, false positive friction, recommendation precision, explanation usefulness. Control: least privilege, redaction, high-risk approval, hallucinated rationale monitoring. Value: faster resolution, loss prevented, lower false positive hold time. Adoption: case explainer -> recommendations -> pre-filled actions.
- 深度复盘结构: Claim: speed-control tradeoff. Evidence: TO-BE workflow, eval matrix, permission matrix, control pack, dashboard. Architecture: action gateway enforces risk tiers. Eval: false negative cost and false positive friction. Control: freeze, denial, legal escalation are gated. Business Value: loss avoided, MTTR, release time. Adoption: compact evidence timeline. Reflection: AI should not hide uncertainty.
- 系统证据材料:
docs/abpa/templates/03-bpmn-pain-metrics.md,docs/abpa/templates/04-requirements-to-eval-matrix.md,docs/abpa/templates/05-ai-control-pack.md,docs/abpa/templates/10-adoption-dashboard.md. - 设计审查问题: Which actions need approval? Cost of false positive vs false negative? How capture analyst feedback? How avoid hallucinated rationale?
- 常见误区: "Fraud 要快, 所以 AI 应该自动冻结和放行."
Case 07: Regulatory Change Impact
- 能力观察点: 业务架构视角, 需求流程视角, Solution Architecture, 产品价值视角.
- 系统经验摘要: Regulatory Change Impact 是展示 业务分析 + architecture 的强案例. AI 不是给法律结论, 而是辅助 clause extraction, obligation mapping, impacted capability/process/system/control analysis, owner assignment 和 executive decision memo. 最终 interpretation 和 approval 由 legal/compliance 完成.
- 复盘主线: Context: 新规则, enforcement action, regulator letter 或 internal policy change 需要判断影响哪些产品, 流程, 系统, 控制. Claim: 文本变化转成 capability, process, system, control, roadmap. Architecture: intake -> clause extractor -> RAG -> BIAN mapper -> impact heatmap -> backlog -> owner review -> memo. Eval: obligation recall, false impact rate, owner accuracy, citation, backlog completeness. Control: legal review, source hierarchy, versioned register, change board. Value: shorter cycle, fewer missed assets, reduced workshops, audit readiness. Adoption: internal policy pilot first.
- 深度复盘结构: Claim: structured impact analysis, not legal judgment. Evidence: heatmap, memo, RACI, control pack. Architecture: obligations map to internal assets, humans validate. Eval: missed obligation more severe than extra false impact. Control: citation, version, owner review, approval status. Business Value: speed, completeness, traceability. Adoption: one impact board. Reflection: connects text, policy, process, system, control.
- 系统证据材料:
docs/abpa/templates/06-executive-decision-memo.md,docs/abpa/templates/02-stakeholder-evidence-map.md,docs/abpa/templates/05-ai-control-pack.md,docs/abpa/templates/09-operating-model-raci.md. - 设计审查问题: How prevent legal advice? What is source hierarchy? How map text to systems? Who owns final approval?
- 常见误区: "AI 可以自动解读监管条文并告诉公司怎么改."
Case 08: Wealth Compliance Guardrail
- 能力观察点: 产品价值视角, EvalOps, Solution Architecture, 需求流程视角.
- 系统经验摘要: Wealth Compliance Guardrail 的好切入点不是让 AI 当投资顾问, 而是在 advisor 发出客户沟通或建议前做 suitability, disclosure, prohibited phrase, product eligibility 和 policy citation 检查. AI 可以做 pre-send review 和 rewrite suggestion, 但客户建议和交易指令仍由 licensed advisor 和 supervisory workflow 负责.
- 复盘主线: Context: advisor 解释产品, 风险, fee, suitability, market commentary, 但沟通不当会带来 mis-selling risk. Claim: AI guardrail 嵌入 workflow. Architecture: draft -> client/product context -> policy/suitability RAG -> prohibited detector -> rewrite -> advisor review -> supervisor sampling. Eval: violation detection, citation precision, false block, unsuitable recommendation detection. Control: suitability gates, templates, no unauthorized advice, supervision evidence. Value: lower defects, faster approval, fewer complaints. Adoption: generic draft checking -> client-specific suitability.
- 深度复盘结构: Claim: compliance-by-design AI product. Evidence: guardrail matrix, control pack, RACI, before/after email. Architecture: checks sit before send. Eval: false negatives breach, false positives kill adoption. Control: human advisor owns message, AI shows policy basis. Business Value: review time, defects, complaints. Adoption: actionable rewrite, not only blocked. Reflection: UX must be strict and useful.
- 系统证据材料:
docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md,docs/abpa/templates/04-requirements-to-eval-matrix.md,docs/abpa/templates/05-ai-control-pack.md,docs/abpa/templates/09-operating-model-raci.md. - 设计审查问题: Hard-block versus warn? How handle jurisdiction rules? How prevent unsuitable advice? What evidence do supervisors review?
- 常见误区: "AI 可以帮理财顾问写投资建议."
Case 09: Product Knowledge RAG
- 能力观察点: 产品价值视角, 需求流程视角, FDE, Solution Architecture.
- 系统经验摘要: Product Knowledge RAG 是最适合系统证据 MVP 的 case, 因为风险相对低, 数据可控, 演示性强. 但重点不是向量库 demo, 而是 content governance, metadata, effective date, source citation, conflict handling, no-answer behavior 和 eval set 的综合系统.
- 复盘主线: Context: 产品规则, fee schedule, eligibility, SOP, campaign terms 经常变化. Claim: 用低风险 case 展示 enterprise-grade RAG. Architecture: approved content -> ingestion -> chunking -> metadata -> hybrid retrieval -> rerank -> answer -> citation -> feedback queue. Eval: curated questions, expected source, citation precision, retrieval recall, stale content rejection, conflict detection. Control: content owner approval, retired content quarantine, RBAC retrieval, source allowlist. Value: lookup time, QA defect, search success, training time. Adoption: one product line, internal users.
- 深度复盘结构: Claim: RAG 从 PoC 转成 governed knowledge product. Evidence: data readiness, eval set, RAG ADR, dashboard, sample question-answer pairs. Architecture: source-of-truth remains document system. Eval: stale document traps and conflict traps. Control: no valid source means refuse or escalate. Business Value: faster lookup only counts with correctness. Adoption: knowledge owners own article fixes. Reflection: fastest credible first evidence narrative.
- 系统证据材料:
docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md,docs/abpa/templates/01-ai-opportunity-canvas.md,docs/abpa/templates/07-data-readiness-pack.md,docs/abpa/templates/04-requirements-to-eval-matrix.md. - 设计审查问题: What metadata is required? How handle effective dates? How choose chunking? What is content governance workflow?
- 常见误区: "把 PDF 放进向量库就能问答."
Case 10: AI Governance / EvalOps Platform
- 能力观察点: EvalOps/Product Ops, 业务架构视角, Solution Architecture, 产品价值视角.
- 系统经验摘要: AI Governance/EvalOps Platform 是把多个 AI case 规模化的中台案例叙事. 它不是审批官僚系统, 而是把 requirements, eval sets, release gates, monitoring signals, incidents, model/prompt versions, control evidence 和 adoption metrics 放进同一条 operating loop, 让 AI 产品持续可控.
- 复盘主线: Context: 多个 AI pilots 后, 质量口径不同, 控制证据分散, model/prompt 变更不可追踪. Claim: enterprise AI operating capability. Architecture: registry -> risk classification -> eval repository -> golden dataset -> eval runner -> release gate -> telemetry -> monitoring -> incident review. Eval: pass rate, regression failure, invalid citation, critical miss, override, latency, cost, feedback. Control: NIST AI RMF lifecycle, ISO/IEC 42001 accountability, OWASP controls, EU AI Act lens. Value: faster controlled rollout, fewer repeat defects, better audit readiness. Adoption: product teams get templates and dashboards.
- 深度复盘结构: Claim: governance should be operational platform, not policy PDF. Evidence: eval architecture, release gate, dashboard, incident review, RACI. Architecture: platform supports cases, business owners keep decision accountability. Eval: domain evals plus shared signals. Control: risk tier determines artifacts and cadence. Business Value: reduce duplicate work, accelerate compliant launch. Adoption: governance must help 产品架构 and engineers. Reflection: governance must be a product.
- 系统证据材料:
docs/AI_ROLE_COMPETENCY_MATRIX_2026.md,docs/AI_ARCHITECTURE_DIAGRAM_PLAYBOOK.md,docs/abpa/templates/04-requirements-to-eval-matrix.md,docs/abpa/templates/05-ai-control-pack.md,docs/abpa/templates/10-adoption-dashboard.md,docs/abpa/templates/12-claim-evidence-map.md. - 设计审查问题: How avoid slowing teams? What is centralized vs product-owned? Minimum artifacts by risk tier? How do incidents update evals?
- 常见误区: "我们需要一个 AI 治理委员会审批所有 AI 项目."
6. 系统证据目录结构
系统证据不是展示页面, 而是一套可被审查的学习资产组织方式。每个核心 case 都应能回答: 问题为什么重要、系统如何工作、质量如何验证、风险如何控制、上线后如何运营、哪些假设仍未证明。
建议目录:
ai-system-evidence/
README.md
cases/
aml-copilot/
case-record.md
process-map.md
requirements-to-eval.md
control-pack.md
architecture-adr.md
operating-metrics.md
reflection.md
product-knowledge-rag/
case-record.md
data-readiness.md
rag-architecture.md
eval-set.md
adoption-metrics.md
regulatory-change-impact/
case-record.md
capability-impact-map.md
decision-memo.md
operating-model-raci.md
standards-map/
nist-ai-rmf-map.md
owasp-llm-top10-controls.md
bpmn-bian-togaf-map.md
review-materials/
system-summaries.md
design-review-questions.md
evidence-index.md
README 只需要证明三件事:
- 你如何从真实业务损耗出发, 而不是从模型功能出发。
- 你如何把系统拆成流程、数据、权限、模型、工具、评测、控制和运营。
- 你如何用证据更新判断, 包括停止、降级、重做和扩展。
单个 case record 建议字段:
Case name:
Business problem:
Baseline and loss:
System boundary:
AI pattern:
Data and context:
Eval contract:
Controls:
Operating metrics:
Adoption path:
Known gaps:
Next evidence to collect:
审查规则:
- 每个 claim 必须链接到至少一个证据材料。
- 每个高风险动作必须有责任人、审批点和回退路径。
- 每个评测指标必须说明数据来源、阈值和失败处理。
- 每个架构图必须能说明被拒绝的替代方案。
- 每个运营指标必须触发某种动作, 不能只是展示数字。
7. 系统复盘演练
每日 45 分钟节奏:
- 选一个 case 和一个能力视角。
- 口头讲清系统摘要: 问题、边界、模式、证据、控制。
- 不看原文复盘主线: context、decision、architecture、eval、risk、operation。
- 回答 3 个设计审查问题。
- 打开一个证据材料, 说明它证明了什么, 不能证明什么。
- 写一条下一轮需要补充的证据。
每周轮换:
| Day | 能力视角 | Case focus | Artifact focus |
|---|---|---|---|
| Monday | 业务分析 | KYC, Regulatory Change, AML | BPMN, stakeholder map, requirements-to-eval |
| Tuesday | 产品架构 | Customer Service RAG, Product Knowledge RAG, Wealth Guardrail | PRD, metric tree, adoption dashboard |
| Wednesday | Solution Architecture | AML, Payments, Lending | C4, ADR, RAG/Agent architecture |
| Thursday | EvalOps / Product Ops | AML, Lending, AI Governance | Eval suite, release gate, monitoring |
| Friday | FDE | Payments, Fraud, Product Knowledge RAG | Prototype README, integration assumptions, pilot report |
| Saturday | 业务架构ure | Regulatory Change, AI Governance, KYC | Capability map, roadmap, operating model |
| Sunday | Evidence governance | Best 3 stories | README, evidence map, gap log |
设计审查问题:
- MVP 里最应该删掉什么, 为什么?
- 最危险的 failure mode 是什么?
- 哪些数据必须具备, 如果拿不到如何降级?
- 模型绝不能做什么?
- 哪个指标会触发暂停或回滚?
- 系统如何改变一线人员的日常流程?
- ROI 如何保守估计, 不夸大收益?
- 合规、风控、工程和运营分别会挑战什么?
- 一线用户不信任系统时, 证据和界面如何设计?
回答质量检查:
- 是否命名了具体业务损耗?
- 是否展示了证据, 而不是只给观点?
- 是否说明人工责任边界?
- 是否描述 eval 和 controls?
- 是否连接到业务价值?
- 是否提到 adoption 和运营节奏?
- 是否包含 trade-off?
- 是否避免把受监管决策描述成全自动系统?
8. 缺口修复计划
| 能力域 | 常见强项 | 常见缺口 | 修复资产 |
|---|---|---|---|
| 业务分析 | 流程、领域、需求思维 | stakeholder evidence 和 exception flow 不够 | 3 张带 pain metric 的 BPMN |
| 产品架构 | case selection 和 business framing | 指标、adoption baseline 和 stop rule 不够 | 3 张 metric tree 与 adoption dashboard |
| Solution Architecture | 架构学习资产丰富 | ADR 缺少取舍、反转条件和 failure mode | 3 份 architecture ADR set |
| EvalOps / Product Ops | eval 概念清楚 | 样本、阈值、release gate 仍偏抽象 | 50 条 golden set sample |
| FDE | 现场问题拆解能力 | prototype 与 integration assumptions 不够 | Product Knowledge RAG demo README |
| 业务架构ure | 金融零售和架构背景 | capability roadmap 与治理节奏需更清楚 | AI capability map 与 12-month roadmap |
最小证据包:
- 3 个完整 case study。
- 1 份 evidence map。
- 每个核心 case 1 张架构图。
- 每个核心 case 1 张 requirements-to-eval matrix。
- 高风险 case 1 份 control pack。
- 1 份 adoption dashboard draft。
- 1 份 business case 或 ROI model。
- 每个 case 1 份 reflection note。
最快补强动作:
- 给每个 case 补系统摘要。
- 给高风险 case 补 decision boundary。
- 给每个 case 补 eval table。
- 给每个 case 补 control table。
- 给每个 case 补 business metric 和 adoption metric。
- 给旧 Web3、架构和 AI 笔记加 freshness label。
Freshness labels:
- Current: 可直接进入系统复盘。
- Needs refresh: 有价值, 但缺 2026 context、metric、diagram 或 eval。
- Historical: 保留为学习记录, 仅作为背景资产。
- Draft: 需要补证据后再使用。
治理规则:
Never delete historical learning assets. Add a freshness label, update note, and evidence angle.
9. 30/60/90 天系统训练路线
Day 1-30: 证据基础
目标: 把原始学习资产整理成可审查的证据。
Weekly focus:
- Week 1: 选择能力主线和 top 10 claims。
- Week 2: 选择 3 个 flagship cases 并映射证据。
- Week 3: 写系统摘要和复盘主线。
- Week 4: 完成第一版 evidence index 与 README。
Concrete outputs:
- 3 条能力主线 positioning statements。
- Claim-to-evidence matrix。
- 10 个 case index。
- 3 个 deep-dive case pages。
- 1 个 standards anchor page。
- 第一轮系统复盘记录。
Quality gate:
- 每个 claim 至少有两个证据资产。
- 每个 flagship case 包含 eval、control、business value、adoption。
- 没有 case 只依赖“我学过某概念”。
Day 31-60: 深度审查能力
目标: 让 case 经得起 design review、risk review 和 architecture review。
Weekly focus:
- Week 5: 为 3 个核心 case 创建架构图。
- Week 6: 创建 requirements-to-eval matrices 和 sample eval cases。
- Week 7: 创建 control packs 和 governance mapping。
- Week 8: 创建 adoption dashboards 和 ROI assumptions。
Concrete outputs:
- AML Copilot deep-dive pack。
- Product Knowledge RAG 或 Customer Service RAG deep-dive pack。
- Regulatory Change 或 AI Governance deep-dive pack。
- 30 个设计审查问题与回答。
- Gap remediation tracker。
Quality gate:
- 能解释每个架构选择的理由。
- 每个 case 至少说明一个被拒绝的替代方案。
- 能说出 release gate 和 stop rule。
- 能明确模型不得执行的动作。
Day 61-90: 复杂系统复盘
目标: 把证据转成稳定的系统判断力, 并根据外部反馈迭代。
Weekly focus:
- Week 9: 按能力视角做系统复盘。
- Week 10: 完成 evidence v1 和 README。
- Week 11: 收集外部评审反馈。
- Week 12: 修复弱 case 并补缺失证据。
- Week 13: 准备最终 deep review。
Concrete outputs:
- Public evidence v1。
- GitHub README 或 evidence README。
- System review question set。
- 外部反馈记录。
- Evidence v1.1 with fixes。
Quality gate:
- 能在 10 秒内打开任一关键证据。
- 能说明自己在 case 中完成的具体工作。
- 能用 failure modes 回答系统如何失败。
- 能用 eval 和 production signals 说明系统是否有效。
- 能把能力 claim 绑定到证据, 而不是绑定到身份标签。
10. 能力视角总结语料
Solution Architecture:
My edge is connecting regulated financial workflows to AI architecture that is measurable and controllable. I can move from BPMN and domain requirements to RAG or agent architecture, ADRs, eval gates, audit logs, observability and rollout controls.
业务架构ure:
My edge is translating AI ambition into capability change. I can map where AI creates value, what workflows and controls must change, which cases should be prioritized, and how governance and adoption scale across a system evidence set.
产品架构:
My edge is productizing AI with eval and adoption built in. I define the user problem, MVP, quality gates, risk controls, launch metrics and feedback loop, so AI becomes a measurable product outcome rather than a demo.
业务分析:
My edge is turning ambiguous AI ideas into testable requirements. I use stakeholder evidence, workflow mapping, decision boundaries, requirements-to-eval and control requirements to make AI delivery concrete.
EvalOps / Product Ops:
My edge is making AI quality operational. I design golden datasets, release gates, monitoring signals, incident reviews and governance dashboards that keep AI systems useful after launch.
FDE:
My edge is working from messy real-world constraints to a working pilot. I can do field discovery, prototype quickly, integrate with systems, capture user feedback, define evals and hand off an operating solution.
11. 最终操作原则
The evidence should not say:
I studied AI, Web3, architecture and product management.
It should prove:
I can identify high-value AI opportunities, turn them into requirements and architecture, evaluate quality, control risk, measure business value, drive adoption, and explain tradeoffs clearly under system review pressure.
The strongest case is the one where a reviewer can see:
- why the business problem matters.
- what evidence you used.
- how the architecture works.
- how quality is evaluated.
- how risk is controlled.
- how value is measured.
- how users adopt it.
- what you learned and would improve next.