AI Procurement Intake:供应商评估沙盒与 Build-Buy 架构
AI procurement intake 是企业 AI 架构的第一道控制点:它把业务 idea、vendor pitch 和本地试点转成可测量的 outcome、workflow、data boundary、risk tier、sandbox evidence 和 build-buy decision。成熟组织不会让供应商 demo 定义问题,而是先定义用例是否值得进入 AI option sp
AI 采购入口 / 供应商评估沙盒 / Build-Buy 决策架构解读 (AI Procurement Intake / Vendor Evaluation Sandbox / Build-Buy Decision Architecture)
配对阅读:本篇的操作手册版(模板/RACI/门禁/runbook)是
docs/AI_PROCUREMENT_INTAKE_VENDOR_EVALUATION_SANDBOX_BUILD_BUY_PLAYBOOK.md。第一遍读本篇建立原理与架构判断;第二遍做案例时再用 playbook 查表落地,两者不需要重复精读。
Date: 2026-06-30
Status: evergreen
Audience: experienced CBAP / 金融零售 AI 产品与架构从业者 / AI governance lead / third-party risk partner
核心导读
AI procurement intake 是企业 AI 架构的第一道控制点:它把业务 idea、vendor pitch 和本地试点转成可测量的 outcome、workflow、data boundary、risk tier、sandbox evidence 和 build-buy decision。成熟组织不会让供应商 demo 定义问题,而是先定义用例是否值得进入 AI option space,再用受控沙盒比较 build、buy、partner、hybrid 或 stop。
问题定义
AI 采购入口不是行政流程, 而是企业 AI 架构的第一道控制点。它决定一个 AI idea 会不会被错误地推入 vendor demo、PoC、采购谈判或生产集成。成熟机构不会先问"哪家供应商最好", 而是先问:
| 控制问题 | 架构含义 | 金融零售后果 |
|---|---|---|
| 这个用例是否值得进入 AI funnel | 先验证 outcome, workflow, data, risk tier, no-AI option | 避免把普通流程问题包装成 GenAI 项目 |
| AI 在流程中扮演什么角色 | read, summarize, recommend, draft, decide, act 分层 | 避免客服、信贷、AML、支付场景中越权决策 |
| 应该 build, buy, partner, hybrid 还是 stop | 将能力差异化、控制权、时间、成本、风险放进同一决策 | 避免因 demo 好看而买入不适合的黑盒 |
| sandbox 评估什么 | 用真实但受控的数据、任务、rubric 和门禁测试供应商 | 避免只看销售演示和通用 benchmark |
| 证据如何进入生产放行 | 评估结果必须连接 ADR, risk acceptance, release gate | 避免 PoC 成功后绕过安全、隐私、模型风险和运营控制 |
上游 procurement intake 的核心价值:
- 把 AI idea 变成可测量的业务和风险假设。
- 把 vendor comparison 变成 architecture comparison。
- 把 PoC 变成受控 sandbox, 不让试点自然漂移成 shadow production。
- 把 build-buy 选择从偏好讨论升级为 evidence-based decision。
- 把后续合同、退出、投资叙事和生产上线建立在证据之上。
边界说明: 本文聚焦 procurement lifecycle 上游的 intake, triage, sandbox, benchmark 和 build-buy decision。合同条款、退出迁移和董事会投资 narrative 属于下游材料, 这里只定义它们需要的输入证据。
架构模型/核心原理
flowchart LR
A[AI idea / vendor pitch / business pain] --> B[Intake funnel]
B --> C{Use-case triage}
C -->|No AI fit| C1[Process / rules / data fix]
C -->|Low value or high risk| C2[Stop or defer]
C -->|Candidate| D[Decision architecture]
D --> D1[Build]
D --> D2[Buy]
D --> D3[Partner]
D --> D4[Hybrid]
D --> E[Sandbox charter]
E --> F[Vendor and internal option benchmark]
F --> G[Evidence pack]
G --> H{Architecture review board}
H -->|No-go| H1[Reject / redesign]
H -->|Limited pilot| H2[Controlled pilot with constraints]
H -->|Production candidate| I[Production promotion gate]
I --> J[Contract, security, privacy, risk, model validation, operating model]
一条实用原则:
Intake owns the question "should this enter the AI option space"; sandbox owns "which option works under controlled evidence"; architecture gate owns "can this be operated safely at production scale".
关键机制
1. Intake Funnel
AI intake 要求每个 idea 先提交最小证据, 而不是直接预约 vendor demo。
| Intake field | 高级要求 | 不合格信号 |
|---|---|---|
| Business outcome | 明确 baseline, target movement, impacted workflow, owner | "提升效率", "智能化", "更懂客户" |
| AI role | read / retrieve / summarize / draft / recommend / decide / act | 直接说"AI 自动处理" |
| Customer or regulatory impact | 是否影响客户承诺、授信、KYC、AML、支付、投诉、收费、适当性 | 只说内部工具所以低风险 |
| Data boundary | 数据源、PII、PCI、账户、交易、文档、语音、日志、跨境、保留 | 不知道会把什么发给供应商 |
| Workflow insertion point | AS-IS / TO-BE 节点, human review, exception path, fallback | 没有流程图, 只有功能清单 |
| No-AI alternative | 流程优化、规则引擎、搜索、RPA、知识治理、报表 | 默认 AI 是唯一方案 |
| Evidence plan | sandbox 数据、rubric、benchmark、control evidence | 只打算看供应商 demo |
2. Triage Gate
把用例分成四类:
| Tier | 典型场景 | 推荐动作 |
|---|---|---|
| T0: Reject / defer | 没有 owner、没有 baseline、数据不可用、风险不可接受 | stop, 先补流程或数据 |
| T1: Learn-only sandbox | 价值假设早期, 数据可脱敏, 不连接生产 | controlled demo and architecture learning |
| T2: Controlled pilot candidate | 明确 workflow, 有评估集, 人类保留最终权责 | sandbox -> limited pilot |
| T3: High-impact candidate | 信贷、AML、KYC、支付、客户承诺、投诉或监管证据 | sandbox 加强, independent challenge, production gate |
3. Sandbox Charter
Sandbox charter 必须在供应商测试前冻结:
| Section | 内容 |
|---|---|
| Scope | 具体流程、用户角色、允许任务、禁止任务 |
| Data | synthetic, masked, historical, gold set, red-team set, access policy |
| Architecture | vendor route, internal baseline, model gateway, logging, retrieval, tool boundary |
| Evaluation | benchmark task, rubric, thresholds, critical failures, slice analysis |
| Controls | privacy, security, HITL, DLP, prompt injection, cost cap, kill switch |
| Evidence | trace, logs, output samples, evaluator notes, cost, latency, defect taxonomy |
| Decision | build / buy / partner / hybrid / stop 的判定规则 |
4. Decision Board
Decision board 不应只由 procurement 或 product 决定。最小构成:
| Role | 负责挑战的问题 |
|---|---|
| Business owner | 业务价值是否真实, 是否愿意承担 adoption 和 residual risk |
| 产品负责人 | 用户场景、MVP、体验、采用、收益假设是否清晰 |
| 业务分析与流程负责人 | 流程、规则、需求、例外、验收和证据是否完整 |
| 解决方案架构能力 | 集成、数据流、RAG、agent、日志、可观测、降级是否可行 |
| 企业架构能力 | 平台复用、能力地图、供应商集中度、目标架构适配 |
| Security / privacy | 数据、身份、权限、日志、DLP、威胁模型是否达标 |
| Risk / compliance / model risk | 风险等级、监管影响、模型验证、人工监督是否充分 |
| Procurement / TPRM | 供应商风险、商业条款、后续合同尽调是否可进入下一阶段 |
Build / Buy / Partner Decision Model
Build-buy-partner 不是三选一口号, 而是一组架构边界决策。
Decision Axes
| Axis | Build 倾向 | Buy 倾向 | Partner 倾向 | Hybrid 倾向 |
|---|---|---|---|---|
| Differentiation | 流程或数据是竞争优势 | 能力通用, 市场成熟 | 需要行业经验转移 | 控制层差异化, 能力层通用 |
| Control need | 数据、模型、策略、审计、工具权限必须内部控制 | 供应商可提供充分控制证据 | 机构缺少短期能力 | 内部保留 policy, eval, gateway, audit |
| Time to value | 可以等待内部能力建设 | 需要快速验证和上线 | 需要加速交付但保留学习 | 先买后抽象, 或买组件建控制面 |
| Scale economics | 用量大, 单位成本可被内部平台摊薄 | 用量不确定或中小规模 | 早期探索 | 高风险部分内部化, 普通能力外部化 |
| Talent readiness | 有 AI platform, data, eval, security, SRE 能力 | 内部团队不足 | 需要 co-build and knowledge transfer | 内部团队能运营控制层 |
| Regulatory evidence | 内部可生成更强证据 | 供应商证据成熟且可导出 | 需要顾问补齐控制设计 | 机构证据层统一, vendor 提供组件证据 |
| Exit optionality | 内部架构可替换 | vendor lock-in 可接受 | 依赖转移需要计划 | 抽象接口降低退出成本 |
Component-Level Decision
不要为整个用例做一个笼统决定。把 AI system 拆到组件层:
| Component | 常见选择 | 判断逻辑 |
|---|---|---|
| Base model | buy or use managed model | 基础模型通常不是金融零售机构差异化来源 |
| Model gateway | build or platform buy | 需要统一 routing, logging, policy, cost, versioning |
| RAG ingestion | hybrid | 文档处理可买, source registry 和 entitlement 应内部控制 |
| Vector store / search | buy managed or internal platform | 取决于数据分类、延迟、成本和地域要求 |
| Prompt / policy registry | build lightweight | 政策、prompt、release evidence 应机构可审计 |
| Eval harness | hybrid | 工具可买, golden set、rubric、门禁阈值必须内部拥有 |
| Agent tool gateway | build | 高风险动作、权限、审批、幂等和审计不宜交给黑盒 |
| Human review workbench | buy, build, or existing workflow | 取决于是否嵌入 AML/KYC/信贷/客服 case system |
| Observability | hybrid | vendor trace 要进入内部 SIEM, audit, cost 和 quality dashboard |
Practical Decision Rule
| Condition | Recommendation |
|---|---|
| 供应商产品强, 但不能导出 trace/eval/log | sandbox 可以学, 不进入生产候选 |
| 供应商质量好, 但工具动作权限不可控 | buy UI/model layer only, build tool gateway |
| 内部模型质量一般, 但证据和控制强 | 可做 high-risk pilot, 因为金融场景安全证据比平均分重要 |
| 多供应商效果接近 | 选择架构适配、证据导出、成本可预测和退出约束更好的方案 |
| 用例不是差异化, 但需要快速 adoption | buy with strong sandbox and production gate |
| 用例是核心风控或客户承诺 | hybrid by default, 内部控制 decision boundary and evidence |
Vendor Sandbox And Benchmark Design
Sandbox Design Principles
- 同一任务, 同一数据, 同一 rubric, 同一 cost and latency measurement。
- 至少比较 vendor option, internal baseline, no-AI baseline。
- 测试 positive cases, negative cases, edge cases, abuse cases, stale-source cases, restricted-data cases。
- 输出必须可追踪到 prompt, model, source, tool call, reviewer, decision。
- sandbox 只能使用批准数据, 不能连接生产写动作。
- benchmark 结论必须包含 architecture fit, not just accuracy。
Benchmark Plan
| Dimension | Measurement | Release implication |
|---|---|---|
| Task quality | groundedness, completeness, policy compliance, extraction accuracy, narrative quality | 决定是否满足 workflow outcome |
| Critical failure | hallucinated commitment, missed red flag, unauthorized advice, PII leakage, wrong adverse action | high-risk 用例通常要求为 0 |
| Retrieval quality | source recall, citation correctness, freshness, entitlement respect | 决定 RAG 是否可用于受控生产 |
| Tool safety | allowed tool choice, argument correctness, approval compliance, idempotency | 决定 agent 是否可启用工具 |
| Human oversight | reviewer agreement, override rate, review time, escalation quality | 决定 HITL 是否真实有效 |
| Cost | cost per case, token variance, document cost, eval cost, monitoring cost | 决定 TCO 和 scale feasibility |
| Latency | p50, p95, timeout, retry, end-to-end workflow time | 决定客户体验和运营队列可用性 |
| Security / privacy | prompt injection result, DLP pass, data retention proof, access control | 决定是否进入 pilot |
| Architecture fit | API, IAM, audit export, SIEM, model gateway, data residency, change control | 决定 build-buy-partner 边界 |
| Evidence maturity | eval export, trace completeness, admin audit, versioning, incident evidence | 决定是否满足金融审计和模型风险 |
Sandbox Evidence Pack
每个 vendor 或 internal option 都要产出:
| Evidence object | 内容 |
|---|---|
| Option card | vendor/internal option, deployment model, model/provider, data route, components |
| Data map | 输入字段、文档、日志、embedding、retention、masking、region |
| Benchmark report | dataset, tasks, rubric, score, confidence, slice failures |
| Failure taxonomy | critical, high, medium, low failure examples and root cause |
| Trace sample | prompt version, retrieval results, model version, output, tool calls, reviewer action |
| Cost and latency sheet | unit economics, p50/p95 latency, rate limit, stress result |
| Architecture fit review | integration, IAM, observability, RAG, agent boundary, platform fit |
| Risk review | privacy, security, third-party, model risk, compliance, operational resilience |
| Recommendation | build / buy / partner / hybrid / stop, with conditions and reversal triggers |
证据与控制
Use a three-layer evidence model: metric proves behavior, control constrains risk, evidence proves the control operated.
| Layer | Examples | Owner |
|---|---|---|
| Outcome metrics | handle time, document review cycle time, AML case completeness, fraud intervention conversion, complaint escalation accuracy | business owner and product owner |
| Quality metrics | groundedness, extraction accuracy, source coverage, critical failure rate, reviewer agreement, override reason | EvalOps and domain SMEs |
| Risk metrics | PII leakage, unauthorized tool call, policy violation, under-escalation, biased slice regression, stale-source answer | risk, compliance, model risk |
| Operational metrics | p50/p95 latency, timeout, fallback, cost per case, rate-limit hit, support ticket volume | platform and operations |
| Adoption metrics | active users, task completion, accepted suggestions, edit distance, trust survey, manual fallback rate | product owner and operations |
Control mapping:
| Risk | Preventive control | Detective control | Corrective control | Evidence |
|---|---|---|---|---|
| Vendor selected before problem clarity | intake completeness gate | funnel review log | reject/defer decision | intake card, decision minutes |
| Demo bias | common benchmark plan | score normalization | re-run benchmark | benchmark report |
| Sensitive data leakage | masking, DLP, approved sandbox data | payload sampling, DLP alert | delete, notify, retrain reviewers | data map, DLP test |
| Prompt injection | red-team cases, tool isolation | injection failure monitor | disable route, update guardrail | red-team report |
| Cost runaway | budget cap, token limits | cost dashboard | throttle, switch model, revise scope | cost sheet |
| Over-reliance | UI uncertainty, mandatory review for high risk | override and review audit | retraining, stricter HITL | reviewer logs |
| Architecture lock-in | gateway, export requirement, component boundaries | dependency review | redesign or reject vendor | ADR, architecture map |
| Production promotion without evidence | release gate checklist | evidence binder completeness check | limited pilot or no-go | gate memo |
金融零售/AI产品场景
1. GenAI Contact Center Copilot
| Intake decision | Sandbox focus | Architecture decision |
|---|---|---|
| AI drafts agent guidance, not customer commitments | policy answer, citation, escalation, vulnerable customer handling | buy copilot UI if strong; build knowledge governance, eval, telemetry export |
Hard failures:
- AI promises fee reversal outside policy。
- AI misses complaint language or vulnerable customer marker。
- AI cites stale product terms。
- AI outputs account data to unauthorized role。
2. KYC Document Intelligence
| Intake decision | Sandbox focus | Architecture decision |
|---|---|---|
| AI extracts and reconciles document facts; final KYC disposition stays human/system controlled | field extraction, document fraud signals, missing document checklist, data lineage | buy OCR/extraction, partner for policy tuning, build case workflow and evidence layer |
Hard failures:
- Wrong identity attribute without confidence flag。
- Document retention exceeds approved period。
- Evidence cannot be exported for audit or regulator inquiry。
- Model cannot handle jurisdiction-specific document rules。
3. AML Investigation Workbench
| Intake decision | Sandbox focus | Architecture decision |
|---|---|---|
| AI summarizes evidence and drafts narrative; no final SAR/no-SAR decision | red flag coverage, source-grounded narrative, analyst override, missed-risk rate | hybrid; internal control over data, RAG, audit, case action boundary |
Hard failures:
- AI omits material suspicious activity。
- AI invents transaction rationale。
- AI suggests final SAR decision as authoritative。
- Case trace cannot reconstruct evidence used。
4. Credit Decision Support
| Intake decision | Sandbox focus | Architecture decision |
|---|---|---|
| AI supports memo drafting and policy retrieval; credit decision remains governed by approved decisioning process | policy retrieval, adverse-action boundary, fair lending slice, explanation evidence | hybrid; build decision boundary and model risk evidence, buy retrieval or document summarization components only |
Hard failures:
- AI uses protected-class proxy or unsupported inference。
- AI drafts adverse action reason not supported by system of record。
- Human reviewers over-rely without challenge。
- Vendor cannot provide versioned evidence for model/prompt changes。
5. Payments Fraud Intervention
| Intake decision | Sandbox focus | Architecture decision |
|---|---|---|
| AI recommends intervention scripts and case prioritization; payment block/release requires deterministic policy and approval | false-positive customer harm, scam typology coverage, latency, tool permissions | hybrid; build tool gateway, approval and audit; consider buy for scam narrative intelligence |
Hard failures:
- AI triggers payment action without authorization。
- AI misses urgent scam indicators。
- Latency breaks real-time intervention window。
- Tool call cannot be replayed or reversed。
反模式
| Anti-pattern | Why it fails | Better pattern |
|---|---|---|
| Vendor-first discovery | Demo defines problem and success criteria | Intake starts from outcome, workflow, risk and baseline |
| Accuracy-only scorecard | Ignores audit, latency, cost, data, tool safety and architecture fit | Multi-dimensional sandbox scorecard |
| PoC using production data without control | Creates privacy and shadow-production risk | Approved sandbox data boundary and DLP |
| One build-buy decision for the whole system | Hides component-level control needs | Component decision matrix |
| Vendor black-box RAG | Cannot prove source, freshness, entitlement or citation | Source registry and retrieval trace |
| Contract promises without technical enforcement | Rights cannot be exercised in operations | Contract-control-evidence mapping |
| Pilot becomes production by adoption pressure | Controls arrive after risk exposure | Production promotion gate with hard stop criteria |
| No no-AI baseline | AI value cannot be defended | Compare process/rules/search baseline |
| Weak cost measurement | Token and eval cost surprise at scale | Cost per case and capacity model |
| Missing exit constraints at intake | Lock-in discovered after integration | Exit constraints and concentration risk before pilot |
最终心智模型
AI procurement 的核心不是“买还是不买”,而是把 option selection 变成一条证据链:intake 判断是否值得进入 AI 空间,sandbox 判断哪个方案在受控条件下有效,architecture gate 判断它能否在生产规模下被安全运营,production promotion gate 判断残余风险是否被授权接受。
| Architecture pattern | Intake question | Sandbox test | Production gate evidence |
|---|---|---|---|
| RAG | 哪些来源权威、最新且有权限边界? | citation correctness、stale-source failure、ACL filtering、retrieval recall | source registry、index version、retrieval eval、access review |
| Agent | 哪些工具可以被调用,权限和副作用是什么? | tool choice accuracy、argument validation、approval path、idempotency | tool policy、audit trace、kill switch、rollback test |
| Copilot | 人类看到、编辑、批准、拒绝和承担什么? | reviewer agreement、override rate、UX trust calibration、escalation | HITL log、training、adoption and quality dashboard |
| Eval | 上线前必须证明哪些行为合同? | golden set、red-team set、slice metrics、critical failures | eval report、threshold decision、exception memo |
| Governance | 谁拥有风险、变更、证据、事故和生命周期? | RACI simulation、gate dry run、evidence completeness | AI inventory、ADR、risk acceptance、operating cadence |
| Model gateway | 哪些 providers 和 versions 可用? | route comparison、fallback、cost/latency benchmark | routing policy、model registry、telemetry export |
| Observability | 单个 case 能否端到端复盘? | trace completeness、log redaction、SIEM export | evidence binder、retention setting、audit export |
最终判断标准:winning option 不是 demo 最流畅的供应商,而是在同一 workflow、同一数据边界、同一 rubric、同一成本与延迟测量、同一控制证据下,最能被集成、监控、审计、退出和持续治理的架构方案。通用能力可以买,政策边界、eval contract、tool gateway、evidence plane 和 residual-risk decision 必须由机构掌握。
Source Anchors
| Anchor | Link | 本文使用方式 |
|---|---|---|
| NIST AI Risk Management Framework | https://www.nist.gov/itl/ai-risk-management-framework | 用 Govern / Map / Measure / Manage 组织 AI risk, evidence, monitoring and management action |
| NIST AI RMF Generative AI Profile | https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence | 用 GenAI risk lens 设计 sandbox red-team, content provenance, data leakage and misuse cases |
| ISO/IEC 42001 AI management systems | https://www.iso.org/standard/81230.html | 用 AI management system 思路定义 accountability, lifecycle, operation, performance evaluation and improvement |
| ISO/IEC/IEEE 29148 Requirements engineering | https://www.iso.org/standard/72089.html | 用 requirements quality, stakeholder concern and validation thinking 支撑 intake and eval contract |
| ISO/IEC/IEEE 42010 Architecture description | https://www.iso.org/standard/74393.html | 用 stakeholder concern, viewpoint and architecture rationale 组织 ADR and architecture fit review |
| Interagency Third-Party Risk Guidance, FDIC FIL-29-2023 | https://www.fdic.gov/news/financial-institution-letters/2023/fil23029.html | 用 third-party lifecycle 思维连接 planning, due diligence, selection, monitoring and termination inputs |
| FFIEC AIO booklet summary, OCC Bulletin 2021-30 | https://www.occ.gov/news-issuances/bulletins/2021/bulletin-2021-30.html | 用 architecture, infrastructure and operations lens 检查 resilience, integration, operations and evidence |
| OWASP Top 10 for Large Language Model Applications | https://owasp.org/www-project-top-10-for-large-language-model-applications/ | 用 prompt injection, sensitive information disclosure, supply chain and excessive agency 设计安全测试 |
SOTA 状态标注 (2026-07-01)
本篇属于第二、三遍深读池(参考架构/深读笔记),未列入 12 周主线必读。时效基线为写作时点;引用前请按 CLAUDE.md 全局时效性硬规则复查最新进展。模块级 SOTA 对照见 docs/AI_SYSTEMATIC_LEARNING_ROADMAP_2026.md 各周「2026 SOTA 对照」行与文末「SOTA 检查」。