AI Technology Radar:情景规划与战略前瞻架构
AI 战略最容易失效的地方,不是少看了某个新模型或新框架,而是把未来当成趋势清单或固定路线图。金融零售机构面对的是一组同时变化的变量:模型能力、推理成本、供应商策略、开源成熟度、agentic workflow 安全、数字身份欺诈、监管预期、劳动力重构、客户信任与竞争对手动作。单点预测没有足够韧性,真正需要的是一套把不确定性转化为可治理决策的架构。
AI Technology Radar / Scenario Planning / Strategic Foresight Architecture
配对阅读:本篇的操作手册版(模板/RACI/门禁/runbook)是
docs/AI_TECHNOLOGY_RADAR_SCENARIO_PLANNING_STRATEGIC_FORESIGHT_PLAYBOOK.md。第一遍读本篇建立原理与架构判断;第二遍做案例时再用 playbook 查表落地,两者不需要重复精读。
Batch 158 foundation note for AI product and architecture strategy in financial retail. Core question: how does a financial retail organization monitor unstable AI technology, regulation, vendors, economics, security, workforce, and customer signals, then convert those signals into portfolio options, architecture runway, experiments, adoption decisions, and executive narratives without chasing hype? Important note: this document is a learning artifact. It is not legal advice, compliance advice, procurement advice, model validation, audit opinion, or investment advice. Formal decisions require review by Legal, Compliance, Risk, Model Risk, Security, Privacy, Procurement, Enterprise Architecture, Business Owners, and accountable executives.
核心导读
AI 战略最容易失效的地方,不是少看了某个新模型或新框架,而是把未来当成趋势清单或固定路线图。金融零售机构面对的是一组同时变化的变量:模型能力、推理成本、供应商策略、开源成熟度、agentic workflow 安全、数字身份欺诈、监管预期、劳动力重构、客户信任与竞争对手动作。单点预测没有足够韧性,真正需要的是一套把不确定性转化为可治理决策的架构。
这篇笔记的核心观点是:AI technology radar、scenario planning 和 strategic foresight 不应被拆成创新团队的展示材料,而应成为产品组合、架构 runway、风险控制和执行叙事之间的决策系统。
external signal
-> source authority and evidence quality
-> assumption ledger
-> trigger indicator
-> radar decision
-> scenario impact
-> portfolio option
-> experiment backlog
-> architecture runway
-> vendor / regulatory watch
-> executive decision record
成熟做法不是声称“知道未来会怎样”,而是持续回答这些问题:
- 哪些关键假设仍然成立,哪些已经变旧;
- 哪些决策可以等待证据,哪些等待会丧失机会;
- 哪些投入是在押注单一未来,哪些投入能保留多个未来下的选择权;
- 哪些架构决策会形成难以逆转的供应商、数据、控制或组织依赖;
- 哪些信号足以改变产品边界、实验优先级、风险控制或执行叙事。
Source Anchors
Access date for source discipline: 2026-06-30. These anchors provide governance language and operating patterns. They do not replace jurisdiction-specific legal, compliance, procurement, or risk interpretation.
| Anchor | Official / primary link | How this note uses it |
|---|---|---|
| NIST AI Risk Management Framework | https://www.nist.gov/itl/ai-risk-management-framework | Uses Govern, Map, Measure, Manage as a risk management backbone for radar-to-control translation |
| NIST AI RMF Core | https://airc.nist.gov/airmf-resources/airmf/5-sec-core/ | Reinforces that AI risk work is a continuous set of functions, not a one-time checklist |
| NIST AI 600-1 Generative AI Profile | https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence | Uses GenAI-specific risk framing for model behavior, content, provenance, incident, and monitoring concerns |
| ISO/IEC 42001 | https://www.iso.org/standard/81230.html | Uses AI management system language for policy, objectives, operation, performance evaluation, and continual improvement |
| Thoughtworks Technology Radar | https://www.thoughtworks.com/radar | Uses radar discipline, rings, and blip movement as inspiration, adapted for enterprise AI governance |
| Thoughtworks Build Your Own Radar | https://www.thoughtworks.com/insights/blog/build-your-own-technology-radar | Uses the mechanics of quadrants and rings while replacing novelty with enterprise decision evidence |
| OECD AI Principles | https://oecd.ai/en/ai-principles | Anchors trust, human rights, democratic values, transparency, robustness, and accountability expectations |
| FFIEC IT Examination Handbook | https://ithandbook.ffiec.gov/ | Anchors financial institution architecture, operations, development, acquisition, maintenance, security, resilience, and third-party technology risk context |
Source-use discipline:
- Treat official sources as anchors, not as complete requirement catalogs.
- Record source date, access date, owner, jurisdiction, authority level, and affected internal artifacts.
- Separate official guidance from analyst commentary, vendor marketing, conference talks, benchmarks, and social media signals.
- Do not convert a source directly into a roadmap item until applicability, materiality, and evidence quality are assessed.
- When a source changes, refresh the assumption ledger and affected decisions rather than rewriting the whole strategy.
1. 问题定义
1.1 为什么 AI 需要 foresight architecture
传统技术规划通常假设技术演进相对连续:年度规划、季度 roadmap、供应商评估、架构评审可以覆盖大多数变化。AI 破坏了这个节奏。一个季度内可能同时出现模型价格下降、供应商数据使用条款改变、官方监管解释更新、新的 prompt injection 攻击、agent framework 能力跃迁、员工信任下降和竞争对手产品上线。
因此,AI foresight 的问题不是“我们应该关注哪些趋势”,而是“哪些外部变化会让内部决策变得过期、昂贵或危险”。它关注三类断裂:
| 断裂 | 典型表现 | 决策风险 |
|---|---|---|
| 证据断裂 | 决策依据来自旧 benchmark、旧成本、旧监管解释或旧攻击模型 | 继续扩张一个已经不成立的路线 |
| 架构断裂 | 数据流、供应商接口、eval、日志、审批链绑定到单一方案 | 后续切换成本过高,议价能力下降 |
| 组织断裂 | 产品、架构、风险、运营、法务各自维护不一致的未来判断 | 投入冲突、控制滞后、执行叙事失真 |
AI 技术雷达、场景规划和战略前瞻架构的共同目标,是让组织知道自己正在赌什么、证据到什么程度、什么条件会改变判断。
1.2 趋势观察和决策架构的区别
趋势观察常见输出是:
New model released.
New regulation proposed.
New agent framework trending.
New vendor partnership announced.
New benchmark improved.
决策架构关心的是:
Which current assumptions changed?
Which portfolio options are now more or less valuable?
Which architecture decisions became more irreversible?
Which experiments should start, stop, or change scope?
Which vendor dependencies require exit planning?
Which controls or evals need refresh?
Which executive decision needs a new narrative?
一个信号重要,不是因为它新,而是因为它改变了某个决策的 expected value、risk、reversibility、timing 或 stakeholder acceptability。
1.3 AI 信号的四层分级
| Category | Definition | Evidence quality | Default action |
|---|---|---|---|
| Hype | 高关注、低决策相关性、企业落地路径不清 | blog posts, social media volume, demo videos, conference claims | 只有映射到命名假设时才进入 watch |
| Weak signal | 某个重要假设可能改变的早期迹象 | early customer behavior, policy drafts, credible research, production anecdotes | 指定 owner,设置触发器 |
| Material signal | 足以改变 backlog、风险姿态、架构方向或供应商动作 | official guidance, contract change, verified incident, production cost delta | 更新 ledger 并进入决策复核 |
| Decision-grade evidence | 足以 fund、stop、scale、hold 或 adopt | controlled experiment, validated eval, legal interpretation, risk acceptance, production telemetry | 触发正式决策或刷新既有决策 |
这套分级的作用是防止组织把精彩 demo、供应商宣传、行业传闻和可审计证据放在同一层级。
1.4 最需要 foresight 的高逆转成本决策
| Decision | Why reversal is hard | Foresight question |
|---|---|---|
| Primary model provider | contracts, data pathways, eval harness, latency tuning, security approvals | Are model economics and provider policies stable enough to commit? |
| Agentic operations platform | workflow redesign, tool permissions, audit trails, SoD controls | Can we constrain autonomy before scaling? |
| Customer identity architecture | regulatory requirements, fraud risk, channel integration, customer friction | Which digital identity standards and fraud patterns are converging? |
| Contact-center GenAI rollout | workforce adoption, QA model, customer harm controls, HR implications | Which tasks should remain human-owned under plausible regulation and quality scenarios? |
| Data residency design | regional deployment, vendor availability, legal interpretation, cost | Will sovereign AI or cross-border data constraints change the architecture? |
| AI platform abstraction | team topology, gateway design, procurement leverage, vendor exit | Which capabilities should be platform-owned versus product-owned? |
2. 架构模型与核心原理
2.1 Strategic Foresight Architecture
战略前瞻架构可以视为三个层次:
Signal layer
technology | regulation | vendor | economics | security | workforce | customer | competitor | data | resilience
|
Interpretation layer
signal quality -> assumption ledger -> trigger indicator -> scenario stress test
|
Decision layer
radar posture -> option portfolio -> experiment backlog -> architecture runway -> governance record
信号层解决“外部发生了什么”;解释层解决“这改变了哪个假设”;决策层解决“我们要 watch、trial、adopt、hold,还是改产品组合和架构投入”。
2.2 Signal Taxonomy
雷达必须覆盖多类信号,因为不同信号改变不同类型的决策。
| Signal class | Example sources | What changes | Financial retail examples |
|---|---|---|---|
| Model capability | model cards, eval results, production benchmarks, internal task evals | use-case feasibility, cost-to-quality curve, human review burden | GenAI contact center improves complaint summarization quality |
| Model economics | provider pricing pages, contract terms, usage telemetry, capacity constraints | unit economics, routing, build-vs-buy, cache strategy | cost per resolved customer call drops below human QA review cost |
| Vendor / platform | release notes, data-use terms, deprecation notices, acquisition news, SLA changes | lock-in risk, procurement posture, exit plan, platform roadmap | provider changes retention policy for prompt logs |
| Regulation horizon | official guidance, consultation papers, supervisory priorities, enforcement actions | control objectives, evidence requirements, release gates | regulator clarifies expectations for AI adverse action explanations |
| Security threat | OWASP-style risks, incident reports, red-team findings, fraud typologies | threat model, controls, monitoring, incident playbook | prompt injection abuses payment exception tool |
| Workforce impact | adoption analytics, HR signals, task redesign data, training evidence | change plan, operating model, role redesign, resistance risk | agents reduce manual case prep but increase quality review burden |
| Customer behavior | complaint trends, trust signals, opt-out rates, channel analytics | product boundaries, transparency, fallback design | customers reject AI-generated debt hardship scripts |
| Competitor / market | product releases, job postings, earnings calls, app flows, public case studies | differentiation, parity requirements, partnership strategy | competitor launches AI-powered open banking cashflow coach |
| Data ecosystem | open banking API changes, data quality shifts, consent patterns | data readiness, lineage, feature availability, privacy posture | new account aggregation coverage improves affordability analysis |
| Operational resilience | outage reports, dependency maps, cloud capacity, fallback telemetry | degraded mode, BCP, kill switch, redundancy | model provider outage affects regulatory report drafting |
2.3 Signal Object Model
每个 material signal 应被记录成稳定对象,而不是会议纪要中的一句话。
| Field | Meaning |
|---|---|
| signal_id | Unique identifier for traceability |
| source | URL, document, system, incident, or internal telemetry source |
| source_authority | official, supervisory, contractual, internal production, research, vendor, market, informal |
| signal_class | technology, regulatory, vendor, economics, security, workforce, customer, competitor, data, resilience |
| affected_assumption | Named assumption in the assumption ledger |
| affected_capability | business capability or architecture capability |
| affected_use_case | product or operational use case |
| confidence | low, medium, high |
| materiality | low, medium, high |
| reversibility | reversible, partially reversible, hard to reverse |
| decision_due | date or event that requires review |
| recommended_action | watch, trial, adopt, hold, stop, refresh decision, open experiment, update control |
| owner | accountable decision owner or forum |
2.4 Signal Quality Ladder
| Level | Evidence | Typical use |
|---|---|---|
| L0 rumor | unsourced claim, social media thread, vendor teaser | ignore unless it maps to a critical assumption |
| L1 directional | credible commentary, early demo, single team anecdote | monitor and assign weak-signal owner |
| L2 corroborated | multiple credible sources, repeated internal observation | add to radar and define trigger |
| L3 measurable | internal eval, telemetry, cost data, red-team finding | open experiment or update controls |
| L4 decision-grade | validated experiment, official obligation, contract change, incident review | change portfolio, architecture, vendor, or release decision |
2.5 Technology Radar Model
企业 AI radar 可以借鉴 Thoughtworks-style radar 的 rings 和 blips,但目标不是判断新技术是否流行,而是让采用姿态变成可追踪的决策。
Radar Rings
| Ring | Meaning | Evidence required | Typical decision |
|---|---|---|---|
| Watch | Monitor because the signal may change a named assumption | source quality, affected assumption, owner, trigger | no build commitment; set review cadence |
| Trial | Run bounded experiment to validate a decision-critical assumption | experiment charter, eval plan, risk controls, success threshold | fund discovery or pilot |
| Adopt | Use in production within defined guardrails | production evidence, controls, support model, exit path, owner | scale or standardize |
| Hold | Do not start new work, or pause expansion, because risk, economics, maturity, or lock-in is unacceptable | documented reason, revisit trigger, affected decisions | stop, defer, or use only for existing estate |
“Assess” 很容易变成停车场;更好的状态是 watch with explicit triggers 或 trial with explicit experiments。
Radar Quadrants
| Quadrant | What belongs here | Example blips |
|---|---|---|
| Models and inference economics | frontier models, small models, distillation, model routing, caching, on-device inference | model commoditization, context window expansion, cost-per-token shifts |
| Agentic operations and workflow | agent frameworks, tool use, state machines, durable workflow, human handoff | agentic banking ops, payment exception agent, regulatory reporting agent |
| Trust, risk, and security | evals, guardrails, identity, provenance, threat detection, red-team methods | prompt injection controls, synthetic identity detection, deepfake voice risk |
| Data, identity, and ecosystem | open banking, digital identity, consent, data quality, financial graph | verifiable credentials, open finance signals, AML typology evolution |
| Governance and regulatory horizon | standards, official guidance, supervisory priorities, audit expectations | NIST AI RMF profile, ISO/IEC 42001, FFIEC handbook changes |
| Workforce and operating model | role redesign, AI academy, human oversight, change saturation | contact-center role shifts, AML investigator task redesign |
Radar Blip Fields
| Field | Example |
|---|---|
| Name | Durable agent workflow state machine |
| Ring | Trial |
| Quadrant | Agentic operations and workflow |
| Decision thesis | Could reduce payment exception handling time if tool permissions, audit trail, and fallback controls are strong |
| Evidence | internal workflow simulation, vendor sandbox, incident review from related agent tools |
| Risks | excessive agency, SoD breach, weak rollback, ambiguous accountability |
| Trigger to move inward | 95 percent task completion in simulation with zero unauthorized tool calls and approved audit evidence |
| Trigger to move outward | policy ambiguity, tool misuse, unacceptable review burden, vendor SLA failure |
| Owner | accountable payments operations and architecture forum |
2.6 Scenario Planning Model
Scenario planning 不负责预测唯一未来,而是用少量高影响不确定性去压测当前假设。
| Uncertainty | Low end | High end |
|---|---|---|
| Model economics | frontier inference remains expensive and concentrated | high-quality inference commoditizes and becomes multi-provider |
| Regulatory intensity | principles-based governance with flexible interpretation | prescriptive controls, transparency, audit, and post-market monitoring |
| Agent autonomy | mostly human-in-the-loop copilots | agentic workflows execute multi-step regulated operations |
| Security threat evolution | known LLM attack patterns remain manageable | synthetic identity, deepfake, prompt injection, and tool misuse accelerate |
| Workforce response | roles adapt through training and redesigned workflows | resistance, capability gaps, and accountability disputes slow adoption |
| Scenario | Narrative | Strategic pressure |
|---|---|---|
| Commodity Intelligence | Frontier quality becomes cheaper and more substitutable across providers | differentiation shifts from model access to data, workflow, evals, UX, governance, and distribution |
| Regulated Intelligence | AI obligations become more prescriptive and evidence-heavy | architecture must support traceability, human oversight, audit evidence, explainability, and release governance |
| Agentic Operations | Banks move from copilots to controlled agents in servicing, operations, and reporting | SoD, delegated authorization, runtime monitoring, tool permission, and kill-switch architecture become mandatory |
| Fraud Acceleration | Deepfakes, synthetic identities, payment scams, and automated social engineering scale | identity, behavioral analytics, customer education, fraud ops, and adaptive controls become differentiators |
| Workforce Recomposition | AI changes work content faster than roles, incentives, and skills can adapt | value realization depends on job redesign, training, trust, and accountability more than model quality |
每个 scenario 应回答:
- 当前战略在哪些条件下会失败;
- 哪些投入是 hedge,哪些投入会 stranded;
- 哪个架构决策会变得过于不可逆;
- 哪个供应商依赖会变危险;
- 哪些控制、客户信任和 workforce 假设会被打破。
2.7 Assumption Ledger
Assumption ledger 是 foresight 和治理之间的桥。它记录组织真正押注的内容。
| Field | Definition |
|---|---|
| assumption_id | Stable ID used across radar, scenarios, experiments, and ADRs |
| assumption_statement | Specific belief that supports a decision |
| decision_supported | Portfolio, product, architecture, vendor, control, or workforce decision |
| confidence | low, medium, high |
| evidence | source links, evals, telemetry, analysis, legal interpretation |
| owner | accountable person or forum |
| review_date | date when evidence must be refreshed |
| trigger_indicator | measurable condition or event that changes the assumption |
| threshold | value or event boundary that requires review |
| impact_if_wrong | cost, risk, delay, control gap, customer harm, workforce impact |
| response | watch, trial, adopt, hold, pivot, stop, escalate |
| Assumption | Supported decision | Trigger | Threshold | Response |
|---|---|---|---|---|
| Multi-provider model routing will reduce lock-in without unacceptable latency | build AI gateway abstraction | p95 latency and cost telemetry | routing adds more than 700 ms or 20 percent cost | narrow routing scope and keep provider-specific path |
| Contact-center agents will trust GenAI summaries if evidence snippets are visible | scale agent assist | adoption and override telemetry | override rate above 35 percent for two review cycles | refresh UX, training, and eval cases |
| Regulatory reporting drafts can be AI-assisted but must remain human-attested | regulatory reporting copilot | control review and audit feedback | evidence lineage incomplete for sampled reports | hold release and improve provenance |
| Payment fraud threat actors will increasingly use deepfake voice in high-risk calls | fraud intervention roadmap | fraud typology reports and case reviews | confirmed cases in target segment or peer incident | accelerate voice risk controls and step-up authentication |
| Model commoditization will make proprietary prompt chains less defensible | platform investment | model benchmark and pricing shift | two providers meet quality threshold at lower cost | invest in eval, data, workflow, and orchestration portability |
好的假设必须 specific、可证伪、链接到决策、有 owner、有证据、有 freshness date,并配套触发阈值。弱假设通常长这样:“AI agents will be the future”“Regulators will probably allow this”“The vendor is enterprise grade”。成熟组织不消灭不确定性,而是让不确定性可治理。
3. 关键机制
3.1 Decision Freshness
每个 radar decision 都应有 freshness date。过期决策不一定错误,但不能被默认信任。
| Decision type | Default freshness window | Refresh trigger |
|---|---|---|
| Model provider choice | 30 to 90 days | price change, data-use term change, major capability release, service outage |
| Regulatory posture | 30 days or event-driven | new official guidance, enforcement action, supervisory priority, legal interpretation |
| Security control | 30 to 60 days | new attack pattern, red-team finding, incident, new tool permission |
| Workforce adoption assumption | 60 to 90 days | adoption telemetry shift, grievance, training failure, role redesign |
| Architecture runway item | quarterly | scenario trigger, platform dependency shift, budget gate |
Decision freshness 让战略可审计:不仅知道当时决定了什么,也知道最近一次证据刷新是什么时候。
3.2 Trigger Indicators
Trigger indicators 定义组织什么时候必须改变 posture。它比“持续关注”更强。
| Trigger type | Example indicator | Decision affected |
|---|---|---|
| Cost trigger | cost per resolved contact falls below target human-assisted baseline | scale GenAI contact center |
| Quality trigger | internal eval pass rate exceeds threshold across critical intents | adopt model for production workflow |
| Risk trigger | red-team finds unauthorized tool execution path | hold agentic workflow release |
| Regulatory trigger | official guidance adds documentation or human oversight expectation | update release gate and control pack |
| Vendor trigger | provider changes data retention or training terms | open vendor review and exit-path analysis |
| Workforce trigger | adoption drops after initial activation or override burden rises | redesign workflow and training |
| Security trigger | prompt injection pattern affects target tool class | refresh threat model and controls |
| Market trigger | competitor launches regulated AI capability with clear customer adoption | reassess differentiation and parity need |
| Indicator type | Useful for | Example |
|---|---|---|
| Leading | early action before impact is visible | vendor roadmap shifts, consultation papers, new attack proof-of-concept, benchmark movement |
| Concurrent | active monitoring during adoption | cost per interaction, human review burden, escalation rate, quality drift |
| Lagging | proof of realized impact or harm | complaint volume, incident loss, audit finding, customer attrition, regulatory finding |
Leading indicator 没有后续验证会变成猜测;只有 lagging indicator 的体系会变成被动响应。
3.3 Option Portfolio
Strategic foresight 的产出应该是 options,而不是单一推荐。Option 是一种有限投入:用较小成本保留未来做更大决策的权利。
| Option type | Description | Financial retail example |
|---|---|---|
| Learning option | small spend to reduce uncertainty | evaluate two model providers on complaint summarization |
| Platform option | invest in reusable capability before full demand is proven | model gateway with logging, routing, eval, and policy enforcement |
| Compliance option | prepare for likely obligations before final interpretation | evidence binder pattern for high-impact GenAI uses |
| Security option | build defense before attack materializes at scale | prompt injection test harness for tool-using agents |
| Vendor option | preserve switching leverage | multi-provider abstraction and contractual exit terms |
| Workforce option | develop role capability ahead of rollout | AI academy path for AML investigators and contact-center supervisors |
| Customer trust option | reduce adoption risk | transparent AI disclosure and human escalation path |
Option 有价值通常满足这些条件:
- uncertainty high;
- payoff 或 avoided loss material;
- 等待会关闭机会;
- 学习成本小于不可逆承诺;
- 同时支持多个 scenario;
- 降低锁定、合规、供应商或组织逆转成本。
示例:
Option: Build a portable evaluation harness for GenAI customer communications.
Value: Works under model commoditization, regulatory tightening, vendor switching, and workforce trust scenarios.
Cost: Moderate platform and governance effort.
Decision unlocked: ability to compare providers, satisfy evidence requests, and refresh release gates.
3.4 Experiment Backlog
Experiment backlog 用来验证关键假设,不是 demo list。
| Field | Meaning |
|---|---|
| experiment_name | concise name |
| assumption_tested | assumption ID |
| scenario_relevance | which scenario this experiment informs |
| decision_unlocked | what decision can be made after evidence |
| method | prototype, eval, simulation, shadow mode, A/B test, red-team, workflow study |
| success_threshold | measurable pass condition |
| risk_controls | privacy, security, human oversight, data minimization, safe sandbox |
| owner | accountable owner or forum |
| timebox | calendar duration |
| evidence_output | report, eval dataset, ADR, decision memo, control update |
| Experiment | Assumption tested | Method | Threshold | Decision |
|---|---|---|---|---|
| Agentic payment exception sandbox | controlled agents can execute low-risk exception steps with audit trail | sandbox simulation and red-team | zero unauthorized tool calls, complete audit trail, human approval on value movement | trial or hold agentic ops |
| Contact-center summary trust study | agents trust summaries with cited source snippets | workflow study and adoption telemetry | override rate below 20 percent for target intents | scale or redesign UX |
| AML typology update retrieval eval | RAG can surface current typology changes without hallucinated citations | eval set from internal cases and typology memos | citation accuracy above threshold with no unsupported high-risk claim | adopt in investigator copilot |
| Digital identity credential verification pilot | verifiable credentials reduce onboarding friction without fraud increase | controlled pilot in low-risk segment | lower manual review rate and no increase in fraud flags | expand or hold |
| Regulatory reporting draft provenance test | report drafting copilot can preserve lineage and attestation | document workflow simulation | every generated claim maps to source evidence | proceed to pilot |
3.5 Architecture Runway
Architecture runway 是当未来更清晰时让产品团队能快速移动的能力集合。它不是提前堆平台,而是保留选择权的架构。
| Capability | Why it matters | Scenario coverage |
|---|---|---|
| AI gateway | routing, policy enforcement, logging, cost control, provider abstraction | commodity intelligence, vendor shocks, security escalation |
| EvalOps platform | repeatable quality, safety, fairness, security, and regression evidence | regulated intelligence, model churn, audit pressure |
| Evidence graph | trace from requirement to data, prompt, model, control, decision, and output | regulated intelligence, audit, reporting |
| Tool permission service | least privilege, SoD, delegated authorization, runtime control | agentic operations, security threat evolution |
| Human oversight workflow | review queues, attestation, escalation, override analysis | regulated intelligence, workforce recomposition |
| Model and vendor inventory | dependency, data use, contract, capability, risk, owner | vendor lock-in, model provider changes |
| Digital identity and fraud intelligence layer | identity proofing, risk signals, behavioral patterns | fraud acceleration, open banking |
| Cost observability | unit economics, routing, budget guardrails, workload forecasting | model economics, portfolio funding |
| Workforce capability telemetry | role readiness, training evidence, adoption, change saturation | workforce disruption |
每个 runway item 都要记录:
| Question | Reason |
|---|---|
| What decision does this make easier later? | avoids infrastructure for its own sake |
| What future scenarios does it hedge? | proves strategic relevance |
| What dependency does it reduce? | exposes vendor lock-in and platform risk |
| What control does it strengthen? | connects architecture to governance |
| What would cause us to stop investing? | protects against sunk-cost behavior |
3.6 Vendor and Regulatory Watch
Vendor watch 和 regulatory watch 要放在同一个 foresight system 中,因为供应商行为与监管预期常常相互影响。
| Vendor watch item | What to monitor | Action trigger |
|---|---|---|
| Data use terms | training use, retention, audit access, regional processing | terms change or exception required |
| Pricing model | token price, committed spend, rate limits, reserved capacity | unit economics or budget threshold breached |
| Model lifecycle | deprecation, replacement, version behavior, backward compatibility | model retirement or behavior drift |
| Enterprise controls | logging, encryption, RBAC, SSO, data residency, SOC reports | control gap blocks regulated use case |
| Integration path | API stability, tool calling, latency, observability | production SLA or resilience risk |
| Exit support | data export, prompt portability, eval portability, contract exit | switching cost becomes unacceptable |
| Regulatory watch item | What to monitor | Action trigger |
|---|---|---|
| AI-specific guidance | risk classification, transparency, human oversight, monitoring | official guidance applies to portfolio use case |
| Banking supervision | model risk, third-party risk, operational resilience, IT governance | exam priority or handbook update affects controls |
| Consumer protection | unfair practices, adverse action, complaints, accessibility | customer-facing AI changes decision or communication |
| Privacy and data | purpose limitation, consent, retention, cross-border | new data source or regional deployment |
| Security and fraud | authentication, identity, cyber, incident reporting | new attack pattern or incident threshold |
| Workforce and employment | monitoring, role impact, fairness, employee rights | AI tool changes work allocation or surveillance |
Watch item 必须落到行动之一:no material impact、continue watch with revised trigger、open experiment、update assumption ledger、update risk/control/eval requirement、update architecture runway、open vendor review、open regulatory interpretation request、escalate to portfolio decision forum。否则 radar 只是仪式。
4. 证据与控制
4.1 Foresight Evidence Map
AI foresight 不是“战略感觉”,它需要 evidence map 把信号、假设和控制连起来。
| Evidence type | Examples | Primary use |
|---|---|---|
| Source authority | official guidance, supervisory priority, signed contract, internal policy | distinguish obligation from commentary |
| Production telemetry | cost, latency, escalation, override, complaints, incident rate | validate adoption and economics assumptions |
| Eval evidence | task eval, red-team result, regression suite, benchmark against internal cases | decide trial/adopt/hold |
| Vendor evidence | release notes, SOC report, data-use terms, SLA, deprecation notice | manage dependency and exit plan |
| Risk evidence | threat intel, incidents, audit finding, model risk review | refresh controls and release gates |
| Workforce evidence | training completion, usage, override, QA burden, role readiness | test realization and operating model assumptions |
| Customer evidence | complaint themes, opt-out, trust signals, conversion, support contacts | adjust product boundary and transparency |
4.2 Governance Cadence
Foresight governance 应轻量但真实:它要能移动 blips、批准实验、刷新 trigger、改变投资和记录责任。
| Cadence | Audience | Purpose | Artifact |
|---|---|---|---|
| Weekly scan | product, architecture, security, risk, operations analysts | triage new signals and assign owners | signal intake log |
| Biweekly radar review | product, architecture, risk, data, operations | move blips, approve experiments, refresh triggers | radar board and decision log |
| Monthly portfolio review | executives, portfolio governance, finance | fund, stop, scale, or hedge options | option portfolio memo |
| Quarterly scenario review | senior leadership, strategy, risk, enterprise architecture | test assumptions and update architecture runway | scenario narrative pack |
| Event-driven review | affected owners | respond to official guidance, vendor change, incident, or major model shift | decision freshness memo |
4.3 Executive Decision Narrative
高层叙事不应是“AI trends update”,而应是可决策的 memo。
What changed:
Why it matters:
Which assumptions are affected:
Which decisions are fresh or stale:
Which options we recommend:
What evidence we have:
What evidence we still need:
What decision is requested:
What would make us change our mind:
这类叙事的关键是承认不确定性,同时明确请求:fund、stop、scale、hold、hedge、refresh controls、open vendor review 或 delay irreversible decision。
4.4 Governance Forums and Decision Rights
| Forum | Decision rights | Inputs | Outputs |
|---|---|---|---|
| Signal triage | classify signal and assign owner | source log, incidents, vendor notices, telemetry | signal object and owner |
| Radar review | move blips and approve trials | radar, evidence, assumption ledger | ring decisions and experiment approvals |
| Architecture council | approve runway and ADR implications | scenarios, dependency map, vendor watch | runway priority and ADR updates |
| Product portfolio board | fund options, stop weak bets, scale strong bets | option portfolio, evidence, risk view | funding and scope decisions |
| Risk and compliance forum | interpret regulatory and control impacts | source anchors, legal interpretation, risk register | control and evidence requirements |
| Executive committee | decide major commitments and narrative | scenario pack, investment memo, residual risk | strategic choice and accountability |
4.5 Control Tests
| Control question | Test method |
|---|---|
| Can each radar blip trace to a named assumption? | sample blips and inspect ledger linkage |
| Are watch items actionable? | verify trigger, owner, review cadence, and escalation path |
| Do trials have measurable thresholds? | inspect experiment charter and stop/scale criteria |
| Are architecture runway items scenario-backed? | map each capability to at least two plausible futures or one high-severity risk |
| Are stale decisions visible? | query decisions past freshness window |
| Are vendor dependencies explicit? | inspect provider inventory, data flows, eval portability, exit triggers |
| Are regulatory signals separated by authority level? | compare official anchors, commentary, and internal interpretation records |
5. 金融零售 / AI 产品场景
5.1 Agentic Banking Operations
Signal:
- Agent frameworks mature and vendors offer tool-using agents for back-office operations.
- Security teams report new excessive-agency and prompt-injection patterns.
- Operations leaders want lower exception handling time.
Decision posture:
- Place agentic payment exceptions in Trial, not Adopt.
- Build runway for tool permission, workflow state, audit trail, human approval, and kill switch.
- Run sandbox simulation before production workflow changes.
- Hold autonomous customer-impacting actions until control evidence is proven.
The product lesson is that the opportunity is not “agents are the future”;the decision is whether controlled agents can handle low-risk exception steps with auditable tool use and human approval before value movement.
5.2 GenAI Contact Center
Signal:
- Model quality improves for summarization and next-best-action suggestions.
- Workforce telemetry shows high activation but inconsistent trust.
- Complaint risk remains high for vulnerable customer journeys.
Decision posture:
- Adopt internal summarization in low-risk intents with citations and QA sampling.
- Trial regulated advice and hardship scripts.
- Hold fully autonomous customer commitments.
Architecture implications:
- source-grounded generation;
- call transcript data controls;
- supervisor review workflow;
- customer harm monitoring;
- quality and trust telemetry.
5.3 Digital Identity and Synthetic Fraud
Signal:
- Synthetic identity and deepfake voice attacks accelerate.
- Verifiable credential ecosystems mature unevenly.
- Open banking and digital wallets change customer onboarding patterns.
Decision posture:
- Watch digital identity standards and ecosystem adoption.
- Trial verifiable credential verification in a low-risk onboarding segment.
- Invest in fraud intelligence runway that can consume identity, device, behavior, and consented data signals.
5.4 AML Typology Evolution
Signal:
- Criminal typologies change faster than static investigation playbooks.
- LLM-based drafting improves narrative quality but may fabricate typology claims.
- Regulators expect explainable and evidence-backed SAR narratives.
Decision posture:
- Trial retrieval-grounded investigator copilot.
- Adopt controlled summarization with source citations after eval threshold.
- Hold unsupported typology generation.
Architecture implications:
- typology knowledge base;
- source lineage;
- human attestation;
- case-level audit trail;
- eval set refreshed from new typologies.
5.5 Open Banking and Personal Finance AI
Signal:
- Open banking improves consented data access.
- Personal finance agents promise cashflow coaching, bill negotiation, and credit optimization.
- Privacy, consumer protection, and suitability concerns increase.
Decision posture:
- Watch customer trust and consent behavior.
- Trial personal finance insights with clear disclosure and human support.
- Hold automated product switching or advice where suitability controls are weak.
5.6 Regulatory Reporting
Signal:
- GenAI can draft regulatory report sections.
- Evidence lineage and attestation remain critical.
- Model provider changes can alter output behavior.
Decision posture:
- Trial drafting assistance in shadow mode.
- Invest in evidence graph, source citation, and human attestation workflow.
- Adopt only for controlled drafting where every claim maps to source evidence.
5.7 Scenario-to-Portfolio Synthesis
| Scenario pressure | Product move | Architecture move | Control move |
|---|---|---|---|
| Model commoditization | differentiate on workflow, trust, and data product quality | gateway, eval portability, cost observability | provider comparison and decision freshness |
| Regulatory tightening | prioritize explainable, auditable use cases | evidence graph, release gates, human attestation | obligation mapping and control evidence |
| Agentic operations | start with low-risk bounded actions | tool permission service, workflow state machine, kill switch | SoD, authorization, runtime monitoring |
| Fraud acceleration | invest in adaptive identity and behavioral signals | fraud intelligence layer and step-up authentication | typology refresh and incident playbooks |
| Workforce recomposition | redesign tasks and review burden before scaling | oversight queues and adoption telemetry | training evidence and hidden-workload monitoring |
6. 反模式
| Anti-pattern | Why it fails | Better pattern |
|---|---|---|
| Trend digest without decisions | creates awareness without action | connect every material signal to an assumption or decision |
| Vendor roadmap as strategy | transfers strategic thinking to supplier incentives | maintain vendor watch, exit path, and internal capability map |
| Benchmark worship | public benchmarks rarely represent regulated workflows | use internal evals and task-specific thresholds |
| Adoption by executive enthusiasm | skips workflow, control, and evidence | require experiment charter and release gates |
| Over-platforming too early | spends before demand and standards clarify | fund runway that preserves options across scenarios |
| Holding everything for perfect certainty | loses learning and timing advantage | use small learning options and controlled trials |
| Treating regulation as a final answer | obligations evolve and require interpretation | run regulatory horizon scanning and decision freshness reviews |
| Ignoring workforce disruption | value leaks when roles, incentives, and trust are not redesigned | track behavior change and capability evidence |
| Single-provider dependency by accident | lock-in emerges through logs, prompts, evals, data flows, and approvals | design portability and exit evidence upfront |
| Scenario theater | produces colorful futures but no decisions | attach every scenario to triggers, options, and architecture runway |
| Assumption ledger as documentation only | assumptions never change decisions | require every stale assumption to resolve into refresh, trial, hold, or escalation |
| Architecture runway without stop rules | platform work becomes self-justifying | define option value, scenario coverage, and conditions to stop investing |
7. 最终心智模型
AI foresight is not trend watching. It is the operating system for decisions under uncertainty.
hype filtered into signals
signals mapped to assumptions
assumptions connected to scenarios
scenarios converted to options
options tested through experiments
experiments shaped into architecture runway
runway and evidence drive watch / trial / adopt / hold decisions
decisions stay fresh through cadence, triggers, and executive narratives
最终要形成的不是“AI 趋势判断力”,而是三种组织能力:
- 证据能力:知道一个信号来自哪里、权威性如何、影响哪个假设、证据是否足以改变决策。
- 期权能力:在不确定性高的时候,用小投入保留未来选择,而不是过早锁死在供应商、平台、流程或组织结构里。
- 架构能力:让产品组合、控制证据、供应商管理、运行时可观测性和执行叙事共享同一套 assumptions、triggers 和 decision records。
成熟组织不靠预测未来取胜。它知道自己假设了什么、什么会改变判断、哪些选项正在被保留、哪些决策太不可逆而不能轻率做出。
SOTA 状态标注 (2026-07-01)
本篇属于第二、三遍深读池(参考架构/深读笔记),未列入 12 周主线必读。时效基线为写作时点;引用前请按 CLAUDE.md 全局时效性硬规则复查最新进展。模块级 SOTA 对照见 docs/AI_SYSTEMATIC_LEARNING_ROADMAP_2026.md 各周「2026 SOTA 对照」行与文末「SOTA 检查」。