AI Requirements Mining / Process Knowledge Extraction Playbook
Requirements mining 不是把文档丢给大模型生成 backlog,而是把多源业务、流程、控制和生产证据转成可追溯、可评估、可验证、可治理、可复用的 requirement and process knowledge assets。AI 可以扩大证据面、发现冲突、聚类噪声和生成候选资产;人类专家负责解释、取舍、授权、治理和上线责任。
AI Requirements Mining / Process Knowledge Extraction Playbook
配对阅读:本手册的原理/架构解读版是
docs/ai-foundations/papers/148-ai-requirements-mining-process-knowledge-extraction-architecture.md。先读 paper 建立机制与取舍,再用本手册落地为模板、RACI 与门禁,两者不需要重复精读。
Requirements mining 不是把文档丢给大模型生成 backlog,而是把多源业务、流程、控制和生产证据转成可追溯、可评估、可验证、可治理、可复用的 requirement and process knowledge assets。AI 可以扩大证据面、发现冲突、聚类噪声和生成候选资产;人类专家负责解释、取舍、授权、治理和上线责任。
在金融零售环境中,PRD、BRD、SOP、政策、工单、会议纪要、通话转写、流程图、API 规格、测试用例、生产日志和控制证据都可能包含需求线索。但这些来源的权威级别、版本、生效日期、权限、隐私、记录保留和业务语义完全不同。高级 requirements mining 的重点,是让系统“广泛挖掘、窄口信任、显式验证、全链路追踪”。
1. Source Anchors
| Anchor | Official link | 本文使用方式 |
|---|---|---|
| ISO/IEC/IEEE 29148 Requirements Engineering | https://www.iso.org/standard/72089.html 和 https://standards.ieee.org/ieee/29148/6937/ | 作为需求质量、生命周期、traceability 和 stakeholder need 管理的标准化锚点。 |
| FFIEC Development, Acquisition, and Maintenance IT Handbook | https://ithandbook.ffiec.gov/it-booklets/development-acquisition-and-maintenance/ | 约束 AI 需求挖掘输出如何进入开发、采购、测试、实施、维护和变更控制。 |
| FFIEC Management IT Handbook | https://ithandbook.ffiec.gov/it-booklets/management/ | 组织治理、风险管理、架构、资源、第三方和监督责任。 |
| NIST SP 800-160 Vol. 1 | https://csrc.nist.gov/pubs/sp/800/160/v1/upd2/final | 把 stakeholder protection needs、security、resilience、assurance 纳入需求抽取和架构评审。 |
| NIST SP 800-218 SSDF | https://csrc.nist.gov/pubs/sp/800/218/final | 将代码、API、测试资产挖掘连接到安全软件开发、漏洞响应和 release evidence。 |
| NIST AI Risk Management Framework | https://www.nist.gov/itl/ai-risk-management-framework | 用 Govern / Map / Measure / Manage 设计 AI 风险识别、评估、门禁和持续改进。 |
| ISO/IEC 42001 AI Management System | https://www.iso.org/standard/81230.html | 用 AI management system 语言设计 policy、role、operation、performance evaluation、internal audit 和 improvement。 |
Source-to-artifact pattern:
official source anchor
-> governance principle
-> mining requirement
-> architecture control
-> evidence artifact
-> owner and metric
2. Executive Framing
高管或业务方常见诉求通常像这样:
我们有很多 PRD、SOP 和 ticket,能不能让 AI 自动生成需求?
我们想把会议记录和客服通话转成 backlog。
我们能不能让 AI 看代码和测试用例,反推出系统需求?
这些诉求需要被改写为系统能力:
Build an evidence-grade requirements and process knowledge extraction capability
that mines candidate needs, rules, events, controls, acceptance criteria and impact links
from governed artifacts,
filters by source authority and permissions,
routes ambiguity to SMEs,
and learns from production feedback without turning AI drafts into approved requirements.
Steering questions:
- 哪些 artifact 是权威来源,哪些只是 pain signal 或 discussion evidence?
- AI 输出进入 baseline 前由谁验证、用什么 rubric、保留什么证据?
- 如何防止越权检索、过期政策引用、PII 泄露和记录处置失控?
- 如何把 mined requirements 连接到 process、API、data、test、control、eval 和 release?
- 如何从 production logs、QA、complaints、incidents 和 human overrides 回流到 portfolio learning?
3. Use Case Boundary
Requirements mining 适合产生候选、冲突、证据和影响分析,不适合直接产生批准结论。
| Use case | Good fit | Boundary |
|---|---|---|
| Requirements discovery | 从多源材料中发现候选需求、冲突、遗漏、重复 | 不自动进入 approved baseline |
| Process knowledge extraction | 从 SOP、流程图、logs、tickets 中抽 activity、event、role、handoff、variant | 不把日志行为直接等同于应然流程 |
| Acceptance criteria drafting | 根据需求和测试资产生成验收候选 | 高影响场景必须 SME 和 QA 验证 |
| Change impact analysis | 政策、API、流程、控制变化后找影响面 | 影响结论必须由 owner 确认 |
| Control linkage | 把 policy、control、test evidence 连接到需求 | 不给法律或合规适用性结论 |
| Portfolio learning | 沉淀复用词汇、模式、eval cases、anti-patterns | 不用未授权 records 或 PII 做无边界训练 |
不适合直接交给 AI 的任务:
| Task | Reason |
|---|---|
| 最终 scope tradeoff | 涉及商业优先级、资源、风险接受和战略选择 |
| 法律或监管解释 | 需要授权职能结合具体事实和管辖范围判断 |
| 高影响客户决策 | 需要授权、控制、解释、申诉和人工责任 |
| records retention 或 legal hold 结论 | 需要 Records、Legal、Compliance 决策 |
| 模型验证结论 | 需要独立模型风险和验证程序 |
4. Target Operating Model
Business / Product Owner
owns outcome, priority, scope, baseline decision
Requirements Architect
owns mining taxonomy, ambiguity workflow, quality rubric, traceability graph
Process Owner / SME
validates process activities, exceptions, variants and operating feasibility
Architecture / Engineering
validates API, data, system, security, performance and integration impact
Risk / Compliance / Control Owner
validates policy/control linkage, risk tier, approval boundary and evidence need
Privacy / Records / Legal
validates data use, records class, hold propagation, access and retention controls
QA / EvalOps
converts mined requirements to tests, eval contracts, thresholds and regression gates
AI Platform / Data Engineering
operates ingestion, retrieval, permission filter, graph, model versioning and monitoring
RACI snapshot:
| Activity | Business owner | Requirements owner | Architect | SME | Risk / Control | Privacy / Records | QA / EvalOps |
|---|---|---|---|---|---|---|---|
| Source inventory | A | R | C | C | C | C | C |
| Authority classification | A | R | C | C | C | C | C |
| Extraction rubric | C | A/R | C | C | C | C | R |
| Requirement validation | A | R | C | R | C | C | C |
| Risk tiering | A | R | C | C | A/R | C | C |
| Eval contract | C | R | C | C | C | C | A/R |
| Release gate evidence | A | R | R | C | C | C | R |
| Portfolio learning | A | R | C | C | C | C | R |
5. Implementation Architecture
Connectors
Confluence / SharePoint / Docs / Jira / Azure DevOps / Git / API gateway
Contact center / CRM / case management / log platform / GRC / test management
|
v
Governed ingestion
artifact id | source type | owner | approval status | version | hash
effective date | data class | record class | permission tag | legal hold flag
|
v
Pre-processing
parsing | OCR/layout | transcript diarization | code/API parsing | test extraction
log event mapping | PII redaction | chunking | metadata enrichment
|
v
Knowledge layer
domain vocabulary | process ontology | authority ladder | policy/control map
requirement graph | event graph | system/API graph | test/eval graph
|
v
AI extraction and reasoning services
candidate requirements | ambiguity | conflict | duplicate | process events
stakeholder concerns | evidence standards | acceptance criteria | impact links
|
v
Human validation workbench
side-by-side source evidence | approve/reject/merge/split/escalate
reason codes | SME comments | decision record | audit trail
|
v
Delivery and governance
backlog sync | requirement baseline | eval contract | test generation
architecture review | control evidence | release gate | learning loop
Architecture non-negotiables:
| Non-negotiable | Why |
|---|---|
| 权限先于检索 | 防止用户通过 AI 摘要看到无权材料 |
| artifact hash and version | 支持复现、审计和变更影响 |
| authority metadata | 防止 ticket、会议纪要和 AI draft 覆盖正式政策 |
| structured output schema | 防止顺滑文本掩盖冲突和不确定性 |
| SME decision log | AI 只生成候选,人负责授权 |
| eval contract | 需求挖掘能力本身也要被评估和门禁 |
| graph traceability | 支持跨需求、流程、系统、测试、控制、release 的 impact analysis |
6. Source Intake and Authority
No artifact enters the mining corpus without owner, version, permission tag and source class. No restricted source enters AI processing without approved purpose and redaction path.
| Source | Required metadata | Extraction focus | Key risk |
|---|---|---|---|
| PRD | owner、version、status、target release、approval | feature、persona、metric、scope、assumption | solution bias |
| BRD | business owner、benefit baseline、decision date | outcome、stakeholder need、policy constraint | vague benefit |
| SOP | process owner、effective date、retired status | activity、role、SLA、exception、evidence | stale process |
| Policy/control | policy owner、effective date、scope、control id | obligation、allowed/prohibited action、approval | misinterpretation |
| Tickets | severity、product、status、linked incident、resolution | pain、defect、workaround、frequency | duplicate noise |
| Work items | workflow status、links、sprint/release、acceptance criteria | backlog、dependency、test/release links | weak traceability |
| Transcripts | consent/notice、channel、QA score、redaction | intent、friction、agent action、complaint signal | PII and transcription error |
| Meeting notes | attendees、roles、decision status、follow-up | decision、assumption、open issue | non-authoritative |
| Process maps | version、notation、owner、scope | intended flow、roles、controls、SLA | idealized flow |
| Code/API specs | repo/version、endpoint、owner、deployment | actual behavior、contract、validation、error | code as false policy |
| Test cases | test owner、result、requirement link、coverage | expected behavior、edge case、regression | happy-path bias |
| Logs | event schema、retention、sampling、data class | variant、latency、failure、handoff、outcome | missing business semantics |
| Controls | control owner、frequency、test result、issue link | control objective、evidence、remediation | control/product disconnect |
Authority decision rules:
| Condition | Decision |
|---|---|
| Approved source, current version, clear owner, permission scoped | Use for extraction and baseline evidence |
| Approved source but expired or superseded | Use only for historical change impact |
| Operational source with high frequency pain signal | Use for discovery, not baseline |
| Meeting note with unapproved decision | Use as clarification prompt and decision candidate |
| Artifact contains restricted data beyond purpose | Exclude or redact before indexing |
| Source owner unknown | Quarantine until ownership is established |
7. Candidate Movement Rules
AI output moves through governed states. It is never approved by generation alone.
| Requirement candidate condition | Backlog action |
|---|---|
| Grounded, no conflict, quality score >= 4, SME approved | Create backlog item with source links |
| Grounded but ambiguous | Create clarification task, not delivery story |
| Conflict between policy and operational practice | Create issue or decision item, not feature story |
| High-impact AI behavior without eval contract | Block from release backlog |
| Source is only ticket or transcript | Convert to problem statement or pain cluster |
| Derived from code or test only | Mark as actual behavior candidate and request owner decision |
Human review level:
| Risk tier | Example | Review requirement |
|---|---|---|
| Low | internal UI label, non-material routing hint | requirements review and sampling |
| Medium | employee workflow recommendation, non-customer-impact field | SME approval and QA test |
| High | customer money, access, eligibility, complaint, regulated communication | product, risk/control, SME and QA approval |
| Restricted | legal hold, sensitive identity, fraud, vulnerability, privileged tool action | specialized owner review and documented gate |
8. Extraction Prompt Contract
Extraction prompts are controlled product artifacts. They must say what to extract, what not to infer, how to expose uncertainty and how to preserve source authority.
Required output schema:
{
"candidate_id": "string",
"candidate_type": "business_requirement | stakeholder_requirement | solution_requirement | transition_requirement | control_requirement | data_requirement | process_rule | acceptance_criterion | risk_issue",
"statement": "string",
"source_refs": [
{
"artifact_id": "string",
"location": "section/page/span/event_id",
"authority_level": "A1|A2|A3|A4|A5|A6",
"effective_date": "YYYY-MM-DD"
}
],
"known_facts": ["string"],
"unknowns": ["string"],
"ambiguity_flags": ["actor_unknown", "decision_boundary_unknown", "data_scope_unknown", "control_owner_missing"],
"conflicts": ["string"],
"quality_score": 0,
"recommended_acceptance_criteria": ["string"],
"validation_owner": "role",
"risk_tier": "low|medium|high|restricted"
}
Prompt rules:
| Rule | Rationale |
|---|---|
| Do not invent missing business rules | 缺证据必须标 unknown |
| Preserve source authority | 不同来源不能被平均化 |
| Separate current behavior from desired behavior | 代码和日志代表事实,不代表应然 |
| Produce questions, not false certainty | 模糊需求需要澄清 |
| Cite exact source spans | 支持 SME 快速验证 |
| Flag policy/control conflicts | 冲突发现是高价值输出 |
| Avoid customer commitments | mined output 不能成为客户可见承诺 |
9. Requirement Quality Gate
| Gate | Pass signal |
|---|---|
| Source grounded | 每个 statement 有 artifact、location、version、owner |
| Authority clear | source level and conflict policy visible |
| Actor clear | customer、employee、system、team、approver 不混淆 |
| Decision boundary clear | read / summarize / recommend / draft / decide / act 已区分 |
| Data boundary clear | source fields、purpose、permission、retention 已定义 |
| Control linkage clear | approval、dual control、review、audit evidence 已连接 |
| Acceptance testable | positive、negative、edge case 和 evidence requirement 已写 |
| Eval ready | dataset、rubric、threshold、critical failure、slice 已定义 |
| Change impact traceable | process、API、test、control、release links 存在 |
| Owner accountable | business owner、SME、architect、QA/EvalOps owner 清楚 |
Scoring interpretation:
| Score | Meaning | Allowed action |
|---|---|---|
| 0 | wrong or unsupported | reject and log reason |
| 1 | discovery note | keep in evidence cluster |
| 2 | grounded but incomplete | send to clarification |
| 3 | clear but not testable/control-linked | improve before backlog |
| 4 | backlog-ready candidate | create item with source links |
| 5 | baseline-ready for high-impact use | release gate eligible after eval |
10. Traceability Graph
Graph structure turns mined text into impact analysis capability.
| Layer | Nodes | Edges |
|---|---|---|
| Strategy | outcome、KPI、benefit hypothesis、risk appetite | justifies、constrains |
| Stakeholder | role、need、concern、decision right | owns、approves、challenges |
| Requirement | candidate、baseline、acceptance criteria | derives_from、verifies |
| Process | activity、event、variant、handoff、exception | precedes、deviates_from、controls |
| System | API、data object、service、UI、code rule | implements、depends_on |
| Quality | test case、eval case、rubric、threshold | verifies、blocks |
| Control | policy、control objective、evidence、issue | constrains、monitors |
| Delivery | backlog、release、change request、incident | delivers、remediates |
Minimum graph queries:
| Query | Why it matters |
|---|---|
| Show all requirements derived from retired SOP sections | 防止过期来源继续驱动 backlog |
| Show requirements without acceptance criteria | 找不可验收需求 |
| Show high-risk requirements without eval contract | 找上线阻断项 |
| Show policy changes impacting AI prompts or retrieval corpus | 防止过期政策输出 |
| Show API schema changes impacting controls and tests | 支持 release impact review |
| Show tickets repeatedly linked to rejected requirements | 识别真实 pain 但方案不对 |
| Show production variants not covered by SOP | 识别流程治理机会 |
11. Process Variant Discovery
生产日志揭示 work-as-done,但需要业务语义化才能成为流程知识。
Input pattern:
case_id, activity, timestamp, resource, lifecycle, channel, product,
risk_tier, amount_band, status, outcome, source_system
Steps:
- 定义 case 粒度:application、dispute、alert、ticket、complaint、service request。
- 标准化 activity:避免把状态码直接当业务活动。
- 生成 top variants:找覆盖 80% 体量的主要路径和高风险长尾。
- 标记 rework、waiting、handoff、skip、loop、override。
- 与 SOP、process map 和 control path 对齐,区分 acceptable exception、control gap、data noise。
- 生成 AI opportunity candidates:summarize、route、draft、validate、retrieve、recommend、tool action。
- 将每个机会连接到 requirement、acceptance criteria、control 和 eval。
Variant interpretation:
| Finding | Product implication | Control implication |
|---|---|---|
| 主路径覆盖低 | 不宜直接自动化,先治理流程和 taxonomy | 例外处理和控制路径需补齐 |
| 高 rework | 改进资料收集、校验、政策解释 | 监控返工原因和 customer harm |
| 多团队 handoff | AI handoff summary 或队列路由 | 责任和 evidence transfer 要清楚 |
| 控制步骤被跳过 | 阻断上线,先修复流程或权限 | control issue and remediation |
| 高等待来自外部资料 | 客户、第三方提醒和 SLA 管理 | 记录通知和暂停计时逻辑 |
| override 集中在某团队 | policy ambiguity 或 training gap | dual control / QA sampling |
12. Eval Contract for the Mining System
Mining system 本身也必须被评估。否则组织会把未经验证的抽取系统当成事实来源。
| Eval area | Metric | Release threshold idea |
|---|---|---|
| Requirement extraction precision | AI candidates accepted as valid by SME | high enough by source class, no critical false positives |
| Critical recall | must-have policy/control/exception requirements found | zero missed critical control in golden set |
| Groundedness | statements fully supported by source refs | unsupported material claim = release blocker |
| Authority classification | source authority correctly ranked | no low-authority source overriding approved source |
| Ambiguity detection | required clarifications correctly flagged | high-risk ambiguity miss = blocker |
| Conflict detection | known conflicts identified | policy/SOP/log conflict misses reviewed |
| Permission safety | no unauthorized source leakage | zero leakage in red-team tests |
| Output schema validity | machine-readable structured output | near-perfect schema compliance |
| SME efficiency | review time per candidate | improves without lowering quality |
| Change impact quality | impacted systems/tests/controls found | validated against known changes |
Critical failures:
| Critical failure | Why it blocks release |
|---|---|
| hallucinated source citation | 破坏证据链 |
| unauthorized PII or restricted source in output | 权限和隐私不可接受 |
| policy/control requirement missed in high-impact workflow | 可能造成控制缺口 |
| low-authority source treated as approved baseline | 需求基线被污染 |
| AI-generated customer commitment | 候选资产越界为外部承诺 |
| hidden conflict between source materials | 冲突被平滑掩盖 |
| output enters backlog without validation evidence | 治理边界失效 |
13. Evidence and Control Checklist
Pre-launch
| Control area | Evidence |
|---|---|
| Source governance | inventory、owner、version、permission、retention、record class |
| Data protection | privacy review where applicable、redaction rules、access matrix |
| Records | record class、legal hold propagation、derived artifact retention |
| Security | connector entitlement、secrets handling、audit logging、vendor boundary |
| Model governance | model card、prompt version、eval results、limitations |
| EvalOps | golden set、rubric、thresholds、critical failures、independent review |
| SME operations | reviewer guide、decision codes、escalation path |
| Traceability | graph schema、source-to-requirement links、impact queries |
| Release | go/no-go memo、exceptions、risk acceptance record |
Production
| Control area | Evidence |
|---|---|
| Usage monitoring | who mined what、source classes、exports、backlog sync |
| Quality monitoring | acceptance rate、reject reasons、ambiguity density、conflict misses |
| Permission monitoring | denied retrievals、redaction events、suspicious access |
| Drift monitoring | source freshness、vocabulary drift、new ticket clusters |
| Change monitoring | policy/API/SOP/model/prompt changes and regression eval |
| Incident handling | leakage、hallucination、wrong baseline、control miss、remediation |
| Portfolio learning | reusable patterns、updated rubrics、added eval cases |
14. 30 / 60 / 90 Roadmap
First 30 days: controlled discovery
| Workstream | Output |
|---|---|
| Select domain | one workflow, e.g., payment dispute, KYC onboarding, fee servicing, AML alert triage |
| Inventory sources | PRD/BRD/SOP/policy/tickets/transcripts/process maps/tests/logs/control evidence |
| Define authority ladder | source classes, approval status, conflict rules |
| Define vocabulary | key terms, role names, activity taxonomy, forbidden ambiguous terms |
| Build small corpus | permission-filtered, redacted, versioned artifact set |
| Design rubric | quality score, ambiguity flags, conflict categories |
| Create golden set | SME-labeled requirements, controls, events, conflicts |
| Run pilot extraction | candidates, source refs, ambiguity questions, initial graph |
Exit criteria:
The team can show source-backed candidates, rejected examples, ambiguity log,
and at least one requirement-to-test-to-control trace for the selected workflow.
Days 31-60: graph, eval and SME workflow
| Workstream | Output |
|---|---|
| Traceability graph | outcome -> requirement -> process -> API/data -> test -> control |
| SME workbench | approve/reject/merge/split/escalate with reason codes |
| Eval contract | dataset, rubric, slices, thresholds, critical failures |
| Process mining link | top variants, rework, handoff, waiting, control gap |
| Backlog integration | only approved candidates sync to delivery system |
| Change impact queries | policy/API/SOP/test changes show impacted assets |
| Control pack | permission audit, evidence pack, release gate memo |
Exit criteria:
The mining system can pass golden-set eval, route ambiguous outputs to SMEs,
and create backlog items only with source refs, quality score and validation evidence.
Days 61-90: production pilot and portfolio learning
| Workstream | Output |
|---|---|
| Production pilot | limited users, limited corpus, high audit logging |
| Monitoring | quality, permission, source freshness, SME decisions, backlog conversion |
| Regression eval | triggered by policy/SOP/API/model/prompt/corpus changes |
| Incident drill | hallucinated source, permission leak, wrong baseline, missed control |
| Portfolio pattern library | reusable requirement patterns, acceptance criteria, eval cases |
| Operating model | RACI, governance cadence, funding and scaling decision |
| Executive review | value evidence, risk issues, expansion roadmap |
Exit criteria:
The organization can demonstrate faster discovery, better traceability,
measurable SME productivity, controlled risk, and reusable portfolio assets.
15. Metrics
| Metric | Meaning |
|---|---|
| discovery cycle time reduction | 从 source intake 到 validated candidate 的时间 |
| validated candidate yield | 每 100 个 artifact 产生的高质量需求数 |
| duplicate reduction | 合并重复 ticket、story、requirement 的比例 |
| clarification throughput | ambiguity 从发现到关闭的时间 |
| backlog quality lift | approved story 的 source refs、acceptance criteria、test link 完整度提升 |
| change impact lead time | 变更影响分析时间 |
| reuse rate | pattern、acceptance criteria、eval case、vocabulary 的复用比例 |
| unsupported claim rate | AI 输出无来源支持的比例 |
| wrong authority rate | 权威等级识别错误或低权威覆盖高权威 |
| critical recall miss | 漏掉高影响政策、控制或例外 |
| permission leakage rate | 未授权信息在输出中出现 |
| ambiguity miss rate | 人工发现但 AI 未标注的关键模糊点 |
| conflict miss rate | 已知冲突未识别 |
| SME disagreement rate | SME 对输出解释不一致 |
| production feedback incorporation | incident、QA、ticket 回流 eval 的速度 |
16. Anti-Patterns and Repairs
| Anti-pattern | Symptom | Repair |
|---|---|---|
| “AI 生成 user story 工厂” | backlog 变多,质量更差 | 强制 source_refs、quality score、SME approval |
| “一个向量库装所有文档” | 权限、版本、记录边界失控 | source registry + retrieval-time ACL + corpus partition |
| “只看文档不看日志” | 自动化理想流程,忽略真实变体 | connect event logs and process mining |
| “只看 ticket 排优先级” | 高频噪声盖过高风险需求 | severity、journey、control、value weighted scoring |
| “让 AI 消除冲突” | 输出顺滑但错误 | show conflicts and route to owner |
| “代码就是需求” | 历史缺陷被产品化 | actual behavior vs desired requirement 分离 |
| “eval 只测摘要质量” | 需求挖掘错误进 backlog | test extraction, authority, conflict, permission and impact |
| “SME review 无结构” | 审过但不可复用 | reason codes and decision log |
| “不治理 derived artifacts” | summary、embedding、graph node 记录风险 | records/privacy controls for all derived artifacts |
| “pilot 后不学习” | 每个团队重复踩坑 | portfolio pattern library and eval expansion |
17. Evidence Artifact Structures
Requirements mining 产物应当以证据字段表达,而不是以可填写表单堆砌。下面三个结构用于约束 intake、candidate review 和 change impact 的最低证据粒度。
17.1 Mining Intake Brief
| Field group | Evidence content | Quality bar |
|---|---|---|
| Workflow and ownership | workflow、business owner、process owner、primary outcome | 能说明为什么选择这个流程,以及谁对结果负责 |
| Risk and source scope | risk tier、source classes included、source classes excluded | 能解释哪些材料进入语料,哪些因权限、记录或质量原因排除 |
| Governance boundary | permission model、record classes、SME reviewers | 能证明检索、抽取和人工验证在授权边界内 |
| Evaluation boundary | quality rubric、eval dataset、release decision owner | 能证明 mining system 本身被评估,输出不会自动进入 baseline |
17.2 Requirement Candidate Review Card
| Field group | Evidence content | Quality bar |
|---|---|---|
| Candidate statement | statement、candidate type、risk tier | 表达应可测试、可追溯,不混合多个需求 |
| Source grounding | source refs、authority level、effective date | 每个 material claim 都能定位到来源和版本 |
| Uncertainty | known facts、unknowns、ambiguity flags、conflicts | 不把缺失信息包装成确定结论 |
| Traceability | linked process activity、API/data/test/control | 候选需求能连接到流程、系统、测试和控制 |
| Review decision | review decision、reason code、owner | 人工验证决定可审计,并能进入后续改进 |
17.3 Change Impact Memo
| Field group | Evidence content | Quality bar |
|---|---|---|
| Change identity | changed artifact、change type、effective date | 说明变化对象、变化性质和生效时间 |
| Impact surface | impacted requirements、workflows、APIs、data objects | 不只列需求,也列系统和流程影响 |
| Assurance impact | impacted tests、evals、controls、records | 说明哪些验证、控制和记录需要更新 |
| Decision path | required approvals、release implication、monitoring update | 明确是否需要 release gate、回归评估或生产监控调整 |
18. Decision Narrative
Requirements mining 的端到端架构从 governed ingestion 开始,对 PRD、SOP、policy、tickets、transcripts、work items、code/API、tests、logs 和 controls 做 source inventory、authority classification、permission filtering 和 versioning。系统用 domain vocabulary 和 process ontology 做结构化抽取,生成 requirement candidate、process event、stakeholder concern、control link、acceptance criteria 和 impact links。所有输出进入 traceability graph,经 quality rubric、eval contract 和 SME validation 后,才进入 backlog 或 baseline。
它不是普通 RAG summarization。RAG 总结回答“这些材料说了什么”,requirements mining 要回答“哪些内容可以成为需求、依据是什么、权威级别如何、哪里冲突、哪里模糊、谁验证、怎么测试、影响哪些系统和控制、上线后怎么监控”。核心资产不是摘要,而是 source-backed graph、quality score、eval contract、SME decision log 和 change impact evidence。
低权威来源应作为 discovery signal,而不是 baseline。通话转写可能错,客户表达可能是情绪或投诉,会议纪要可能不是批准决定,ticket 可能是重复噪声。进入 baseline 需要更高权威来源或 owner 追认。
冲突图是高价值发现。政策和控制是约束,SOP 是 intended process,生产日志是真实行为。系统不应让 AI 平滑合并冲突,因为冲突本身可能代表流程绕行、控制缺口、SOP 过期、系统缺陷或政策解释不清。
生产日志揭示 work-as-done。文档挖掘告诉我们设计意图,日志挖掘告诉我们真实行为。两者结合,才能决定该做 AI assistant、流程重构、数据质量改进、API 集成,还是控制修复。
隐私、records 和权限必须先于检索。raw artifacts、chunks、embeddings、summaries、review notes 和 graph nodes 都可能成为 derived artifacts,必须有 retention、legal hold propagation、purpose limitation 和 audit log。
核心原则:
Mine broadly, trust narrowly, validate explicitly, trace everything,
and let production evidence improve the operating knowledge base.