Compliance · intermediate · 2026

NIST AI RMF Implementation: Complete Enterprise Checklist (2026)

How to implement the NIST AI Risk Management Framework for LLMs and autonomous agents — the four core functions mapped to AIZA-HexTyx controls, with a complete enterprise checklist.

17 min read
In This Guide
1. What the NIST AI RMF Is — and Why It Matters Now 2. GOVERN — Establishing Accountability and Policy 3. MAP — Identifying Risk Context 4. MEASURE — Quantifying AI Risk 5. MANAGE — Implementing and Updating Controls 6. Implementation Checklist

What the NIST AI RMF Is — and Why It Matters Now

The NIST AI Risk Management Framework is not a compliance checkbox. It is a structured methodology for making AI systems more trustworthy — covering the full lifecycle from design through deployment, monitoring, and retirement. Unlike the EU AI Act, which is a legal requirement with specific obligations and penalties, the AI RMF is a voluntary framework. But "voluntary" understates its real-world importance: it is increasingly the framework procurement teams, government agencies, and enterprise security reviewers use to evaluate AI vendor trustworthiness.

In 2026, organizations deploying LLMs, autonomous agents, and RAG systems face two converging pressures: regulatory requirements that increasingly reference NIST AI RMF controls, and enterprise buyers who ask for AI RMF alignment documentation before signing contracts. The framework covers ground that neither OWASP LLM Top 10 nor EU AI Act fully addresses — particularly the GOVERN and MAP functions, which deal with organizational accountability and risk context rather than technical controls.

The four functions — GOVERN, MAP, MEASURE, MANAGE — are not sequential phases. They are concurrent, continuous activities that reinforce each other. GOVERN creates the policy context that makes MEASURE meaningful. MAP identifies the risks that MANAGE must mitigate. MEASURE provides the evidence that GOVERN uses to adjust policies. Understanding this cycle is the key to implementing the framework rather than just documenting it.

AIZA-HexTyx maps to 15 NIST AI RMF controls across all four functions. The compliance_map.md file in the deployment package documents each control with its corresponding evidence artefact. See the governance framework: AI Agent Governance Framework: Complete 2026 Checklist →

GOVERN — Establishing Accountability and Policy

GOVERN is the foundation. Without it, MEASURE produces numbers nobody acts on, and MANAGE implements controls nobody owns. The five GOVERN controls that AIZA-HexTyx addresses:

GOV-1.1 — Policies, processes, and procedures are in place. The Aegis policy engine provides programmable, versioned, hot-reloadable rules. Every rule change is logged with a timestamp. The policy snapshot (GET /gateway/policy/rules) is the GOV-1.1 evidence artefact — it shows what policies are active at any given moment and when they changed.

GOV-1.2 — Accountability and responsibility are defined. The AEGIS_TENANT_ID environment variable namespaces all audit records to a specific organizational unit. Webhook registration to your SIEM assigns operational responsibility — the team that receives the webhook alert is the team accountable for the event.

GOV-1.3 — Organizational teams understand their AI roles. The Aegis dashboard and AXIOM analyst provide role-appropriate access: security leads see raw findings and calibration history; AI engineers see read-only scan results and remediation guidance; executive stakeholders see the PDF report and Executive Markdown summary.

GOV-4.1 — Risk management processes include AI-specific considerations. The calibration loop — where HexTyx scan findings automatically update Aegis detection rules — is the operational expression of GOV-4.1. It is not a policy document about risk management; it is a running system that executes risk management automatically.

GOV-6.1 — Policies are reviewed and updated. Policy hot-reload (POST /gateway/policy/reload) without gateway restart means policy updates can be deployed immediately when new attack patterns emerge. The calibration history shows when updates occurred and why.

The most common GOVERN failure: policies exist in documents but are not enforced in running systems. If your Aegis policy engine has not been tested with POST /gateway/calibrate and the calibration loop is not connected, your GOVERN documentation is aspirational, not operational.

MAP — Identifying Risk Context

MAP is where you understand what you are actually protecting and what can go wrong. Four MAP controls:

MAP-1.1 — Context is established. System personas in AIZA-HexTyx define the operational context for each AI endpoint: support_agent, it_helpdesk, finance_advisor. Each persona has a known vulnerability profile, expected tool access, and tenant scope. This structured context definition is the MAP-1.1 artefact — it documents what each AI component is supposed to do and for whom.

MAP-2.1 — AI system stakeholders are identified. The tenant_id and user_id fields in every audit record link every AI decision to a specific organizational entity and user. This creates the stakeholder accountability chain MAP-2.1 requires.

MAP-3.1 — AI system capabilities and limitations are understood. HexTyx scan output documents exactly what each endpoint's limitations are — which attack categories succeed, at what bypass rate, with what confidence. This is not a marketing claim about model capability; it is measured evidence of actual behavior under adversarial conditions.

MAP-5.1 — Likelihood and impact are evaluated. The risk score (0-100) and severity classification (critical/high/medium/low) in every HexTyx finding operationalize MAP-5.1. Risk scores above 70 trigger automatic policy escalation in the calibration loop — turning the MAP assessment into an automated MANAGE action.

For the complete checklist approach: LLM Security Checklist: 40 Things to Check Before Launch →

MEASURE — Quantifying AI Risk

MEASURE is where most organizations fail first. They have governance policies (GOVERN) and risk awareness (MAP), but no quantified measurement system. The five MEASURE controls:

MEASURE-1.1 — Risk evaluation methods are established. HexTyx's dual-layer evaluator (regex pattern matching + LLM-as-judge + ground truth verification) is the risk evaluation method. The three confidence tiers — possible, probable, confirmed — map directly to MEASURE-1.1's requirement for graduated risk assessment rather than binary pass/fail.

MEASURE-2.1 — AI risks are assessed. Every HexTyx scan produces a JSON risk assessment: target risk score, bypass rate per mutation strategy, confirmed leaks, cascade amplification factor. These are the MEASURE-2.1 artefacts — not qualitative assessments but quantified measurements with reproducible methodology.

MEASURE-2.6 — AI system performance is monitored. Continuous Aegis monitoring tracks live system performance against the baseline established in MEASURE-2.1. When runtime behavior diverges from the scan baseline — when a new bypass pattern appears in production that HexTyx did not test for — the calibration loop captures it and updates the detection library.

MEASURE-3.1 — Risk metrics are tracked. The Memory Graph provides the longitudinal risk tracking MEASURE-3.1 requires. Attack cluster growth, cross-model transfer rates, and bypass rate trends over multiple scan runs give you the time-series data that shows whether your AI security posture is improving or degrading.

MEASURE-4.1 — Risk measurement results are documented. Seven report formats — JSON, SARIF, STIX 2.1, MITRE ATT&CK Navigator layer, Executive Markdown, PDF, JUnit — provide the documentation artefacts MEASURE-4.1 requires in whatever format your compliance review demands.

MANAGE — Implementing and Updating Controls

MANAGE is where policy becomes action. Five MANAGE controls:

MANAGE-1.1 — Response plans are established. The Aegis policy engine's action taxonomy — block, redact, alert, flag — is the response plan. Every finding type in every HexTyx scan maps to a corresponding Aegis action. The policy rules snapshot shows which actions are mapped to which finding types.

MANAGE-1.3 — Monitoring occurs continuously. Aegis runs five checkpoints on every request in under 40ms. This is not periodic monitoring — it is continuous, request-level governance. The audit log accumulates evidence of every decision, every block, every redaction.

MANAGE-2.2 — Risks are prioritized for treatment. The calibration loop prioritizes risk treatment automatically: high-bypass attack patterns are promoted to the injection library first; risk score thresholds are tightened where evasion is detected; policy rules are escalated when bypass rates exceed 60%. Human prioritization is a fallback, not the primary mechanism.

MANAGE-3.1 — Risks are mitigated. The complete AIZA-HexTyx mitigation stack: Aegis CP1 (injection detection), CP2 (RAG chunk validation), CP3 (tool call interception), CP4 (output secret redaction), CP5 (async audit). TaintTracker for agent cascade. MultiTurnTracker for session-wide risk. Dead Man's Switch for canary protection.

MANAGE-4.1 — Lessons are incorporated. Every CalibrationNode in the Memory Graph is a documented lesson incorporated: what attack pattern was discovered, when it was detected, what rule was created, what bypass rate changed. This is the MANAGE-4.1 evidence trail — not a post-incident retrospective, but a continuous learning record.

For the full automated testing approach: Automated LLM Security Testing: The Complete 2026 Guide →

The existing NIST AI RMF overview and control mapping reference: NIST AI RMF Implementation Guide for LLM Products →

Implementation Checklist

For the EU AI Act compliance checklist that complements NIST AI RMF: EU AI Act Requirements & Checklist: Complete Enterprise Guide →

GOVERN checklist:

MAP checklist:

MEASURE checklist:

MANAGE checklist:

Generate Your NIST AI RMF Evidence Package

Run a full HexTyx scan and get OWASP mapping, MITRE ATT&CK Navigator layer, SARIF output, calibration history, and policy snapshot — covering GOVERN, MAP, MEASURE, and MANAGE artefacts in one run.

Start NIST AI RMF Assessment →