AI risk policies often sound convincing until a team has to decide whether a system is ready for real people. The US National Institute of Standards and Technology designed its voluntary AI Risk Management Framework, or AI RMF, to help organisations bring trustworthiness considerations into the design, development, use and evaluation of AI systems. Its practical core is organised around four functions: Govern, Map, Measure and Manage.
These functions are not a one-time sequence and they are not a certification label. They are connected kinds of work. Governance shapes every decision; mapping gives measurements context; measurement supplies evidence; and management turns that evidence into priorities and action.
Govern: make responsibility visible
Govern is the foundation. A team needs named owners, escalation routes, policies, documentation expectations and a way to include affected perspectives. It should know who can approve deployment, who can stop it, who investigates an incident and which evidence must be retained.
A useful governance record is specific. Instead of saying “humans remain in control”, identify the decision a human reviews, the information available to that reviewer, the time allowed and what happens when confidence is low. Governance also covers suppliers: a purchased model does not transfer responsibility away from the organisation using it.
Map: understand the system in context
Map asks what the AI system is for, where it will operate and who may experience its benefits or harms. The same model can carry very different risk in a private drafting tool, a hiring screen or a clinical workflow. Teams should describe intended use, foreseeable misuse, affected groups, dependencies, data provenance and the consequences of failure.
Good mapping also records assumptions. Is the input language supported? Will a user know that output is machine-generated? Can a person challenge a result? Which environment was used during testing, and how does production differ? An explicit assumption can be tested; a hidden assumption becomes a surprise.
Measure: collect evidence that answers the mapped questions
Measure is broader than reporting one accuracy score. Depending on the context, evidence may include reliability under changing inputs, subgroup performance, privacy and security tests, harmful-content rates, calibration, accessibility, human-factors studies and post-deployment observations. The metric must connect to the risk described during mapping.
Teams should define thresholds before results arrive, document limitations and keep failed tests rather than selecting only favourable numbers. Some harms are difficult to reduce to a single metric, so qualitative review and feedback from affected people can complement quantitative testing. Measurement reduces uncertainty; it does not eliminate it.
Manage: prioritise, respond and keep watching
Manage turns the risk picture into decisions. A team may mitigate a risk, restrict the use case, add human review, monitor a leading indicator, prepare a rollback or decide not to deploy. Priorities should reflect likelihood, severity and the organisation’s risk tolerance—not merely which fix is easiest.
Management continues after launch. Models, data, user behaviour and external conditions can change. A practical plan therefore includes monitoring, incident handling, change control and a trigger for reassessment. If a model update changes behaviour, an earlier test report is historical evidence, not automatic approval of the new version.
Use profiles to make the framework local
NIST describes profiles as implementations of the framework’s functions, categories and subcategories for a particular setting. A current profile can describe how an organisation manages a use case now; a target profile can describe the desired state. The gap between them creates a concrete improvement backlog.
This tailoring matters because the Playbook is not a checklist that every organisation must follow in full. NIST presents suggested actions that users can select according to their use case, resources and interests. Copying every suggestion without connecting it to a real decision can create paperwork without safer outcomes.
A small-team starting pattern
- Govern: name the owner, reviewer, stop authority and evidence location.
- Map: write the intended use, affected people, main failure modes and assumptions.
- Measure: choose tests and thresholds that address those failure modes.
- Manage: decide what must change before launch and what will trigger rollback or review afterward.
Run the four functions again when the use case, model, data or operating environment changes. The value is not the number of documents produced. It is whether the organisation can explain what it knew, what it tested, what it decided and who remains accountable.
What the framework can—and cannot—prove
Using the AI RMF can improve discipline and shared language, but it does not by itself prove that a system is safe, lawful or suitable. NIST describes the framework as voluntary, and its supporting Playbook as a flexible resource. Legal obligations, sector rules and independent assurance may still apply.
The responsible claim is therefore modest: the four functions help teams ask better questions and connect principles to evidence-backed decisions. They are a structure for ongoing risk management, not a badge that ends the conversation.
Primary sources
NIST AI Risk Management Framework and NIST AI RMF Playbook. Accessed 11 September 2026.