Key takeaways
- The NIST AI RMF organises AI risk work into four functions: Govern, Map, Measure and Manage.
- ISO/IEC 42001, published in December 2023, is a certifiable standard for an AI management system.
- Agents that send, pay, post or delete need action-level controls, not only model-level reviews.
- Risk tiers let low-risk use cases move fast while high-risk ones get deeper review.
What an AI governance framework is, and how NIST and ISO 42001 fit
An AI governance framework answers four questions: what AI do we use, what could go wrong, who decides, and how do we know it is working. In practice that means an inventory of AI systems, a risk assessment for each, a named owner, and monitoring after launch. The NIST AI Risk Management Framework, released in January 2023, structures this as Govern, Map, Measure and Manage, and NIST added a Generative AI Profile in July 2024. It is voluntary and free to use.
ISO/IEC 42001 takes a management-system approach, similar to ISO 27001 for security, with policies, objectives, risk treatment and continual improvement, and organisations can be certified against it. Many enterprises use NIST as the practical risk method and ISO 42001 as the management wrapper, especially if customers ask for certification. If you operate in the EU, map the framework to the EU AI Act's risk categories too. None of this is legal advice, and regulated firms should involve counsel.
How to govern AI agents that take actions
Traditional model governance focuses on bias, accuracy and explainability. An AI governance framework for agentic AI also has to govern actions. An agent that reads invoices and posts them to the ledger, or replies to customers, carries operational risk that a chatbot does not. The key controls are scoped permissions, so the agent can only reach the systems and records it needs, and an action policy that lists which steps run automatically and which wait for a person.
A practical rule: anything irreversible or external, such as sending an email, making a payment, posting a journal entry or deleting data, needs human approval until the agent has a long, measured track record on that exact task. Every action should be logged with the input, the decision and who approved it. Add evaluation before each release, spend limits, and a kill switch the business owner can use without calling engineering. These controls fit neatly under the NIST Manage function.
Do you need an AI center of excellence?
An AI center of excellence helps when several teams are building or buying AI at once and nobody can see the whole picture. A small group of three to eight people, usually from IT, data, security, legal and a couple of business units, can own the inventory, run risk reviews, share reusable patterns and keep a list of approved vendors. The risk is that it becomes a queue that slows every request.
The better model is a hub that sets standards and a spoke in each business unit that delivers. Use risk tiers so a low-risk internal drafting assistant is approved in days, while a customer-facing agent or anything touching hiring, credit or health data gets a full review. AI change management matters as much as the controls: publish decisions, explain the rules in plain language, and give teams an AI governance framework template they can fill in themselves.
How it works
- 1
Inventory and risk tiers
We list every AI system and agent in use or planned, and sort them into risk tiers based on data, users and the actions they can take.
- 2
Map to NIST AI RMF and ISO 42001
We map your policies and controls to the frameworks you need, and mark the gaps that matter for your sector and customers.
- 3
Define agent action policies
For each agent we set permission scopes, logging, evaluation thresholds and which actions run automatically versus wait for approval.
- 4
Set up the review process
We create a lightweight intake form, a review cadence and a small governance group, so low-risk requests clear in days.
- 5
Pilot the framework on live use cases
We run two or three real agents through the process, with a person approving every send, payment, posting or deletion, then adjust the framework from what we learn.
Before and after
Typical ranges from comparable deployments. Your baseline is measured before anything is built.
Tools it works with
- OneTrust
- Credo AI
- ServiceNow
- Microsoft Purview
- Okta
- Langfuse
- AWS Bedrock Guardrails
- Azure AI Foundry
- Jira