Governance for AI agents in production
In short
AI agent governance decides what an agent may see, what it may do, who signs off, and how you prove what happened. On Dynamiq those controls are part of the platform: detectors screen input for PII and prompt injection, validators check output, approval gates pause consequential steps for a person, and every run is traced and replayable. Real-time evals score live traffic and an AI agent surfaces issues from production traces, so oversight continues after launch.
The problem
Most organizations can build an agent. Fewer can answer the questions that come before it goes live: what data can it reach, what can it change, who approved the risky step, and how will anyone know if its answers get worse next month. A policy document does not stop a model from ignoring an instruction, and controls each team builds on its own end up different for every agent. Risk, compliance and security teams need the same enforceable controls on every agent, and evidence they can pull without asking engineers to reconstruct a run.
What it moves
- Time to approve a new agent
- How long risk and security review takes when the controls and evidence already exist.
- Flag rate
- How often detectors flag PII, prompt injection or policy violations, by agent and channel.
- Approval turnaround
- How long consequential steps wait for a person, and how often reviewers edit or reject.
- Quality drift
- Changes in real-time eval scores and AI-found issues after a model, prompt or data change.
How the agent works
Step 1
Access decided up front
SSO on Enterprise, organization roles, private projects and scoped access keys decide who can build, deploy and run agents, and credentials can resolve per end user where access must follow each person's own permissions.
Step 2
Input screened
Detectors flag PII, prompt injection and content policy violations before the model sees a message, and a Choice node routes flagged input to a refusal or a person.
Step 3
Output checked
Validators confirm the output is valid JSON, an allowed value or a matching pattern before anything downstream uses it, and fail the step or retry when it is not.
- Human review
Step 4
Approval on consequential steps
A payment, a database write, an email or a filing pauses for a person to approve the exact payload, edit the fields you allow, or reject with feedback the agent acts on.
- Human review
Step 5
Evaluated before release
A new version is scored against datasets built from real traces and compared with the version in production, and a reviewer decides whether to deploy it or roll back.
Step 6
Monitored after launch
Real-time evals score a sample of live traffic, an AI agent reviews production traces for issues and regressions, and every run stays replayable for audit.
Controls
Run governance where the agents run: on Dynamiq Cloud, or self-hosted in your own AWS, Azure, GCP, IBM Cloud, OpenShift or any Kubernetes 1.32+ cluster. Dynamiq holds SOC 2, and signs a BAA for HIPAA workloads and a DPA under GDPR. Our engineers define the success criteria, evals and approval points with your risk and compliance teams before the first agent ships.
- Enforced, not prompted
- Detectors, validators and approval gates are workflow nodes the platform runs every time, not instructions a model can talk its way around.
- Detect, never rewrite
- Detectors classify and flag; they do not redact or rewrite text, so what happens to a flagged message is a routing rule you own.
- Least-privilege access
- Private projects, scoped access keys that can expire, and per-user credentials keep an agent's reach no wider than the person or service it acts for.
- Change control
- Every save is a new version. Deploy a specific version, compare it with the last one and roll back in one step.
- Evidence on demand
- Every run is traced and replayable, including detector verdicts and approval decisions, and traces export in bulk as JSON.
- Continuous evaluation
- Real-time evals and AI-found issues keep scoring an agent after it ships, so drift shows up in a chart before it shows up in a complaint.
Systems it connects to
- Identity providers over OIDC (Okta, Microsoft Entra ID)
- Slack and Microsoft Teams for approvals
- Jira and ServiceNow for review queues
- Models from 29 model providers, the most-used through the AI Gateway
- Databases and warehouses, under per-user credentials
- Trace exports as JSON for your archive or SIEM
- Evaluation datasets built from production traces
Built with
- Guardrails and approvalsDetect PII and prompt injection before a model sees them, enforce content policies, and pause tool steps for a person to approve.
- EvalsScore an agent on a dataset before you ship, then keep scoring a sample of its live traffic after you deploy.
- ObservabilityEvery run is traced and replayable, node by node, with the cost, latency and tokens each step used.
- Apps and deploymentsDeploy a saved workflow version as an App with its own hostname. Call it over HTTP, embed a chat widget, or trigger it on a schedule or an event.
Industries
- Financial servicesBack-office automation, customer service by phone and chat, and KYC and AML work, on one governed platform your compliance team can audit.
- InsuranceAnswer policyholders by phone and chat, prepare claims and underwriting files, and keep an adjuster or underwriter approving every decision.
- HealthcarePatient intake, referral and document review, and clinical knowledge assistants, with PII and prompt injection screened before a model reads anything.
- Government and public sectorCitizen services, benefits and permit review, and program reporting, with a person in the loop and your agency in control of the infrastructure.
- TelecommunicationsAnswer subscriber calls and chats in their own language, reconcile billing, and give support and network teams answers with a citation attached.
Questions and answers
What is AI agent governance?
The controls that decide what an agent may see and do, who approves its consequential actions, and how you prove what happened afterward: access control, input and output checks, approvals, audit trails and ongoing evaluation.
How is governing an agent different from governing a model?
A model produces text; an agent calls tools that change records, send messages and move money. Governance has to cover the actions, so approval gates and per-user credentials matter as much as content checks.
Does Dynamiq redact PII?
No. Detectors flag PII and prompt injection so you can route a message to a refusal or a person; they do not redact or rewrite it.
Can we show an auditor what an agent did?
Yes. Every run is traced and replayable node by node, including each detector verdict and approval decision, and traces export as JSON in bulk.
Does governance stop at launch?
No. Real-time evals score a sample of live traffic, and an AI agent reviews production traces and surfaces issues and regressions for a person to check.
Does this make us compliant with the EU AI Act?
It supports record-keeping and human oversight with traces and approval gates, but Dynamiq does not certify compliance. Whether a system is high-risk, and what that requires, depends on your own assessment.

See this agent on your data.
Talk to our team. We will walk through how it works in your environment, with your systems.