Governance for AI agents in production

Screen what reaches the model, approve what the agent does, and keep a replayable record of every run, with the same controls on every agent.
Run 4821Completed
loan-file-review1.9 s
Detect PIISSN blocked before the model120 ms
Extract fields740 ms
Credit policy6 passages180 ms
Underwriter approvalapproved
Update loan system90 ms
Faithfulness 0.96Policy check passedReplay
Illustration of a run trace: each step of a loan-file review with its timing, a PII guardrail event, a knowledge lookup and the underwriter's approval.

In short

AI agent governance decides what an agent may see, what it may do, who signs off, and how you prove what happened. On Dynamiq those controls are part of the platform: detectors screen input for PII and prompt injection, validators check output, approval gates pause consequential steps for a person, and every run is traced and replayable. Real-time evals score live traffic and an AI agent surfaces issues from production traces, so oversight continues after launch.

The problem

Most organizations can build an agent. Fewer can answer the questions that come before it goes live: what data can it reach, what can it change, who approved the risky step, and how will anyone know if its answers get worse next month. A policy document does not stop a model from ignoring an instruction, and controls each team builds on its own end up different for every agent. Risk, compliance and security teams need the same enforceable controls on every agent, and evidence they can pull without asking engineers to reconstruct a run.

What it moves

Time to approve a new agent
How long risk and security review takes when the controls and evidence already exist.
Flag rate
How often detectors flag PII, prompt injection or policy violations, by agent and channel.
Approval turnaround
How long consequential steps wait for a person, and how often reviewers edit or reject.
Quality drift
Changes in real-time eval scores and AI-found issues after a model, prompt or data change.

How the agent works

Every step is traced. The steps marked for review wait for a person.
  1. Step 1

    Access decided up front

    SSO on Enterprise, organization roles, private projects and scoped access keys decide who can build, deploy and run agents, and credentials can resolve per end user where access must follow each person's own permissions.

  2. Step 2

    Input screened

    Detectors flag PII, prompt injection and content policy violations before the model sees a message, and a Choice node routes flagged input to a refusal or a person.

  3. Step 3

    Output checked

    Validators confirm the output is valid JSON, an allowed value or a matching pattern before anything downstream uses it, and fail the step or retry when it is not.

  4. Step 4

    Approval on consequential steps

    A payment, a database write, an email or a filing pauses for a person to approve the exact payload, edit the fields you allow, or reject with feedback the agent acts on.

    Human review
  5. Step 5

    Evaluated before release

    A new version is scored against datasets built from real traces and compared with the version in production, and a reviewer decides whether to deploy it or roll back.

    Human review
  6. Step 6

    Monitored after launch

    Real-time evals score a sample of live traffic, an AI agent reviews production traces for issues and regressions, and every run stays replayable for audit.

Controls

Run governance where the agents run: on Dynamiq Cloud, or self-hosted in your own AWS, Azure, GCP, IBM Cloud, OpenShift or any Kubernetes 1.32+ cluster. Dynamiq holds SOC 2, and signs a BAA for HIPAA workloads and a DPA under GDPR. Our engineers define the success criteria, evals and approval points with your risk and compliance teams before the first agent ships.

Enforced, not prompted
Detectors, validators and approval gates are workflow nodes the platform runs every time, not instructions a model can talk its way around.
Detect, never rewrite
Detectors classify and flag; they do not redact or rewrite text, so what happens to a flagged message is a routing rule you own.
Least-privilege access
Private projects, scoped access keys that can expire, and per-user credentials keep an agent's reach no wider than the person or service it acts for.
Change control
Every save is a new version. Deploy a specific version, compare it with the last one and roll back in one step.
Evidence on demand
Every run is traced and replayable, including detector verdicts and approval decisions, and traces export in bulk as JSON.
Continuous evaluation
Real-time evals and AI-found issues keep scoring an agent after it ships, so drift shows up in a chart before it shows up in a complaint.

Systems it connects to

  • Identity providers over OIDC (Okta, Microsoft Entra ID)
  • Slack and Microsoft Teams for approvals
  • Jira and ServiceNow for review queues
  • Models from 29 model providers, the most-used through the AI Gateway
  • Databases and warehouses, under per-user credentials
  • Trace exports as JSON for your archive or SIEM
  • Evaluation datasets built from production traces

Questions and answers

What is AI agent governance?

The controls that decide what an agent may see and do, who approves its consequential actions, and how you prove what happened afterward: access control, input and output checks, approvals, audit trails and ongoing evaluation.

How is governing an agent different from governing a model?

A model produces text; an agent calls tools that change records, send messages and move money. Governance has to cover the actions, so approval gates and per-user credentials matter as much as content checks.

Does Dynamiq redact PII?

No. Detectors flag PII and prompt injection so you can route a message to a refusal or a person; they do not redact or rewrite it.

Can we show an auditor what an agent did?

Yes. Every run is traced and replayable node by node, including each detector verdict and approval decision, and traces export as JSON in bulk.

Does governance stop at launch?

No. Real-time evals score a sample of live traffic, and an AI agent reviews production traces and surfaces issues and regressions for a person to check.

Does this make us compliant with the EU AI Act?

It supports record-keeping and human oversight with traces and approval gates, but Dynamiq does not certify compliance. Whether a system is high-risk, and what that requires, depends on your own assessment.

See this agent on your data.

Talk to our team. We will walk through how it works in your environment, with your systems.