# How to evaluate an AI agent platform.

URL: https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit

> A vendor-neutral checklist for choosing an enterprise AI agent platform: deployment, security, governance, evals, models, integrations and exit. Free scorecard.

19 questions to ask every vendor, why each one matters, and a scorecard to compare the answers side by side.

[Download the Excel scorecard](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit/scorecard.xlsx) · [Ask us these questions](https://www.getdynamiq.ai/book-a-demo)

[Deployment and data residency](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#deployment) · [Security and compliance](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#security) · [Governance in production](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#governance) · [Quality and evaluation](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#quality) · [Models and integrations](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#models) · [Build and ownership](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#build) · [Commercials](https://www.getdynamiq.ai/resources/agent-platform-evaluation-kit#commercials)

## Deployment and data residency

1.  ### Runs in your environment
    
    Can the whole platform run in our cloud account or data center, not only the model?
    
    Why it matters: Regulators and security teams ask where prompts, documents and traces live. A managed-only platform makes that answer someone else's.
    
2.  ### Isolated operation
    
    Can it run with no outside access, and who sets that up?
    
    Why it matters: Some workloads can never reach the internet. Find out early whether that is a product capability or a custom project.
    
3.  ### Data location
    
    Where are prompts, retrieved documents, traces and credentials stored, and for how long?
    
    Why it matters: Retention and residency rules apply to agent traces as much as to source data.
    

## Security and compliance

1.  ### Certifications you can verify
    
    Which attestations do you hold, and will you share the reports under NDA?
    
    Why it matters: A logo on a web page is not evidence. Ask for the report and the scope it covers.
    
2.  ### Identity and access
    
    Do you support SSO with our identity provider, projects and role-based access?
    
    Why it matters: Agents act on real systems. Access to build and deploy them should follow your identity controls.
    
3.  ### Secrets handling
    
    How are credentials for connected systems stored, rotated and revoked?
    
    Why it matters: An agent platform holds keys to your CRM, core systems and data stores.
    

## Governance in production

1.  ### Human approval
    
    Can any step wait for a person to approve, and can the reviewer edit the values first?
    
    Why it matters: Most regulated workflows need a person to sign off on the steps that move money, data or decisions.
    
2.  ### Input screening
    
    Are PII and prompt injection detected before data reaches a model?
    
    Why it matters: Screening at the platform level beats relying on every agent builder to remember.
    
3.  ### Traceability
    
    Is every run recorded step by step, with inputs, outputs, cost and latency, and can we replay it?
    
    Why it matters: When an auditor or customer asks why an agent did something, you need the full record.
    

## Quality and evaluation

1.  ### Evals before launch
    
    Can we score agents against our own datasets and metrics before they go live?
    
    Why it matters: Success criteria agreed up front are the difference between a demo and a production agent.
    
2.  ### Evals after launch
    
    Are live deployments evaluated continuously, and can production traces become test cases?
    
    Why it matters: Models, data and prompts drift. Quality has to be measured every day, not once.
    

## Models and integrations

1.  ### Model choice
    
    Which model providers are supported, can we run open models ourselves, and can we switch without rewriting agents?
    
    Why it matters: The best model for a task changes every few months. Lock-in at the model layer is expensive.
    
2.  ### Connections to your systems
    
    How does an agent reach our systems: prebuilt connectors, APIs, databases, MCP servers?
    
    Why it matters: An agent that cannot act in your systems is a chatbot.
    
3.  ### Per-user access
    
    Can a deployed agent act with each end user's own permissions?
    
    Why it matters: Retrieval and actions should respect the permissions people already have.
    

## Build and ownership

1.  ### Two ways to build
    
    Can business teams build visually and engineers build in code, on the same runtime?
    
    Why it matters: Separate tools for each audience mean two platforms to govern.
    
2.  ### Openness
    
    Is the core open source, and under which license?
    
    Why it matters: An open core lowers switching costs and lets your engineers read what runs in production.
    
3.  ### Delivery help
    
    Will your engineers work with ours to reach production, and who owns what they build?
    
    Why it matters: The first agent is the hardest. Make sure the result, including the code, is yours.
    

## Commercials

1.  ### How you buy
    
    Can we buy through our cloud marketplace, and how is usage priced?
    
    Why it matters: Marketplace procurement can shorten purchasing and count toward existing cloud commitments.
    
2.  ### Exit
    
    If we leave, what do we take with us: agents, prompts, data, traces?
    
    Why it matters: A clean exit path is the best test of how much you depend on a vendor.
    

## See how Dynamiq answers every question.

Book a working session and bring your checklist. Our engineers will answer on your use case.

[Start free](https://app.getdynamiq.ai/signup) · [Talk to the team](https://www.getdynamiq.ai/book-a-demo)
