How to evaluate an AI agent platform.

19 questions to ask every vendor, why each one matters, and a scorecard to compare the answers side by side.

Deployment and data residency

  1. Runs in your environment

    Can the whole platform run in our cloud account or data center, not only the model?

    Why it matters: Regulators and security teams ask where prompts, documents and traces live. A managed-only platform makes that answer someone else's.

  2. Isolated operation

    Can it run with no outside access, and who sets that up?

    Why it matters: Some workloads can never reach the internet. Find out early whether that is a product capability or a custom project.

  3. Data location

    Where are prompts, retrieved documents, traces and credentials stored, and for how long?

    Why it matters: Retention and residency rules apply to agent traces as much as to source data.

Security and compliance

  1. Certifications you can verify

    Which attestations do you hold, and will you share the reports under NDA?

    Why it matters: A logo on a web page is not evidence. Ask for the report and the scope it covers.

  2. Identity and access

    Do you support SSO with our identity provider, projects and role-based access?

    Why it matters: Agents act on real systems. Access to build and deploy them should follow your identity controls.

  3. Secrets handling

    How are credentials for connected systems stored, rotated and revoked?

    Why it matters: An agent platform holds keys to your CRM, core systems and data stores.

Governance in production

  1. Human approval

    Can any step wait for a person to approve, and can the reviewer edit the values first?

    Why it matters: Most regulated workflows need a person to sign off on the steps that move money, data or decisions.

  2. Input screening

    Are PII and prompt injection detected before data reaches a model?

    Why it matters: Screening at the platform level beats relying on every agent builder to remember.

  3. Traceability

    Is every run recorded step by step, with inputs, outputs, cost and latency, and can we replay it?

    Why it matters: When an auditor or customer asks why an agent did something, you need the full record.

Quality and evaluation

  1. Evals before launch

    Can we score agents against our own datasets and metrics before they go live?

    Why it matters: Success criteria agreed up front are the difference between a demo and a production agent.

  2. Evals after launch

    Are live deployments evaluated continuously, and can production traces become test cases?

    Why it matters: Models, data and prompts drift. Quality has to be measured every day, not once.

Models and integrations

  1. Model choice

    Which model providers are supported, can we run open models ourselves, and can we switch without rewriting agents?

    Why it matters: The best model for a task changes every few months. Lock-in at the model layer is expensive.

  2. Connections to your systems

    How does an agent reach our systems: prebuilt connectors, APIs, databases, MCP servers?

    Why it matters: An agent that cannot act in your systems is a chatbot.

  3. Per-user access

    Can a deployed agent act with each end user's own permissions?

    Why it matters: Retrieval and actions should respect the permissions people already have.

Build and ownership

  1. Two ways to build

    Can business teams build visually and engineers build in code, on the same runtime?

    Why it matters: Separate tools for each audience mean two platforms to govern.

  2. Openness

    Is the core open source, and under which license?

    Why it matters: An open core lowers switching costs and lets your engineers read what runs in production.

  3. Delivery help

    Will your engineers work with ours to reach production, and who owns what they build?

    Why it matters: The first agent is the hardest. Make sure the result, including the code, is yours.

Commercials

  1. How you buy

    Can we buy through our cloud marketplace, and how is usage priced?

    Why it matters: Marketplace procurement can shorten purchasing and count toward existing cloud commitments.

  2. Exit

    If we leave, what do we take with us: agents, prompts, data, traces?

    Why it matters: A clean exit path is the best test of how much you depend on a vendor.

See how Dynamiq answers every question.

Book a working session and bring your checklist. Our engineers will answer on your use case.