# Online evaluation

URL: https://www.getdynamiq.ai/glossary/online-evaluation

> Automatically scoring a sampled share of a deployed system's live production traffic on an ongoing basis, not a one-time test.

Automatically scoring a sampled share of a deployed system's live production traffic on an ongoing basis, not a one-time test.

An online evaluation attaches a metric to a system that is already live and scores a sampled share of its real traffic continuously, rather than a fixed dataset run once before a release. Nothing is re-executed for this: the metric reads what the system already produced.

A model, a prompt or an upstream data change can degrade quality after launch in ways a pre-release test never surfaces. A regulated operation needs continuing evidence that a live system still performs the way it did when it was approved, not only a report from the day it launched.

Sampling a share of a deployed agent's real conversations and scoring each one for policy adherence as they happen, catching a drift in quality quickly instead of waiting for a periodic manual audit to find it.

**In Dynamiq**, attach a metric to a deployed App's evaluations with a sample rate and a path that pulls the metric's input from each recorded trace, and every scored trace links back to its full execution trace, so a low score is one click from the run that produced it.

## See it in Dynamiq

-   [Evals](https://www.getdynamiq.ai/product/evaluations)

## Related terms

-   [LLM evaluation](https://www.getdynamiq.ai/glossary/llm-evaluation)
-   [LLM as a judge](https://www.getdynamiq.ai/glossary/llm-as-a-judge)
-   [AI observability](https://www.getdynamiq.ai/glossary/ai-observability)

## See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.

[Start free](https://app.getdynamiq.ai/signup) · [Talk to the team](https://www.getdynamiq.ai/book-a-demo)
