# LLM evaluation

URL: https://www.getdynamiq.ai/glossary/llm-evaluation

> Scoring an LLM or agent's outputs against test data with defined metrics, so changes in quality are measured rather than judged by feel.

Scoring an LLM or agent's outputs against test data with defined metrics, so changes in quality are measured rather than judged by feel.

LLM evaluation is built from three pieces: a metric that defines how to score, a dataset that defines what to score against, and a run that puts the two together and produces results. Comparing a candidate version against what is already in production, on the same dataset, with the same metric, turns a subjective impression into a number you can actually compare.

Regulated teams cannot ship a prompt or model change on a hunch. They need a documented, reviewable comparison of before and after, on a dataset that does not change out from under them.

Comparing a new prompt against the current one on the same released dataset of real questions, scored by the same metric, before deciding whether to promote the new version.

**In Dynamiq**, Evals combine versioned metrics, an LLM-as-a-judge rubric, a predefined RAG evaluator such as Faithfulness or Context Recall, or your own Python function, with versioned datasets you can build directly from production traces, and evaluation runs that score a workflow's fresh outputs or a dataset as recorded. Every run pins the exact metric and workflow version it used, so it can be rerun and trusted later.

## See it in Dynamiq

-   [Evals](https://www.getdynamiq.ai/product/evaluations)

## Related terms

-   [LLM as a judge](https://www.getdynamiq.ai/glossary/llm-as-a-judge)
-   [Online evaluation](https://www.getdynamiq.ai/glossary/online-evaluation)
-   [AI observability](https://www.getdynamiq.ai/glossary/ai-observability)

## See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.

[Start free](https://app.getdynamiq.ai/signup) · [Talk to the team](https://www.getdynamiq.ai/book-a-demo)
