# LLM inference

URL: https://www.getdynamiq.ai/glossary/llm-inference

> Running a trained model to produce an output for a given input, as opposed to training the model in the first place.

Running a trained model to produce an output for a given input, as opposed to training the model in the first place.

Inference is what happens every time a model actually answers something: tokenizing the input, generating output tokens one at a time, and returning the result. It is distinct from training, which happens once, ahead of time, to produce the model's weights in the first place. Inference is also where ongoing cost and latency come from, since it scales with every call, not with a one-time training run.

Where inference happens, on a shared public endpoint or on infrastructure a regulated enterprise controls, is as much a governance question as it is a performance one, alongside the more familiar concerns of cost and response time.

An agent's every reasoning step and tool decision is a separate inference call, so one long agent run can mean many calls to a model before it produces a final answer.

**In Dynamiq**, call the most-used models through one AI Gateway endpoint, use any of 29 model providers from the SDK, or run inference on your own infrastructure by self-hosting an open model on the vLLM engine behind an OpenAI-compatible endpoint, with autoscaling replicas and no call ever leaving the cluster it runs in.

## See it in Dynamiq

-   [Models and AI Gateway](https://www.getdynamiq.ai/product/models)

## Related terms

-   [AI gateway](https://www.getdynamiq.ai/glossary/ai-gateway)
-   [Model routing](https://www.getdynamiq.ai/glossary/model-routing)
-   [Open-weight model](https://www.getdynamiq.ai/glossary/open-weight-model)
-   [Context window](https://www.getdynamiq.ai/glossary/context-window)

## See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.

[Start free](https://app.getdynamiq.ai/signup) · [Talk to the team](https://www.getdynamiq.ai/book-a-demo)
