Prompt injection
A prompt injection tries to make a model ignore its actual instructions and follow different ones smuggled in through the input instead. A direct injection is typed straight into a chat, telling the model to ignore its previous instructions. An indirect one is hidden somewhere the agent will read it anyway, inside a scraped web page or an uploaded document, where the agent never expected instructions to appear.
An agent with tools that can send a message, query a database, or take an action is a much bigger target than a chatbot, because a successful injection does not just produce an unwanted sentence, it can trigger a real action in a real system.
A scraped web page contains hidden text instructing the agent to reveal its system prompt or call a tool it was never asked to use, buried in content the agent was only supposed to summarize.
In Dynamiq, a prompt injection detector node classifies a message before it reaches an agent, backed by a dedicated classifier, and flagged input can route to a refusal or a human review branch instead of the agent, with every verdict recorded in the run's trace.

See an agent on your own workflow.
Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.