Retrieval-augmented generation (RAG)

Retrieving relevant passages from your own documents and giving them to a model as context, so its answer is grounded in that material.

Retrieval-augmented generation, RAG, has two halves. Indexing converts documents into chunks, embeds each chunk, and writes the result into a vector store, done once, in advance. Retrieval embeds an incoming query, searches that store for the closest chunks, and hands them to a model as context for its answer, done on every request.

A general-purpose model does not know a firm's internal policies, contracts or case history, and a regulated answer needs to be grounded in the firm's own current documents, not in whatever the model happened to learn during training, which may be outdated or simply not about this firm at all.

An agent asked about a firm's travel policy on business-class flights retrieves the actual current policy passage and answers from it, instead of producing a plausible-sounding guess drawn from general knowledge of how travel policies typically work.

In Dynamiq, a Knowledge Base bundles conversion, chunking, embedding and vector storage into one managed pipeline with a search endpoint, or the same components can be wired by hand with vector store search and writer nodes against 8 vector stores. Either way, the result attaches to an agent as a retrieval tool.

See an agent on your own workflow.

Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.