Document parsing
Document parsing turns a file format into text a model can actually reason over, while preserving the structure that gives the content its meaning: headings, tables, lists, emphasis. That is different from plain OCR, which recovers characters but not the relationships between them, so a table becomes a row of numbers with no columns.
A contract or a policy document is not just words, it is words in a specific structure, clause numbers, defined terms, tables of figures. Flattening that structure into a wall of text loses information a regulated reader, human or model, actually depends on.
A scanned invoice is parsed into Markdown with its line-item table intact, so an amount can still be matched to the specific item it belongs to, instead of a stream of numbers with no columns to anchor them.
In Dynamiq, Document Parse in the AI Gateway runs a vision-capable model over a PDF or image and returns Markdown that keeps headings, tables and lists, while skipping running headers, footers and blank pages, callable from a playground or a single API endpoint.

See an agent on your own workflow.
Bring a process and its documents. Our engineers will show you how Dynamiq runs it, in your environment or ours.