Reading documents
How Merchaint turns a PDF into data, and what to do when it reads it wrong.
Understand it first
How extraction works
Between a document arriving and data appearing, Merchaint does three things in order. It works out what the document is, picks a schema to read it with, then reads the fields.
How Merchaint reads different documents
A PDF made by accounting software already has its text inside it. A scan or a phone photo doesn't — it's a picture that has to be worked out. Merchaint handles both, and that's why results differ.
The screens you use
The document inbox
Inbox is where every document lands, whichever way it arrived, and where you watch it being read. Most people have this page open all day.
Check what was extracted
Open a document and you see the original page next to the fields Merchaint read from it. This is where you confirm it got things right — and the one screen that decides whether you can trust everything after it.
Schema Studio
A schema is the list of fields Merchaint looks for in a kind of document. Schema Studio is where those lists live. It's the one part of reading documents that you actually set up.
When it goes wrong
Fix a document that was read wrong
Sometimes Merchaint reads a document as the wrong type, or misses a value. Reprocessing gives it another go — and lets you tell it which schema to use, so it doesn't make the same mistake twice.
The exceptions queue
Exceptions collects documents Merchaint couldn't read, grouped by what went wrong. It's a short daily list — and it can stay empty even when documents were missed.