Blog Document AI No. 02
How to extract data from purchase order PDFs
If your customers send purchase orders as PDFs, someone or something has to turn each one into data your order system can use. There are four ways to do it. Each has a place, and each fails in its own way.
1. Typing them in
The default. A person opens the PDF and keys the order.
It works on any layout, needs no software, and a good order desk catches odd things a machine would not, such as a price that does not match the quote.
It is slow, and it does not scale with volume. Errors are rare per field but there are a lot of fields: an order with 20 lines has well over 100 values to key. The mistakes that get through tend to be the costly kind, such as a transposed digit in a quantity or a delivery sent to the invoice address.
2. Templates
A template tells the software where each field sits on one customer's layout: the order number is in this box, the lines are in this table between these columns.
For a handful of large customers who send the same layout every time, templates are accurate and cheap to run.
They break when the layout changes, and layouts change whenever a customer updates their ERP, edits their form or exports from a different system. Every new customer needs a new template. With a long tail of customers, maintaining templates becomes a job of its own.
3. OCR plus rules
OCR (optical character recognition) turns the page image into text. Rules or regular expressions then look for fields: a number after "PO No.", a date near "Order date".
OCR is mature and reads clean print very accurately. The trouble is structure. Tables come out as lines of words with the columns lost, and two address boxes side by side are often read across, so their lines interleave. The rules that find fields are, in effect, templates written in a different way, and they fail in the same places.
If your PDFs come from software they usually have a text layer already, and you can skip OCR and read the text directly. The structure problem stays.
4. Vision language models
A vision language model takes the page image and an instruction, and answers in text. Asked to fill a fixed JSON structure, it reads the page much as a person does: it sees that a table has columns, that one box is labelled "Deliver to" and another "Invoice to", and that "Q-55120" next to "Your quote" is a quote reference.
The advantage is that there is no template. A layout the model has never seen is handled the same way as a familiar one.
The risks are different from OCR's:
- It can be wrong with confidence. A model may put a value in the wrong field, or occasionally produce a value that is not on the page. Instructions help ("use null when a value is not present") but do not remove the risk.
- Formats drift. Left to itself, a model writes dates and amounts however they appeared. You need a fixed schema and a step afterwards that normalises dates, numbers and codes.
- Cost and privacy. Hosted model APIs charge per token, and send your customers' orders to a third party.
The fix for the first two is to treat the model's output as a reading to be checked, not as the answer. Purchase orders are good for this because they check themselves: quantity times unit price should equal each line amount, the lines should sum to the subtotal, and the subtotal plus tax and carriage should equal the total. A misread digit almost always breaks one of those sums.
Which to use
- A few customers, steady layouts, low volume: typing, or templates.
- Many customers, changing layouts: a vision model with a fixed schema, normalisation and arithmetic checks, and a person reviewing only the orders that fail a check.
That second approach is what VioPO does, with the model running on hardware we operate rather than a third party's API. It is not open yet, but you can see its output on three sample orders.