Maciej Nuzia · Documents

I automate document reading and pass the data into your system

Invoices, delivery notes and service reports use different layouts, but they contain fields needed later in the process. I build automation that reads a number, date, amount or line item and sends it to the right system. The accuracy of that reading depends on the real files, including scans and phone photos.

Photos and scans are a different job from files with the text already inside, and it shows in the estimate.

Day to day

Today, somebody reads the document and retypes the data

Open the file, find the number, the date and the amount, switch to the other window and key the same thing in again. With a few documents a day that is nothing. With a full inbox it eats a slice of somebody's day, every day, and has no name in any job description.

  • the same number keyed in twice, because the second system cannot see the first
  • a phone scan, crooked and half in shadow
  • a document from a new supplier looks different and has to be learned
  • line items you cannot see on the first page, because the table runs on

Who it's for

Start with a document type that keeps returning

You can spot a candidate by whether the document comes back in the same shape. An invoice from a regular supplier does, even though every issuer draws it differently. A letter that arrives once and never again has too short a life to build anything around, and that is a job I turn down. These are the places where the same layout keeps landing:

  • Accounts payable keying in other people's invoices
  • Freight forwarding, where every carrier has its own note
  • Law firms and contract teams
  • Procurement buried in quotes
  • Field service: repair reports
  • Property managers: tenancy agreements and meter readings

Kinds of document

What can be read and where the data can go

A supplier invoice

Number, date, seller, amounts and line items. The automation lines them up against the order the invoice refers to, and where the two drift apart it sets the document aside for review.

A supplier with a layout of their own

The fields are named differently, the table sits somewhere else, part of the data hides in the footer in small print. So I describe the fields by what they mean, not by where they sit on the page.

Working out what the document is

The post arrives mixed together: an invoice, a proof of delivery, a complaint. The automation sorts them by kind and drops each file into the right queue along with the fields it read.

A photo taken on a phone

A driver photographs a signed delivery note in a car park and the shot comes in at an angle. Files like that get straightened before reading, and which of them make it through is something I check on your photos.

Two documents side by side

The order next to the invoice, the delivery note next to the goods receipt. The automation surfaces the lines that disagree and leaves the call to a person.

The entry on the far side

The fields that were read go where they belong: into the accounting package, into the ERP, into the spreadsheet somebody opens every morning anyway. This step is often the longest one, because the receiving system takes data on its own terms.

If what you want is to ask questions of those documents, that is a different road: RAG for business. And when the reading is one step inside a larger back-office task, I start from the page on AI process automation.

On your files

A real sample shows how difficult the reading will be

A pile with the worst copies taken out of it gives a picture that falls apart in the first week of work.

01

One kind of document to begin with

I take the one there is most of, and write down with you the fields that are genuinely needed further on. Some boxes on today's form drop out, because somebody fills them in out of habit and nobody ever looks at them again.

02

Field by field

I set out which fields to take and how to recognize each one: by the heading, by what sits next to it, by the shape of the number. A fixed layout needs no more than an ordinary rule for that, while paperwork with a shifting layout needs a model.

03

A confidence threshold with a person behind it

Every field comes back with a note on how sure the reading was. We agree which fields are the expensive ones to get wrong. When one of those falls below the threshold, the document waits for human eyes.

04

Going live, and the count of held documents

The automation starts on that one kind of document. After that I look at what it held back and why, and the list of fixes comes straight out of that. The count of held documents is worth a glance every so often: when it jumps, a supplier has usually changed their layout.

Technologies

A text file and a photo need different reading methods

The first one carries text inside and it only has to be lifted out. The second has to be recognized from the image first, and everything then rests on how much of it is legible.

  • OpenAI API
  • Claude / Anthropic
  • OCR
  • Vision models
  • Python
  • PostgreSQL
  • Accounting integrations
  • Webhooks
  • AWS

Weak spots

A misread value can be passed to the next system

A file the amount cannot be read from

A photo shot into the light, a fax, a copy of a copy. Some of that is fixed by preparing the image before reading, and some stays unreadable for anyone who looks at it. How big that second group is in your inbox is something I check on a sample before promising anything.

A layout changed mid-year

A supplier shifts their template and tells nobody. The reading step then lands on the neighboring row and passes on a number that draws no attention. So I watch the things that have to agree: the line items against the total at the bottom, the number against what is already in the system.

The files you keep under lock

Personnel files, medical records, contracts under a confidentiality clause. Before a file travels to a model provider, someone with authority over that data has to say what may be sent out at all. For some documents the answer is nothing at all.

FAQ

Documents, extraction and where the data goes

How is this different from plain OCR?

OCR turns an image into text and stops there. What is left is a wall of characters, and somebody still has to fish the number and the amount out of it. A model adds that second step: it works out that the digits next to the word "total" are the amount due, and that out of the several dates on an invoice only one belongs in the due date field.

Which documents can be read?

Invoices, delivery notes, purchase orders, contracts, service reports, forms. What counts is repeatable content: the same fields come back in every copy, even when the layout differs from one supplier to the next. One-off letters and free text with no fixed fields drop out.

What happens to a field the automation is unsure of?

The field comes back empty or flagged as uncertain, and the document waits for a person. We agree which fields have to be right and which ones can be corrected later. Without that list the automation either stops everything or lets through things it should have held back.

Where do the scans go while they are being read?

That depends on where the reading runs. With a model provider the file leaves your network, so what the provider keeps on its side and for how long starts to matter. Reading placed inside your own infrastructure keeps the files in the building. That brings its own questions: where it stands and who looks after it. That choice is yours, and it is made before the project starts.

What drives the effort up the most?

The number of different layouts to support, and the share of scans among them, because an image has to be recognized first. Next comes the system on the receiving end: one takes data through an API, another wants a batch file prepared exactly the way it likes.

A sample

Show me one type of document

Pick the one you get most of and write down what somebody retypes out of it today, and where it goes. Both the conversation and the report from the estimation tool rest on that.