"How do I get the data in these PDFs into Excel?" is one of the questions I hear most often in companies. Invoices, orders, bank statements, price lists: the data already exists, but it is locked in a format designed to be read by a person, not by a spreadsheet. So someone retypes it, line by line.
The first distinction is between digital and scanned PDFs. A digital PDF, exported from business software or an invoicing tool, already contains the text: you just extract it. A scan or a photo is an image, and needs OCR to become text. The difference matters, because with crooked pages, stamps and handwriting, OCR quality decides much of the final accuracy.
Extracting the text is not enough, though. The text of an invoice is not a table: columns break, totals sit at the bottom, and the same information has different names from one supplier to the next. This is where AI comes in: a language model reads the text the way a person would and returns structured fields, for example document number, date, supplier, line items with quantities and amounts, subtotal, VAT and total.
The weak point is that a model can be confidently wrong, or write a plausible value that is not in the document. That is why, in the systems I build, every extracted value carries the piece of text it came from, and the software checks that this text really exists in the document. If it does not, the field is flagged instead of landing in Excel.
Then come the rules, which are cheap and never hallucinate: an Italian VAT number has a check digit, line items must add up to the subtotal, subtotal plus VAT must equal the total, dates must be plausible. An invoice with a wrong total is stopped here, before it reaches your books.
Only then does the data go to Excel. The simplest format is a CSV that Excel opens directly; for continuous use it is better to write straight into your business software or a database, so nobody even has to open the file. What you must not do is mix verified and doubtful data: uncertain fields should be shown to a person, with the document beside them, and confirmed with one click.
My website has a free demo that does exactly this: you paste the text of an invoice or upload a digital PDF, the model extracts the fields, the browser checks every value against the text and applies the rules, and you download the result as a CSV for Excel. One of the samples is an invoice with a deliberately wrong total, so you can see how it gets flagged.
A real project adds OCR for scans, matching against the records the company already has, and a log of corrections, which becomes the most valuable data for improving the system. The principle stays the same: extract, verify, and only then export. If you want to apply it to your own documents, the details are on the document data extraction page.