How do you pull data out of hundreds of PDFs?
You decide the exact fields you need, collect every PDF under a stable name, extract the text (with OCR only for scanned pages), and pull each field into a table. Then you check every value against the document it came from before anything is loaded. The checking is what makes the table usable; extraction alone is the easy half.
Most projects go wrong on the inputs rather than the reading: a file saved under the wrong name, a document from the wrong year, a scan too poor to read, or a blank field that someone filled with a guess.
How it works
Write the field list first
Name each field, its format and where it appears in the document. A field nobody can point to in a real sample does not belong in the first version.
Collect and name the files
Save every PDF under a stable key, such as an account number or ID plus the year or period. The name is how a later run knows what it already has and what is new.
Extract the text and the fields
Digital PDFs give up their text directly; scanned pages need OCR. A form or table parser, or a language model given the field list, then maps the text to your fields.
Check each value against its source
Confirm the document is the one it claims to be, keep the page reference, and leave a missing value empty and marked. Anything uncertain goes to a person for review.
Load with a dry run
Show what would change before writing: new rows, filled gaps and any value that would overwrite an existing one. Then load, and keep the before-and-after files.
What it costs to run
OCR and document parsers are billed per page. The list below is list prices from Google and Amazon read today; a language model, storage and the tool that runs the steps are billed separately by their providers.
| Item | List price | Source |
|---|---|---|
| Google Document AI, Enterprise Document OCR | First 1,000 pages listed as free; then US$1.50 per 1,000 pages, up to 5 million pages | Google Cloud, Document AI pricingRead on Sep 28, 2026 |
| Google Document AI, Form Parser (fields and tables) | US$30 per 1,000 pages, up to 1 million pages (Google's own example: 100 pages cost US$3) | Google Cloud, Document AI pricingRead on Sep 28, 2026 |
| Amazon Textract, Detect Document Text (OCR) | US$0.0015 per page for the first million pages (US West, Oregon) | Amazon Web Services, Amazon Textract pricingRead on Sep 28, 2026 |
| Amazon Textract, Analyze Document with Tables | US$0.015 per page for the first million pages (US West, Oregon) | Amazon Web Services, Amazon Textract pricingRead on Sep 28, 2026 |
| Amazon Textract free tier (new AWS customers, first three months) | 1,000 pages a month of Detect Document Text; 100 pages a month with Forms, Tables and Layout | Amazon Web Services, Amazon Textract pricingRead on Sep 28, 2026 |
Third-party list prices, read on the date shown. They are not Betterlane prices.
What we learned building it
From the college database in Betterlane's own product, built from public Common Data Set PDFs that colleges publish. Not a client system.
- Save every file under an ID and a year
- Each PDF was stored under the school's ID plus the year. The database was then joined to the federal IPEDS and College Scorecard data on the federal school ID.
- Check the name inside the file before loading
- Each document must name the school it was filed under. The check caught files saved for the wrong school (Bard filed as Barnard, WPI as Whitman). Files we could not confirm were left out.
- Blank stays blank
- A value that is not in the PDF is stored empty and marked as not found. Nothing is guessed to make the table look complete.
- Dry run first, and fill gaps instead of overwriting
- Every write script runs as a dry run by default and saves before-and-after files. New data fills empty fields; it does not replace values already there.
- Say how far the coverage goes
- The database holds 1,959 colleges, and admission factors are live for 314 of them. One field that did not hold up, the GPA figures, was withdrawn from the product rather than shown.
When it is not worth it
- If you receive a few documents a month, typing the fields in by hand is cheaper and easier to trust.
- If every document has a different layout and the fields you need are not written anywhere consistent, start by agreeing what the documents must contain.
- If the numbers feed a payment or a legal decision, extraction can prepare them, but a person still has to approve each one.
Questions
Do I need OCR for every PDF?
No. A PDF created by software already contains its text. OCR is needed for scanned pages and photos, and it is where most reading errors come from.
Can a language model do the extraction?
It can map messy text to your fields, but it should not be trusted blindly. Keep the page reference, check each value against the document and send uncertain ones to a person.
What happens when a value is missing?
It stays empty and is marked as not found. Filling it with a likely value makes the whole table harder to trust.
Where does the result go?
Usually a spreadsheet or the system that uses the data, with a column pointing back to the source file and page.
Have a pile of documents you need as a checked table?
See the service: Document processingLet’s talkAll guides
- What should an inbound voice agent answer, and what must it hand to a person?
- What does a website chatbot do for a small business, and when is it worth it?
- How do you connect WhatsApp to an assistant with Meta's Cloud API?
- How do you automate a support inbox without sending a wrong reply?
- n8n, Zapier or Make: which one should a small business choose?
- How do you build a report that updates itself when the data changes?
- How does an internal knowledge assistant answer from your documents without making things up?
- How do you automate a content workflow and keep an editor in control?
- How do you build a sourced lead list without buying one?
- What does a small business website need to be found by Google and ChatGPT?
- When does a business need a customer portal instead of email?
- When should a business move a process off spreadsheets into an internal tool?
- Why does my n8n workflow run twice for the same thing?
- My n8n workflow worked, and now it fails with 401 or 403. Why?
- How do I find out when an n8n workflow fails silently?