How do you pull data out of hundreds of PDFs?

You decide the exact fields you need, collect every PDF under a stable name, extract the text (with OCR only for scanned pages), and pull each field into a table. Then you check every value against the document it came from before anything is loaded. The checking is what makes the table usable; extraction alone is the easy half.

Most projects go wrong on the inputs rather than the reading: a file saved under the wrong name, a document from the wrong year, a scan too poor to read, or a blank field that someone filled with a guess.

Updated Sep 28, 2026

How it works

  1. Write the field list first

    Name each field, its format and where it appears in the document. A field nobody can point to in a real sample does not belong in the first version.

  2. Collect and name the files

    Save every PDF under a stable key, such as an account number or ID plus the year or period. The name is how a later run knows what it already has and what is new.

  3. Extract the text and the fields

    Digital PDFs give up their text directly; scanned pages need OCR. A form or table parser, or a language model given the field list, then maps the text to your fields.

  4. Check each value against its source

    Confirm the document is the one it claims to be, keep the page reference, and leave a missing value empty and marked. Anything uncertain goes to a person for review.

  5. Load with a dry run

    Show what would change before writing: new rows, filled gaps and any value that would overwrite an existing one. Then load, and keep the before-and-after files.

What it costs to run

OCR and document parsers are billed per page. The list below is list prices from Google and Amazon read today; a language model, storage and the tool that runs the steps are billed separately by their providers.

ItemList priceSource
Google Document AI, Enterprise Document OCRFirst 1,000 pages listed as free; then US$1.50 per 1,000 pages, up to 5 million pagesGoogle Cloud, Document AI pricingRead on Sep 28, 2026
Google Document AI, Form Parser (fields and tables)US$30 per 1,000 pages, up to 1 million pages (Google's own example: 100 pages cost US$3)Google Cloud, Document AI pricingRead on Sep 28, 2026
Amazon Textract, Detect Document Text (OCR)US$0.0015 per page for the first million pages (US West, Oregon)Amazon Web Services, Amazon Textract pricingRead on Sep 28, 2026
Amazon Textract, Analyze Document with TablesUS$0.015 per page for the first million pages (US West, Oregon)Amazon Web Services, Amazon Textract pricingRead on Sep 28, 2026
Amazon Textract free tier (new AWS customers, first three months)1,000 pages a month of Detect Document Text; 100 pages a month with Forms, Tables and LayoutAmazon Web Services, Amazon Textract pricingRead on Sep 28, 2026

Third-party list prices, read on the date shown. They are not Betterlane prices.

What we learned building it

From the college database in Betterlane's own product, built from public Common Data Set PDFs that colleges publish. Not a client system.

Save every file under an ID and a year
Each PDF was stored under the school's ID plus the year. The database was then joined to the federal IPEDS and College Scorecard data on the federal school ID.
Check the name inside the file before loading
Each document must name the school it was filed under. The check caught files saved for the wrong school (Bard filed as Barnard, WPI as Whitman). Files we could not confirm were left out.
Blank stays blank
A value that is not in the PDF is stored empty and marked as not found. Nothing is guessed to make the table look complete.
Dry run first, and fill gaps instead of overwriting
Every write script runs as a dry run by default and saves before-and-after files. New data fills empty fields; it does not replace values already there.
Say how far the coverage goes
The database holds 1,959 colleges, and admission factors are live for 314 of them. One field that did not hold up, the GPA figures, was withdrawn from the product rather than shown.

When it is not worth it

  • If you receive a few documents a month, typing the fields in by hand is cheaper and easier to trust.
  • If every document has a different layout and the fields you need are not written anywhere consistent, start by agreeing what the documents must contain.
  • If the numbers feed a payment or a legal decision, extraction can prepare them, but a person still has to approve each one.

Questions

Do I need OCR for every PDF?

No. A PDF created by software already contains its text. OCR is needed for scanned pages and photos, and it is where most reading errors come from.

Can a language model do the extraction?

It can map messy text to your fields, but it should not be trusted blindly. Keep the page reference, check each value against the document and send uncertain ones to a person.

What happens when a value is missing?

It stays empty and is marked as not found. Filling it with a likely value makes the whole table harder to trust.

Where does the result go?

Usually a spreadsheet or the system that uses the data, with a column pointing back to the source file and page.

Have a pile of documents you need as a checked table?

See the service: Document processingLet’s talk

All guides