Data extraction · Burbank, California

Documents go in.
A spreadsheet you can actually use comes out.

Invoices, filings, statements, reports, and authorized public records arrive in different layouts. You get one table, in the columns you asked for, with the source of every row recorded so you can check any line against the original.

714documents attempted
639parsed
5,872rows extracted
75failures, logged with reasons

From a single run. The failures are listed with the reason each one failed — they are not quietly dropped from the count.

One filing, six rows, and the page each one came from

A published document on the left; the spreadsheet rows it became on the right.

A government filing on the left with the filer name and filing ID masked, and on the right the six spreadsheet rows extracted from it, each carrying owner, ticker, transaction type, date, amount, source document and page number.
The filer's name and the filing ID are both masked because this is a sample, not a delivery. Masking the name alone would be cosmetic — the ID leads straight back to it. This particular document is RC4-encrypted with an empty password: it opens fine, but nothing inside decompresses until it is decrypted, so ordinary parsers find no text, assume it is a scan, and reach for OCR. It is not a scan, and OCR would have put errors into financial figures that had none.

Three layouts in. One table out.

Invoices, statements and confirmations agree on nothing. The file you get back does.

Three example source documents — a ruled invoice table, a running-balance account statement, and an order confirmation with no table at all — and the single output table that all three resolve into.
These are example layouts, not a client's files. The third one has no table in it at all, which is the case most converters fail on: they look for ruled rows, find none, and hand back a wall of text. Reading a document is not the same as detecting a grid in it.

A real, public-source sample you can open

A four-row extraction from an official Ohio Legislative Service Commission PDF. It includes the original source link, page citations, normalized amounts, and a total that reconciles to the source.

Grant awards PDF extraction

The workbook shows the exact delivery format: clean rows, clear source references, and a stated limitation instead of an invented claim about automation.

Download .xlsx

A recurring register, not another one-off conversion

This two-period demonstration shows what happens after the next batch arrives: append it to the register, reconcile totals, identify new and changed records, preserve the source trail, and isolate anything that needs a decision.

Summary sheet showing prior and current record counts, new and changed records, exception count, and reconciled current amount.
Control summary. Counts and amounts reconcile before the file is treated as complete.
Change log identifying new, changed, and missing-current-period grant records with source references.
Period changes. New, changed, and missing-current-period records are separated for review.
Exception report showing a conflicting duplicate amount, a missing award date, and a record absent from the current period.
Fail-closed exceptions. Conflicting or missing values stay blank or flagged instead of being guessed.

Demonstration boundary: every record and source filename in this workbook is fictional and was created to exercise the workflow. It proves the delivery structure and error-handling approach; it does not claim customer demand or unattended accuracy.

Recurring grant-record workflow

Five tabs: buyer summary, current register, change log, source data, and exceptions. The workbook is designed for reviewable recurring work, not as a black-box automation claim.

Download .xlsx

How I work

Four rules, agreed before anything starts.

Schema first

You tell me the columns you need before I begin. Every document is made to fit them, and anything that will not fit is reported to you rather than guessed at.

Every row cites its source

Each row carries the file and the page it was read from, so you can open the original and verify any line without taking my word for it.

Nothing is silently dropped

Any line that cannot be read with confidence is flagged in its own column with the reason. A short list of known problems beats a clean-looking file that is quietly wrong.

Normalised on arrival

Dates come back as ISO dates whatever format they went in as. Currency comes back as a number you can add up, not text with a symbol stuck to it.

Who you would be working with

Fox Parker

Data Science student · Pasadena City College

Second year, working toward an Associate in Data Science. Two internships at StoryArc, a UX/UI company — graphic design first, then data entry.

I deliver finished files against a deadline I have agreed to. I am not available for live calls during the working day, and I would rather say that up front than discover it is a problem later.