Our role
Designed and developed for the client: extraction pipeline, bulk processing and the export and API surface.
What it is
Document data extraction
Industry
Finance operations
Live at
pdfdata.co
Public traction
About $5.3K/month in reported revenue; plans from about $29/month
Indie Hackers; product site, as of September 2026
Upload invoices and receipts and get structured data back: line items, totals, dates and vendors, exported to Excel, CSV or JSON, or delivered through an API at volume.
The problem
Somebody types invoices into a spreadsheet. Line items, totals, dates, vendors, one document at a time, in a month-end week that always runs longer than planned. The error rate is not the real cost — the real cost is that nothing downstream can start until the typing is finished, so reporting is always behind the business.
The first call
We started from the documents rather than the software: how many arrive, in what formats, from how many vendors, and what happens to the data once it is out. That decided the shape of the build — a pipeline, not a viewer — and made the export format the first thing to get right instead of the last.
What we built
An extraction pipeline that reads invoices and receipts and returns structured records: line items, totals, dates and vendors, exported to Excel, CSV or JSON. Bulk processing for the month-end pile, and an API for the teams that would rather never open the product at all, so one extraction serves a person and a system equally.
How they run today
Documents go in as they arrive and the structured file is waiting; at volume the same thing happens through the API without anyone opening a browser. Month end is a review of exceptions rather than a week of typing.
What changed
The bottleneck moved from data entry to decisions. Reporting describes where the business is rather than where it was when the typing finished, and the people who were typing are reading the numbers instead, which is the job they were hired for.
Screens
Services behind it
- AI SaaS / AI Product DevelopmentComplete AI-native products: architecture, frontend, backend, auth, billing, agents, pipelines, dashboards and cloud.Product strategy, AI architecture, Billing and subscriptions, Dashboards
- Data & AI InfrastructureMake company data usable by AI: ingestion, extraction, retrieval pipelines, vector search and AI-ready APIs.Document ingestion, Extraction, Vector search, Knowledge bases
Not sure where to start?
Describe the problem in a few lines. You get a straight answer on what we would build and how long it takes.