OCR in-process
Scanned PDFs read without sending documents to a third party, on the API hosts themselves.
Uploads and scans arrive, text is extracted or OCR'd in-process, keywords and labels are found, and a rules engine files, renames, labels and expires documents on triggers. Every rule can be previewed before it runs, and conflicting rules are caught.
Document capture fails in two ways: the OCR is wrong and nobody notices, or the automation does something nobody expected. We built against both: every stage is versioned so it reruns when improved, and every rule shows what it would do before it does it.
The same parts make an AP invoice or a bill of lading readable and routable in a distributor's back office, which is the service this case study underwrites.
What we built
Scanned PDFs read without sending documents to a third party, on the API hosts themselves.
Extraction, OCR and labelling each carry a version; improving one reruns only that stage.
Text, keyword, file name, type, folder, size, dates, label and OCR state, with six actions from label to delete-after-N-days.
Any rule can be tested against the current documents and explained for one document.
Rules that fight each other are found before they run.
What a rule did, when, and the way back.
Why it matters to you
The same pipeline reads a supplier invoice; the rules match it to a purchase order and route it to an approver; the archive keeps it. The missing piece is your ERP's PO and receipt data, which is an integration we also do.
Not unless you want it to. OCR runs in-process on the hosts we control, and nothing is sent to an outside AI service unless that trade is worth it to you.
A person, always. The preview and explain tools exist so the person sees what the machine saw before anything is posted.
Thirty minutes, no obligation. We reply within one business day.