Invofox Self Serve
Invofox Self Serve is an automated document processing API that parses PDFs and images into structured JSON and Markdown, featuring classification, table extraction, and zero-retention compliance.
Invofox Self Serve is an automated document processing API that parses PDFs and images into structured JSON and Markdown, featuring classification, table extraction, and zero-retention compliance.
What the product does and how it is positioned
Invofox provides a unified document extraction API designed to convert unstructured PDFs, scanned pages, and images into normalized JSON datasets and Markdown for downstream LLM and database applications.
The system handles the entire ingestion lifecycle, including document deskewing, dual-pass OCR, document splitting, classification, tabular reconciliation, and automated confidence scoring.
Source-supported ways to use the product
Extracting vendor details, tax identifiers, dates, and line-item totals from supplier invoices into structured JSON payloads.
Parsing borrower information, loan terms, and settlement values from mortgage forms and Closing Disclosures.
Automatically separating batch-uploaded PDFs containing mixed files such as bank statements and payslips into discrete classified records.
Extracting employer data, employee identification numbers, and remuneration totals from standard employee payslips.
The documented workflow, where available
Create an account and generate the necessary API keys.
Select the required standard document models or establish a custom schema.
Transmit documents via a POST request to the extraction endpoint.
Receive the validated, schema-mapped JSON or Markdown output directly or via webhook.
Invofox processes unstructured document uploads through a sequential technical pipeline. When files are received via the REST API endpoint, integrity checks handle password-protected or corrupted files before routing them into pre-processing modules for deskewing, denoising, and sharpening.
Text and structural elements are interpreted using a dual-pass OCR system that isolates layout geography while transcribing text. Multi-page batches are divided by an automated page splitter, after which classification models categorize the document type. Extracted tables, currencies, and dates are then normalized, cross-checked against business rules, and returned via webhook or polling along with confidence scores and region provenance.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
Invofox supports invoices, receipts, delivery notes, purchase orders, payslips, bank statements, utility bills, ID documents, contracts, tax forms, and US mortgage forms such as closing disclosures.
Users can submit corrections back to the feedback endpoint, which updates the pipeline configuration to help reduce the same error from repeating across future documents.
Processing occurs by default within the European Union, with US processing available upon request. Scale and Enterprise plans can opt into zero-retention mode, while on-premise deployments are restricted to Enterprise contracts.
The platform accepts PDF files, scanned documents, and image formats including PNG and JPG for extraction.
The platform natively supports Latin-script languages such as English, Spanish, Portuguese, French, Italian, and German, while non-Latin scripts require a separate configuration window.