DataTorch is the review-and-AI layer for document work. It reads the scans your digitization pipeline already produces, lets your own extraction models take the first pass, and routes what is left to the person qualified to judge it. INPI runs document analysis on DataTorch, and the same workflow appears in published work from the University of Sydney and TU Wien.
Teams extract and classify information from forms, filings, and records by hand. Getting models to do the first pass, reviewing their output in the same place, and feeding corrections back rarely fits legacy tooling.
DataTorch does not scan your documents and does not hold your system of record. It works on the files your digitization pipeline already writes to storage you control, and hands the reviewed result back — so adding it does not disturb the scanners, the records system, or the processes built around them.
Yours
Your scanning and digitization
01
Connect
Point DataTorch at the storage your scans already land in — S3, Azure, or Google Cloud. Files are read in place.
02
Model first pass
Your own OCR, layout, or NLP models run the first extraction pass from a pipeline step, on your compute.
03
Expert review
Extractions land beside the page against your schema. A reviewer confirms or corrects in one click.
04
Export
Confirmed fields leave through the API or a structured export — into your records, and into your next training run.
Yours
Your records system
Review any document type in one annotation workflow.
Confirm and correct extracted fields in place, against your schema.
Reviewers claim work from a task inbox, and every run keeps its log.
First-pass OCR, layout, or NLP from your own models, on your compute.
01 / Expert review
Key–value fields and classification schemas match your work, and the model’s extraction lands beside the page so your reviewer confirms or corrects in one click. Reviewers claim assigned work from a task inbox, and each step records who submitted what, with its run log retained.
02 / Your models
Wrap your own OCR, layout, or NLP models as Agents and wire them into a pipeline for first-pass extraction. They run on your compute and their predictions surface in the normal review flow — your weights and your documents never go to a vendor model. Every page a reviewer confirms adds to the labeled set your next training run uses.

Data control
The questions your IT and quality teams will ask first, answered without hedging.
Your bucket, your credentials
Connect your own S3, Azure Blob, or Google Cloud Storage. DataTorch reads the images in place — they are never copied into a vendor bucket.
Your models, your hardware
Model agents run as Python on compute you control. Your weights and your images never reach a third-party model.
Hosted, your cloud, or no network at all
Run on our hosted cloud, deploy the same product into your own cloud, or install it fully air-gapped via Helm with an offline signed license.
Your data stays portable
Annotations export to COCO, YOLO, or JSON, and the GraphQL API and Python SDK reach every project — so the work is yours to take elsewhere.
DataTorch is not a certified or validated system, and it does not replace your quality process. It is built to run inside the environment your quality team already controls — on your hardware, against your storage, with no outbound path for your images. We’ll go through the specifics with your team on a call.
INPI uses DataTorch for document analysis, and the same workflow shows up in peer-reviewed and preprint research:
Form-NLU: Dataset for the Form Language Understanding
Ding et al. · University of Sydney · SIGIR 2023
Synthetic Data for Applications in Document Analysis
Muth · TU Wien · 2023
Start free on our hosted cloud, private project included. Or book 30 minutes to scope an on-prem or bring-your-own-cloud deployment — no slides either way.
datatorch@datatorch.io · we reply within a business day
The annotation platform for specialized imagery — review, score, and share datasets your team works from.
© 2026 DataTorch. All rights reserved.