Why a generation of polished SaaS labeling platforms doesn't fit regulated life-sciences, and what the on-prem alternative needs to look like.
DataTorch · March 26, 2026
A note before the argument: the cloud-native annotation platforms (Labelbox, Scale, V7, Sama, and friends) are good products. The engineering teams behind them ship faster than we do. The interfaces are polished, the integrations are real, and for many ML teams they're the obviously-right choice.
This post is about the labs they don't fit. There are more of them than the SaaS market has noticed.
The labs we work with share four constraints. Any one of them complicates a cloud-only labeling tool. Together, they make it a non-starter.
Data cannot leave the environment. Patient samples, proprietary imaging assays, IP-sensitive datasets: depending on which one applies, the operative regulation is HIPAA, CLIA, GxP, FDA QSR, or a pharma NDA. The detail varies; the conclusion is the same: the data and its derivatives stay inside.
Models cannot leave the environment. Cytogenetics labs that have invested in proprietary scoring models, pharma R&D groups with imaging models trained on internal data, specialty diagnostics shops with hard-won classifiers: none of them want their model weights sitting on a vendor's GPU. The model is the asset; the cloud tool wants to host it.
The workflow has structure the vendor doesn't model. Assay-specific fields, range checks, reviewer sign-off chains, cytogenetic LUTs, lot tracking. Generic labeling tools render this as "custom metadata," which works until QC asks you to reproduce the exact reviewer state for a case from nine months ago.
The buyer is not the engineer. In the labs we serve, the person making the decision is a Lab Director, Director of Cytogenetics, or VP of Clinical Informatics. They aren't evaluating annotation tools on developer experience. They're evaluating on whether their compliance team will sign off and whether their SMEs will tolerate the change.
A SaaS tool optimized for a Series-B AI startup is solving a different problem.
If you accept the four constraints above, the requirements on the alternative tool are concrete:
That set of requirements is approximately what DataTorch is. We didn't invent these criteria; we listened to the labs we wanted to work with and built around what they couldn't have.
If your data and models can sit in someone else's cloud, the cloud-native tools are great and you should use them. We don't compete for that buyer.
If they can't, the on-prem option needs to be just as serious as the cloud option. That's the bar we've been trying to clear.
Book a call if you want to compare against a specific cloud tool you've already evaluated.
Tools and community for image annotation: label, review, and share datasets, from everyday photos to the imagery only specialists can read.
© 2026 DataTorch. All rights reserved.