Cloud annotation tools and the labs that can't use them

Why a generation of polished SaaS labeling platforms doesn't fit regulated life-sciences, and what the on-prem alternative needs to look like.

DataTorch · March 26, 2026

A note before the argument: the cloud-native annotation platforms (Labelbox, Scale, V7, Sama, and friends) are good products. The engineering teams behind them ship faster than we do. The interfaces are polished, the integrations are real, and for many ML teams they're the obviously-right choice.

This post is about the labs they don't fit. There are more of them than the SaaS market has noticed.

The shape of "doesn't fit"

The labs we work with share four constraints. Any one of them complicates a cloud-only labeling tool. Together, they make it a non-starter.

Data cannot leave the environment. Patient samples, proprietary imaging assays, IP-sensitive datasets: depending on which one applies, the operative regulation is HIPAA, CLIA, GxP, FDA QSR, or a pharma NDA. The detail varies; the conclusion is the same: the data and its derivatives stay inside.

Models cannot leave the environment. Cytogenetics labs that have invested in proprietary scoring models, pharma R&D groups with imaging models trained on internal data, specialty diagnostics shops with hard-won classifiers: none of them want their model weights sitting on a vendor's GPU. The model is the asset; the cloud tool wants to host it.

The workflow has structure the vendor doesn't model. Assay-specific fields, range checks, reviewer sign-off chains, cytogenetic LUTs, lot tracking. Generic labeling tools render this as "custom metadata," which works until QC asks you to reproduce the exact reviewer state for a case from nine months ago.

The buyer is not the engineer. In the labs we serve, the person making the decision is a Lab Director, Director of Cytogenetics, or VP of Clinical Informatics. They aren't evaluating annotation tools on developer experience. They're evaluating on whether their compliance team will sign off and whether their SMEs will tolerate the change.

A SaaS tool optimized for a Series-B AI startup is solving a different problem.

What the alternative needs to look like

If you accept the four constraints above, the requirements on the alternative tool are concrete:

  • Deploys inside the customer's environment. Not "VPC-tenanted on the vendor cloud", actually inside.
  • Supports air-gapped installs, because some labs need that.
  • Treats the customer's ML as the primary model, not a fallback to a vendor model.
  • Treats the workflow's structure (schemas, sign-offs, lookup tables) as first-class, not as a metadata blob.
  • Doesn't ship telemetry, license check-outs, or support backdoors. The audit you sign off on once is the audit that holds.

That set of requirements is approximately what DataTorch is. We didn't invent these criteria; we listened to the labs we wanted to work with and built around what they couldn't have.

A pragmatic note

If your data and models can sit in someone else's cloud, the cloud-native tools are great and you should use them. We don't compete for that buyer.

If they can't, the on-prem option needs to be just as serious as the cloud option. That's the bar we've been trying to clear.

Book a call if you want to compare against a specific cloud tool you've already evaluated.

Tools and community for image annotation: label, review, and share datasets, from everyday photos to the imagery only specialists can read.

Platform

PlatformImaging labsExplore datasetsEnterprisePricing

© 2026 DataTorch. All rights reserved.