Why we built Agents and Pipelines for regulated labs

Cloud labeling tools assume you'll use a vendor model. Regulated labs can't. Here's the architecture choice we made instead.

DataTorch · July 15, 2025

Most cloud annotation platforms ship with a built-in model: you upload your data, their model pre-labels it, your reviewers correct, training data comes out. It's a clean story. It also doesn't work if you can't put your data in their cloud, or if your lab already has its own ML stack that took years to validate.

When we started talking to regulated life-sciences labs (cytogenetics groups, pathology teams, specialty diagnostics), we kept hearing the same shape of problem. Their SMEs were the bottleneck on scoring throughput. They had, or wanted, in-house models to take pressure off those SMEs. But every option on the market either required moving data to a vendor cloud, or required throwing out the model they'd built and using the vendor's.

That's not a feature gap. That's an architectural assumption. So we built around the opposite assumption: your model is the model.

Agents and Pipelines, briefly

In DataTorch, a Pipeline is a workflow that runs against incoming data. The interesting unit inside a Pipeline is a Job Agent: a container, service, or runtime that emits structured outputs. You wrap your inference (whatever it is: Python service, Docker container, ONNX runtime) as a Job Agent and call it from a Pipeline. The pipeline runs your inference against new data, predictions land in the SME review surface, corrections feed back as labeled training data.

The agent owns the model. We own the orchestration, the review UI, and the audit trail. The split is deliberate.

What this means for your lab

A few things flow from the split.

Your ML team keeps their stack. No rewriting inference to fit a vendor SDK. If your model runs as a Docker container today, that's the artifact you point a Job Agent at.

Your data stays inside. The Pipeline runs inside your environment; your model runs inside your environment; the SMEs review inside your environment. There is no point in the loop where your data has to leave the perimeter for the workflow to function.

Your corrections compound. Every SME correction is a labeled training example. You can export them on whatever schedule fits your retraining cycle: weekly, monthly, never. We don't see any of it.

What it doesn't do

Worth being honest: Agents and Pipelines isn't a magic-model menu. We don't ship pre-trained models for cytogenetic scoring or digital pathology. If you don't have a model and don't want to build one, DataTorch is the wrong tool. The product assumes you've already decided to own your ML, or are actively heading that way.

If that sounds like your lab, we should talk.

Tools and community for image annotation: label, review, and share datasets, from everyday photos to the imagery only specialists can read.

Platform

PlatformImaging labsExplore datasetsEnterprisePricing

© 2026 DataTorch. All rights reserved.