Event Banner

When the Pipeline Breaks: Building ML Infrastructure for Biotech R&D

How biotech ML teams scale reproducible, auditable workflows as data and compute demands grow

Description

AI-native biotech teams have made machine learning central to how they do discovery and diagnostics, and the work behind it keeps getting heavier: terabyte imaging, multi-step genomics, molecular simulation, GPU-heavy training, and experiments that can take days or months to finish. More of it runs on its own now, with longer chains of steps and less hands-on supervision. Most teams start with their own scripts or a basic scheduler, and that works for a while. Then something shifts. A new data type, a bigger team after a raise, or the first real conversation with regulators. Suddenly you can't re-run an experiment from six months ago without guessing at what changed. Compute bills go up and no one can explain where the money went. The audit trail that used to be fine isn't anymore, and for teams working with patient data, privacy and residency rules like HIPAA and GDPR start shaping how the infrastructure gets built. The team has to decide whether to keep extending what they have, standardize on a platform, or rethink the architecture. This roundtable puts ML and engineering leaders from AI-native biotech and life-sciences companies in the same room to talk through that shift, from reproducibility and compute cost to compliance and keeping research teams focused on the science as the work scales.

Date: 2026-07-22

Time (ET): 2:00 PM EDT, Jul 22, 2026

Time (Local): 6:00 PM UTC, Jul 22, 2026

Location: online

Speakers

Jasmin Bharadiya

Jasmin Bharadiya

Senior Data Engineer, Character Bio

Joshua Paul

Joshua Paul

Head of Compute, Data Science, Software, and Operations, Octant

Niels Bantilan

Niels Bantilan

Chief Machine Learning Engineer, Union.ai

Vivek Mittal

Vivek Mittal

Partner and Managing Director, Head of Biopharma, Health Advances

Nathan Silberman

Nathan Silberman

Chief Technology Officer, Artera AI

Sivan Bercovici

Sivan Bercovici

Chief Technology and Business Officer, Karius

Guided Questions

Vivek Mittal

Vivek Mittal

You evaluate platform technologies at Health Advances and advise companies from discovery through to market, so you've had a close look at a lot of teams from the outside. When you assess one, how much does the state of its data and ML infrastructure actually move the needle, and where have you seen teams look further along than they really are?

Jasmin Bharadiya

Jasmin Bharadiya

At Character you're keeping AI research reproducible across genomics, imaging, and clinical data at once, each with its own quirks. In precisionmedicine work where a biomarker result has to hold up months later, which of those data types quietly causes the most reproducibility trouble, and what did you have to build to keep them trustworthy together instead of one at a time?

Joshua Paul

Joshua Paul

At Octant you've focused on getting data science and computational chemistry into scientists' hands through AI and LLM tooling, sitting on top of a platform generating high-throughput screening data. When those tools go to chemists and biologists instead of staying with your team, what has to hold underneath for the results to stay reliable and reproducible as more people lean on them, and where does that turn out to be harder than it looks from the outside?

Niels Bantilan

Niels Bantilan

You maintain Flyte and you've spent years on the infrastructure that keeps ML pipelines reproducible and runs heavy genomics or imaging workloads across different kinds of compute. You've watched a lot of teams reach the point where a homegrown setup or a general-purpose tool starts to break as they scale. When does that usually happen, and when is it actually not worth switching yet, when are teams better off with the simple thing a while longer?

Nathan Silberman

Nathan Silberman

At Artera you've built an FDA-authorized diagnostic that reads gigapixel pathology slides through a chain of AI models to guide how an oncologist treats a patient. Clearing that regulatory bar asks things of the infrastructure that a research pipeline never has to. What did meeting it force you to lock down that a research team would never think about, and where does that clinical-grade rigor get hardest as you scale to larger foundation models?

Sivan Bercovici

Sivan Bercovici

At Karius you built the platform behind a clinical diagnostic that has to pull real pathogen signal out of millions of DNA fragments per sample, where a false positive isn't a bug, it's a patient consequence. As you add new tests, what's the hardest part of keeping that signal trustworthy, and how do you decide when the evidence is strong enough to report versus hold back?