Share this job
Senior Data/ML Engineer
Apply for this job

About the company

Early-stage privacy-tech startup that makes sensitive data (health records, financial transactions, clinical trials) safely usable through synthetic data generation and privacy-safe computation on real data. Selling into some of the largest US health systems.

The role: Senior Data/ML Engineer (100% Remote)

Co-own how the company decides data is safe to release, and build the engineering that makes that decision real at customer scale. You'll be the third technical person on the team, splitting time roughly evenly between privacy/disclosure methodology and the pipeline that produces it.

  • Extend the privacy methodology for synthetic data and decide which real-data aggregates can be released, coarsened or suppressed
  • Own the synthesis pipeline (a typed DAG on Ray): cohort selection, schema reshaping, encoding, distributed training of a database-level generative model, generation and evaluation
  • Make it correct and fast against real customer databases, including debugging data that synthesizes badly
  • Work directly with customers, explaining privacy results to compliance teams

What they're looking for

  • ~5+ years of applied ML or data science at a product or data-infrastructure company
  • Strong Python and production-quality code
  • Comfortable working directly with customers and defending methodological choices
  • Helpful: healthcare or regulated data (EHR, claims, OMOP, HIPAA de-identification), Ray or Spark at scale, generative modeling and evaluation

Tech: Python, Ray, PyTorch, Node, TypeScript, React, Kubernetes on all major clouds

Compensation: $150K-$170K base + equity

Location: 100% Remote (US)

Company name shared after an intro call with HSF.

Apply for this job
Powered by