The Data Backbone for Medical AI

Senior Data Engineer — Medical AI Data Platform (m/f/d)

Full time · 40h · Dresden (hybrid) · remote within EU possible · €70,000–95,000

Why this role exists

Cancilico is a Dresden-based, seed-funded health-tech company, spun out of TU Dresden and University Hospital Dresden. Our first product, MyeloAID, uses AI to automate the morphological differentiation of bone marrow smears. No one else in the EU is doing AI-driven bone marrow analysis at this depth, and we are expanding into new diagnostic verticals and, over time, multimodal models that combine morphology with clinical, laboratory, and molecular data.

Those ambitions depend on more than training better models. We need to know where every image, annotation, and clinical attribute came from; how it was transformed; whether two records represent the same sample or patient; and which exact data supported an experiment, evaluation, or release.

We have built the foundations of that system: an internal data registry combining PostgreSQL, object storage, DVC, and orchestrated processing pipelines. It manages raw whole-slide and microscope images, metadata history, patient and sample identity, deduplication, protected holdouts, and curated datasets for model development.

Your job is to own and evolve this platform into the dependable data backbone for our computer-vision products and future multimodal AI. You will make complex medical data discoverable, reproducible, secure, and usable by the people building and validating our models.

What you’ll do

This is a deeply hands-on engineering role. You will personally write, test, review, deploy, debug, and maintain production Python and SQL—not only design the system or coordinate its implementation.

  • Own the application and data architecture of our internal data registry, from immutable raw assets through enriched and deduplicated records to curated, versioned datasets for training and evaluation.
  • Build reliable ingestion pipelines for data from hospitals, scanners, microscope cameras, annotation systems, and product workflows. Define contracts and validation at each boundary so partial uploads, retries, and malformed metadata fail safely and visibly.
  • Evolve our PostgreSQL data model and interfaces for images, annotations, patients, samples, clinical metadata, dataset membership, and provenance. Design migrations that preserve history and keep a live system trustworthy as its schema changes.
  • Make data quality operational. Own the platform workflows around identity resolution and deduplication, including checksum, metadata, and perceptual-hash controls. Work with the AI/Data Science team, who own the advanced image-similarity models, to integrate their methods into those workflows. Detect inconsistencies and drift, protect patient-level holdouts from leakage, and give operators clear tools to inspect and resolve ambiguous cases.
  • Publish reproducible datasets that AI engineers can consume confidently. Connect database state, object-store artifacts, code, and manifests so a training run or released model can always be traced back to the exact data it used.
  • Prepare the platform for multimodal data: clinical and laboratory values, molecular findings, additional imaging modalities, and changing annotation taxonomies, without turning the core model into an unmaintainable collection of one-off fields.
  • Operate the platform as a production system. Improve observability, reconciliation, recovery, access boundaries, documentation, and automated tests while working with our infrastructure engineers on the services underneath it.
  • Work with QM/RA to translate data-governance and change-control needs into practical technical controls and evidence. You will contribute the traceability and validation records behind compliance; you will not be expected to own regulatory submissions yourself.

You will collaborate closely with AI engineers, clinical experts, QM/RA, and infrastructure engineers. You will not be expected to own production model development, scientific performance evaluation, or the underlying Kubernetes and networking platform.

What we’re looking for

Must have

  • Strong hands-on software and data-engineering experience. You currently write production-quality Python and SQL and have personally implemented, tested, deployed, debugged, and operated data systems—not only designed or managed them.
  • Strong Python and SQL, including practical PostgreSQL data modelling, query design, and schema evolution.
  • Experience building idempotent and recoverable batch or incremental pipelines, with deliberate validation, retries, observability, and failure handling.
  • A good understanding of data contracts, provenance, versioning, and the trade-offs between relational metadata and large assets in object storage.
  • Strong testing and debugging discipline. You look for silent data loss, duplication, leakage, and partial state—not just whether the happy path completed.
  • Product-minded ownership. You can turn the needs of AI engineers, clinical colleagues, and data operators into a platform that is coherent rather than a sequence of special cases.
  • Confidence using modern AI coding tools critically and effectively as part of your development workflow.

Nice to have

  • Experience building data systems in a regulated, safety-critical, or highly controlled environment. Healthcare, life sciences, GxP, finance, aerospace, and similar backgrounds are all relevant, and we value this experience highly.
  • Experience with medical-device development or verification, ISO 13485, IVDR/MDR, IEC 62304, computerized-system validation, or clinical-performance studies.
  • Experience with medical imaging, computer vision, or other large unstructured datasets. Whole-slide images, microscopy, DICOM, tiling, and image annotations are particularly relevant.
  • Hands-on experience with S3-compatible object storage, DVC or another data-versioning system, and Prefect or a comparable workflow orchestrator.
  • Experience with entity resolution, record linkage, similarity search, or human-in-the-loop deduplication and conflict review.
  • Familiarity with annotation workflows, label taxonomies, COCO-style datasets, and inter-rater disagreement.
  • Experience handling pseudonymised health data, role-based access, retention rules, and auditable change histories.
  • Familiarity with how ML teams consume datasets for training, evaluation, and experiment tracking. You do not need to be a model researcher.

You do not need a medical background, but you should be genuinely curious about the clinical problem and comfortable learning the language needed to model it accurately.

Our stack & setup

  • Data platform: Python, PostgreSQL, SQLAlchemy, Alembic, S3-compatible object storage, and DVC.
  • Pipelines and operations: Prefect for orchestration, with automated health, reconciliation, and data-quality checks.
  • Downstream consumers: PyTorch training repositories, MLflow experiments, annotation tooling, and FastAPI-based product services consume versioned data from the registry.
  • Engineering: GitHub for code, CI/CD, and tasks. The underlying platform is operated with our infrastructure engineers rather than being the primary responsibility of this role.

You’ll report to our CAIO and work alongside AI engineers, infrastructure and product engineers, QM/RA, and ~10 clinical annotation consultants (hematologists and medical technical assistants).

What we offer

  • Salary: €70,000–€95,000 depending on experience and fit.
  • 30 days paid leave, flexible core hours (10:00–15:00).
  • 40-hour weeks.
  • Pick your machine: Mac, Windows, or Ubuntu.
  • Flexible German benefits, negotiable based on your needs.
  • Working language is English. No German required.
  • Hematologists and medical technical assistants are colleagues you can walk over and ask.
  • Hybrid in Dresden preferred. Fully remote within EU time zones possible for the right candidate. You’ll need an existing right to work in the EU; we can’t sponsor relocation or visas at this stage.
  • We welcome applications regardless of background, parental status, disability, or neurodivergence.