Principal Data Engineer

Careers

Requisition ID

549681

Category

Research & Development

Location

United Kingdom, London

Apply

Location: This is a hybrid remote/in-office role.

About our Company:

Medidata is powering smarter treatments and healthier people through digital solutions to support clinical trials. Celebrating over 25 years of ground-breaking technological innovation across more than 38,000 trials and 12 million patients, Medidata offers industry-leading expertise, analytics-powered insights, and one of the largest clinical trial data sets in the industry. More than 1 million registered users across approximately 2,300 customers trust Medidata's seamless, end-to-end platform to improve patient experiences, accelerate clinical breakthroughs, and bring therapies to market faster. A Dassault Systèmes brand (Euronext Paris: FR0014003TT8, DSY.PA), Medidata is headquartered in New York City and has been recognised as a Leader by Everest Group and IDC. Discover more at www.medidata.com. Listen to our latest podcast, from Dreamers to Disruptors, and follow us at @Medidata.

Our team

At the heart of Medidata's ecosystem, the Data Platform Team powers and connects every application, serving as the engine for enterprise data convergence. Every interaction across our global platform generates critical data—and our mission is to transform that raw information into high-impact insights.

Operating on a Data-as-a-Product philosophy, our team builds the foundational infrastructure. This infrastructure includes high-throughput streaming and cloud warehousing to automated data governance. It fuels clinical analytics, AI/ML innovations, and global data sharing.

Why Join Us?

  • Direct Impact at Scale: Promote the central data engine behind every Medidata application, directly accelerating clinical trials and life-saving operational outcomes globally.
  • Modern Distributed Stack: Build at the intersection of real-time event streaming pipelines, scalable cloud data warehouses, and enterprise-grade automated data security.
  • Product-Minded Engineering: Treat data as a first-class product, transforming static databases into high-value, reusable assets for internal AI/ML teams and external partners.
  • Fuel Advanced AI/ML: Promote next-generation predictive modelling and clinical analytics by standardising and safeguarding complex healthcare datasets.

What will you do

Reporting to a Director of Engineering, as a Principal Data Engineer / Architect, you will lead the strategic vision and hands-on execution of our next-generation Object-Centric Data Fabric. You will transition traditional application-centric architectures into a centralised semantic layer that seamlessly unifies multi-stream operational data—including Electronic Data Capture (EDC), patient telemetry, and real-world health datasets. In this role, AI augmentation is natively woven into your workflow. It acts as a force multiplier to automate routine mapping, query optimization, and regulatory documentation. This allows you to focus on driving high-impact platform architecture.

  • Data Fabric & Lake Architecture: Architect and evolve the enterprise semantic data fabric, converting multi-stream clinical execution datasets into an object-centric model. Design and execute a modern Data Lake strategy centred on Apache Iceberg as the core storage format, ensuring high-performance querying and seamless interoperability with Snowflake and heterogeneous compute engines.
  • AI-Accelerated Schema & Pipeline Engineering: Develop and maintain end-to-end multi-stream ingestion pipelines for complex clinical trial schemas. Use AI-driven schema inference and ontology alignment tools to auto-draft mapping artifacts, dramatically reducing integration timelines across different life science datasets.
  • High-Throughput Streaming & Backend Services: Build scale, fault-tolerant real-time ingestion pipelines using Kafka, AWS, and Snowflake. Write robust enterprise services in Java or Scala, leveraging AI coding assistants for rapid code generation, refactoring, and performance tuning.
  • Technical Strategy & Database Optimization: Promote technical direction and engineering best practices across teams for Change Data Capture (CDC), clustering, data migration, and aggregation. Use AI query-optimization tools to analyse execution plans, auto-tune complex Snowflake/Iceberg SQL workloads, and eliminate performance bottlenecks.
  • Intelligent Telemetry & Closed-Loop Reasoning: Integrate automated AI inferencing and reasoning layers directly into data pipelines to detect telemetry anomalies in real-time and automatically map safety signals back to operational trial datasets.
  • Quality & AI-Driven Compliance: Lead Test-Driven Development (TDD) and Behavior-Driven Development (BDD) initiatives. Use AI test generators to produce HIPAA/GxP-compliant synthetic clinical trial datasets for automated validation. Leverage AI tools to auto-draft validation artifacts, data lineage manifests, and audit documentation required under GxP, HIPAA, and GDPR standards.
  • System Resilience: Troubleshoot complex production issues across distributed data environments and implement resilient, self-healing pipeline architectures.

Key Business Value & Strategic Impact

Your leadership will directly advance life sciences technology. By pairing modern lakehouse architecture (Snowflake, Apache Iceberg, Kafka) with AI-augmented workflows, you will empower our platform to process critical safety signals in hours rather than weeks. This enables adaptive clinical trial execution, guarantees zero-loss data integrity, and significantly accelerates regulatory submission timelines for life-saving therapies.

Qualifications

  • Education & Experience: Bachelor's or Master's degree in Computer Science, Data Science, Software Engineering, or equivalent practical experience, alongside proven years of dedicated professional experience in enterprise data engineering and architecture.
  • Data Warehousing & Data Lake Mastery: Expert-level mastery of SQL (OLAP/OLTP) and enterprise cloud data platforms (specifically Snowflake), combined with practical experience architecting open data lake structures using Apache Iceberg.
  • Software Engineering & Streaming: Demonstrated expertise building production backend services in Java or Scala on AWS, alongside deep familiarity with real-time streaming architectures (e.g., Apache Kafka) and modern data engineering design patterns.
  • AI-Augmented Engineering Proficiency: Active, practical experience integrating AI developer tools (e.g., GitHub Copilot, Cursor) into daily workflows to accelerate SQL query generation, code refactoring, automated testing, and technical documentation drafting.
  • Engineering Rigor & Methodologies: Proven command of Git revision control, CI/CD pipeline automation, and TDD/BDD practices, augmented by AI-driven test case generation and quality checks.
  • Domain & Regulatory Awareness: Strong foundational understanding of clinical trial data workflows (e.g., EDC architectures), healthcare data models, and life science compliance standards (GxP, HIPAA, GDPR), with the ability to apply AI/ML tools for schema mapping and zero-loss data integrity verification.

Base pay is one part of the Total Rewards that Medidata provides to compensate and recognise employees for their work. Most sales positions are eligible for a commission on the terms of applicable plan documents, and many of Medidata's non-sales positions are eligible for annual bonuses. Medidata believes that benefits should connect you to the support you need when it matters most and provides benefits, including medical, dental, life and disability insurance; a generous pension; and 25+ paid holidays per year.

We will accept applications on an ongoing basis until we fill the position.

#LI-Hybrid

#LI-AB1

Note: Please be on the lookout for job scams. Medidata recruiters will never ask applicants for monetary compensation, credit card, or banking details.