Skip to content

Senior Data Engineer / Databricks Expert

  • Remote
    • Wrocław, Dolnośląskie, Poland
  • PLN 20,000 - PLN 30,000 per month

Job description

Stermedia.ai is looking for an experienced Senior Data Engineer with strong Databricks expertise to join a data platform team building a global analytics environment for a leading pharmaceutical organization.

The platform integrates data from more than 80 facilities worldwide, including laboratories, manufacturing systems, logistics platforms, and operational environments. Its goal is to transform distributed datasets into reliable, analytics-ready data that supports reporting, operational monitoring, and data-driven decision-making.

You’ll contribute to the full delivery lifecycle — from direct client discussions and business requirements analysis to solution design, implementation, and production deployment. Working alongside international data engineering, analytics, and machine learning teams, you’ll focus on building scalable data pipelines and transformation layers using Databricks, Apache Spark, Python, SQL, and Delta Lake.

About the project

The project focuses on building a scalable data platform for operational data from pharmaceutical laboratories, production systems, and global supply chains.

Data from multiple regions is consolidated into a centralized analytics environment. The role focuses on using Databricks and a lakehouse architecture to ingest, process, and organize this data into reliable datasets for downstream analytics.

As a Databricks expert, you’ll design and maintain data pipelines, optimize distributed processing, and support data quality, governance, and production reliability. DBT is a complementary skill for SQL-based transformations and analytics engineering workflows.


Job requirements

  • Design, develop, and maintain data pipelines in Databricks using Python, PySpark, and SQL

  • Build and maintain Delta Lake tables and transformation layers following bronze, silver, and gold architecture

  • Integrate data from multiple source systems into consistent, analytics-ready datasets

  • Implement incremental processing, change data capture, and schema evolution where required

  • Develop and orchestrate production workflows, including dependencies, scheduling, retries, and monitoring

  • Optimize Spark workloads, SQL queries, and compute usage to improve performance and manage costs

  • Implement automated data quality checks, validation, and pipeline tests

  • Support data access management, governance, and lineage using Unity Catalog

  • Maintain reusable code, technical documentation, and Git-based development workflows

  • Collaborate with business stakeholders, data engineers, analysts, and machine learning specialists to translate requirements into reliable data solutions

Required skills

  • Strong commercial experience with Databricks, including developing and operating production data pipelines

  • Advanced knowledge of Python, PySpark, and SQL

  • Hands-on experience with Apache Spark, distributed data processing, and performance troubleshooting

  • Practical experience with Delta Lake, including incremental loads, merge operations, and schema management

  • Strong understanding of lakehouse architecture, ETL/ELT patterns, and data modeling

  • Experience orchestrating and monitoring workflows in Databricks

  • Practical knowledge of Unity Catalog, including permissions, data organization, and lineage

  • Experience optimizing pipeline performance and compute resource usage

  • Familiarity with cloud storage and services in at least one major cloud environment: Azure, AWS, or GCP

  • Experience with Git, code reviews, automated testing, and CI/CD workflows

  • Strong analytical and communication skills, with the ability to work directly with international stakeholders

  • Good command of English

Nice to have

  • Experience with DBT, including models, tests, macros, and integration with Databricks

  • Experience with Structured Streaming and Databricks Auto Loader

  • Familiarity with Terraform and infrastructure as code

  • Experience supporting datasets and workflows used by machine learning teams

  • Experience building global analytics platforms or integrating data from multiple facilities

  • Familiarity with pharmaceutical, manufacturing, laboratory, or supply chain data

  • Databricks certifications in data engineering

Technology focus

  • Data platform: Databricks

  • Processing: Apache Spark, PySpark

  • Languages: Python, SQL

  • Storage and table format: Delta Lake, cloud object storage

  • Architecture: Lakehouse, bronze / silver / gold layers

  • Governance: Unity Catalog

  • Orchestration: Databricks jobs and workflows

  • Development and delivery: Git, CI/CD, automated testing

  • Complementary analytics tooling: DBT

We offer you

  • Opportunities to work with modern data engineering and machine learning technologies

  • An annual self-development budget

  • The opportunity to contribute to a variety of interesting projects

  • Internal workshops and knowledge-sharing sessions

  • Support for personal branding through articles, conference talks, and leading internal workshops

  • Flexible working hours

  • The possibility of remote work

  • A chillout room, free beverages, and team and company events

  • A friendly atmosphere

  • MultiSport

  • LuxMed

Salary

20,000–30,000 PLN + VAT (B2B)


or