
Senior Data Engineer / Databricks Expert
- Remote
- Wrocław, Dolnośląskie, Poland
- PLN 20,000 - PLN 30,000 per month
Job description
Stermedia.ai is looking for an experienced Senior Data Engineer with strong Databricks expertise to join a data platform team building a global analytics environment for a leading pharmaceutical organization.
The platform integrates data from more than 80 facilities worldwide, including laboratories, manufacturing systems, logistics platforms, and operational environments. Its goal is to transform distributed datasets into reliable, analytics-ready data that supports reporting, operational monitoring, and data-driven decision-making.
You’ll contribute to the full delivery lifecycle — from direct client discussions and business requirements analysis to solution design, implementation, and production deployment. Working alongside international data engineering, analytics, and machine learning teams, you’ll focus on building scalable data pipelines and transformation layers using Databricks, Apache Spark, Python, SQL, and Delta Lake.
About the project
The project focuses on building a scalable data platform for operational data from pharmaceutical laboratories, production systems, and global supply chains.
Data from multiple regions is consolidated into a centralized analytics environment. The role focuses on using Databricks and a lakehouse architecture to ingest, process, and organize this data into reliable datasets for downstream analytics.
As a Databricks expert, you’ll design and maintain data pipelines, optimize distributed processing, and support data quality, governance, and production reliability. DBT is a complementary skill for SQL-based transformations and analytics engineering workflows.
Job requirements
Design, develop, and maintain data pipelines in Databricks using Python, PySpark, and SQL
Build and maintain Delta Lake tables and transformation layers following bronze, silver, and gold architecture
Integrate data from multiple source systems into consistent, analytics-ready datasets
Implement incremental processing, change data capture, and schema evolution where required
Develop and orchestrate production workflows, including dependencies, scheduling, retries, and monitoring
Optimize Spark workloads, SQL queries, and compute usage to improve performance and manage costs
Implement automated data quality checks, validation, and pipeline tests
Support data access management, governance, and lineage using Unity Catalog
Maintain reusable code, technical documentation, and Git-based development workflows
Collaborate with business stakeholders, data engineers, analysts, and machine learning specialists to translate requirements into reliable data solutions
Required skills
Strong commercial experience with Databricks, including developing and operating production data pipelines
Advanced knowledge of Python, PySpark, and SQL
Hands-on experience with Apache Spark, distributed data processing, and performance troubleshooting
Practical experience with Delta Lake, including incremental loads, merge operations, and schema management
Strong understanding of lakehouse architecture, ETL/ELT patterns, and data modeling
Experience orchestrating and monitoring workflows in Databricks
Practical knowledge of Unity Catalog, including permissions, data organization, and lineage
Experience optimizing pipeline performance and compute resource usage
Familiarity with cloud storage and services in at least one major cloud environment: Azure, AWS, or GCP
Experience with Git, code reviews, automated testing, and CI/CD workflows
Strong analytical and communication skills, with the ability to work directly with international stakeholders
Good command of English
Nice to have
Experience with DBT, including models, tests, macros, and integration with Databricks
Experience with Structured Streaming and Databricks Auto Loader
Familiarity with Terraform and infrastructure as code
Experience supporting datasets and workflows used by machine learning teams
Experience building global analytics platforms or integrating data from multiple facilities
Familiarity with pharmaceutical, manufacturing, laboratory, or supply chain data
Databricks certifications in data engineering
Technology focus
Data platform: Databricks
Processing: Apache Spark, PySpark
Languages: Python, SQL
Storage and table format: Delta Lake, cloud object storage
Architecture: Lakehouse, bronze / silver / gold layers
Governance: Unity Catalog
Orchestration: Databricks jobs and workflows
Development and delivery: Git, CI/CD, automated testing
Complementary analytics tooling: DBT
We offer you
Opportunities to work with modern data engineering and machine learning technologies
An annual self-development budget
The opportunity to contribute to a variety of interesting projects
Internal workshops and knowledge-sharing sessions
Support for personal branding through articles, conference talks, and leading internal workshops
Flexible working hours
The possibility of remote work
A chillout room, free beverages, and team and company events
A friendly atmosphere
MultiSport
LuxMed
Salary
20,000–30,000 PLN + VAT (B2B)
or
All done!
Your application has been successfully submitted!
You've already applied for this job
Thank you for your interest - we've already received your application, so this new submission can't be accepted. Your previous application is on file.
