Data/Databricks Engineer
$$$$
We're looking for a Middle+ Data/Databricks Engineer to join our data team and take ownership of pipelines built on Databricks. You'll design and run production-grade ETL/ELT workloads, work with the lakehouse (Delta Lake) architecture, and partner closely with analysts, data scientists, and business stakeholders who rely on the data you deliver. "Middle+" here means roughly 3โ5 years of hands-on data engineering experience: you can own a pipeline end to end with minimal supervision, make sound design decisions, and are starting to mentor more junior engineers โ without yet carrying full architectural ownership of the platform.
Contract & Logistics
- Contract type: B2B (valid Czech IฤO / trade license required).
- Location: Remote, candidate must be based in the Czech Republic; occasional visit to Prague office.
- Start date: ASAP
- English: Fluent
- Czech: Native/Fluent
What You'll Do
- Design, build, and maintain scalable ETL/ELT pipelines on Databricks using PySpark and SQL.
- Implement and evolve Delta Lake tables following a medallion (bronze/silver/gold) architecture.
- Develop batch and, where needed, streaming data processing jobs (Structured Streaming).
- Build and maintain CI/CD for Databricks notebooks and jobs (Databricks Asset Bundles, Repos, Git-based workflows).
- Orchestrate pipelines with Databricks Workflows, Azure Data Factory, Airflow, or similar tools.
- Monitor job performance and cluster costs; troubleshoot failures and optimize for reliability and spend.
- Implement data quality checks, testing, and monitoring for critical datasets.
- Collaborate with analysts, data scientists, and business stakeholders to translate requirements into data models.
- Participate in code reviews, document technical decisions, and help shape engineering best practices.
- Provide guidance and informal mentorship to junior team members.
What We're Looking For
- 3+ years of hands-on experience with Apache Spark, ideally on Databricks specifically.
- Strong PySpark and SQL skills, including performance tuning of Spark jobs.
- Practical experience with Delta Lake and lakehouse architecture concepts.
- Solid grasp of data modeling (dimensional modeling, star schema) and ETL/ELT design patterns.
- Experience with at least one major cloud platform โ Azure, AWS, or GCP (Azure Databricks is a plus).
- Familiarity with pipeline orchestration tools (Databricks Workflows, ADF, Airflow, or equivalent).
- Working knowledge of Git and CI/CD practices for data engineering.
- Understanding of data governance, access control, and security basics (Unity Catalog is a plus).
- Good written and spoken English; Czech or Slovak is a plus but not required.
- Comfortable working independently and communicating clearly in a remote, distributed team.
Nice to Have
- Databricks certification (Data Engineer Associate or Professional).
- Experience with streaming technologies such as Kafka or Azure Event Hubs.
- Experience with dbt for transformation and testing.
- Solid Python software-engineering practices โ testing, packaging, typing.
- Infrastructure-as-code experience (Terraform).
- Exposure to MLOps or ML pipeline support.
- Experience in a regulated or large enterprise data environment.
What We Offer
- Fully remote work with flexible hours, based anywhere in the Czech Republic.
- Long-term B2B cooperation with a stable pipeline of work.
- A modern, actively evolving data stack rather than legacy maintenance work.
- Direct collaboration with senior engineers and architects โ real input on design decisions.
- Budget for training, certifications, and conference attendance.
- A flat structure with fast decision-making and minimal bureaucracy.
Required languages
English
C1 - Advanced
Czech
C1 - Advanced
Published 12 August
18 views
ยท
0 applications
๐
Average salary range of similar jobs in
analytics โ
Loading...