Middle Software Engineer (Java/C++) โ Query Engine / Data Platform
About the Client:
Our client is a leading enterprise data platform company building an open, high-performance data lakehouse for AI and analytical workloads. The platform combines an intelligent SQL query engine, an AI-ready semantic layer, and an open catalog built on Apache Iceberg โ enabling Fortune 500 companies across finance, energy, manufacturing, and logistics to unify, query, and govern data at massive scale across cloud and on-premise sources.
About the Role:
We are looking for a Middle-level Software Engineer to work on the core query engine of a large-scale distributed data platform. You will develop features across query planning, optimization, and execution, contribute to performance-critical components, and help evolve the platform's ability to process massive analytical workloads across heterogeneous data sources.
This is a systems-level, backend engineering role focused on distributed data processing internals โ not application development or CRUD services.
Responsibilities:
- Develop and maintain features across the query engine โ planning, optimization, execution, and data access layers.
- Write high-quality, performance-conscious code in Java and/or C++.
- Work with SQL semantics, query plans, and execution operators over large-scale distributed data.
- Integrate with columnar formats, open table formats, and connectivity drivers used across the platform.
- Contribute to CI/CD pipelines and automated testing โ participate in build, test, and release workflows in Jenkins.
- Deploy and validate work in Kubernetes (GKE / EKS / AKS) across GCP, AWS, and Azure using Docker.
- Debug complex issues spanning query planning, distributed execution, memory management, and I/O.
- Collaborate closely with US-based engineering teams on design reviews, code reviews, and technical decisions.
Required Qualifications:
- Education: B.S. or M.S. in Computer Science, Computer Engineering, or a related technical field.
- Programming: Strong proficiency in Java or C++, with solid OOP and software design fundamentals. Candidates comfortable with both are especially valued.
- SQL & Data: Strong SQL skills and solid understanding of relational and analytical data systems, including query execution concepts.
- Data processing systems experience: Hands-on experience with data processing systems โ query engines, distributed databases, ETL/ELT frameworks, or analytical platforms.
- Engineering experience: 3+ years in backend/systems software engineering.
- CI/CD & DevOps basics: Practical experience with Jenkins pipelines and modern development workflows.
- Containers: Working knowledge of Docker and basic Kubernetes (running workloads, debugging pods, kubectl fluency).
- Cloud: Comfortable operating in at least one major cloud (GCP, AWS, or Azure).
- Version Control: Confident with Git / GitHub workflows.
- English: Upper-Intermediate or higher (B2+) โ daily written and verbal communication with a US-based engineering team.
- Availability: Able to work EU hours with a 2โ3 hour shift toward US West Coast time to ensure daily overlap with the client team.
Desired Skills:
- Experience with Apache Arrow (columnar in-memory format).
- Familiarity with Apache Calcite (SQL parsing, planning, optimization framework).
- Exposure to Gandiva or LLVM-based expression compilation.
- Experience with open table formats โ Apache Iceberg, Delta Lake, or Hudi.
- Familiarity with distributed computing frameworks (e.g., Apache Spark, Kafka) and MPP SQL query engines (e.g., Presto, Trino, or similar).
- Understanding of query planning, optimization, and execution internals.
- Experience with data connectivity drivers: JDBC, ODBC, Arrow Flight.
- Kubernetes on managed services (GKE / EKS / AKS) and multi-cloud exposure.
- IaC tools such as Terraform.
- Experience with performance profiling and optimization of latency-sensitive systems.