Senior Software Engineer (Java/C++) โ Query Engine / Data Platform
About the Client:
Our client is a leading enterprise data platform company building an open, high-performance data lakehouse for AI and analytical workloads. The platform combines an intelligent SQL query engine, an AI-ready semantic layer, and an open catalog built on Apache Iceberg โ enabling Fortune 500 companies across finance, energy, manufacturing, and logistics to unify, query, and govern data at massive scale across cloud and on-premise sources.
About the Role:
We are looking for a Senior Software Engineer to design, build, and own significant components of the core query engine of a large-scale distributed data platform. You will lead work across query planning, optimization, and execution โ driving architectural decisions, mentoring middle engineers, and pushing the platform's performance and scalability boundaries across all major clouds.
This is a systems-level, backend engineering role focused on distributed data processing internals โ not application development or CRUD services.
Responsibilities:
- Architect, design, and implement significant components of the query engine โ planner, optimizer, execution operators, memory management, and data access layers.
- Write and review high-quality, performance-critical code in Java and/or C++.
- Own end-to-end delivery of features from technical design through production rollout.
- Drive performance optimization work โ profiling, benchmarking, and eliminating bottlenecks in latency- and throughput-sensitive paths.
- Integrate deeply with columnar formats, open table formats, and connectivity drivers.
- Mentor middle engineers, lead code reviews, and set technical standards for the team.
- Contribute to CI/CD architecture and quality processes across Jenkins, containerization, and multi-cloud deployment (GCP, AWS, Azure via Kubernetes / Docker).
- Debug complex, cross-layer issues spanning query planning, distributed execution, memory management, and I/O.
- Partner with US-based engineering leads on design reviews, architectural decisions, and roadmap execution.
Required Qualifications:
- Education: B.S. or M.S. in Computer Science, Computer Engineering, or a related technical field.
- Programming: Strong proficiency in Java or C++, with deep understanding of OOP, systems design, memory management, and concurrency. Candidates strong in both are especially valued.
- SQL & Data: Advanced SQL skills and strong understanding of query execution internals โ planning, optimization, and operator implementation.
- Data processing systems experience: Demonstrated experience building or extending data processing systems โ query engines, distributed databases, ETL/ELT frameworks, or analytical platforms.
- Engineering experience: 5+ years in backend/systems software engineering, with a track record of owning significant components.
- CI/CD & DevOps: Solid experience with Jenkins and modern engineering workflows.
- Containers & Orchestration: Strong working knowledge of Docker and Kubernetes.
- Cloud: Hands-on experience with at least one major cloud (GCP, AWS, or Azure); exposure to more than one is a strong plus.
- Version Control: Confident with Git / GitHub workflows and rigorous code review practices.
- English: Upper-Intermediate or higher (B2+) โ daily written and verbal communication with a US-based engineering team.
- Availability: Able to work EU hours with a 2โ3 hour shift toward US West Coast time to ensure daily overlap with the client team.
- Leadership: Prior experience mentoring engineers, leading technical designs, or owning components end-to-end.
Desired Skills:
- Deep experience with Apache Arrow (columnar in-memory format).
- Hands-on experience with Apache Calcite (SQL parsing, planning, optimization framework).
- Experience with Gandiva or LLVM-based expression compilation and code generation.
- Experience with open table formats โ Apache Iceberg, Delta Lake, or Hudi.
- Deep familiarity with distributed computing frameworks (e.g., Apache Spark, Kafka) and MPP SQL query engines (e.g., Presto, Trino, or similar).
- Deep understanding of query planning, optimization (cost-based, rule-based), and execution internals.
- Experience with data connectivity drivers: JDBC, ODBC, Arrow Flight.
- Kubernetes on managed services (GKE / EKS / AKS) and multi-cloud exposure.
- IaC tools such as Terraform.
- Track record of performance optimization at the systems level โ profiling, vectorization, cache/memory-conscious design.
- Prior experience contributing to open-source data infrastructure projects.