Senior Software Engineer (Java/C++) — Query Engine / Data Platform Offline
About Our Client:
Our client is a leading enterprise data platform company building an open, high-performance data lakehouse for AI and analytical workloads. The platform combines an intelligent SQL query engine, an AI-ready semantic layer, and an open catalog built on Apache Iceberg — enabling Fortune 500 companies across finance, energy, manufacturing, and logistics to unify, query, and govern data at massive scale across cloud and on-premise sources.
About the Role:
We are looking for a Senior Software Engineer to work on the core query engine of a large-scale distributed data platform. You will own significant components across query planning, optimization, and execution — both driving improvements and acting as a trusted owner of production quality: diagnosing complex customer-reported issues and delivering safe, well-tested fixes across supported release lines. Engineers with staff-level scope and ambitions are very welcome.
This is a systems-level, backend engineering role focused on distributed data processing internals — not application development or CRUD services.
Responsibilities:
- Architect, design, and implement query engine components — planner, optimizer, execution operators, memory management, and data access layers
- Diagnose and resolve complex production and customer-reported issues across the full query path — correctness, performance, memory, concurrency, resiliency, and compatibility
- Develop, review, and backport fixes across multiple supported release lines; strengthen regression prevention
- Write and review high-quality, performance-critical code in Java and/or C++
- Analyze query profiles, execution plans, logs, and telemetry to drive performance optimization — profiling, benchmarking, and eliminating bottlenecks in latency- and throughput-sensitive paths
- Integrate with columnar formats, open table formats, and connectivity drivers
- Improve regression tests, diagnostics, and observability; contribute to CI/CD architecture and quality processes in Jenkins
- Mentor middle engineers, lead code reviews, and help set technical standards
- Debug complex cross-layer issues spanning query planning, distributed execution, memory management, and I/O, partnering with US-based engineering leads.
Required Qualifications:
- B.S. or M.S. in Computer Science, Computer Engineering, or a related field
- 5+ years in backend / systems software engineering with a track record of owning significant components end-to-end
- Strong proficiency in Java or C++ (comfort with both is especially valued), with deep understanding of OOP, systems design, memory management, concurrency, and asynchronous programming
- Advanced SQL and solid understanding of query execution internals — planning, optimization, operator implementation
- Experience building or extending data processing systems: query engines, distributed databases, ETL/ELT, or analytical platforms
- Ability to assess change risk in a mature production system; experience supporting software across multiple versions, branches, or customer environments
- Hands-on experience with Jenkins and modern CI/CD workflows; Docker and Kubernetes
- Hands-on experience with at least one major cloud (GCP, AWS, or Azure)
- Confident Git/GitHub workflows and rigorous code-review habits
- Comfortable with AI-assisted development workflows — using modern AI tools for code comprehension, debugging, and test generation, and critically validating their output
- Prior experience mentoring engineers, leading technical designs, or owning components end-to-end
- English Upper-Intermediate or higher (B2+) — daily written and verbal communication with a US-based engineering team
- Availability to work EU business hours shifted 2–3 hours later for daily overlap with US West Coast mornings.
Desired Skills:
- Apache Arrow (columnar in-memory format) and SQL planner/optimizer frameworks such as Apache Calcite
- LLVM-based runtime expression compilation / code generation
- Open table formats — Apache Iceberg, Delta Lake, or Hudi — and columnar file formats such as Parquet, Avro, or ORC
- MPP query engines (Presto, Trino, or similar); background with distributed data platforms such as Spark, Snowflake, or Databricks
- Deep understanding of query optimization (cost-based, rule-based), vectorized execution, memory management, spilling, and caching
- Connectivity drivers: JDBC, ODBC, Arrow Flight
- Managed Kubernetes (GKE/EKS/AKS) and multi-cloud exposure; Terraform; cloud object storage (AWS S3, Azure Data Lake Storage, Google Cloud Storage)
- SaaS operations experience: incident response, profiling, observability, release and backport processes
- Systems-level performance work and open-source contributions to data infrastructure projects.
Details:
- Engagement: Long-term contract
- Location: Europe (EU / EEA / UK), remote
- Working hours: EU business hours, shifted 2–3 hours later for daily overlap with the US West Coast team