Principal Data Infrastructure/Scala Engineer (IRC303896)

GlobalLogic ๐Ÿ”ฅ
$$$$

Job Description
8+ years of software development experience, including experience working on complex backend, distributed systems, data infrastructure, or data-intensive projects.
Proven experience designing and developing scalable backend platforms, distributed services, query engines, data-processing systems, or infrastructure components handling large volumes of data or high-throughput workloads.
Strong hands-on experience with Scala, Java, or another JVM-based language, with the ability to work effectively in Scala-based production codebases.
Hands-on experience with Apache Spark, preferably with Scala, for large-scale distributed data processing.
Experience building or working with distributed query engines, SQL engines, or data access/query platforms, such as Trino, Presto, Spark SQL, Hive, or similar technologies.
Hands-on experience with modern data lake and table formats, such as Apache Iceberg, Delta Lake, Apache Hudi, or similar technologies.
Strong understanding of data lake architecture, including table formats, partitioning, schema evolution, snapshots, versioning/time travel, data consistency, transactions, and compaction.
Experience working with large-scale production data platforms, including environments with large numbers of tables, partitions, or high-volume data processing workloads.
Deep understanding of distributed systems, data processing, storage engines, or data infrastructure, including scalability, fault tolerance, concurrency, and performance optimization.
Experience developing and troubleshooting high-throughput, low-latency, production-grade systems.

The role is primarily focused on backend, distributed systems, data infrastructure, and system-level engineering rather than full-stack application development.
Job Responsibilities
Architect & Build: Lead the architecture, design, and implementation of highly scalable, complex backend systems and robust data infrastructure using Scala.
Data Lake Engineering: Design and manage enterprise-grade data lake solutions leveraging Apache Iceberg or equivalent table formats, ensuring optimal data consistency, transactions, schema evolution, partitioning, and compaction.
Performance Optimization: Maximize system throughput, minimize latency, and optimize resource utilization across large-scale distributed data processing systems.
Production Excellence: Systematically diagnose, troubleshoot, and resolve complex, high-impact production and customer-facing infrastructure issues.
Codebase Modernization: Champion the gradual improvement and modernization of a mature, production-critical codebase without causing disruption to active customers or services.
Technical Leadership & Collaboration: Act as a technical authority, fostering strong communication and collaboration across engineering teams to align on technical objectives.
Navigate Ambiguity: Exercise a strong ownership mindset to operate effectively and deliver high-quality outcomes in areas with limited documentation or historical knowledge.
Department/Project Description
We are seeking an expert Data/Backend Infrastructure Engineer with strong distributed systems experience, large-scale data processing, cloud-native architecture, data lake/table formats, query engines, and production systems engineering.

Required skills experience

Scala 5 years
AWS, Azure, Apache Spark, Apache Flink, Presto

Required languages

English B2 - Upper Intermediate
Ukrainian Native
Published 4 September ยท Updated 7 October
Last 30 days: 44 views ยท 4 applications
All time: 74 views ยท 8 applications
Response activity: Medium
Last responded yesterday
See stats of candidates who applied for this job ๐Ÿ‘€
To apply for this and other jobs on Djinni login or signup.
Loading...