Site Reliability Team Lead
We are looking for an SRE Team Lead to join our global SRE organization. In this role, you'll lead a team of SREs while partnering with engineering teams globally to improve the reliability, scalability, and operational excellence of our cloud platform.
As an SRE Team Lead, you'll combine hands-on technical leadership with people management, driving engineering-focused reliability initiatives while growing and mentoring a team of SREs across automation, deployment processes, observability, and operational excellence in production.
Responsibilities
• Team Leadership: Manage, mentor, and grow a global team of SRE engineers across Israel and Ukraine, running regular 1:1s, setting goals, and supporting career development. Own hiring,
onboarding, and performance management for the team.
• Distributed Team Operations: Keep the team working as one unit across sites, with shared standards, consistent handoffs, and clear ownership so that reliability work is not fragmented by location or timezone.
• Technical Direction: Set the technical roadmap for the team's reliability, automation, and observability initiatives, and stay hands on enough to guide design decisions and unblock complex problems.
• Reliability Engineering: Guide the design and implementation of solutions that improve the reliability, availability, and scalability of our production platform, and ensure the team proactively identifies and eliminates operational risks.
• Automation and Platform Engineering: Prioritize and oversee the build of internal tools and automation that eliminate manual operational work, improve engineering productivity, and streamline production workflows.
• Observability: Drive the team's roadmap for monitoring, alerting, dashboards, and production visibility, reducing alert fatigue and strengthening operational insight across services.
• Production Rollouts: Oversee the team's work on deployment processes using modern release strategies such as Canary, Blue/Green, and Feature Flags, ensuring safe and reliable releases.
• Production Reliability and Incident Response: Act as an escalation point for critical production incidents, guide root cause analysis, and ensure long term preventive improvements are implemented and tracked.
• On-Call and Operational Excellence: Own the team's on-call rotation and coverage across sites,participate as needed, and drive continuous improvements that reduce operational toil and prevent future incidents.
• Cloud and Infrastructure: Oversee the team's work on our Kubernetes based cloud platform,CI/CD pipelines, and production infrastructure running on GCP and AWS.
• Cross Team Partnership: Represent the SRE team in planning and decision making with Software Engineering, DevOps, DBA, and Product leadership, and align the team's priorities with broader engineering goals.
Requirements
• 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Infrastructure Engineering, including some experience leading or mentoring other engineers.
• Hands on experience operating Kubernetes in production environments.
• Strong experience working with public cloud platforms (GCP or AWS).
• Strong programming and scripting skills (Python preferred, Go or Bash are a plus).
• Experience designing and building automation and internal engineering tools.
• Experience working with CI/CD pipelines and modern deployment methodologies.
• Hands on experience with observability platforms such as Datadog, Prometheus, or Grafana.
• Strong understanding of Linux, networking, distributed systems, and cloud native architectures.
• Excellent troubleshooting, debugging, and root cause analysis skills.
• Strong communication skills and proven experience working with remote colleagues across sites and timezones, with a genuine interest in growing into people leadership.
• Fluent English, written and spoken, as the team works across multiple countries
Advantages
• Prior formal people management or team lead experience.
• Experience leading or coordinating engineers who are not co-located.
• Experience with Infrastructure as Code (Terraform, Ansible, etc.).
• Experience with messaging and distributed technologies such as Kafka, Pub/Sub, or Redis.
• Experience supporting modern deployment strategies such as Canary, Blue/Green, or Feature Flags.
• Familiarity with OpenTelemetry and modern observability tooling.
• Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
• Experience working in large scale SaaS production environments.
• Relevant cloud or Kubernetes certifications (GCP, AWS, CKA, CKAD).