Senior Site Reliability Engineer, (Ukraine)
$$$
Product
Are you passionate about building reliable, scalable, and highly available production systems? Do you enjoy solving complex engineering challenges through automation, observability, and software engineering?
Optimove is looking for a Senior Site Reliability Engineer (SRE) to join our global SRE organization. In this role, you'll partner with engineering teams globally to improve the reliability, scalability, and operational excellence of our cloud platform.
As a Senior SRE, you'll lead engineering-driven reliability initiatives by building automation, improving deployment processes, advancing observability, and driving operational excellence across our production environment.
Responsibilities
- Reliability Engineering β Design and implement solutions that improve the reliability, availability, and scalability of our production platform while proactively identifying and eliminating operational risks.
- Automation & Platform Engineering β Build internal tools and automation that eliminate manual operational work, improve engineering productivity, and streamline production workflows.
- Observability β Improve monitoring, alerting, dashboards, and production visibility while driving initiatives that reduce alert fatigue, improve incident detection, and strengthen operational insights.
- Production Rollouts β Design, improve, and support production deployment processes using modern release strategies such as Canary, Blue/Green, and Feature Flags to ensure safe and reliable releases.
- Production Reliability & Incident Response β Lead the investigation and resolution of critical production incidents, perform root cause analysis, and implement long-term preventive improvements that improve platform resilience and reduce recurring issues.
- On-Call & Operational Excellence β Participate in the team's on-call rotation, ensuring reliable production operations. Respond to critical incidents, collaborate with engineering teams during service disruptions, and drive engineering improvements, automation, and operational best practices that reduce operational toil and prevent future incidents.
- Cloud & Infrastructure β Design, maintain, and improve our Kubernetes-based cloud platform, CI/CD pipelines, and production infrastructure running on GCP and AWS.
- Engineering Partnership β Collaborate closely with Software Engineers, DevOps, DBAs, and Product teams to improve reliability across the software development lifecycle and promote SRE best practices.
Requirements
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Infrastructure Engineering.
- Hands-on experience operating Kubernetes in production environments.
- Strong experience working with public cloud platforms (GCP or AWS).
- Strong programming and scripting skills (Python preferred; Go or Bash are a plus).
- Experience designing and building automation and internal engineering tools.
- Experience working with CI/CD pipelines and modern deployment methodologies.
- Hands-on experience with observability platforms such as Datadog, Prometheus, or Grafana.
- Strong understanding of Linux, networking, distributed systems, and cloud-native architectures.
- Excellent troubleshooting, debugging, and root cause analysis skills.
- Strong communication skills and experience collaborating with globally distributed engineering teams.
Advantages
- Experience with Infrastructure as Code (Terraform, Ansible, etc.).
- Experience with messaging and distributed technologies such as Kafka, Pub/Sub, or Redis.
- Experience supporting modern deployment strategies such as Canary, Blue/Green, or Feature Flags.
- Familiarity with OpenTelemetry and modern observability tooling.
- Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
- Experience working in large-scale SaaS production environments.
- Relevant cloud or Kubernetes certifications (GCP, AWS, CKA, CKAD).
Why Join Us?
Join a team where Site Reliability Engineering is treated as an engineering disciplineβnot just an operations function. You'll work on large-scale cloud infrastructure serving enterprise customers while building automation, improving observability, and enabling engineering teams to deliver software safely and reliably. - You'll collaborate with a global SRE organization using modern cloud-native technologies, Kubernetes, distributed systems, and large-scale SaaS workloads. You'll have the opportunity to lead impactful reliability initiatives, influence engineering best practices, and solve complex production challenges that directly improve the experience of both customers and engineering teams.
Required languages
English
B2 - Upper Intermediate
Ukrainian
B1 - Intermediate
GCP/AWS, Kubernetes, Python, CI/CD, Datadog, Kafka, Terraform
Published 30 July
19 views
Β·
1 application
π
$3500-5000
Average salary range of similar jobs in
analytics β
Loading...