Site Reliability Engineer (SRE)
We are looking for a Site Reliability Engineer (SRE) to join a global engineering team building and operating a large-scale cloud SaaS platform. In this role, you'll work closely with experienced SREs and software engineers to improve platform reliability, scalability, and operational excellence.
This position is a great fit for an engineer with experience in SRE, DevOps, or Cloud Operations who wants to deepen expertise in Kubernetes, cloud infrastructure, automation, and production engineering.
Responsibilities
- Production Reliability โ Help maintain the availability, stability, and performance of our production platform while proactively identifying opportunities to improve system reliability.
- Monitoring & Observability โ Improve monitoring, dashboards, and alert quality to increase production visibility and reduce alert fatigue.
- Automation โ Build scripts and automation solutions that reduce manual operational work, improve engineering efficiency, and enhance production reliability.
- Production Operations & Incident Response โ Participate in troubleshooting production issues, perform root cause analysis, and contribute to long-term reliability improvements.
- Cloud Platform Operations โ Support and improve our Kubernetes-based cloud platform, CI/CD pipelines, and production infrastructure running on GCP and AWS.
- Engineering Collaboration โ Work closely with Software Engineers, DevOps, DBAs, and Product teams to improve production reliability and operational excellence across the platform.
On-Call & Operational Excellence โ Participate in the team's on-call rotation, supporting reliable production operations. Respond to production incidents alongside experienced SREs while continuously improving operational processes and reducing manual effort through automation.
Requirements
- 2โ3 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or a similar role.
- Hands-on experience with Kubernetes and containerized environments.
- Familiarity with public cloud platforms (GCP or AWS).
- Experience with Linux systems and basic networking concepts.
- Experience with scripting or programming (Python, Bash, or similar).
- Familiarity with monitoring and observability platforms (Datadog is an advantage).
- Strong analytical and troubleshooting skills.
- Excellent communication skills and the ability to collaborate with globally distributed engineering teams.
- A strong desire to learn, take ownership, and continuously improve systems and processes.
A proactive mindset, strong work ethic, and the hunger to grow as a Site Reliability Engineer.
Nice to Have
- Experience with CI/CD pipelines.
- Familiarity with Infrastructure as Code tools such as Terraform or Ansible.
- Experience with messaging technologies such as Kafka, Pub/Sub, or Redis.
- Exposure to Canary, Blue/Green, or Feature Flag deployment strategies.
- Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
Experience working in cloud-native or SaaS production environments.