Shift Site Reliability Engineer
Stape, Inc
๐ฅ
Responds Quickly
$$$$
Product
Key tasks:
- On-call rotation for 24/7 support of the main products and services.
- Document issues and remediation steps.
- Uphold SLAs and SLOs by applying SRE best practices, including incident response, post-mortem analysis, and the creation of operational playbooks.
- Enhance infrastructure health by implementing checks and scripts to address known issues.
- Proactively create monitors within the GKE/K8s ecosystem.
Your background:
- At least 3 years of experience as a CloudOps/SRE Engineer.
- 2+ years of extensive experience with Kubernetes (deployment, scaling, networking, troubleshooting).
- Experience with monitoring tools like Prometheus, Grafana, and logging solutions like Grafana Loki Stack/Promtail or analogs.
- Hands-on experience with high-load applications in production will be a big plus.
- Practical experience in supporting business applications (like NodeJS, PHP, GoLang, Python).
- Strong understanding of networking concepts, protocols, and microservices architecture.
- Proficiency in Git & GitHub or other version control.
- Experience with issue processing (RCA, Postmortems).
- Familiarity with incident response and management tools like BetterStack, PagerDuty, or others.
Will be a plus:
- Familiarity with GCP/Scaleway, Terraform, Docker, Linux, CI/CD, PostgreSQL/MySQL.
- Proficiency in at least one scripting language (e.g., Bash, Python).
- Solid spoken and written English skills (ideally Upper-Intermediate level or higher).
- Familiarity with Cloudflare (Workers, WAF, managing DNS, and Cloudflare for SaaS).
- Google Cloud Certifications (or at least desire to obtain one).
Required skills experience
Kubernetes
2 years
Required languages
English
B2 - Upper Intermediate
Published 21 September
35 views
ยท
1 application
๐
$2700-5000
Average salary range of similar jobs in
analytics โ
Loading...