Stape, Inc

Shift Site Reliability Engineer

Stape, Inc ๐Ÿ”ฅ Responds Quickly
$$$$
Product

Key tasks:
 

  • On-call rotation for 24/7 support of the main products and services.
  • Document issues and remediation steps.
  • Uphold SLAs and SLOs by applying SRE best practices, including incident response, post-mortem analysis, and the creation of operational playbooks.
  • Enhance infrastructure health by implementing checks and scripts to address known issues.
  • Proactively create monitors within the GKE/K8s ecosystem.

 

Your background:

  • At least 3 years of experience as a CloudOps/SRE Engineer.
  • 2+ years of extensive experience with Kubernetes (deployment, scaling, networking, troubleshooting).
  • Experience with monitoring tools like Prometheus, Grafana, and logging solutions like Grafana Loki Stack/Promtail or analogs.
  • Hands-on experience with high-load applications in production will be a big plus.
  • Practical experience in supporting business applications (like NodeJS, PHP, GoLang, Python).
  • Strong understanding of networking concepts, protocols, and microservices architecture.
  • Proficiency in Git & GitHub or other version control.
  • Experience with issue processing (RCA, Postmortems).
  • Familiarity with incident response and management tools like BetterStack, PagerDuty, or others.

 

Will be a plus:

  • Familiarity with GCP/Scaleway, Terraform, Docker, Linux, CI/CD, PostgreSQL/MySQL.
  • Proficiency in at least one scripting language (e.g., Bash, Python).
  • Solid spoken and written English skills (ideally Upper-Intermediate level or higher).
  • Familiarity with Cloudflare (Workers, WAF, managing DNS, and Cloudflare for SaaS).
  • Google Cloud Certifications (or at least desire to obtain one).


 

Required skills experience

Kubernetes 2 years

Required languages

English B2 - Upper Intermediate
Published 21 September
35 views
ยท
1 application
To apply for this and other jobs on Djinni login or signup.
Loading...