Senior DevOps Engineer โ Relocation to Dubai
We are a product company building AI-powered mobile applications used by over 10 million people in 150 countries. We own the entire product lifecycle โ from the initial idea to large-scale global implementation. We thrive in a fast-paced environment, making data-driven decisions without unnecessary bureaucracy. If you value autonomy, rapid releases, and seeing the direct impact of your work, you will fit right in.
The Role
We are looking for a Senior DevOps Engineer to take full ownership and leadership of the infrastructure for our ambitious project in the Cloud Gaming & Streaming space.
In this role, you will architect, scale, and maintain a high-performance, fault-tolerant system provisioning and streaming content across a server fleet growing from hundreds toward thousands of machines. You will lead infrastructure decisions, set technical standards, build key systems (like observability and bare-metal orchestration) from scratch, and mentor a Mid DevOps Engineer on the team.
Work Format & Relocation
- Location: Dubai, UAE (Office).
- Trial Period: First 2โ3 months are fully remote for a smooth onboarding.
- Relocation Support: Full corporate visa sponsorship and relocation package provided after the trial period.
Key Responsibilities
- Fleet Architecture & Provisioning: Architect and automate end-to-end bare-metal and virtualized host provisioning โ storage, GPU passthrough (vfio), and VM lifecycle across a large, heterogeneous server fleet.
- Orchestration & Scale: Design and maintain fleet-scale orchestration systems: parallel deployments, health checks, and self-healing routines ensuring consistency across thousands of machines.
- Infrastructure as Code & Configuration: Lead IaC practices using Terraform and Ansible to maintain reproducible host configurations, deterministic automation, and VM images.
- Advanced Virtualization & OS Administration: Deeply manage Linux/KVM virtualization (libvirt, QEMU) and Windows guest systems, including complex GPU passthrough and low-latency virtual networking.
- Greenfield Observability: Architect and implement the observability stack from scratch (Prometheus, Grafana, Loki/ELK) across the entire fleet for deep performance monitoring and alerting.
- Network & Traffic Management: Design and optimize network configurations at scale โ NAT, port forwarding, routing, and traffic paths for thousands of concurrent, low-latency streaming clients.
- CI/CD & Deterministic Automation: Build robust CI/CD pipelines (GitLab CI / GitHub Actions) for infrastructure code, images, and deterministic scripting (Bash, Python, PowerShell).
- Security & Storage: Implement enterprise-grade secrets management (Vault / SOPS) and manage large-scale content distribution via S3-compatible object storage (e.g., Cloudflare R2).
- Leadership & Incident Response: Set reliability engineering standards, drive incident response/on-call processes, and collaborate closely with core C++/C# infrastructure engineers.
Requirements
- Experience: 4+ years in a DevOps / Infrastructure role with proven experience building and managing large-scale, high-load systems or large server fleets (dozens/hundreds+ of nodes).
- OS & Virtualization Expertise: Deep administration skills in both Linux and Windows; practical expertise with KVM/libvirt, QEMU, and GPU passthrough / vfio.
- IaC & Configuration Management: Strong hands-on experience with Terraform and Ansible for host consistency, image management, and fleet-wide automation.
- Scripting Mastery: High proficiency in Python, Bash, or PowerShell with the ability to write deterministic, idempotent, and restart-safe automation scripts across heterogeneous hardware.
- Networking & Observability: Strong grasp of Linux networking (NAT, iptables, routing, traffic shaping) and experience standing up monitoring/logging stacks (Prometheus, Grafana, Loki/ELK) from scratch.
- CI/CD & Security: Experience automating infrastructure via GitLab CI/GitHub Actions and managing secrets at scale (Vault, SOPS).
- Language: English (Upper-Intermediate / B2 or higher โ required for technical collaboration).
Nice-to-have
- Direct experience in Cloud Gaming, real-time data streaming, or high-load gaming environments.
- Hands-on experience with GPU-heavy workloads and large-file distribution via object storage (S3 / Cloudflare R2).
- Basic knowledge of C++ or C# to collaborate effectively with core streaming infrastructure developers.
What We Offer
- Architectural Greenfield & Autonomy: Full ownership to design the fleet observability, orchestration, and automation architecture from zero.
- Leadership & Real Impact: Direct influence on the tech stack of a Cloud Gaming platform targeting a global, multi-million user audience.
- Relocation Package: Full financial and organizational support for your move to Dubai (corporate visa, flight tickets, initial housing) after the trial period.
- Modern Office: Work from a vibrant, tech-driven workspace in Dubai alongside a strong engineering team.
- Fast-Paced Culture: A product-driven environment with quick decision-making, high engineering freedom, and zero bureaucracy.
- Compensation: Highly competitive USD-pegged salary matching your senior technical background.