Senior Infrastructure Engineer (GPU Platform)
On behalf of our client, we are looking for a Senior Infrastructure Engineer (GPU Platform)
Requirements:
- Production experience with multi-node GPU training infrastructure
- Strong Linux, containers, CUDA, and NVIDIA GPU stack knowledge
- Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
- Deep experience with Kubernetes or Slurm
- Experience with infrastructure automation and observability
- Experience diagnosing issues across training workloads, networking, storage, and GPU hosts
- Strong incident leadership and provider-facing communication skills
- English โ Upper-Intermediate or higher
Would be a plus:
- Experience in an AI lab, HPC environment, or specialist GPU cloud
- PyTorch, Megatron, DeepSpeed, or other distributed-training frameworks
- Experience with parallel storage and checkpoint optimization
- Experience working with multi-provider GPU platforms
Company offers:
- Long-term employment with possibilities for professional growth
- Fully remote work
- Reasonably flexible schedule
- 15 days of paid vacation
- Regular performance reviews