GPU Infrastructure Engineer
We’re looking for an experienced Infrastructure Engineer to work on a project of our Customers. They are building a multi-node GPU platform for large-scale model training and inference using managed GPU providers. In this role, you’ll have to own the platform performance, reliability, and the technical relationship with the Client’s providers.
Our Customers provide SaaS solutions that help companies to optimize their business. These solutions include business planning to automate and optimize business, delivery, and workflow solutions. The platform leverages industry-leading Artificial Intelligence (AI) and Machine Learning (ML) for better prediction and prevention of disruptions across business.
Candidate’s location – Europe.
Responsibilities:
- Design and operate distributed GPU training infrastructure
- Validate cluster topology, RDMA/InfiniBand, and NCCL performance
- Standardize Kubernetes or Slurm scheduling, GPU images, and software versions
- Diagnose issues across training workloads, networking, storage, and GPU hosts
- Build monitoring, benchmarks, runbooks, and reliability standards
- Work with GPU providers to resolve incidents and define technical requirements
Support SFT, DPO, RL, and large-scale inference workloads
Requirements:
- Location in Europe
- Production experience with multi-node GPU training infrastructure
- Strong Linux, containers, CUDA, and NVIDIA-GPU-stack knowledge
- Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
- Deep experience with Kubernetes or Slurm
- Infrastructure automation and observability experience
- Strong skills in incident leadership and provider-facing communication
- English level – Upper-Intermediate or higher
Will be a plus:
- Experience in an AI lab, HPC environment, or specialist GPU cloud
- Knowledge of distributed-training frameworks such as PyTorch, Megatron, or DeepSpeed
- Experience with parallel storage, checkpoint optimization, and multi-provider platforms