AI / ML Infrastructure Engineer
$$$
Location: New York, NY or San Francisco, CA (Hybrid / On-site)
Employment Type: W-2 Contract or Full-Time (Must be authorized to work on W-2 in the US without sponsorship)
Experience: 3+ Years
Job Summary
We are seeking an AI / ML Infrastructure Engineer to join our client's high-performance engineering team. You will design, scale, and optimize the underlying infrastructure powering large-scale generative AI and machine learning models in production.
Key Responsibilities
- Architect and scale high-performance GPU clusters (NVIDIA hardware) for model training and low-latency inference.
- Optimize model serving pipelines using Triton Inference Server, vLLM, and TensorRT.
- Collaborate with data scientists and MLOps engineers to streamline distributed training workflows.
- Monitor, troubleshoot, and optimize cloud infrastructure costs and performance bottlenecks.
Requirements
- Minimum 3 years of hands-on experience in ML infrastructure, systems engineering, or DevOps roles.
- Strong proficiency in PyTorch, CUDA ecosystem, and container orchestration (Kubernetes, Docker).
- Deep experience with cloud providers (AWS, GCP) and Infrastructure as Code (Terraform).
- Valid W-2 work authorization in the United States (no C2C or cross-border arrangements).
Sound like a fit?
Apply now and our recruiting team will get back to you.
Required skills
PyTorch, DevOps, AWS
Required languages
English
B2 - Upper Intermediate
Published 7 October
18 views
ยท
0 applications
๐
Average salary range of similar jobs in
analytics โ
Loading...