AI Platform Engineer (Strong Senior/Lead DevOps/MLOps)
We are seeking a hands-on AI Platform Engineer to join our team. The primary mission of this role is to operate, automate, manage, and evolve
our enterprise AI Platform so that internal development teams can rely on a stable, secure, scalable, and well-governed environment for building AI solutions.
This is a platform engineering and operations role, not an application development position. The ideal candidate is proactive, autonomous, understands cloud platform engineering principles, and can execute technical initiatives with minimal supervision.
The role spans infrastructure, CI/CD automation, API management, platform governance, and AI services, with a strong emphasis on Infrastructure as Code and operational excellence.
Requirements:
- Hands-on experience operating enterprise platforms in Microsoft Azure environments
- Strong experience with Terraform and Infrastructure as Code practices
- Experience designing and maintaining CI/CD pipelines using Azure DevOps
- Practical knowledge of Azure API Management (APIM)
- Experience with Azure AI Foundry, Azure OpenAI, or related Azure AI services
- Understanding of cloud networking, security, identity management, and access control
- Strong troubleshooting and problem-solving skills.
- Proven ability to work autonomously and manage tasks end-to-end
Nice-to-Have Skills:
- Experience with Portkey AI for LLM routing, observability, guardrails, caching, or governance
- Knowledge of multi-model AI platforms and model gateway architectures
- Familiarity with AWS Bedrock or other cloud AI platforms
- Experience implementing platform observability solutions and operational dashboards
Knowledge of enterprise AI governance, compliance, and security controls
Candidate Profile:
- Operates independently while keeping stakeholders informed and aligned
- Has a platform-first mindset and understands how developers consume shared services
- Is comfortable working across infrastructure, DevOps, API management, and AI services
- Takes ownership of operational excellence, automation, and continuous improvement
- Is proactive, reliable, and capable of driving technical tasks from design through implementation
- Learns quickly and adapts to an evolving cloud and AI technology landscape
Values documentation, standardization, and long-term maintainability of platform solutions
Job responsibilities:
AI Platform Operations
- Troubleshoot platform-level issues and coordinate resolution across multiple Azure services
- Support platform scalability, reliability, availability, and operational excellence
- Manage lifecycle operations for LLMs, embedding models, AI services, and platform environments across development, non-production, and production
- Administer and maintain Azure AI Foundry workspaces, projects, model deployments, connections, and platform configuration
Infrastructure as Code & Automation
- Develop, maintain, and enhance Infrastructure as Code using Terraform
- Build and manage reusable infrastructure modules, templates, and deployment patterns
- Automate environment provisioning, configuration management, and platform operations
- Ensure consistency and repeatability across environments through automation
- CI/CD & DevOps
Design, implement, and maintain CI/CD pipelines using Azure DevOps - Manage build, release, deployment, validation, and rollback processes
Improve deployment reliability through testing, automation, and release governance
Promote DevOps best practices across the platform ecosystem
API Platform & Integration Management
- Operate and maintain Azure API Management (APIM) as a core component of the AI Platform
- Configure APIs, policies, authentication, authorization, throttling, and routing mechanisms
- Support secure exposure of AI capabilities to internal development teams
- Assist in troubleshooting API connectivity, authentication, and integration challenges
Observability & Monitoring
- Configure and maintain platform-wide logging, monitoring, tracing, and telemetry
- Build dashboards and reporting for platform usage, performance, adoption, cost, and operational health
- Monitor platform capacity, quotas, consumption, and service reliability
- Implement alerting and diagnostics to ensure stable platform operations