Senior MLOps Engineer
Who we are:
Adaptiq is a technology hub specializing in building, scaling, and supporting R&D teams for high-end, fast-growing product companies in a wide range of industries.
About the Product:
Our client is a leading online gaming operator developing a new-generation transactional platform for user account management. This greenfield solution is being built for a highly regulated, high-traffic environment and is expected to support millions of concurrent users.
The platform will handle real-time financial transactions and requires strong availability, sub-second performance, data integrity, security, and complete auditability. A major engineering challenge is ensuring the system remains reliable and performant during significant traffic surges and peak events.
The team is looking for a hands-on MLOps Engineer who will contribute to designing and building scalable, reliable ML infrastructure, creating and maintaining production-grade data and ML pipelines, and supporting the complete ML lifecycle. Particular attention will be given to data quality, monitoring, governance, and overall system reliability.
About the Role:
We are looking for a Senior MLOps Engineer to build and evolve scalable cloud infrastructure, real-time data pipelines, and production ML workflows.
You will work with Infrastructure as Code, event-driven systems, and end-to-end ML processes, covering model deployment, monitoring, data governance, and reliability. Working closely with Data Engineers and Data Scientists, you will help shape the new platform, ensure reliable data and ML workflows from end to end, and understand the impact of upstream data changes on downstream models.
This position offers significant technical ownership and the opportunity to contribute to architectural decisions, technology selection, and the long-term evolution of the platform.
Key Responsibilities:
- Design, provision, and maintain scalable cloud infrastructure using Infrastructure as Code.
- Develop and operate robust real-time, event-driven data pipelines with fault tolerance and sub-second latency.
- Design and optimize storage systems and data lake infrastructure supporting both transactional and analytical workloads.
- Build high-availability networking solutions, API gateways, automated failover mechanisms, and zero-downtime routing.
- Embed data validation, observability, and monitoring into pipelines to maintain reliability and quickly detect potential issues.
- Build and maintain infrastructure supporting ML model training, validation, deployment, and inference.
- Automate deployment and testing processes for containerized workloads running on Kubernetes.
- Partner with cross-functional teams to establish ML governance practices and dependable end-to-end data flows.
- Contribute to long-term cloud scalability, FinOps and cost optimization initiatives, as well as disaster recovery planning.
- Provide technical guidance, contribute to architecture reviews, and mentor team members on engineering best practices.
Required Competence and Skills:
- 5+ years of experience in cloud infrastructure, data engineering, or distributed systems architecture.
- Strong hands-on experience with infrastructure provisioning tools such as Terraform, container technologies including Docker and Kubernetes, and cloud environments such as AWS or GCP.
- Practical experience with distributed streaming technologies such as Apache Kafka and Change Data Capture (CDC) approaches.
- Strong knowledge of object storage architectures such as S3, distributed data lakes, and SQL/NoSQL databases.
- Strong Python skills and scripting experience, including Bash, for automation and pipeline orchestration.
- Good understanding of ML workflows, data governance, and model lifecycle management.
- Ability to turn high-level architectural concepts into practical implementations and optimize them for performance.
- Strong communication skills and experience working closely with Data Scientists and Engineering teams.
Nice to Have:
- Previous experience in fintech, iGaming, or other high-throughput, real-time transactional environments.
Why Us:
We provide 20 days of vacation leave per calendar year (plus official national holidays of a country you are based in).
We provide full accounting and legal support in all countries we operate.
We utilize a fully remote work model with a powerful workstation and co-working space in case you need it.