Platform / SRE Architect - multi-tenant SaaS, stateful workloads on Kubernetes
START HERE
Most of what we need is ordinary platform work. One part is not.
We run a multi-tenant SaaS platform on Kubernetes with a MongoDB data tier, and we are designing it for its enterprise stage. The hardest problem in the engagement is reshaping stateful workloads in production without losing data, and designing tenant isolation so that one customer's load cannot degrade another's.
If reshaping a StatefulSet's storage in a live cluster is a problem you have actually solved rather than read about, keep reading. If it is not, the rest of this list will look familiar and you will probably be bored.
WHAT YOU WILL OWN
The data tier. Design and implement high availability for MongoDB: replication topology, migration paths that are reversible at every step, and backup and restore that is verified rather than assumed.
Tenant isolation. Design multi-tenant resource isolation so workloads stay predictable as customer count and data volume grow. Noisy-neighbor behavior is the failure mode we care about most.
Release integrity. Mature the deployment pipeline so releases are deterministic and auditable, and so what is running matches what was declared.
Observability that means something. Define the service-level indicators that reflect real customer experience, and build the measurement behind them. We are not looking for more dashboards.
Async and background processing. Architect the model for long-running, data-heavy operations.
Capacity. Right-size resources against measured usage, introduce autoscaling, and load test at realistic data volumes.
Cloud strategy. Give us your long-term view. We are on Azure and are not planning to migrate, but we want an architect's assessment to plan against.
WHAT WE NEED TO SEE
You have done stateful surgery in production. Reshaping storage, promoting a standalone database to a replicated one, or migrating a data tier with a rollback path at each step. This is the one thing we cannot backstop with review.
You have solved tenant isolation for real. Not "we used namespaces." You have hit a noisy-neighbor problem, diagnosed it, and have opinions about which mitigations actually held.
You debug by measurement. You instrument and observe rather than reason from intuition, and you have been wrong about a bottleneck before and can say so.
You can read Node.js well enough to reason about where memory and time go. You do not need to write it.
You will disagree with a proposed design and explain why. You will be the most experienced infrastructure person here, and deference would waste the engagement.
5+ years in platform engineering, SRE or cloud architecture, with real design ownership on at least one production system.
English sufficient for written technical discussion and calls.
THE STACK, FOR MATCHING
Kubernetes and Helm in production, MongoDB, Azure and AKS preferred (equivalent AWS or GCP depth is fine), Terraform, Prometheus, Grafana, Alertmanager, Loki, Node.js services.
Nice to have: GitOps (Argo CD or Flux), k6 or JMeter, exposure to blockchain infrastructure or carbon markets. The domain is genuinely optional, we will teach it.
HOW TO APPLY
Send your CV and answer one question. It matters more to us than the CV does.
Describe a time you changed the shape of a stateful workload in production. A storage migration, a topology change, a database promotion, anything where the data had to survive. What was your rollback at each step, and what did you check before you started?
Answers that describe the happy path tell us less than answers that describe what almost went wrong. We are not looking for a keyword list.
ABOUT CLIMISSION
Climission is a climate technology company on a mission to empower stakeholders across carbon markets and sustainable value chains with state-of-the-art tools and platforms, enabling a seamless transition to a more sustainable and transparent global ecosystem. Founded on the principles of innovation, transparency, and sustainability, Climission builds on the groundbreaking work of Envision Blockchain and the open-source Guardian platform. The company develops and operates a suite of products for the carbon economy, including a Managed Guardian Service (enterprise SaaS for carbon market operations), an AI Toolkit for climate data, and an Indexer for digital environmental assets, alongside its contributions to open-source infrastructure. Climission works with leading registries, enterprises, and international bodies to make the verification and tokenization of environmental assets trustworthy, scalable, and intelligent.
Company website: https://climission.com/