Our client builds an enterprise LLM gateway. It connects customers’ applications and AI agents to multiple model providers through an OpenAI-compatible API. Customers register their own provider API keys. The gateway routes requests, reduces model costs, enforces policies, and tracks cost and usage per agent.
We are looking for a Senior Full-Stack Engineer with strong Python and TypeScript skills. You will build production SaaS features across the backend and frontend. Your work will cover database design, FastAPI endpoints, React screens, tests, and deployment. Features can also include Python middleware in the gateway.
Two areas will be your primary focus:
- Payments and billing — building the commercial layer on top of the platform’s usage metering:subscription plans, usage-based billing, invoicing, credits, and entitlement enforcement.
- Gateway guardrails and telemetry — building the policy enforcement and observability layer that makes the gateway safe, measurable, and auditable for enterprise customers.
You will join a small, senior, distributed team. You will own features through deployment and take part in architecture decisions. When needed, you will work directly with customers to define, build, deploy, and improve features that use AI.
What’s interesting about this project:
- Own features from database design to the customer interface and production deployment.
- Work with Python, FastAPI, React, TypeScript, Kubernetes, and Terraform.
- Help decide how the product works and how the team builds it.
- Build a gateway that connects customer applications and AI agents to multiple model providers.
- Measure how gateway changes affect model costs, response quality, and latency.
What you’ll work on
- Payment system integration (primary focus): Design and build the billing layer end to end — integrate Stripe (or an equivalent provider) for subscription plans, metered/usage-based charging, invoicing, credits and prepaid balances, and payment lifecycle handling.
- LLM gateway guardrails (primary focus): Design and implement the policy layer that inspects traffic flowing through the gateway.
- Telemetry and observability (primary focus): Extend the request telemetry pipeline that powers customer-facing analytics — enriching per-request records with cost, token, cache, routing, and guardrail metadata, building efficient time-series aggregations.
- Gateway middleware: Build Python hooks inside the LLM proxy. Work includes prompt compaction, model routing based on request complexity, semantic caching, and response enrichment. Keep request handling correct and latency low.
- Backend API development: Build and evolve FastAPI services with async SQLAlchemy and PostgreSQL, including multi-tenant data isolation, RBAC, session-based authentication, encrypted secret storage, and Alembic migrations.
- Frontend development: Build customer-facing features in React and TypeScript — billing and plan management, guardrail configuration, telemetry and spend analytics views, admin/provisioning screens, and the LLM playground with streamed responses.
- Multi-provider model management: Extend provider integrations, model catalogues, budgets, and rate limits so tenants can bring their own models and providers safely.
Technology Stack
- Backend: Python 3.13, FastAPI, Uvicorn, Pydantic, Poetry
- Data: PostgreSQL 16, SQLAlchemy 2 (async, asyncpg), Redis Stack
- LLM Gateway: OpenAI-compatible proxy (LiteLLM) with custom Python middleware, multi-provider model routing
- Frontend: React, TypeScript, Vite, Tailwind CSS, Chart.js
- Payments: Stripe or equivalent — subscriptions, metered billing, webhooks, invoicing
- Testing: pytest, Vitest, Testing Library
- Infrastructure: Google Cloud Platform, Kubernetes (GKE), Helm, Terraform, Docker
- CI/CD: GitHub Actions
- Observability: Prometheus, Grafana, distributed tracing
- Code quality: black, isort, flake8, ESLint, Prettier
Responsibilities
- Own features end to end across the gateway middleware, backend API, and frontend — including schema design, API contracts, UI, tests, and rollout.
- Build and operate the subscription and usage-based billing system, including payment provider integration, webhook reliability, usage-to-invoice mapping, and entitlement enforcement.
- Design and implement guardrails that inspect and enforce policy on LLM prompts and completions, with per-tenant configuration and full auditability.
- Extend the telemetry pipeline and analytics aggregations that power customer-facing cost and usage reporting, and add tracing, metrics, and alerting across services.
- Write middleware that runs in the live request path, treating latency, failure modes, and graceful degradation as first-class design concerns.
- Design multi-tenant data models with correct isolation between organizations, workspaces, and keys.
- Write meaningful automated tests for the code you ship.
- Participate in architecture decisions, code reviews, and planning; document decisions for a distributed team.
- Handle sensitive data responsibly — secret encryption at rest, least-privilege access, and audit logging.
Requirements
- 5+ years of professional full-stack experience. You have built and shipped production SaaS features across the backend and frontend.
- Strong Python skills and experience building production APIs with FastAPI (or Django REST / Flask, with a willingness to work in FastAPI), including async programming
- Strong React and TypeScript skills — comfortable with component-driven UI, SPA state management, routing, and integrating REST APIs including streamed responses
- Practical familiarity with LLM integrations, retrieval-augmented generation (RAG), AI agents, or AI tools such as Cursor, Claude Code, or Copilot. You do not need an AI research or specialist agent infrastructure background.
- Experience with Stripe or a similar payment provider. Relevant work includes subscriptions, webhooks, idempotency, proration, refunds, and usage-based billing in production.
- Solid PostgreSQL skills: schema design, migrations, indexing, and query optimization — including aggregations over large event/log tables
- Experience with multi-tenant architectures, authentication and authorization, and tenant data isolation
- You can use Docker and run services in Kubernetes. You can read a Helm chart and debug a failing pod.
- Experience with Redis or a comparable cache, and with designing for cache correctness and invalidation
- You work well in a startup with limited structure. You clarify unclear problems, choose practical solutions, and deliver them with little oversight.
- Clear written and spoken English for collaboration and documentation
Nice to Have
- Experience working directly with customers to define, build, and improve solutions.
- Experience with LLM API streaming, token accounting, rate limits, retries, and cost tracking.
- Experience building guardrails, content moderation systems — or working on trust & safety, compliance, or security tooling
- Familiarity with LLM gateways or proxies or with building an internal LLM platform
- Experience with LLM observability and evaluation tooling and with building evaluation harnesses
- Experience with usage metering and rating systems — turning high-volume event streams into accurate, auditable billable quantities
- FinOps or cloud-cost-optimization background; experience showing customers where their money goes
- Terraform and infrastructure-as-code experience on GCP (GKE, Cloud SQL, Secret Manager) or an equivalent cloud
Our benefits:
- No micromanagement
- Freedom to engage in decision-making and implementation
- Ability to work in a team of professionals (the ratio of middle and above specialists 80/20)
- Participation in the development of high-quality products
- Direct communication with clients on a partnership level
- Health insurance
- A $1,000 flexible benefits budget per individual year
- 20 paid working days off and 10 days sick leave
- Opportunity to work remotely
Join us and be among those who care!