Senior Data Engineer
## About Us
Smirnov Labs is a fast-growing engineering company founded by an ex-Googler. We help product teams launch and scale quickly by leveraging modern technology and AI-assisted development.
Our projects are diverse โ we partner with clients at the earliest stages and own the technical delivery end-to-end.
We're a small, senior-heavy team where every engineer has real impact. No bureaucracy, no hand-holding โ just sharp people shipping fast.
This is an outstaff role: you'll be embedded directly with a client's engineering team as part of our delivery group.
## The Role
The client runs a custom AI system, and it's only as good as the data reaching it. Your job is to get that data in: connect to whatever third-party API holds it, pull it reliably, shape it, and keep it flowing.
That means a lot of integration work โ every source has its own auth, its own pagination, its own rate limits, and its own idea of what a schema is. You'll own those connectors and the pipelines behind them end-to-end.
## What You'll Do
- Build and own integrations with third-party APIs โ REST, GraphQL, webhooks, the occasional CSV drop or legacy SOAP endpoint. Auth flows, pagination, rate limits, retries, backfills.
- Design and run the pipelines that feed the client's AI system: ingestion, normalisation, enrichment, and delivery into the stores it reads from.
- Build the ingestion path for AI workloads โ chunking, embeddings, and keeping vector indexes fresh as source data changes, without full re-indexing every time.
- Model the data. Design schemas that hold up as sources are added and change underneath you.
- Make pipelines idempotent, observable, and recoverable. Things will break upstream; the system should degrade predictably and tell you why.
- Own data quality: validation, schema-drift detection, reconciliation, alerting on the failures that matter rather than all of them.
- Ship and run your own work โ Docker, CI/CD, cloud environments. You deploy what you build.
- Work directly with the client's engineering and AI teams on what the models actually need from the data โ no layers in between.
- Use AI-assisted development tools (Claude Code, Cursor, GitHub Copilot) as part of your daily workflow.
## What We're Looking For
- 5+ years building and running production data pipelines
- Strong Python and strong SQL โ both at a level you'd defend in review
- Real depth in third-party API integration: OAuth and token refresh, pagination strategies, rate limiting, retries and backoff, incremental syncs, and handling APIs whose documentation is wrong
- Experience with a workflow orchestrator (Airflow, Dagster, Prefect, Temporal, or similar) and an opinion about it
- Solid Postgres: schema design, query performance, migrations. Warehouse experience (BigQuery, Snowflake, Redshift) a plus
- Batch and streaming ingestion patterns, and a sense of when each is the right call
- Comfortable owning your own deployments โ Docker, CI/CD, cloud (AWS or GCP)
- A track record of shipping without a detailed spec and without process handed to you
- Comfort working in the US timezone (overlapping working hours required)
- Upper-Intermediate or higher English
- Ability to own and drive features independently with minimal supervision
## Nice to have
- Hands-on experience feeding LLM or RAG systems โ chunking strategies, embedding pipelines, vector stores (pgvector, Qdrant, Pinecone, whatever you've used). Side projects count; commercial experience is not required here.
- dbt, or another transformation layer you've run in anger
- Data quality and observability tooling (Great Expectations, Monte Carlo, or your own)
- Experience with managed connectors (Fivetran, Airbyte) โ including knowing when to stop fighting them and write your own
- Early-stage or agency background: greenfield work, shifting requirements, direct client contact
## What We Offer
- Competitive salary above market average
- Fully remote work
- Flat structure โ work directly with the client's engineering team, no middle management
- Real technical ownership: you make the architecture decisions on the pipelines you build
- Diverse and technically challenging projects
- Modern AI-powered development workflow
- Small team culture โ your voice matters, your code ships
- Paid vacation and sick leave
- Flexible schedule within the US timezone overlap
- Professional and career growth through real ownership, not courses and certificates
## How to Apply
Send your CV with a brief note about the nastiest API you've had to integrate โ what made it hard, and how you made it reliable.