Data Engineer Scraping, ETL, LLM Pipelines

$$$$

We're building a real-time startup & investor data platform โ€” the data is the product, served to customers via API, a BigQuery feed, and MCP. 

The core challenge
Many of our pipeline processes are currently coupled and run sequentially not as parallelizable as we need. As we scale into millions of records and add tons of new data points, we want an architecture that lets us add sources and processes independently, run them in parallel, and scale cleanly. That's the heart of this role.

What you'll do

  • Re-architect coupled, sequential pipelines into modular, parallelizable ones that scale to millions of records.
  • Design scalable pipeline architecture on GCP-based services.
  • Add new datasets end to end: structured extraction, LLM-based parsing, clean storage in Supabase / BigQuery with well-defined schemas.
  • Build eval frameworks for our LLM calls so we can route the right model to the right use case mixing OpenAI with open-source models (e.g. Gemma) deliberately, not by default.
  • Integrate scraping providers (Scrapfly, Bright Data) and handle the occasional hard source (Cloudflare, captchas) as we approach new datasets.

What we're looking for

  • 4+ years building data pipelines, backend services, and automated data processing at scale.
  • Strong track record designing scalable, parallelizable pipeline architecture โ€” decoupling dependent processes, orchestration, throughput.
  • Deep ETL and web-scraping experience (e.g. Scrapy, Playwright for dynamic sources); comfort working through providers like Scrapfly / Bright Data.
  • Hands-on with GCP (BigQuery a plus), Airflow, PostgreSQL, FastAPI, and Docker.
  • Practical LLM-pipeline experience building evals and choosing models for cost/quality per use case.
  • Fluent English.

We care more about depth in scalable pipeline architecture, ETL, and scraping than about an exact stack match if you've built these systems well with adjacent tools, we want to talk.

Details

  • Engagement: [contract / full-time]
  • Location: [remote]

Required skills experience

Data Processing 2 years
Web Scraping / Scraping 2 years
Pipeline/CRM hygiene 1.5 years
Scrapy 1.5 years
Playwright 1.5 years
RabbitMQ 1.5 years
PostgreSQL 2 years
Supabase 1.5 years
BigQuery 1.5 years
Airflow 2 years
FastAPI 1.5 years
Docker 1.5 years
AWS 1.5 years
GCP BigQuery 1.5 years

Required languages

English C1 - Advanced
Ukrainian Native
Published 7 July
53 views
ยท
7 applications
See stats of candidates who applied for this job ๐Ÿ‘€
To apply for this and other jobs on Djinni login or signup.
Loading...