Senior Data Ingestion Engineer (AWS / Document Extraction / AI-RAG Prep)
ZentixSoft is looking for two Senior Data Ingestion Engineers! ๐
Format: Direct Contract Engagement (1-Year Contract, 100% Full-Time Remote).
๐ก Why ZentixSoft Partner Network?
๐ธ Transparent Compensation: Direct payroll arrangement with $40 USD/hour base compensation.
โ๏ธ Work-Life Balance: Sustainable engineering workflows with predictable long-term deliverables.
๐ฆพ Trust & Transparency: Zero micromanagement and full autonomy over your technical pipeline architecture.
๐ Culture & Growth: Long-term project stability within large-scale financial and insurance data domains.
๐งฉ Responsibilities:
- Pipeline Design & Engineering: Design and deploy scalable data ingestion pipelines for high-volume structured/unstructured documents (PDFs, scans, emails, Word, Excel, PowerPoint).
- Document Extraction & OCR: Implement OCR and document processing workflows using AWS Textract (or equivalent) for text extraction, cleaning, normalization, and metadata tagging.
- RAG & Vector Storage Prep: Execute semantic chunking, metadata extraction, vector storage schema design, and retrieval mechanism preparation for downstream AI models.
- Integrations & Connectors: Build robust connectors with enterprise sources, including SharePoint, email servers, and public cloud repositories.
- Validation & Observability: Implement automated error monitoring, OCR extraction validation, and CI/CD automated testing using AWS Step Functions and CloudWatch.
๐ Our Perfect Match (Requirements):
- Senior Data Engineering: 7+ years of commercial Data Engineering experience, with strong proficiency in Python and SQL.
- AWS Expertise: 5+ years of hands-on experience in AWS environments, specifically AWS S3, Step Functions, CloudWatch, and public cloud data processing.
- Unstructured Data & OCR: Proven background in building document extraction pipelines (handling PDFs, scanned images, emails, Office documents) and utilizing OCR technology (AWS Textract or similar).
- Source Connectors & Pipeline Operations: Practical experience integrating sources like SharePoint and email, performing text normalization, metadata tagging, and maintaining CI/CD/Git testing standards.
- EU Residency & Location: Candidates MUST reside in an EU member country (with active legal residency; citizenship can be non-EU/global).
- Language & Communication: C1 Advanced English (verbal and written) for direct technical collaboration.
โ Good to Have:
- Experience with Vector Databases, RAG architectures, semantic chunking, and vector retrieval mechanisms.
- Background in Insurance or Financial Services industries handling enterprise security standards.
- Exposure to Azure or Databricks platform components.
๐ Project Details:
โณ Duration: 1 Year (Full-time contract).
๐ Target Start: End of October.
๐ Location: 100% Remote (Must be physically located in an EU member state with valid residency).