ZentixSoft

Senior Data Ingestion Engineer (AWS / Document Extraction / AI-RAG Prep)

$$$$

 

ZentixSoft is looking for two Senior Data Ingestion Engineers! ๐Ÿš€

Format: Direct Contract Engagement (1-Year Contract, 100% Full-Time Remote). 
 

๐Ÿ’ก Why ZentixSoft Partner Network?
 ๐Ÿ’ธ Transparent Compensation: Direct payroll arrangement with $40 USD/hour base compensation. 
โš–๏ธ Work-Life Balance: Sustainable engineering workflows with predictable long-term deliverables. 
๐Ÿฆพ Trust & Transparency: Zero micromanagement and full autonomy over your technical pipeline architecture. 
๐ŸŽ Culture & Growth: Long-term project stability within large-scale financial and insurance data domains.
 

๐Ÿงฉ Responsibilities:

  • Pipeline Design & Engineering: Design and deploy scalable data ingestion pipelines for high-volume structured/unstructured documents (PDFs, scans, emails, Word, Excel, PowerPoint).
  • Document Extraction & OCR: Implement OCR and document processing workflows using AWS Textract (or equivalent) for text extraction, cleaning, normalization, and metadata tagging.
  • RAG & Vector Storage Prep: Execute semantic chunking, metadata extraction, vector storage schema design, and retrieval mechanism preparation for downstream AI models.
  • Integrations & Connectors: Build robust connectors with enterprise sources, including SharePoint, email servers, and public cloud repositories.
  • Validation & Observability: Implement automated error monitoring, OCR extraction validation, and CI/CD automated testing using AWS Step Functions and CloudWatch.
     

๐ŸŽ“ Our Perfect Match (Requirements):

  • Senior Data Engineering: 7+ years of commercial Data Engineering experience, with strong proficiency in Python and SQL.
  • AWS Expertise: 5+ years of hands-on experience in AWS environments, specifically AWS S3, Step Functions, CloudWatch, and public cloud data processing.
  • Unstructured Data & OCR: Proven background in building document extraction pipelines (handling PDFs, scanned images, emails, Office documents) and utilizing OCR technology (AWS Textract or similar).
  • Source Connectors & Pipeline Operations: Practical experience integrating sources like SharePoint and email, performing text normalization, metadata tagging, and maintaining CI/CD/Git testing standards.
  • EU Residency & Location: Candidates MUST reside in an EU member country (with active legal residency; citizenship can be non-EU/global).
  • Language & Communication: C1 Advanced English (verbal and written) for direct technical collaboration.
     

โž• Good to Have:

  • Experience with Vector Databases, RAG architectures, semantic chunking, and vector retrieval mechanisms.
  • Background in Insurance or Financial Services industries handling enterprise security standards.
  • Exposure to Azure or Databricks platform components.
     

๐Ÿ“ Project Details:

โณ Duration: 1 Year (Full-time contract).

๐Ÿš€ Target Start: End of October.

๐Ÿ“ Location: 100% Remote (Must be physically located in an EU member state with valid residency).

Required languages

English B2 - Upper Intermediate
Published 2 October
12 views
ยท
0 applications
To apply for this and other jobs on Djinni login or signup.
Loading...