Middle QA Automation Engineer ( Python)
Description
We are building an AI powered document search assistant for a large enterprise client in Japan. Project design documentation accumulated over roughly twenty years currently sits on an isolated file server with no search capability at all, which forces engineers to rely on subject matter experts to locate specifications manually. The solution migrates this content into SharePoint within the client internal network and delivers natural language search across it. Roughly half of the corpus is in legacy Excel format where critical information lives inside shapes, callouts and embedded screenshots rather than in cells, so the pipeline combines format conversion, drawing layer parsing and OCR before indexing. The architecture is built on Azure (Python-based processing, Azure AI Search, Azure AI Document Intelligence) with Google Cloud considered as an alternative track. The current stage is a Proof of Concept focused on two highest-risk areas, unstructured data extraction and incremental index updates, with production rollout to follow.
Requirements
- Possessing 4+ years of dedicated professional experience in quality assurance and automation engineering, demonstrating a strong track record of successful project contributions.
- Demonstrated advanced proficiency in Quality Assurance methodologies and practices, ensuring high standards of software reliability and performance.
- Possessing intermediate hands-on experience in QA Automation, including the design, development, and maintenance of automated test suites.
- Strong working knowledge and intermediate-level experience with Python programming for scripting, test automation, and data analysis.
- A foundational understanding and beginner-level exposure to Artificial Intelligence concepts, with an eagerness to explore its applications in QA.
Preferred Skills:
- Familiarity with Microsoft Azure cloud services at a beginner level, particularly concerning its potential use in test environments or automation pipelines.
Job responsibilities
- Define and own the evaluation framework for the PoC, including extraction accuracy and search relevance metrics
- Validate extraction of unstructured content from legacy Excel files, covering text inside shapes, callouts, and embedded images processed through OCR
- Verify that content survives format conversion without loss, comparing source documents against indexed output
- Test incremental index update logic, confirming that only changed and new files are reprocessed and that deletions propagate correctly
- Measure and report index update latency and delta correctness
- Design test cases for edge scenarios (corrupted legacy files, macro-heavy workbooks, scanned images with low quality, near duplicate and outdated documents)
- Build a ground truth dataset together with the client subject matter experts and use it to score natural language search results
- Perform integration and regression testing across the ingestion, extraction, and indexing pipeline
- Support client validation sessions and prepare evaluation results for stakeholder review