AI-First Senior QA Engineer (Manual)
AI-First Senior QA Engineer (Manual)
Ukraine, Kyiv · On-Site · Full-Time
Company: Orbox.ai - AI-Native PropTech Operating System for Real Estate
Location: Ukraine, Kyiv on-site, office in the city centre
Employment: Full-time, permanent
Seniority: Senior
Reports to: Technical Lead
Works with: Product Analyst, Product Designer, Engineering team, Fleet of Agents
Mobilization reservation: Full reservation (deferment) from mobilization for male employees
PropTech experience: Not essential; very good if you have direct or indirect experience
English: Advanced (B2+)
About
We are building an AI-native operating system for real estate. The product is a PropTech platform where LLMs and AI agents are not bolted on top - they are the architecture. Owners, operators and guests interact with the AI system through natural language. AI autonomously generates compliance documents, routes staff, answers guest queries from a RAG-indexed knowledge base in real time and orchestrates financial split-payments.
But the product is only half of it.
We are equally an AI-first delivery organization. The way we discover, define and ship is as differentiated as what we ship. Specs, stories and designs are machine-readable artifacts our coding and review agents consume directly — and quality is verified by a human who knows exactly where agents fail, where specs lie and where the product breaks in real hands.
The role
This is the role that decides whether what we ship actually works for humans and for the agents that built it.
You will report directly to the Head of Delivery & Product, with short decision loops to founders and direct access to engineering.
This is a hands-on, manual-first QA role at the core of the delivery machine. Most of our code is written by AI agents against machine-readable specs. That makes human quality judgment more valuable, not less: agents pass their own checks and still ship broken flows, wrong edge-case behaviour and UX that technically matches the spec but fails the user. You are the last honest reader of the product before it reaches staging review and then customers on prod.
You will own manual testing end to end: exploratory testing, spec-based test design, UAT coordination, regression passes and release sign-off. You will test not just deterministic UI flows, but AI-driven behaviour agent responses, RAG answers, document generation where “expected result” is a judgment call, not an assert.
You will work side by side with the AI-First Product Analyst (who writes the specs you test against) and engineering (whose agents implement them). You will use our shared Claude Team / Cowork with premium seating along with the company's AI tools daily, deeply and natively — not as a novelty, but as the default way you work.
We operate in a fast-moving startup environment where speed and precision both matter. You must be opinionated about which QA tasks AI does well today (test-case generation, coverage analysis, log triage), which it does not (exploratory judgment, UX intuition, risk feel) — and how to combine both deliberately.
What you will own & do
AI-First Way of Working
• Daily, deep, native use of AI tools for test-case generation from specs, coverage gap analysis, bug report drafting, log and session triage, regression checklist upkeep.
• Treat AI tools the way an engineer treats their IDE: opinionated, tuned, integrated into every step of your workflow.
• Contribute to internal evals from the QA side: define what “good” looks like for AI-generated code, AI-generated test artifacts and our user-facing AI features.
Manual Testing & Test Design
• Own manual testing end to end: exploratory sessions, spec & story-based test design, edge cases, negative paths, cross-role and multi-tenant scenarios.
• Convert machine-readable specs & stories with acceptance criteria into structured, traceable test cases and checklists — every requirement maps to a verification.
• Run deep exploratory testing where scripted coverage ends: real-world flows, messy data, weird devices, impatient users.
• Verify complex B2B2C surfaces: dashboards, role-based access, real-time data, payment flows, document generation, back-office tooling.
Testing AI-Driven Behaviour
• Test the non-deterministic layer: agent conversations, autonomously generated compliance documents — correctness, groundedness, tone, failure modes.
• Design and run structured evaluation of LLM outputs with the Analyst and engineering: treat AI output as something to be measured, not trusted.
• Hunt for hallucinations, prompt-injection surfaces, permission leaks across tenants and roles, and silent agent failures.
Quality Process & Release Ownership
• Own release readiness: regression scope, smoke passes, UAT coordination with stakeholders, GO/NO-GO input for every release.
• Write bug reports engineers (and agents) can act on without a follow-up question: environment, steps, expected vs actual, evidence, severity.
• Feed spec & story defects back upstream: when the work item is wrong or ambiguous, it gets fixed at the spec/story layer, not patched in QA.
• Build the QA practice from scratch: process, tooling, test data, defect workflow in Jira — lightweight enough for startup speed, rigorous enough to trust.
What we are looking for
Must have
4+ years of hands-on manual QA experience on web products, with senior-level ownership of quality for a product or major product area.
• 0 → 1 builder: comfortable being the first and only QA, setting up the practice yourself, working directly with founders and engineering — not joining an already-scaled QA org.
• Experience testing LLM-based features, chatbots, RAG systems or agent workflows — including designing evals or output rubrics.
• Daily, deep, native use of AI tools in your QA workflow: test generation, triage, analysis, documentation. You have a real, opinionated workflow built around them and can defend your choices.
• Excellent test design craft: from acceptance criteria to test cases, edge cases, negative scenarios and risk-based prioritization — structured, traceable, unambiguous.
• Strong exploratory testing instincts: you find the bugs no script would — and can explain how you found them.
• Sharp written bug reporting and strong verbal communication in English (C1+), under startup timelines, without losing rigor.
• Comfortable with complex B2B / multi-sided product surfaces: multi-tenant SaaS, role-based access, real-time data, payments, back-office tooling.
Working knowledge of APIs and databases for verification: reading API responses (Postman or similar), basic SQL to check data behind the UI, reading logs.
Strong advantage (Nice to have)
• Background in PropTech, hospitality tech, Fintech or multi-sided marketplace platforms.
• Exposure to payment flows, compliance / regulatory workflows or document-generation systems.
• Test automation literacy (Playwright, Cypress or similar) — enough to direct agents that write automation, even if you don't write it daily.
Why this role
You will help build the first AI-native operating system for real estate in Dubai and define what quality means when AI agents write most of the code. Both are differentiators, both are hard and both are yours to own at the senior level.
• Genuine quality ownership — you build the QA practice from zero and hold the release bar. This is not a “execute someone else's test cases” role.
• Full reservation from mobilization for male employees — officially arranged by the company, so you can focus on the work.
• Clear growth trajectory — this role evolves into AI-First QA Lead / Head of Quality as the team scales.
• Direct access to founders and Head of Delivery & Product — short decision loops, no middle management.
• Technically ambitious product surface AI-native, Agentic system, RAG, IoT, real-time systems, government and regulatory APIs and eventually tokenization infrastructure.
AI-first by default — you will co-design how quality assurance works inside an AI-first delivery machine.
• Great office in the centre of Kyiv — one on-site team, short feedback loops, fast decisions.