Orbox

AI-First Senior QA Engineer (Manual)

$$$$

AI-First Senior QA Engineer (Manual)

Ukraine, Kyiv · On-Site · Full-Time

 

Company: Orbox.ai - AI-Native PropTech Operating System for Real Estate

 

Location: Ukraine, Kyiv  on-site, office in the city centre

Employment: Full-time, permanent

Seniority: Senior

Reports to: Technical Lead 

Works with: Product Analyst, Product Designer, Engineering team, Fleet of Agents

Mobilization reservation: Full reservation (deferment) from mobilization for male employees

PropTech experience: Not essential; very good if you have direct or indirect experience

English: Advanced (B2+)

 

About

We are building an AI-native operating system for real estate. The product is a PropTech platform where LLMs and AI agents are not bolted on top - they are the architecture. Owners, operators and guests interact with the AI system through natural language. AI autonomously generates compliance documents, routes staff, answers guest queries from a RAG-indexed knowledge base in real time and orchestrates financial split-payments.

 

But the product is only half of it.

 

We are equally an AI-first delivery organization. The way we discover, define and ship is as differentiated as what we ship. Specs, stories and designs are machine-readable artifacts our coding and review agents consume directly — and quality is verified by a human who knows exactly where agents fail, where specs lie and where the product breaks in real hands.

 

The role

This is the role that decides whether what we ship actually works  for humans and for the agents that built it.

 

You will report directly to the Head of Delivery & Product, with short decision loops to founders and direct access to engineering.

 

This is a hands-on, manual-first QA role at the core of the delivery machine. Most of our code is written by AI agents against machine-readable specs. That makes human quality judgment more valuable, not less: agents pass their own checks and still ship broken flows, wrong edge-case behaviour and UX that technically matches the spec but fails the user. You are the last honest reader of the product before it reaches staging review and then customers on prod.

 

You will own manual testing end to end: exploratory testing, spec-based test design, UAT coordination, regression passes and release sign-off. You will test not just deterministic UI flows, but AI-driven behaviour  agent responses, RAG answers, document generation  where “expected result” is a judgment call, not an assert.

 

You will work side by side with the AI-First Product Analyst (who writes the specs you test against) and engineering (whose agents implement them). You will use our shared Claude Team / Cowork with premium seating along with the company's AI tools daily, deeply and natively — not as a novelty, but as the default way you work.

 

We operate in a fast-moving startup environment where speed and precision both matter. You must be opinionated about which QA tasks AI does well today (test-case generation, coverage analysis, log triage), which it does not (exploratory judgment, UX intuition, risk feel) — and how to combine both deliberately.

 

What you will own & do

 

AI-First Way of Working

• Daily, deep, native use of AI tools for test-case generation from specs, coverage gap analysis, bug report drafting, log and session triage, regression checklist upkeep.

• Treat AI tools the way an engineer treats their IDE: opinionated, tuned, integrated into every step of your workflow.

• Contribute to internal evals from the QA side: define what “good” looks like for AI-generated code, AI-generated test artifacts and our user-facing AI features.

 

Manual Testing & Test Design

• Own manual testing end to end: exploratory sessions, spec & story-based test design, edge cases, negative paths, cross-role and multi-tenant scenarios.

• Convert machine-readable specs & stories with acceptance criteria into structured, traceable test cases and checklists — every requirement maps to a verification.

• Run deep exploratory testing where scripted coverage ends: real-world flows, messy data, weird devices, impatient users.

• Verify complex B2B2C surfaces: dashboards, role-based access, real-time data, payment flows, document generation, back-office tooling.

 

Testing AI-Driven Behaviour

• Test the non-deterministic layer: agent conversations, autonomously generated compliance documents — correctness, groundedness, tone, failure modes.

• Design and run structured evaluation of LLM outputs with the Analyst and engineering: treat AI output as something to be measured, not trusted.

• Hunt for hallucinations, prompt-injection surfaces, permission leaks across tenants and roles, and silent agent failures.

 

Quality Process & Release Ownership

• Own release readiness: regression scope, smoke passes, UAT coordination with stakeholders, GO/NO-GO input for every release.

• Write bug reports engineers (and agents) can act on without a follow-up question: environment, steps, expected vs actual, evidence, severity.

• Feed spec & story defects back upstream: when the work item is wrong or ambiguous, it gets fixed at the spec/story layer, not patched in QA.

• Build the QA practice from scratch: process, tooling, test data, defect workflow in Jira — lightweight enough for startup speed, rigorous enough to trust.

 

What we are looking for

 

Must have

 4+ years of hands-on manual QA experience on web products, with senior-level ownership of quality for a product or major product area.

• 0 → 1 builder: comfortable being the first and only QA, setting up the practice yourself, working directly with founders and engineering — not joining an already-scaled QA org.

• Experience testing LLM-based features, chatbots, RAG systems or agent workflows — including designing evals or output rubrics.

• Daily, deep, native use of AI tools in your QA workflow: test generation, triage, analysis, documentation. You have a real, opinionated workflow built around them and can defend your choices.

• Excellent test design craft: from acceptance criteria to test cases, edge cases, negative scenarios and risk-based prioritization — structured, traceable, unambiguous.

• Strong exploratory testing instincts: you find the bugs no script would — and can explain how you found them.

• Sharp written bug reporting and strong verbal communication in English (C1+), under startup timelines, without losing rigor.

• Comfortable with complex B2B / multi-sided product surfaces: multi-tenant SaaS, role-based access, real-time data, payments, back-office tooling.

 Working knowledge of APIs and databases for verification: reading API responses (Postman or similar), basic SQL to check data behind the UI, reading logs.

 

Strong advantage (Nice to have)

• Background in PropTech, hospitality tech, Fintech or multi-sided marketplace platforms.

• Exposure to payment flows, compliance / regulatory workflows or document-generation systems.

• Test automation literacy (Playwright, Cypress or similar) — enough to direct agents that write automation, even if you don't write it daily.

 

Why this role

You will help build the first AI-native operating system for real estate in Dubai and define what quality means when AI agents write most of the code. Both are differentiators, both are hard and both are yours to own at the senior level.

 

• Genuine quality ownership — you build the QA practice from zero and hold the release bar. This is not a “execute someone else's test cases” role.

• Full reservation from mobilization for male employees — officially arranged by the company, so you can focus on the work.

• Clear growth trajectory — this role evolves into AI-First QA Lead / Head of Quality as the team scales.

• Direct access to founders and Head of Delivery & Product — short decision loops, no middle management.

• Technically ambitious product surface  AI-native, Agentic system, RAG, IoT, real-time systems, government and regulatory APIs and eventually tokenization infrastructure.

 AI-first by default — you will co-design how quality assurance works inside an AI-first delivery machine.

• Great office in the centre of Kyiv — one on-site team, short feedback loops, fast decisions.

Required skills experience

QA 4 years
Test cases 4 years
Test Planning 4 years
AI Automations 1 year
SQL 1 year
API Testing 1.5 years
Bug Reporting 3 years
Exploratory Testing 3 years

Required domain experience

Machine Learning / Big Data 6 months

Required languages

English B2 - Upper Intermediate
Ukrainian Native
Published 26 September
42 views
·
2 applications
To apply for this and other jobs on Djinni login or signup.
Loading...