Senior Solo AI Engineer (Python, RAG, Conversational and Voice AI)

$$$$
Product

We build Arabic and English chat and voice assistants for UAE government services. They answer questions from official documents with citations, and they complete real transactions: traffic fines, licence renewals, payments, identity verification, all through SOAP and REST integrations with government systems.

Our work spans two worlds. Some projects run in the cloud with hosted models and managed GPUs. Others run fully on-premise inside air-gapped government networks: local models on our own GPU servers, no internet, no external APIs. You will work across both, and you will ship the same quality in each.

We are a small team and we deliver fast. A full system typically goes from requirements to production in under 30 days. That pace is possible because each engineer owns their work end to end: read the documents, build the feature, integrate the services, test the complete journey in chat and voice, package the release, and support it live.

What you will build

  • Chat applications: multi-step citizen journeys with intent detection, tool calling, confirmations, and clean handling of records like vehicles, fines, and receipts. English and Arabic, including right-to-left layouts.
  • Voice: streaming speech recognition and synthesis, barge-in, interruption handling, and smooth handoff from voice to visual UI for forms, payments, and OTP.
  • RAG pipelines: ingestion of PDF, DOCX, scanned and bilingual documents; hybrid retrieval; reranking; citations that point to the exact source; evaluation so we know retrieval actually works.
  • Agentic workflows: model reasoning combined with deterministic controls, so the assistant completes real services and never invents fees, eligibility rules, or records.
  • Model serving: hosted model APIs and cloud GPUs on some projects; open models served locally with vLLM or similar on others, including quantization, latency, and throughput tuning on government hardware.
  • Integrations: SOAP 1.1/1.2, WSDL, XML, and REST services; UAE PASS and OAuth identity flows; payment gateways with callbacks, receipts, idempotency, and audit logs.
  • Deployment: containers, systemd, reverse proxies, TLS. For air-gapped sites, complete offline packages: Python wheels, model weights, container images, migrations, and rollback, with nothing silently downloaded at install time.

How we work

This part matters as much as the technical list. You will often be the only engineer on a delivery, working directly from client documents and existing code.

  • You start from the documentation and the codebase, not from a task-by-task checklist.
  • You search the docs, code, logs, WSDLs, and captured traffic before asking for help, and you escalate only when something genuinely blocks you: a missing credential, an external contract, a business decision.
  • You fix root causes, not symptoms, and you check whether the same defect exists in the paths nobody reported.
  • You test your own work across the whole journey before calling it done: chat, voice, retrieval, integrations, UI, deployment, and the scenarios your change might have broken. Done means verified, not "the model answered correctly once."
  • You give short, honest progress updates with realistic estimates, especially when something slips.
  • You stay accountable for the outcome the citizen sees, not just the component you wrote.

What we need from you

  • 4+ years of production Python: async services, APIs, background workers, real error handling.
  • LLM applications you actually shipped: RAG, tool calling, or agents doing useful work in production.
  • Experience with hosted model APIs, and either experience running open models on GPUs or the clear ability to learn it fast.
  • Solid Linux: containers, networking, TLS, and debugging a server that has no internet access.
  • Integration skill: you can take a WSDL or API spec, get a working integration, and diagnose it when the live service disagrees with the documentation.
  • The independent working style described above, for real, not as an interview answer.
  • Based in the UAE or able to relocate quickly; the role is primarily on-site.

Strong pluses (we will teach what you are missing)

  • Arabic, especially Gulf dialect, and English-Arabic code-switching
  • Streaming speech systems: ASR, TTS, WebSockets, WebRTC
  • GPU serving and quantization (vLLM, TensorRT-LLM, or similar)
  • UAE PASS, UAE government services, or payment gateway integrations
  • Air-gapped packaging: offline repositories, checksums, SBOMs
  • SQL Server or comparable enterprise databases

 

Required skills experience

LangGraph 2 years
Voice AI Agent 6 months
On-Premise Infrastructure 2 years
Agentic AI 2 years
Python 5 years

Required languages

English B2 - Upper Intermediate
Published 7 September
26 views
ยท
6 applications
Last responded 49 minutes ago
See stats of candidates who applied for this job ๐Ÿ‘€
To apply for this and other jobs on Djinni login or signup.
Loading...