KAWORKS

Senior Voice AI Engineer, Evaluation and Reliability

$$$$

We are currently looking for a Senior Voice AI Engineer, Evaluation and Reliability to join a professional German-based product company team that designs innovative eHealth solutions for hospitals.

You will work on a stable healthcare product.

 

Senior level, and we mean it. You would own the evidence that this system is safe to put in front of

patients, on your own judgement, and tell us when the answer is no. That isn't a job we can hand to

someone still learning what to measure.

 

You would not be building a voice stack. We use a third-party voice platform that already handles the

telephony, the speech recognition and synthesis, and the knobs for end-of-speech detection and

interruption. What it does not do is tell us whether any of that is working well enough for patients, and it

gives us no way to check that a change hasn't made things worse. That's the job.

 

So you'd own the conversation and the evidence about it: how the assistant behaves when a caller

interrupts, corrects themselves or won't answer; how fast the whole thing responds once a round trip to

our backend is in the loop; and the body of test cases we run before every change.

 

You need

  • To have run a phone or voice assistant that took real calls from real users, and to be able to tell us what its response time was and how you measured it. Having built the audio pipeline yourself is fine but not the point โ€” we care that you know what to measure and what good looks like.
  • To have tuned a voice system rather than only shipped one: end-of-speech detection, barge-in, how long to wait before assuming the caller has finished. You'd be doing this against a vendor's settings rather than your own code, and their own documentation says clean barge-in can't be guaranteed on phone audio.
  • To know what happens to speech recognition accuracy when the audio comes down a phone line instead of a good microphone, and to have measured it rather than read about it.
  • A way of testing conversation changes that doesn't rely on listening to a few calls and hoping. This is the centre of the role. Whatever form it took โ€” a replay harness, a scored scenario set, anything โ€” we want to hear how it worked and where it let you down.
  • Enough Python to build that harness and the tooling around it.
  • We are considering candidates who work through Ukrainian FOP.

 

Nice to have

  • Having worked against a closed voice platform, with the frustrations that brings, rather than only your own stack.
  • Telephony background: SIP , WebRTC, Asterisk, Twilio or similar. Useful for knowing why calls sound the way they do, though you would not be running any of it.
  • Work with health data, or any other field where getting it wrong has consequences.
  • German, though we're not counting on it.

 

You don't need

  • To speak German.
  • To have built streaming audio infrastructure, a turn-taking system or a speech pipeline. The platform does that, and rebuilding it isn't on the table.
  • To know PHP or anything about our existing backend. Another team builds the endpoints you'd call.
  • To have trained a model from scratch.

Required skills experience

AI/ML 4 years
ip-telephony 2 years
Voice AI Agent 2 years

Required domain experience

Healthcare / MedTech 3 years

Required languages

English C1 - Advanced
Ukrainian Native
Published 28 August
18 views
ยท
4 applications
To apply for this and other jobs on Djinni login or signup.
Loading...