Senior Prompt Engineer (Onsite, Dubai)
Europe, Ukraine HOTAbout the project
Our client, a flagship airline from the GCC region, is scaling its conversational AI platform from text to voice. Text-based agents are live; the next step is a production-grade voice experience over IVR. You will be embedded onsite with the client’s AI, backend, and telephony teams as the reference person for everything spoken, backed by our voice AI platform team.
Engagement terms
- Onsite in Dubai, client office, 5 days/week (Mon–Fri)
- Initial phase: approx. 4 months (August – December)
- Possible extension into further phases (4 / 6 / 12 months) subject to contract renewal and performance, with home leave between phases
- We cover: work visa, housing, round-trip flights, medical insurance
- Overtime is possible and paid at the standard hourly equivalent
- Possible performance bonus based on client’s end-of-phase feedback (“Exceeded Expectations”)
What you will do
- Design and iterate voice-agent prompts for an LLM-based conversational stack: latency, barge-in and interruption handling, turn-taking, confirmation patterns, graceful fallback to DTMF or a human agent
- Build multilingual voice flows (English and Arabic): you own the engineering side – prompt structure, SSML, custom lexicons, pronunciation handling; Arabic linguistic quality is covered by dedicated linguists and native testers
- Solve voice-specific slot capture: emails, guest names, PNRs, phone numbers — spell-back flows, per-character confirmation, ASR-error recovery
- Work within a multi-LLM pipeline — answer generation, response validation (hallucination checks), guardrails, PII masking, routing — and tune each layer without degrading latency
- Drive ASR/TTS quality: dual-STT configurations, provider selection and tuning (Azure AI Speech, Deepgram, Google, ElevenLabs, OpenAI Realtime), barge-in behavior, endpointing
- Instrument voice KPIs on the analytics dashboard — containment, task completion, success/error rate, abandonment, misrecognition on critical slots — and run A/B experiments on prompt variants
- Make prompt and rule changes safe, auditable, and reversible in production
Must have
- 4+ years building production conversational systems
- Proven LLM agent prompting experience (OpenAI, Anthropic or equivalent) with structured tool use and streaming — be ready to show a prompt library or before/after metrics
- Hands-on experience with at least one ASR engine (Azure AI Speech, Google STT, Deepgram, Whisper) and one TTS engine (Azure, ElevenLabs, Google)
- Shipped a voice slot-capture flow (email or name capture) and understand why it is hard
- Working knowledge of SIP and telephony basics: signaling, RTP/SRTP, DTMF modes, codecs, call routing via SIP trunk to a bot platform
- Experience with validation/guardrail LLM layers (hallucination checking, PII masking)
- Fluent English
Strong plus
- LiveKit or similar low-latency voice-agent frameworks (Pipecat, Vocode), OpenAI Realtime API
- Hands-on production work on the voice channel; Asterisk experience
- Multi-agent orchestration or custom agent frameworks
- CCaaS platforms: Genesys Cloud, NICE CXone, Amazon Connect, Twilio Flex
- Airline/travel domain terminology (PNR, ticket numbers, airport codes) and its phonetic handling
- Open-source voice-agent contributions; experience running voice user-testing with real customers
What success looks like
- Voice agent live in production in English and Arabic with a defined containment target (Arabic linguistic QA by our native testers)
- Misrecognition on the top three high-value slots (email, PNR, phone number) cut by 30%+ vs. launch baseline
- Repeatable evaluation harness so every prompt change is measured before it ships
Ready to rumble?
Send your CV or contact us here.