Voice AI Engineer

NeuroDrift

IndiaFull-timeToday
Fresher / Entry-levelRemoteIT Services and IT Consulting

Key skills

PythonAsyncioFastAPIWebRTCSIPRTPLLM IntegrationKubernetesReal-time Audio ProcessingTelephony IntegrationTool CallingLatency Optimization

What you’ll do

  • Build and operate end-to-end real-time voice agents, focusing on telephony, audio pipelines, and LLM integration. The role involves optimizing latency and managing stateful audio workers in production while presenting technical decisions to clients.

What they’re looking for

  • Requires 3+ years of experience with Python and async frameworks, along with a proven track record of shipping SIP-trunk and real-time audio integrations. Candidates must be proficient in Kubernetes and have experience with LLM voice tools like OpenAI Realtime or ElevenLabs.

Job description

NeuroDrift is a US-based AI voice and enterprise software company. We build real-time voice AI agents that run live on real phone lines, at scale, for enterprise customers. Our platform (CallDash) handles production call traffic every day — so latency, telephony quirks, and audio edge cases are our daily reality, not a research problem. We're looking for a strong real-time voice engineer with broad expertise across the stack — someone who has shipped voice into production and debugged it when it broke under load. What you'll do Build and operate real-time voice agents end to end — telephony, audio pipeline, LLM integration, tool calling Own latency: chase every millisecond from end-of-utterance to time-to-first-audio Integrate and harden SIP-trunk connections on real call platforms Deploy and scale stateful audio workers in production Take part in client meetings — present your approach and defend your technical decisions What we need (3–4 years experience) 3+ years Python, including async (asyncio, FastAPI or similar) Real-time audio in production — WebRTC, SIP, RTP Production telephony integration — you've shipped at least one SIP-trunk integration on a real call platform and debugged it under load LLM voice integration in production — Gemini Live, OpenAI Realtime, ElevenLabs Conversational, Retell, or comparable Function / tool calling in voice contexts — you understand how tool-call scheduling affects perceived latency and barge-in behavior Kubernetes — deploying stateful audio workers with autoscaling A latency mindset — you instinctively reach for end-of-utterance, time-to-first-token, and time-to-first-audio metrics before the user complains The setup Fully remote High-intensity role: 50–60 hours/week — this is startup pace, not a 9-to-5 Working hours primarily IST, with availability for client meetings in US time zones You'll thrive here if you like owning hard problems end to end, you're comfortable in front of clients, and "it works on my machine" isn't in your vocabulary.

To apply and see any additional details, continue to the original posting on the employer’s site.

Before you apply — beat the ATS

Most fresher applications are filtered by software before a human sees them. Build a clean, keyword-matched resume free, or check your current one.