Production Voice Infrastructure
High-throughput Voice AI SaaS platform handling over 1,000 production calls daily in the Mr. LADs app with sub-300ms p95 latency.
- Infrastructure: Self-hosted LiveKit on DigitalOcean (India) + UAE local LAN SIP modem deployment (SIM -> modem -> SIP gateway -> LiveKit on local LAN behind NAT).
- Worker Scaling: VM-hosted API & batch manager paired with 3 Cloud Run auto-scale workers. LiveKit agents expect always-on connections, yet run on scale-to-zero Cloud Run through a keep-alive, hold, and polling layer (about 30% lower infrastructure cost).
- Multi-Provider LLM/TTS: Dynamic routing across Gemini Realtime, Sarvam, and Ultravox; Cartesia as the primary TTS layer, with Cartesia and Fish Audio voice cloning for brand-matched voices.
- Fire-and-Forget RAG: Google File Search RAG returning context without stalling streaming audio loops.
- Data Layer: Relational SQL schema for the platform, designed end to end by Sahil.
- Call Evaluation: Recorded call audio and transcripts scored by an LLM-as-a-judge plus human review for task completion, grounding, and word-level language confusion across Indic languages; reviewer-annotated error spans feed a correction lexicon applied from the next call onward.