AI Solutions Engineer building production AI systems, not prototypes. Leading TechieMaya's Voice AI vertical (VOAG): 1,000+ calls/day at sub-300ms p95 latency on self-hosted LiveKit and SIP telephony, with multi-provider model routing and per-tenant cost tracking.
Client-facing from scoping calls to go-live, shipping into constrained environments — client VMs, on-premise networks, strict NAT gateways — and staying accountable after launch, from evals and monitoring to cost tuning, with no ops team.
Equally hands-on with multi-agent systems (Google ADK, LangGraph), RAG over private data, and real-time speech pipelines in Python and Go on GCP.
Lead the Voice AI vertical: scaled to 1,000+ calls/day at sub-300ms p95 latency; kept LiveKit agents — built for always-on connections — running on serverless Cloud Run through a keep-alive, hold, and polling layer, cutting infrastructure cost 30%.
Hybrid cloud SIP: engineered a UAE telephony bridge (SIM -> modem -> SIP -> LiveKit) deployed inside a client's own VM behind a strict NAT gateway; self-hosted LiveKit across India and UAE with tenant-aware routing and GitHub Actions-driven CI/CD.
Agentic workflows & multi-provider routing: built a tenant-aware tool registry over OAuth 2.0 for live calendar scheduling, omnichannel messaging, and mid-call human handoff; added fire-and-forget RAG for long-document reference with no conversational dead air, routed across Gemini Live, Sarvam, and Ultravox for cost/language optimisation, with Cartesia as the primary TTS layer (plus Fish Audio) for brand-matched voice cloning.
Call evals: recorded audio and transcripts scored by an LLM-as-a-judge plus human review for task completion, grounding, and word-level language confusion across Indic languages; reviewer-annotated error spans feed a correction lexicon applied from the next call onward.
Ownership: voice AI lead in client scoping meetings; administer GCP for VOAG and the wider company (least-privilege IAM, VM hardening, cost control); designed VOAG's complete relational SQL schema.
WhatsApp Dispatch Automation — B2B Luxury Ground Transport (Dubai)
Built a Go gateway (whatsmeow) for group-chat messaging unsupported by Meta's official API, paired with a Google ADK multi-agent system (orchestrator, booking, support agents) resolving flight, maps, fleet-availability and fare data via tools — cutting booking turnaround from ~30 minutes to under 1 minute; extended the gateway to a second client for message-monitoring and forwarding across 1,000+ daily messages.
MAGe — Multi-Agent Media Generation Engine
Hierarchical agent pipeline (creative director -> scriptwriter -> reference selector -> generator -> reviewer) producing on-brand video ads from a Playwright-scraped brand profile, using Google Veo with frame-carry continuity and FFmpeg assembly; tracks per-session API cost via Python contextvars across concurrent async workers.
Privacy-First On-Premise RAG (A2A Protocol)
Inverted RAG topology for compliance-sensitive clients — the reasoning agent deploys onto client infrastructure over an Agent-to-Agent protocol, queryable by any external agent with zero data exfiltration; hybrid dense-vector and keyword index for mixed document types.
Freelance Voice AI Consultant — Quantashift Consultancy Services (contracted to MGS Technology)
Jan. 2026 – Mar. 2026 · Remote (Pune, India)
Hireups — Real-Time AI Interviewer with Photorealistic Video Avatar (Hireups)
Brought in after the internal team was blocked for months on LiveKit real-time video; delivered a live photorealistic avatar interviewer on Gemini's realtime audio-native model with sub-second, interruptible multilingual dialogue (English/Arabic in production).
Shipped three production deployments — GCP VM, Cloud Run (non-trivial: LiveKit workers require persistent connections), and the client's private server, which needed a custom LiveKit build to satisfy ZeroSSL on an untrusted IP range.
Rearchitected a failing WebRTC monolith into scalable GCP microservices; added a multi-LLM failover mesh (Gemini/Groq/Ultravox) for uninterrupted sessions, and moved proctoring inference in-browser (TensorFlow.js, MediaPipe, COCO-SSD) to eliminate server video compute.
Privacy-First On-Premise RAG: Inverted RAG topology over A2A protocol for compliance-sensitive clients with zero external document exfiltration.
WhatsApp Dispatch Automation: B2B group dispatch for Dubai luxury transport company (`whatsmeow` Go + Google ADK) plus independent musician community vertical.
VOAG — Enterprise Voice AI SaaS: 1,000+ daily calls in Mr. LADs app, sub-300ms latency, UAE LAN SIP modem architecture, fire-and-forget RAG.
UniBias — Live Attention Tracker: Vision Transformers, FastAPI, OpenCV. Privacy-first tracker analysing periodic webcam and screen frames to detect distraction and trigger alerts, with no recording.
Blood Bank Demand-Forecasting System: IEEE-published SARIMA/XGBoost dynamic micro-expiry simulation reducing platelet wastage from 11.2% to 2.5%.
Technical Skills Inventory
AI & Agents: Google ADK, LangGraph, LiteLLM, LightRAG, Multi-Agent Systems, A2A Protocol, MCP, RAG, Vector DBs, LLM-as-a-judge
Co-author: Led problem formulation, forecasting-model development (SARIMA/XGBoost; SARIMA lowest MAE at 5.85), and simulation. Reduced simulated wastage 11.2% -> 2.5% while holding 99.1% fulfillment, statistically validated across 30 iterations (paired t-tests).
Education & Extracurricular Leadership
Ajay Kumar Garg Engineering College (AKGEC)
B.Tech in Computer Science · CGPA: 8.0 | Sept. 2022 – Jun. 2026 · Ghaziabad, India
Coordinator, Cloud Computing Cell: Organised ML and cloud workshops for 200+ students.
Coordinator, Centre of Metaverse: Led "Rescue X", a VR first-responder platform, to Top 5 at the National IDE Boot-camp 2025.
Professional Certifications
Building AI Voice Agents for Production, LiveKit / DeepLearning.AI (Sep. 2025)