Voice AI & Telephony Engineering
LiveKit (WebRTC), SIP / VoIP bridging, WebSockets, Gemini Live API, Sarvam, Ultravox, Groq, streaming Speech-to-Text (STT) and Text-to-Speech (TTS) low-latency pipelines.
LiveKit (WebRTC), SIP / VoIP bridging, WebSockets, Gemini Live API, Sarvam, Ultravox, Groq, streaming Speech-to-Text (STT) and Text-to-Speech (TTS) low-latency pipelines.
Google Agent Development Kit (ADK), LangGraph, Multi-Agent Systems, A2A Protocol, Model Context Protocol (MCP), Retrieval-Augmented Generation (RAG), LiteLLM, LightRAG, Vector DBs (ChromaDB, Qdrant, Pinecone).
LLM-as-a-judge scoring, human review workflows, span-level error annotation, word-level language-confusion tracking for multilingual (Indic) voice calls, agent test playgrounds.
Google Veo (video generation), Imagen (image generation), Cartesia (primary TTS and voice cloning in VOAG), Fish Audio (voice cloning), FFmpeg video stitching, Playwright (headless scraping & visual brand extraction).
Python (`asyncio`, `contextvars` session metering), Go (`whatsmeow` session library), SQL (relational schema design), JavaScript / Next.js (front ends for AI apps), FastAPI, REST APIs, OAuth 2.0, Redis, NGINX, Docker, Microservices Architecture.
Google Cloud Platform (Cloud Run, Compute Engine, least-privilege IAM, VM provisioning & hardening, cost control), DigitalOcean, ZeroSSL, Git, GitHub Actions (CI/CD), NAT Gateways.
SARIMA, XGBoost, ARIMA, Prophet, Predictive Modeling, Time-Series Forecasting, Computer Vision Inference.