Voice AI · AWS Infrastructure
Voice Agents, Owned End to End
A template-driven voice AI platform, migrated off managed cloud onto self-hosted AWS EC2 — with per-call cost telemetry built into the code and validated on real calls.
The Economics
Why I Took It Off Managed Cloud
Voice AI wrappers — Retell, Vapi, Bland — sell the same stack at $0.10-0.15/min: LiveKit media plus the same OpenAI and Deepgram APIs you can call directly. Managed LiveKit Cloud is fairer at $0.07/min, but it still scales linearly with every minute sold.
Self-hosting flips the model: one ~$30/mo EC2 instance carrying unlimited agents, plus API costs at source rates. The break-even lands around 1,000 minutes a month — after that, every minute sold is margin the wrappers would have kept.
So the platform was migrated: LiveKit Cloud is no longer used. Same agents, same code — owned infrastructure underneath.
Self-hosted vs LiveKit Cloud — validated projections
| Monthly volume | Self-hosted | LiveKit Cloud |
|---|---|---|
| 1,000 min/mo | ~$74 | ~$70 |
| 10,000 min/mo | ~$474 | ~$700 |
| 50,000 min/mo | ~$2,100 | ~$3,500 |
At 50,000 minutes a month the owned stack costs $2,100 against $3,500 — and the gap compounds with every client added to the same box.
The Architecture
One Box, Everything On It
The real deployed call path, exactly as it runs — verified against the production repos, not a diagram of aspirations.
Widget — Vercel / Next.js
iframe-embeddable frontend with a server-side token route: Cloudflare Turnstile verification and per-IP rate limiting (5 req/min) before a LiveKit token is ever issued.
nginx → LiveKit Server
WSS terminates at nginx (SSL, port 443) and proxies to a Dockerized LiveKit Server on :7880 — TURN over 3478/5349, WebRTC media on UDP 50000-60000.
Python worker — systemd
One systemd service per agent (Restart=always, Docker-gated): Deepgram nova-3 STT → GPT-4.1 → Deepgram Aura-2 TTS, Silero VAD, streaming throughout.
The world outside
n8n webhooks for lead capture, calendar booking, and call-completed events; Cloudflare R2 for call recordings. AI APIs stay direct — no wrapper markup.
Validated per-minute cost — real test calls
| Deepgram STT — nova-3per audio second | $0.0077 |
| Deepgram TTS — Aura-2per 1K characters | $0.0237 |
| OpenAI GPT-4.1per token in/out | $0.0367 |
| AWS EC2 — amortized~$30/mo flat | ~$0.001 |
| Total, validated | $0.068 |
The LLM dominates — a ~4,000-token system prompt rides every turn, which is exactly the kind of finding you only get from measuring your own calls.
Cost Telemetry
Every Call, Priced by Its Own Code
Provider dashboards can't isolate a single call's cost — so cost tracking was built into the agents themselves. An llm_node() override captures LLM prompt and completion tokens and TTS character counts into per-call state, and the call-completed webhook carries the raw usage out.
An n8n Code node then computes STT, TTS, LLM, and total cost from the raw data — rates declared in one place — and writes every call to Google Sheets and the CRM. No dashboard dependency: code-level tracking is authoritative.
The whole pattern ships as a reusable template doc, so any new agent gets instrumented the day it's cloned.
Hosting changes the transport, not the model. The self-hosting penalty is 100-300ms — imperceptible next to a 500-1500ms LLM turn.
Singapore (ap-southeast-1) with a permanent Elastic IP for Asia; us-east-1 beside OpenAI and Deepgram for US callers.
Fargate was evaluated and rejected — LiveKit requires host networking. The honest answer was a real box, run like one.
The Fleet
Four Agents, One Template
Every agent is cloned from the same template — customized per client in under 30 minutes: clone, fill the env, edit the system prompt, deploy.
“May”
Reference deploymentThe portfolio receptionist — first agent migrated to EC2, and the birthplace of the cost-tracking system. Her repo carries the reusable template doc.
“Sarah”
Commercial cleaning clientAnswers questions, qualifies leads (business name included), and books appointments against the client's calendar — timezone-correct through DST.
Bilingual agent
Filipino-EnglishSpeaks Taglish with polite po/nyo markers and a custom Filipino voice persona — running as its own systemd service on the same box.
Agency agent
Marketing agency siteLive receptionist on the agency's contact page — one of the two public deployments you can call yourself, right now.
Production Scars
What Live Traffic Teaches
Every system below earned its keep by failing first — and being fixed with the evidence in hand.
The timezone bug
Bookings used a hardcoded UTC offset — 1 hour off in winter — and the availability check was built in UTC, ~7 hours off. Rebuilt on luxon with America/Los_Angeles zoning: correct through every DST transition.
Token endpoint hardening
The token route is the front door, so it got a door: Turnstile verification plus per-IP rate limiting before any LiveKit credential is minted.
Lead state in every payload
A contextvars pattern carries lead and booking fields into the call-completed webhook, so n8n routes booked calls and missed calls to the right sheet — no post-hoc matching.
Two of these agents are live on public contact pages — call one and hear the stack for yourself.