DEVELOPER-FIRST VOICE AI
Build voice AI agents with 3 lines of code.
Low-latency, carrier-grade infrastructure. Trigger calls via REST, stream WebRTC in browsers, receive signed webhooks, and bring your own carriers or models.
<150ms
Turn-taking latency
99.99%
Carrier SLA
100+
Dialing countries
BYOC
Zero telecom markup
5-MINUTE QUICKSTART
First AI phone call in under 60 seconds.
Place an autonomous call to any mobile number with one authenticated request.
- 1
Obtain your API key
Generate a secret from the Vaani console. Keys look like vaani_{32_hex}.
- 2
Make the API call
POST to /api/trigger-call/ with agent, number, and metadata.
curl -X POST "https://app.vaanivoice.ai/api/trigger-call/" \
-H "X-API-Key: vaani_••••••••••••••••••••••••••••••••" \
-H "Content-Type: application/json" \
-d '{
"agent_id": "0df81e22-8321-4ba2-b5e1-8843c08b8b09",
"phone_number": "+919876543210",
"metadata": {
"customer_name": "Aarav Sharma",
"order_id": "ORD-88219"
}
}'CORE REST API
Dispatch, observe, and manage agents from code.
/api/trigger-call/
Telephony & call dispatch
Fire a call with live customer context in one request—metadata weaves straight into dialogue.
- Dynamic variable injection from metadata
- Calling-hour guardrails with emergency / BYOC bypasses
- Atomic concurrency so trunks never oversubscribe
/api/transcript/{id} · /api/stream/{id}
Telemetry & transcripts
Pull speaker-diarized transcripts, dual-channel recordings, and per-turn latency when you need them.
- Token-level AGENT / USER timestamps
- Secure streaming URLs for stereo recordings
- ASR · LLM TTFT · TTS breakdown per turn
/api/agents/*
Agent management
Create, clone, and publish agents from code across multi-tenant workspaces.
- Programmatic CRUD and publish
- Phonetic dictionaries and keyword boosts
- No console-only lock-in
REAL-TIME WEBHOOKS
Event-driven voice. No polling.
Stream telephony signals, transfers, and post-call analytics to your endpoints as they happen. Every payload is signed with X-Webhook-Signature.
- call_started
- user_picked_up_at
- human_transfer_initiated
- human_transfer_successful
- call_ended
- call_postprocessing
VERIFY · HMAC SHA-256
expected = "sha256=" + hmac.new(
secret.encode(), request_body, hashlib.sha256
).hexdigest()
if not hmac.compare_digest(signature_header, expected):
raise HTTPException(403, "Invalid signature")SAMPLE · call_postprocessing
application/json
{
"event": "call_postprocessing",
"call_id": "outbound-1783515029-ea856e02",
"timestamp": "2026-09-23T10:45:12.000Z",
"to": "+919876543210",
"from": "+918035315360",
"data": {
"call_duration": 42.15,
"end_reason": "Normal Clearing",
"summary": "Customer confirmed delivery reschedule for Friday at 4 PM.",
"entities": {
"reschedule_date": "2026-09-25",
"slot": "16:00",
"confirmed": true
},
"dispositions": {
"Final Outcome": "Rescheduled",
"Sentiment": "Positive"
},
"recording_url": "https://app.vaanivoice.ai/api/stream/…",
"transcript": "[10:44:30] AGENT: Namaste Aarav ji…\n[10:44:35] USER: Haanji, Friday ko shift kardo."
}
}BRING YOUR OWN
Carriers and models—your stack, not a walled garden.
BYOC
Bring your own carrier
- Connect existing SIP trunks—keep negotiated rates
- Retain brand caller IDs without porting
- Zero per-minute telecom markup on your trunks
- Warm-transfer into your PBX or human queue
BYOL
Bring your own models
- Use your own LLM API keys or self-hosted endpoints
- Choose among 8+ TTS engines and regional voices
- Route to private inference for data residency
- Full control over model mix per agent
IN-APP WEBRTC
Voice agents in any browser or app.
Embed bidirectional agents in React, Next.js, Flutter, iOS, or Android without PSTN costs. Create a short-lived session, connect the realtime room, and stream transcripts plus agent state on the data channel.
- Session · POST /api/webrtc/session/create
- Stream · waveforms, tokens, Listening / Thinking / Speaking flags
CONNECT
const room = new Room();
await room.connect(wsUrl, sessionToken);
room.on(RoomEvent.DataReceived, (payload) => {
const message = JSON.parse(new TextDecoder().decode(payload));
console.log("Transcript:", message.text);
});DEVELOPER EXPERIENCE
Tooling that keeps you shipping.
Voice sandbox
Test with laptop mic or one-click dial to your handset.
Webhook simulator
Fire sample payloads at local or staging receivers.
Teacher AI diffs
Surgical prompt edits with git-style change tracking.
SDKs & OpenAPI
Python, TypeScript/JavaScript, and OpenAPI v3 specs.
FAQ
Common questions from engineering teams.
Vaani averages glass-to-glass turn-taking from under 150ms to about 450ms depending on model tier. That covers the full roundtrip: caller stops speaking (VAD), streaming STT, LLM time-to-first-token, neural TTS chunk synthesis, and the first audio packet on RTP/WebRTC. Per-turn millisecond benchmarks show up in your latency telemetry dashboard.
Continuous client- and server-side voice activity detection listens while the agent speaks. When the customer talks over the AI, the engine cancels the LLM stream and flushes the playback buffer in under 150ms. The agent stops, keeps context, and listens—without acoustic feedback or robotic overlap.
Register any existing carrier contract or on-premise PBX. Enter your SIP server URI, credentials, and transport (UDP, TCP, or TLS). Vaani routes on your rate cards and verified brand caller ID with zero per-minute telecom markup from the platform.
Yes. Use built-in tools (transfer, book appointment, sheets sync, and more) or custom OpenAPI/REST tools with a JSON schema and dynamic parameters—e.g. fetch_order_status(order_id). While the tool runs, Vaani masks latency with conversational fillers before speaking the result naturally.
Native bilingual and code-mixed speech models follow natural colloquial switches. If someone starts in English and mid-sentence moves to Hindi—"Can you check my balance, kal credit hua tha ya nahi?"—recognition tracks the transition without a keypress or restart.
Webhooks retry with exponential backoff. Non-2xx responses or timeouts over 10 seconds re-queue the event. Every payload is signed with HMAC SHA-256 in the X-Webhook-Signature header so you can verify origin with zero trust.
Multi-tier automatic failover via provider resolution. Each agent has primary and fallback providers—if the primary degrades or fails, Vaani promotes the fallback or next-best tier in real time so live conversations stay connected.
In India, an automated calling-hour enforcer blocks commercial promotional calls outside 08:00–20:00 IST, with bypasses for exempt transactional BYOC trunks and emergency alerts. Campaigns can pre-flight scrub against national Do-Not-Disturb registries and org-wide blacklists.
Yes. The WebRTC in-app voice SDK opens bidirectional sessions in React, Next.js, iOS, Android, and Flutter. Users talk over their device mic on secure realtime media—zero PSTN fees—plus waveforms and live captions if you want them.
Audio in transit uses SRTP / WebRTC DTLS; data at rest uses AES-256. For banking, healthcare, and payments you can enable zero-retention mode or selectively disable recording and transcript logging to meet HIPAA, PCI-DSS, and GDPR needs.
An atomic Redis-backed concurrency engine reserves lines safely. When you hit capacity, single API calls queue FIFO; bulk campaigns back off automatically so trunks don’t return 503s.
Opus (48kHz) for WebRTC, and G.711 μ-law / A-law (8kHz) for standard PSTN trunks.
Yes—private cloud and hybrid VPC via Kubernetes (EKS, GKE, AKS, or bare metal), including on-premise PBX handoff.
NEXT
Ready to build the future of voice?
Start with 50 free minutes. No credit card required.