How to Set Up an ElevenLabs Voice Agent for Lead Capture

ElevenLabs Conversational AI is a hosted voice-agent runtime that wires speech-to-text, an LLM brain, tool calls, and ElevenLabs' own text-to-speech into a single low-latency pipeline you can attach to a phone number. The promise is "ship a voice agent in ten minutes." The reality, if you point it at real inbound leads, is that the demo agent breaks on the first call.
The thesis of this piece: most ElevenLabs tutorials end at "your agent picks up the phone." That's the easy 20%. The remaining 80% — structured qualification, real CRM write-back with idempotent fallback, human handoff before the lead bounces, latency budget you can actually defend — is what separates a working demo from an agent that earns its $0.10/minute. We'll walk the full lead-capture stack as it stands in May 2026: ElevenLabs Agents Platform + Twilio for the phone layer + HubSpot for the destination, with the design decisions named at each fork.
The throughline: every minute of voice latency above ~900ms is a lead you've already lost, and every CRM write that isn't idempotent is a duplicate you're paying a SDR to dedupe on Monday.
The quick answer
What is an ElevenLabs voice agent?
A hosted pipeline that takes an inbound phone call, transcribes the caller in real time, sends turn-by-turn context to an LLM you pick (GPT-4o, Claude, Gemini, or your own server), runs whatever tool calls the LLM asks for, and speaks the response back in an ElevenLabs voice — all in under a second per turn.
Why does it matter for lead capture?
Inbound voice is the highest-intent surface most SaaS companies still ignore because hiring a 24/7 SDR is uneconomic at < $50M ARR. A voice agent at $0.08–$0.12 per minute (ElevenLabs Agents pricing) makes the math work — but only if the agent qualifies, writes back, and hands off cleanly.
The non-obvious truth: the LLM model you pick matters less than the prompt structure and the tool contract. Founders obsess over GPT-5 vs Claude. The lead-quality difference comes from how cleanly your qualify_lead tool schema constrains the conversation.
The foundations (everyone explains these, we'll be quick)
The ElevenLabs Conversational AI stack has four layers, stitched together by the platform:
- Speech-to-text. ElevenLabs uses its own STT (Scribe v2 baseline) plus the option to swap in Deepgram or AssemblyAI. Inference target is ~150ms for short utterances.
- LLM brain. You pick. Native options include GPT-4o, GPT-4o-mini, Claude (Sonnet/Haiku), Gemini Flash, and a custom server endpoint. Per ElevenLabs' own latency guide, Gemini Flash 1.5 returns first tokens in under 350ms; GPT-4o and Claude land in the 700–1000ms range.
- Tools. Webhook tools that hit your endpoints, MCP server connections, or native integrations (HubSpot, Salesforce, Zapier, n8n).
- Text-to-speech. ElevenLabs Flash v2.5 or Multilingual v2, ~75ms model inference for typical short outputs.
The four layers chain together. End-to-end pipeline latency lands at 450–750ms for native models and pushes past 800ms once you wire in a heavier external LLM. That number is the budget. Spend it carelessly and the caller hangs up.
Phone connectivity is solved one of two ways: (a) buy a number directly through ElevenLabs' native Twilio integration (one-click, less control), or (b) own the Twilio number, then attach it via account SID + auth token in the ElevenLabs dashboard. Option (b) is what production teams pick because it preserves your Twilio call recording, your existing compliance posture, and the option to bypass ElevenLabs if you ever need to. The official Twilio + ElevenLabs guide covers the click path. Five minutes to a ringing agent.
That's the demo. Now the part the tutorials don't write.
Where the surface tutorials break: 1 — the qualification prompt
The single biggest production failure mode is a system prompt that reads like "you are a friendly assistant for Acme Corp." That works for the demo call. It does not work for a real inbound where the caller has 90 seconds of patience and you have one shot to capture: who they are, what company, what pain, what budget tier, what timeline. Five fields.
The fix is structural, not stylistic. Don't ask the LLM to "qualify the lead." Give it a tool. The tool's schema becomes the qualification rubric:
{
"name": "qualify_lead",
"description": "Called once at the end of a qualified conversation to persist the lead.",
"parameters": {
"type": "object",
"required": ["full_name", "email", "company", "pain_point", "budget_band", "timeline"],
"properties": {
"full_name": { "type": "string" },
"email": { "type": "string", "format": "email" },
"company": { "type": "string" },
"pain_point": { "type": "string", "maxLength": 280 },
"budget_band": { "type": "string", "enum": ["under_5k", "5k_25k", "25k_100k", "over_100k", "unknown"] },
"timeline": { "type": "string", "enum": ["now", "this_quarter", "this_year", "exploring"] }
}
}
}Two things happen. First, the LLM will not call qualify_lead until it has all six fields, so the system prompt's "ask, listen, capture" instruction becomes load-bearing instead of decorative. Second, the enums make budget and timeline gradeable downstream — your sales team can sort by over_100k + now without parsing free text.
System prompt then becomes one paragraph: "You're qualifying inbound leads for {product}. You have ~3 minutes. Ask one thing at a time. Capture the six fields the qualify_lead tool needs. When you have all six, call the tool, confirm the email back to the caller, and offer to either book a meeting via book_meeting or text the caller a link. Hang up warmly."
That prompt + that tool schema is 80% of the agent's lead quality. The model choice is the remaining 20%. Premium tier ($0.12/min, gpt-4o) is the right default; Turbo (gpt-4o-mini, $0.10/min) is fine if your agent's job is purely "intake" and a downstream human re-reads the transcript.
Where the surface tutorials break: 2 — CRM write-back you can actually trust
The tutorials all show "now connect HubSpot." The HubSpot integration is real — ElevenLabs ships a native template that exposes four webhook tools (search_contact, create_contact, update_contact, create_deal) hitting the HubSpot CRM v3 API with Bearer auth. Wire those up and the agent can read and write contacts mid-call.
What the tutorials skip: idempotency, retry, and the dirty-data fallback.
Three failure modes you will hit in week one:
- Duplicate contacts. Caller's email is already in HubSpot from a year-old form fill. Agent creates a new contact. Your SDR sees two records. Fix: every
create_contactcall must firstsearch_contactby email and route toupdate_contactif found. Don't trust the LLM to remember this — bake it into a singleupsert_contactwebhook on your side that wraps both HubSpot calls and returns one contact_id. - Webhook timeouts mid-call. HubSpot's API p99 is 2–3 seconds. If your webhook tool waits synchronously, the agent goes silent for those seconds. Caller assumes they got cut off. Fix: respond to the tool call in under 500ms with
{"status": "queued", "lead_id": "uuid-here"}and process the actual HubSpot write in a background job. The agent thanks the caller while the queue drains. - The "I'll text you a link" handoff that never sends. If the conversation closes with the agent promising a calendar link, the SMS is a separate Twilio API call your webhook must trigger. Test this with your own number weekly — silent failures here look like agent success in the dashboard and silence in the lead's inbox.
The lesson is bigger than HubSpot. Voice agents look stateless to the LLM but are extremely stateful to your CRM, your SMS provider, and your calendar. Each external system has its own failure mode, and the voice surface gives you exactly zero seconds to recover gracefully. Treat every tool call like a payment: idempotent, retried, logged, alerted.
Where the surface tutorials break: 3 — the latency budget is a real budget
Here's the math on a single conversational turn:
| Stage | Native pipeline | With GPT-4o brain | With Claude brain |
|---|---|---|---|
| STT (transcribe last utterance) | 150ms | 150ms | 150ms |
| LLM first token | 350ms (Flash) | 700–900ms | 800–1000ms |
| Tool call round trip (if any) | +300–800ms | +300–800ms | +300–800ms |
| TTS to first audio | 75ms | 75ms | 75ms |
| Total (no tool) | ~575ms | ~925ms | ~1125ms |
| Total (with tool) | ~875ms–1375ms | ~1225ms–1725ms | ~1425ms–1925ms |

Sources: ElevenLabs' published latency benchmarks and the model-by-model comparison from TokenMix's 2026 voice API breakdown.
900ms is the perceptual cliff for phone conversation. Past that, callers think they've been disconnected. Past 1500ms, they hang up.
What this means in practice:
- For lead capture (3-minute calls, mostly Q&A, occasional tool call): use GPT-4o-mini or Gemini Flash, not GPT-4o. The intelligence ceiling is high enough; the latency floor is the constraint.
- Cache the system prompt and the tool catalog at the LLM provider. OpenAI's prompt caching takes ~50ms off first token after the first call.
- Make every webhook tool respond in under 500ms. Background the heavy work. This is non-negotiable.
- Do not stack tool calls in one turn. If
search_contactreturns "found," the LLM should respond conversationally next turn and only callupdate_contactafter another caller utterance.
The latency budget is not a "nice to have." It is the difference between a $20-CAC lead and a hung-up call.
Where the surface tutorials break: 4 — handoff before the lead bounces
The most common failure I see in production agents: a high-intent caller asks something the agent can't answer ("can you do SOC2 reports for our auditor?") and the agent says "I'll have someone reach out." The lead is gone. Asynchronous follow-up converts at a fraction of the rate of "let me grab them right now."
Build a transfer_to_human tool. Two implementations work:
- Warm transfer via Twilio. Use Twilio's Dial TwiML verb to bridge the caller to a sales rep's cell phone. The agent says "putting you through now"; Twilio rings the rep; the rep picks up. ~5 seconds end-to-end. Caller does not redial.
- Slack notify + callback within 60s. Webhook hits a Slack channel with the lead's transcript-so-far. First rep to claim it dials the caller back. The agent says "we'll be back to you in under a minute" and stays on the line keeping the caller engaged with company context until the callback lands.
The first is harder to wire and converts ~2x higher. Pick it.
Trigger conditions for transfer_to_human: any caller-stated budget over your top band, any compliance/legal/security question the agent isn't authorized to answer, any explicit "can I talk to a human." Code the triggers into the system prompt as a hard rule, not a suggestion.
How does this compare to OpenAI's Realtime API?
OpenAI shipped GPT-Realtime-2 on May 7, 2026 as a single-model voice pipeline — STT, reasoning, and TTS all inside one model with no chaining. We covered the architectural shift in OpenAI just collapsed the voice stack. For lead-capture specifically, Realtime-2 wins on raw latency (sub-500ms end-to-end is achievable) but loses on three things that matter more than latency:
- Voice quality. ElevenLabs' Flash v2.5 is still the most human-sounding TTS in production; Realtime-2's voices are improving but read as synthetic on a 4G phone speaker.
- Tool ecosystem. ElevenLabs' Agents Platform has named integrations for HubSpot, Salesforce, n8n, Zapier, and an MCP bridge. Realtime-2's tool surface is the OpenAI Functions API — flexible but you're rebuilding plumbing.
- Pricing. Realtime-2 bills per audio token; for a 3-minute call that's typically more expensive than ElevenLabs' flat $0.10/min Turbo tier.
For a SaaS team prioritizing lead intake quality over absolute lowest latency, ElevenLabs + GPT-4o-mini brain is still the right pick in May 2026. For a team building a voice product where every millisecond matters and the voice is the product, Realtime-2 has the edge.
The same architectural pattern also shows up on the support side. Anthropic's Claude Agent SDK has reshaped triage flows for support teams, and a voice agent is functionally the same animal: a tool-using LLM with a strict turn structure and a stateful external system to write to.
The edge cases that will bite you
A few things production agents hit that the docs underplay:
- Background noise from car/cafe calls. ElevenLabs Scribe handles this reasonably but not perfectly. Add a system prompt rule: "If you can't make out the caller, ask them to repeat once, then ask if they'd prefer a callback." Don't loop on "sorry, can you say that again."
- Numbers and emails over voice. Caller says "ypratham at gmail dot com." STT hears "why pratham at gmail dot com." Add a confirmation step: every captured email gets read back letter-by-letter before the
qualify_leadtool fires. - Pricing questions. Inbound voice attracts price-shoppers. Decide upfront: will the agent state pricing or always punt to "let me put you with our team"? Stating pricing wins on conversion but only if your pricing is simple. Hide the decision behind a
pricing_policyfield in the system prompt — flip it without retraining. - TCPA, GDPR, recording disclosure. Every inbound call needs a "this call may be recorded" preamble in jurisdictions that require two-party consent. ElevenLabs does not auto-inject this. Wire it into the first agent utterance.
- The 95% silence discount. ElevenLabs gives a 95% billing discount for periods of silence longer than 10 seconds. This matters when the caller puts you on hold to grab their credit card. Don't fight the silence; the meter respects it.
The Honest Take
What this article assumes, that could turn out wrong:
ElevenLabs' lead in voice quality is the load-bearing reason to pick this stack. If OpenAI's TTS catches up in the next six months — and Realtime-2 is already 80% of the way there on the synthetic-vs-human axis — the calculus flips. The pricing math also assumes Twilio is in your stack already. If you're starting from zero, ElevenLabs' native phone numbers are a one-click path that ships faster, at the cost of long-term portability.
I haven't published a private benchmark of HubSpot webhook p99 latency over a sustained call volume — the 2–3 second p99 number is HubSpot's own published target, and I treated it as the planning constraint. Real-world p99 under burst load may be worse, which would push the "respond in 500ms with a queued status" pattern from "recommended" to "mandatory."
The "warm transfer converts ~2x higher" claim is from sales operator anecdote and small-sample HubSpot data, not a clean industry study. Treat the magnitude as directionally right and the exact multiplier as uncertain.
One framing limit: this piece is calibrated for SaaS inbound where the call is the first interaction. For agents that operate inside a longer multi-channel funnel (email → SMS → call), the qualification rubric changes and the tool surface needs to read prior interaction history. That's a separate architecture and a different article.
The 2026 voice infrastructure market is moving fast. ElevenLabs locked in distribution via the Spotify partnership, and as we argued at the time, ElevenLabs has become a B2B infrastructure company, not a creator tool. Six months from now the integration list will look different and at least one of the brains in the table above will be retired. The pattern — qualify-via-tool-schema, idempotent CRM write, warm transfer, defended latency budget — will not.

FAQ
Do I need to host my own LLM endpoint or can I use ElevenLabs' built-in models?
Use the built-in models for v1. Custom LLM server adds 100–200ms of round-trip and a deployment surface you don't need until you're past 10k minutes/month or have a finetuned model. Default to GPT-4o-mini (Turbo tier, $0.10/min). Upgrade only if call quality data tells you to.
Can ElevenLabs voice agents handle outbound calling for cold sales?
Technically yes, with the outbound call API. Legally and ethically: be very careful. US TCPA and equivalents in EU, India, and Australia require explicit prior consent for automated outbound to consumers. Most production teams use ElevenLabs only for inbound and warm callbacks to leads who initiated contact.
What's the difference between ElevenLabs Agents Platform and a custom build with Twilio + Deepgram + OpenAI + ElevenLabs TTS?
About 200–400ms of latency and 6 weeks of engineering time. The custom build wins on pricing at very high volume (>500k minutes/month) and on flexibility. The hosted platform wins on every other axis at every other volume. Start hosted; rebuild only when the cost math forces it.
How do I handle multilingual inbound calls?
Set the agent's first-language detection to "auto" in the platform settings, use ElevenLabs Multilingual v2 voice (slightly slower than Flash but worth it), and write your system prompt in English — the LLM will respond in whichever language the caller starts in. Tested production languages: English, Spanish, German, French, Hindi, Portuguese, Japanese.
What's the realistic lead conversion rate for a voice agent versus a contact form?
Operator data from 2025 puts voice-agent qualified leads at 2–4x the conversion rate of form-fill leads to booked meeting, primarily because the voice agent surfaces and addresses objections in real time. The catch: voice agents see only the callers willing to dial — typically 10–20% of your inbound traffic. Net effect on pipeline depends on whether you can route warm leads to the voice surface deliberately.
Will Google penalize my site for hosting AI-generated voice transcripts?
No, as of the December 2025 Core Update. Google's stated position is that transcripts of real conversations between users and AI agents are first-party content. Index them with noindex if they contain PII; index them with consent if they're useful (FAQ-style transcripts from sales calls perform well as long-tail SEO landing pages).
The bottom line
A working ElevenLabs voice agent for lead capture is four moving parts, not one. Get the system prompt thin and the tool schema thick. Treat every webhook like a payment: idempotent, queued, retried. Defend the 900ms latency budget like it's a SLO — because it is, and the SLA is the caller hanging up. Build the warm-transfer path on day one, not week eight.
The platform itself is mature. The reason most lead-capture agents fail in production is not ElevenLabs, OpenAI, or Twilio. It's that the team treated the agent as a chatbot with a microphone instead of as a distributed system with a 900ms total latency budget. Build it like the latter and the $0.10/min math becomes the cheapest SDR you've ever hired.