CashNow Voice Lab

ElevenLabs Speech Engine Β· ElevenLabs APIs Β· Pipecat β€” on the test line
Call Center β†—

Setup

idle
Use headphones so the agent doesn't hear itself. Your browser will ask for the microphone.
Advanced settings

Live conversation

Start a call β€” the two-party transcript, tool calls and interruptions appear here live.

Latency

per turn Β· caller stops β†’ agent audio
β€”
time to first word
β€”
voice-to-voice p50
β€”
last turn
0
turns
0
tool calls
β€”
barge-in stop
hear (end-of-turn + STT)think (LLM)toolsspeak (TTS β†’ first audio)T3* = caller talked over the agent Β· T3† = agent re-engaged after silence (not counted)

Your calls

phone and browser calls, newest first Β· click a call for its transcript and timings
WhenNumberEngine Β· modelLengthGreetingFirst replyTypical replySlowestRepliesTalk-oversHow it ended

Launch concurrent calls

Type any number (971…). New numbers are added to the lab's allowlist when you press Start. Each call needs its own phone with someone answering. Mode B = the agent looks things up with live tool calls during the call.
Browser-mic calls stay on the Live call tab (one mic per page). Everything here runs at the same time β€” the summary shows how latency holds up while they all talk.
0
live now
0
calls
β€”
V2V p50
β€”
V2V avg
β€”
V2V max
0
tool calls
0
errors
Pick phone numbers and/or synthetic callers, then Start all.

Run a comparison

scripted synthetic callers Β· same script, same LLM, same tools
Load test = raise Concurrent calls. ElevenLabs trial concurrency is limited, so start small (1 β†’ 3 β†’ 5). Each synthetic call uses real ElevenLabs + LLM minutes.

Past comparisons

WhenEnginesScriptRunsΓ—ConcStatus

Results

hear = caller stops β†’ transcript ready (turn detection + STT) Β· think = β†’ LLM first token (incl. tools) Β· speak = β†’ first agent audio. All measured at the same point on our side for every engine.

Tools the agent can call

shared by all three engines
OnNameTypeKindDelayFail%
Mode A offers only action + system tools (data is preloaded). Mode B and Hybrid offer everything. Add a delay to see how each engine handles slow lookups, or a fail % to test error handling.

Edit tool

Pick a tool on the left.

Synthetic customers

fictional β€” no real customer data in the lab
IdNameLangOwes (AED)OverdueScenario

Customer record (what Mode A preloads / what tools return)

Agent prompt

used by all engines

Voices

ElevenLabs EU workspace

Synthetic caller scripts

Each turn: {"say": "...", "bargeInAfterMs": 1400} β€” bargeInAfterMs makes the caller talk over the agent's previous answer. The caller voice is "Aryan" (Indian male).

Phone allowlist (test numbers)

One number per line, digits only. The lab refuses to dial anything else.

Calls

WhenEngineSourceModeCustomerDurTurnsV2V p50V2V avg1st wordToolsBargeEnd

System status

How it's wired

Speech Engine β€” our bridge streams caller audio (PCM 16 kHz) to ElevenLabs' conversation WebSocket (EU). ElevenLabs does speech-to-text, turn detection and interruptions, then connects back to this lab's brain (/se/ws) with the transcript; the brain calls the LLM + tools and streams text back; ElevenLabs speaks it.

ElevenLabs APIs (DIY) β€” our bridge runs the loop itself: Scribe v2 Realtime (its VAD ends the turn) β†’ the same brain β†’ Flash TTS over the multi-context WebSocket. Barge-in = partial transcript while the agent speaks.

Pipecat β€” a Python pipeline on this server: Silero VAD + Smart Turn decide the turn, ElevenLabs realtime STT/TTS, the same LLM; tools call back into the same tool runner.

Phone β€” the VoiceGateway test line (not the prod line). Its allowlist admits the office network, so the lab reaches it through a relay on the Mac; if that relay is down, phone calls fail (browser + synthetic still work).

Latency is measured identically for every engine at our side: the moment the caller's audio stops β†’ the first audible agent audio we receive. The phone network's own delay is the same for all engines and not included.