In-house voice infrastructure

Real-time voice.
At infrastructure cost.

QuickVoice is a real-time text-to-speech and speech-to-text API for voice agents and live applications. Stream with WebSocket or call plain REST. Pricing starts at $0.0018 per 1,000 characters.

QuickVoice 1M characters $2.80
Versus
ElevenLabs Multilingual v2 / v3 $100.00
Engine online
TTS streaming output
Input

“I found three options that match your request.”

Azelma24 kHz mono
0 msfirst audio <75 msstreaming
TransportWebSocket
FormatOpus / WAV
Request cost$0.0028 / 1K
TTS
$0.0028/1K chars
STT
$0.0018/1K chars
First audio
<75ms
Free
100,000chars
Native adapters
01 Hear the model

One line. Many voices.
Hear the range.

Six featured voices from the wider roster read the same words every time. Hear the model directly, with no montage, music, or post-production hiding the output.

Recorded from the live engine
GET /v1/voices
Azelma
English, Female
Alba
English, Male
Heart
English, Female
Michael
English, Male
Eve
English, Female
George
English, Male

Every voice reads the same line, so you can compare them directly.

Live Your input

Now make it yours.

Generate speech from your own copy, record a line, or upload audio to transcribe. The signal above switches to the real generated audio while it plays.

REST fallbackWebSocket first300 char demo
Live voice workspace
REST and WebSocket
0 / 300 characters
02 One key, both directions

Text becomes voice.
Voice becomes data.

The same bearer token authenticates every endpoint. Use REST when the input is ready, or WebSocket when text and audio arrive a chunk at a time.

One bearer key QUICKVOICE_KEY
01
Text to speech

Type it. Hear it.

$0.0028 / 1K
TEXT“Hello, Maya.”
TTS

Twelve natural English voices. Opus for the wire; lossless WAV when you need it.

REST
POST /v1/tts
Stream
WSS /v1/tts/stream
Returns
audio/opus or audio/wav
Text-to-speech reference
02
Speech to text

Speak it. Read it.

$0.0018 / 1K
STT
TEXT00:01.2 Hello, Maya.

Multilingual transcription with word-level timestamps for complete files and live audio.

REST
POST /v1/stt
Stream
WSS /v1/stt/stream
Accepts
wav, mp3, m4a, flac, ogg
Speech-to-text reference
Plain HTTP from any language. Quickstart WebSocket protocol Errors & limits
03 Integrations

Your stack.
Our voice layer.

No proprietary client and no new orchestration layer. Use a native adapter where it helps, or call the same REST and WebSocket surface from any language.

VOICE LAYERQuickVoice
RESTWebSocket
Native adapter

LiveKit Agents

Python v0.1.4Node v0.1.4

Drop real-time STT and TTS into a LiveKit agent.

pip install livekit-plugins-quickdial View plugin on PyPI
Native adapter

Pipecat

Python v0.1.0

QuickVoice services for Pipecat voice pipelines.

pip install pipecat-quickdial View plugin on PyPI
Universal

Plain HTTP

Any languageNo SDK required

Bearer auth, JSON in, audio or transcript out.

curl https://api.quickdial.ai/v1/tts Open quickstart
04 Production path

One process.
Fewer places to wait.

Each request stays co-located from authentication through model execution and output. The shorter path reduces network variance and keeps metering tied to the request.

Request path One worker
01
Authenticate Bearer key
Key scoped
02
Prepare Text or audio
Input ready
03
Run model Co-located
In process
04
Stream Bytes or text
Incremental
No network hop mid-request. Error codes, rate-limit headers, and retry behavior remain explicit at the edge. Read errors & limits
01

No channel reservations

Stateless workers scale horizontally. You do not reserve concurrent channels up front.

02

Per-key isolation

Request rate is governed per key, so separate workloads can be isolated cleanly.

03

Request-level metering

Every request reports usage, metered to four decimal places for exact reconciliation.

05 Transparent economics

Pay for characters.
Keep the margin.

TTS bills the characters you send. STT bills the characters returned. No seats, monthly minimum, or rounding every request up to the next minute.

Text to speech Opus01
$0.0028/ 1K characters

Opus. WAV (lossless) at $0.0045 / 1K.

Start generating
Speech to text02
$0.0018/ 1K characters

Billed on timestamped transcript characters returned.

Start transcribing
Cost proof for text to speech

Hear the difference. Then see it.

Compare the same character volume at published list prices. Every rate is shown on one scale with its source beside it.

Published TTS list pricesCost at 1M characters
QuickVoiceOpus$0.0028 / 1KCurrent public price$2.80
ElevenLabsMultilingual v2 / v3Source $0.1000 / 1K36× price gap$100.00
DeepgramAura-2Source $0.0300 / 1K11× price gap$30.00
Sarvam AIBulbul v3Source $0.0312 / 1K11× price gap$31.17
OpenAItts-1-hdSource $0.0300 / 1K11× price gap$30.00
Google CloudNeural2Source $0.0160 / 1K5.7× price gap$16.00

Directional comparison using published list prices per 1,000 characters, checked August 14, 2026. Enterprise discounts, taxes, and unrelated platform fees are excluded. Sarvam lists ₹30 per 10,000 characters; the USD estimate uses ₹96.24 per $1. Deepgram Aura-2 is compared because Nova-3 is speech-to-text and billed per audio minute.

Speech-to-text metering

Usage you can reconcile.

Audio vendors meter differently. QuickVoice keeps the unit explicit instead of forcing a misleading apples-to-oranges multiplier.

Billable unitTranscript charactersreturned
TimestampsWord levelincluded
RoundingNo minute blocksexact
ReportingPer request4 decimals
06 Before you build

Clear answers.
No fine print.

The questions developers ask before they put a voice API into a real product.

Browse all documentation
Which characters am I actually billed for?

For text-to-speech, the characters in the text you send, so you can price a request exactly before you make it. For speech-to-text, the characters of the transcript returned, since the input is audio and has no character count. Both are metered to four decimal places and reported per request.

Can I clone a voice, or synthesise a language other than English?

Not today. Text-to-speech ships twelve natural English voices and no cloning. Speech-to-text is the multilingual half. It auto-detects and transcribes several languages including French, German, Spanish, Italian and Portuguese.

Do I have to use an SDK?

No. It's plain REST and WebSocket, so curl or your language's HTTP client is enough. Official plugins exist for LiveKit Agents and Pipecat only because those frameworks want a native adapter. They're a convenience, not a requirement.

When should I use WebSocket instead of REST?

Use REST when you have the whole input up front; the response already streams back. Use WebSocket when the input itself arrives incrementally: piping LLM tokens into speech before the sentence finishes, or transcribing a live microphone.

What happens when I run out of free characters?

It becomes pay-as-you-go at the rates above, with no monthly minimum and no seat count. Requests fail with a distinct error rather than silently billing you if there's no payment method on file.

Ready First request in about a minute

Put voice in production.
Keep the economics.

Create a key, send text or audio, and stream the result back into your product.