Type it. Hear it.
Twelve natural English voices. Opus for the wire; lossless WAV when you need it.
- REST
- POST /v1/tts
- Stream
- WSS /v1/tts/stream
- Returns
- audio/opus or audio/wav
QuickVoice is a real-time text-to-speech and speech-to-text API for voice agents and live applications. Stream with WebSocket or call plain REST. Pricing starts at $0.0018 per 1,000 characters.
“I found three options that match your request.”
Six featured voices from the wider roster read the same words every time. Hear the model directly, with no montage, music, or post-production hiding the output.
Every voice reads the same line, so you can compare them directly.
Generate speech from your own copy, record a line, or upload audio to transcribe. The signal above switches to the real generated audio while it plays.
The same bearer token authenticates every endpoint. Use REST when the input is ready, or WebSocket when text and audio arrive a chunk at a time.
Twelve natural English voices. Opus for the wire; lossless WAV when you need it.
Multilingual transcription with word-level timestamps for complete files and live audio.
No proprietary client and no new orchestration layer. Use a native adapter where it helps, or call the same REST and WebSocket surface from any language.
Drop real-time STT and TTS into a LiveKit agent.
pip install livekit-plugins-quickdial View plugin on PyPI QuickVoice services for Pipecat voice pipelines.
pip install pipecat-quickdial View plugin on PyPI Bearer auth, JSON in, audio or transcript out.
curl https://api.quickdial.ai/v1/tts Open quickstart Each request stays co-located from authentication through model execution and output. The shorter path reduces network variance and keeps metering tied to the request.
Stateless workers scale horizontally. You do not reserve concurrent channels up front.
Request rate is governed per key, so separate workloads can be isolated cleanly.
Every request reports usage, metered to four decimal places for exact reconciliation.
TTS bills the characters you send. STT bills the characters returned. No seats, monthly minimum, or rounding every request up to the next minute.
Opus. WAV (lossless) at $0.0045 / 1K.
Start generatingBilled on timestamped transcript characters returned.
Start transcribingCompare the same character volume at published list prices. Every rate is shown on one scale with its source beside it.
Directional comparison using published list prices per 1,000 characters, checked August 14, 2026. Enterprise discounts, taxes, and unrelated platform fees are excluded. Sarvam lists ₹30 per 10,000 characters; the USD estimate uses ₹96.24 per $1. Deepgram Aura-2 is compared because Nova-3 is speech-to-text and billed per audio minute.
Audio vendors meter differently. QuickVoice keeps the unit explicit instead of forcing a misleading apples-to-oranges multiplier.
The questions developers ask before they put a voice API into a real product.
Browse all documentationFor text-to-speech, the characters in the text you send, so you can price a request exactly before you make it. For speech-to-text, the characters of the transcript returned, since the input is audio and has no character count. Both are metered to four decimal places and reported per request.
Not today. Text-to-speech ships twelve natural English voices and no cloning. Speech-to-text is the multilingual half. It auto-detects and transcribes several languages including French, German, Spanish, Italian and Portuguese.
No. It's plain REST and WebSocket, so curl or your language's HTTP client is enough. Official plugins exist for LiveKit Agents and Pipecat only because those frameworks want a native adapter. They're a convenience, not a requirement.
Use REST when you have the whole input up front; the response already streams back. Use WebSocket when the input itself arrives incrementally: piping LLM tokens into speech before the sentence finishes, or transcribing a live microphone.
It becomes pay-as-you-go at the rates above, with no monthly minimum and no seat count. Requests fail with a distinct error rather than silently billing you if there's no payment method on file.
Create a key, send text or audio, and stream the result back into your product.