Every PlayHT feature you can click through in the dashboard — text to speech, voice cloning, dubbing, voice isolation — is backed by a REST API you can call directly from your own code. This guide walks through the fastest path from “no API key” to “playable audio file,” using the text-to-speech endpoints as the example. The same patterns apply across the rest of the platform.
1. Get your API key
Every request to PlayHT's API needs credentials tied to your account. You'll find them in your account settings:
- Log in to your PlayHT account.
- Open Settings → API Access.
- Copy your key and store it somewhere safe — a secrets manager or environment variable, not a committed file.
If you don't have an account yet, you can create one and generate a key in a couple of minutes. New accounts include enough credit to work through this entire guide without adding a card.
2. Authenticate every request
PlayHT's API uses a bearer token over HTTPS. Set your key in the Authorization header on every call:
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json
Requests without a valid header return a 401. There's no session or cookie state to manage — every request is authenticated independently, which makes it straightforward to call the API from a server, a serverless function, or a background worker.
3. Make your first request
The core text-to-speech endpoint takes a block of text, a voice, and an output format, and returns generated audio. Here's a minimal example with curl:
curl -X POST https://api.playht.example/v2/tts \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Welcome to PlayHT. This is your first generated voice.",
"voice": "en-US-cora",
"output_format": "mp3",
"speed": 1.0,
"sample_rate": 24000
}'
A few fields worth knowing:
- text — plain text or SSML, depending on the endpoint. Long-form content can be sent as a single request; the API handles chunking internally.
- voice — the identifier of a stock or cloned voice available on your account. A separate endpoint lists every voice you have access to.
- output_format —
mp3,wav, andoggare typically supported; choose based on what your playback pipeline expects.
4. Handling the response: polling vs. streaming
PlayHT exposes text-to-speech in two shapes, and picking the right one matters more than almost any other decision you'll make while integrating the API.
Job-based generation (polling)
A standard request like the one above kicks off a generation job and returns a job ID right away. Poll a status endpoint until the job reports completed, then fetch the resulting audio URL:
async function generateAndWait(text) {
const job = await fetch("https://api.playht.example/v2/tts", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
body: JSON.stringify({
text,
voice: "en-US-cora",
output_format: "mp3"
})
}).then(r => r.json());
let status = job;
while (status.status !== "completed") {
await new Promise(res => setTimeout(res, 1000));
status = await fetch(`https://api.playht.example/v2/tts/${job.id}`, {
headers: { "Authorization": "Bearer YOUR_API_KEY" }
}).then(r => r.json());
}
return status.output_url;
}
This model is ideal when you're generating audio in bulk — narrating articles, building an audiobook, producing ad variations — and don't need the first byte of audio the instant you send the request.
Streaming generation
For anything interactive — a voice agent, a live assistant, a real-time dubbing preview — waiting on a full job round-trip adds latency your users will notice. PlayHT's streaming endpoint opens a connection and returns audio chunks as they're generated, so playback can begin before generation finishes:
const response = await fetch("https://api.playht.example/v2/tts/stream", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
body: JSON.stringify({
text: "This plays back as soon as the first chunk lands.",
voice: "en-US-cora",
output_format: "mp3"
})
});
const reader = response.body.getReader();
// Feed each chunk straight into your audio player or a socket
// as it arrives, rather than waiting for the response to close.
5. Choosing the right endpoint for your use case
A rough rule of thumb we give developers:
- Batch content generation (articles, e-learning modules, IVR prompts, audiobooks) — use the job-based endpoint. Fire off requests in parallel, poll or use webhooks, and optimize for throughput rather than time-to-first-byte.
- Real-time voice agents and live interaction (conversational agents, live captions read aloud, in-call voice changer) — use the streaming endpoint and start playback on the first chunk. Every millisecond between “user finishes speaking” and “agent starts responding” is felt directly by the person on the other end.
Mixing both in one product is common: a voice agent might stream its live responses in the moment, then use the job-based endpoint overnight to pre-render a set of standard prompts.
6. A note on rate limits
Rate limits are enforced per API key and vary by plan — check the response headers on any call for your current limit and remaining quota. If you exceed your limit, the API returns a 429 with a Retry-After value. A few practical habits:
- Back off and retry using the
Retry-Afterheader rather than a fixed delay. - Batch and queue bulk jobs instead of firing large numbers of requests at once.
- Reuse a single streaming connection for a conversation turn instead of opening a new one per sentence.
If you're building something with sustained high volume — a call center integration, a media pipeline processing hundreds of hours of content — it's worth reviewing a plan with higher limits before you hit a wall in production. See pricing for current tiers.
Next steps
That's enough to go from a fresh account to streaming or job-based audio in production code. From here:
- Browse the full API reference for every endpoint, parameter, and error code.
- Read more about the underlying text-to-speech engine and what shapes how natural a voice sounds at this level of detail.
- If you're still deciding on a plan, compare limits and included credits on the pricing page.
Once your first request returns audio, the rest is just building — swap in your own voices, wire up webhooks, and scale from a demo script to a production pipeline.