Skip to main content

How to Use Text to Speech

Neuphonic's primary offering is its text-to-speech technology, which serves as the foundation for various features, including Agents. Visit our Playground to experiment with different models and voices, and then continue reading below to learn how to implement this in code.

note

Don't forget to visit the Quickstart guide to obtain your API Key and get your environment set up.

Speech Synthesis​

You can generate speech using the API in two ways: Server Side Events (SSE) and WebSockets.

SSE is a streaming protocol where you send a single request to our API to convert text into speech. Our API will then stream the generated audio back to you in real-time, providing the lowest possible latency. Below are some examples of how to use this endpoint.

The SDK examples demonstrate how to send a message to the API, receive the audio stream from the server, and play it through your device's speaker.

# Replace <API_KEY> with your actual API key.
# To switch languages, replace the lang_code in the path parameter (e.g., /en) with the desired language code.
curl -N --request POST \
--url https://api.neuphonic.com/sse/speak/en \
--header 'Content-Type: application/json' \
--header 'X-API-KEY: <API_KEY>' \
--header 'Accept: text/event-stream' \
--data '{
"text": "Hello, world!"
}'
warning

The chosen voice needs to be available for the chosen language.

Text-to-Speech Configuration​

The settings for Text-to-Speech generation can include the following parameters.

NameTypeDescription
lang_codestringLanguage code for the desired language. See the full Languages page for all supported codes. Common values: 'en' (English), 'es' (Spanish), 'zh' (Chinese)
voice_idstringThe voice ID for the desired voice. Based on what voice_id you chose different models will be leveraged. Examples: '8e9c4bc8-3979-48ab-8626-df53befc2090'
speedfloatPlayback speed of the audio. Supported values snap to 0.7 (slow), 1.0 (normal), or 1.5 (fast). Any value <1.0 is treated as slow, >1.0 as fast.
sampling_rateintSampling rate of the audio returned from the server. Options: 8000, 16000, 22050, 24000. Default: 24000
encodingstringEncoding of the audio returned from the server. Options: 'pcm_linear', 'pcm_mulaw'
timestampsbooleanReturn word timings alongside the audio, for highlighting text as it is spoken. See Word Timings. Default: false
ipa_keystringUse a regional variant's sounds for <phoneme:...> tags. See Pronunciation. Options: 'es-castilian', 'pt-br'

Word Timings​

Set timestamps=true to get the start time and length of every word alongside the audio, so you can highlight text as it is spoken.

wss://api.neuphonic.com/speak/en?api_key=<API_KEY>&timestamps=true

Messages then carry an alignment next to the audio:

{
"data": {
"audio": "<audio_data_string>",
"sampling_rate": 24000,
"alignment": {
"units": ["The", "lens", "costs", "$42", "today."],
"unitStartTimesMs": [100, 240, 560, 900, 1780],
"unitDurationsMs": [140, 300, 340, 880, 380],
"charStartOffsets": [0, 4, 9, 15, 19],
"charEndOffsets": [3, 8, 14, 18, 25]
}
}
}

The five arrays line up: index i of each describes the same word.

FieldMeaning
unitsThe word, as you sent it
unitStartTimesMsWhen it starts, in milliseconds from the beginning of the audio
unitDurationsMsHow long it lasts, in milliseconds
charStartOffsetsWhere the word starts in your text
charEndOffsetsWhere it ends (exclusive)

Reading them​

Collect entries as they arrive, keyed by index across the five arrays, and highlight whichever one covers the current playback position (`ms >= unitStartTimesMs[i] && ms < unitStartTimesMs[i]

  • unitDurationsMs[i]) by selecting myText.slice(charStartOffsets[i], charEndOffsets[i])`. Times are measured from the start of the stream, so they stay correct whatever your text's length, and some messages carry no entries while others carry several — just keep appending.

Things to expect​

  • Offsets point at your original text, not at expanded speech — $42 comes back as one entry spanning its whole read-out, and 東京 stays 東京, not its pronunciation.
  • Not every character is covered: spaces and most punctuation have no sound of their own.
  • Some voices have no timings. alignment is then null, with alignment_unavailable explaining why, so you can tell "none coming" from "not yet".
  • Every language is supported; for scripts written without spaces (Japanese, Chinese, Thai) a "word" is the smaller unit natural to that script.

Pronunciation​

If a word is said wrongly, look it up in the pronunciation tool on our dashboard to get its correct IPA transcription, then wrap it in a <phoneme:...> tag inline in your text:

"Please welcome <phoneme:ˈdeɪtə>, our guest."

Only the tagged word changes; the rest of the sentence is spoken as normal. Bare, /slashed/, or [bracketed] IPA all work, and ˈ marks the stressed syllable.

Available in ten languages (English, German, Dutch, French, Italian, Spanish, Portuguese, Russian, Hindi and Korean) — use the phonemes of the language you're synthesising, or the request is rejected. For Castilian Spanish or Brazilian Portuguese, add ipa_key=es-castilian or ipa_key=pt-br to use that variant's sounds.

Inline Pauses​

Insert silence anywhere in your text for pacing, using <pause:500ms>, <pause:2s>, or SSML's <break time="500ms"/>:

"Welcome to Neuphonic. <pause:1s> Your audio is being generated."
"And the winner is ... <break time="2s"/> you!"

Billed at 1 character per second of silence.

Providing a context ID​

You can provide a context ID to uniquely identify the audio chunks generated for a given text input. This ID is returned with the corresponding audio, allowing you to track, interrupt, or associate the audio with its original request.

client = Neuphonic()

ws = client.tts.AsyncWebsocketClient()
audio_bytes = bytearray()

async def on_message(message: APIResponse[TTSResponse]):
nonlocal audio_bytes
audio_bytes += message.data.audio
print(message.data.context_id) # here the context_id will be "1"

ws.on(WebsocketEvents.MESSAGE, on_message)

await ws.open()

await ws.send(
{"text": "This is message one example.<STOP>", "context_id": "1" }
)

Resetting the Stream​

Send <RESET> as a message to discard everything currently buffered — any text not yet spoken and the voice's carried-over prosody — without closing the connection. The next text you send starts a fresh utterance, and its word timings (if timestamps=true) start again from zero:

await ws.send('<RESET>')

Use it to interrupt what's currently being generated, e.g. on barge-in, without paying the cost of reopening the websocket.

More Examples​

To see more examples, check out our Python SDK examples and JavaScript SDK examples on GitHub.