Text To Speech with Timestamps
Generate speech from text and return word/character-level timestamps aligned with the audio. Useful for subtitle sync, per-character highlight animation, and speech-region visualization.
The request body matches the standard /v1/text-to-speech endpoint (voice_id, text, model, language, prompt, output). Instead of raw audio bytes, this endpoint returns a JSON object containing base64-encoded audio plus words and characters arrays.
Use the optional granularity query parameter to return only word-level or only character-level timestamps and reduce payload size.
Language note. For languages that do not use whitespace between words — such as Japanese (
jpn) and Chinese (zho) — word-level alignment collapses the entire sentence into a single "word". For those languages, always requestgranularity=charto receive usable per-character timestamps.
See Listing all voices for available voices.
/v1/text-to-speech/with-timestampsAuthorizations
X-API-KEYstringheaderrequiredAPI key for authentication. You can obtain an API key from the Typecast API Console.
Query parameters
Filter for which timestamp arrays to return.
- Omitted: returns both
wordsandcharacters. word: returnswordsonly (charactersis null).char: returnscharactersonly (wordsis null).
Languages without whitespace (e.g., jpn, zho): word alignment yields a single segment covering the whole sentence, so use char to obtain meaningful timestamps.
wordcharText-to-speech request parameters
Text to convert to speech. Minimum 1 character, maximum 2000 characters. Credits consumed based on text length. Supports multiple languages including English, Korean, Japanese, and Chinese. Special characters and punctuation are handled automatically.
Voice model to use for speech synthesis.
- ssfm-v30: Latest model with improved prosody and additional emotion presets (recommended)
- ssfm-v21: Stable production model with reliable quality
ssfm-v30ssfm-v21Audio output settings including volume (0-200), pitch (-12 to +12 semitones), tempo (0.5x to 2.0x), and format (wav/mp3) for controlling the final audio characteristics
Use remove_silence_ms (integer, 0–1000 ms) to shorten detected silence.
Show nested model
Adjusts the relative volume of the output audio: 0 (completely silent), 50 (half volume), 100 (standard volume, default), 150 (50% louder than standard), 200 (maximum volume, twice as loud as standard).
Since this only scales the existing volume, using volume can amplify the loudness differences between voices if they have different baseline levels. For consistent output across all clips, use target_lufs instead.
- Note: This parameter cannot be used simultaneously with the
target_lufsparameter.
Required range: 0 <= x <= 200
Adjusts the pitch in semitones to affect perceived gender and age: -12 (one octave lower, deeper voice), -6 (half octave lower), 0 (original pitch, default), +6 (half octave higher), +12 (one octave higher, higher voice)
Controls speech speed: 0.5 (half speed, very slow and clear), 0.75 (slightly slower than normal), 1.0 (normal speaking speed, default), 1.5 (50% faster than normal), 2.0 (double speed, very fast speech)
Sets the target absolute loudness (LUFS) for the output audio. This normalizes all generated voices to a consistent volume level, regardless of the original source's loudness. Values closer to 0 are louder, while values closer to -70 are quieter.
- Required range: -70 <= x <= 0
- Recommended values: -14 (common streaming standard), -23 (broadcast standard)
- Note: This parameter cannot be used simultaneously with the
volumeparameter. Usetarget_lufsfor consistent absolute loudness across different clips, or usevolumefor traditional relative scaling.
Output audio format.
WAV format:
- Uncompressed PCM audio
- 16-bit depth, mono channel, 44100 Hz sample rate
- Higher quality, larger file size
- Recommended for professional audio production
MP3 format:
- Compressed MPEG Layer III audio
- 320 kbps bitrate, 44100 Hz sample rate
- Smaller file size
- Recommended for web streaming and distribution
wavmp3When enabled, shortens detected silences longer than the specified duration to that duration. The value is in milliseconds (ms). The value is the duration of silence to retain, not the amount to remove.
Accepted values:
- An integer from 0 to 1000, with a recommended range of 0 to 200.
- Omitted or
null: silence removal is disabled. 0: removes detected silence longer than 0 ms. This does not disable the feature.- Booleans, strings, fractional values, and out-of-range values are invalid.
Example: 100 shortens detected silences longer than 100 ms to 100 ms. Lower values include shorter silence segments for removal and leave less silence in each affected segment.
Emotion and style settings for the generated speech, including emotion type (happy/sad/angry/normal) and intensity (0.0 to 2.0) to control the emotional expression
SmartPrompt (ssfm-v30)
Text that comes AFTER the text field in TTSRequest. Provides forward context for emotion inference.
The model analyzes the flow: previous_text → text (synthesized) → next_text
- Maximum 2000 characters
- Helps the model anticipate emotional transitions
- Leave empty if no following context is available
Discriminator field to identify the prompt type. Must be set to "smart" for context-aware emotion inference.
Text that comes BEFORE the text field in TTSRequest. Provides backward context for emotion inference.
The model analyzes the flow: previous_text → text (synthesized) → next_text
- Maximum 2000 characters
- Helps the model understand emotional build-up and context
- Leave empty if no preceding context is available
PresetPrompt (ssfm-v30)
Discriminator field to identify the prompt type. Must be set to "preset" for preset-based emotion control.
Emotion preset to apply to the generated speech.
Supported emotions: normal, happy, sad, angry, whisper, toneup, tonedown
Check available emotions for each voice in the models field of the GET /v3/voices API response.
normalsadhappyangrywhispertoneuptonedownPrompt (ssfm-v21)
Emotion preset to apply.
Supported emotions for ssfm-v21: normal, happy, sad, angry
Check available emotions for each voice in the models field of the GET /v3/voices API response.
Language code following ISO 639-3 standard. Case-insensitive (both "ENG" and "eng" are accepted). If not provided, will be auto-detected based on text content.
ssfm-v30 Supported Languages (37)
| Code | Language |
|---|---|
| ARA | Arabic |
| IND | Indonesian |
| POR | Portuguese |
| BEN | Bengali |
| ITA | Italian |
| RON | Romanian |
| BUL | Bulgarian |
| JPN | Japanese |
| RUS | Russian |
| CES | Czech |
| KOR | Korean |
| SLK | Slovak |
| DAN | Danish |
| MSA | Malay |
| SPA | Spanish |
| DEU | German |
| NAN | Min Nan |
| SWE | Swedish |
| ELL | Greek |
| NLD | Dutch |
| TAM | Tamil |
| ENG | English |
| NOR | Norwegian |
| TGL | Tagalog |
| FIN | Finnish |
| PAN | Punjabi |
| THA | Thai |
| FRA | French |
| POL | Polish |
| TUR | Turkish |
| HIN | Hindi |
| UKR | Ukrainian |
| VIE | Vietnamese |
| HRV | Croatian |
| YUE | Cantonese |
| ZHO | Chinese |
| HUN | Hungarian |
ssfm-v21 Supported Languages (27)
| Code | Language |
|---|---|
| ARA | Arabic |
| IND | Indonesian |
| RON | Romanian |
| BUL | Bulgarian |
| ITA | Italian |
| RUS | Russian |
| CES | Czech |
| JPN | Japanese |
| SLK | Slovak |
| DAN | Danish |
| KOR | Korean |
| SPA | Spanish |
| DEU | German |
| MSA | Malay |
| SWE | Swedish |
| ELL | Greek |
| NLD | Dutch |
| TAM | Tamil |
| ENG | English |
| POL | Polish |
| TGL | Tagalog |
| FIN | Finnish |
| POR | Portuguese |
| UKR | Ukrainian |
| FRA | French |
| HRV | Croatian |
| ZHO | Chinese |
Timestamp endpoint note. For languages without inter-word whitespace — Japanese (
jpn) and Chinese (zho) — word-level alignment collapses the whole sentence into a single segment. Always pair these languages withgranularity=charto receive usable per-character timestamps.
Voice identifier. Two prefixes are supported:
tc_— Built-in Typecast voices (e.g.,tc_60e5426de8b95f1d3000d7b5). See Listing all voices for available IDs.uc_— Custom voices created via Instant cloning (e.g.,uc_64a1b2c3d4e5f6a7b8c9d0e1). Only the owner of a cloned voice can use it.
Case-sensitive: must use lowercase prefix.
Response
200Success - Returns base64 audio and timestampsapplication/json
Word-level timestamps (with attached punctuation). null when the request uses granularity=char.
Character-level timestamps (including punctuation and whitespace). null when the request uses granularity=word.
Audio encoding format of the bytes in audio — either wav or mp3, mirroring the request's output.audio_format.
wavmp3