The CLI can call Typecast Timestamp TTS and save alignment data alongside the generated audio. Use this when an agent needs subtitles for Shorts, caption timing for social video, karaoke-style highlights, or lip-sync metadata.
Generate subtitles
# Save audio and SRT subtitles
cast "Hello, world. This is a test." \
--out hello.wav \
--timestamp-out hello.srt
# Save audio and WebVTT subtitles
cast "Hello, world. This is a test." \
--out hello.wav \
--timestamp-out hello.vtt \
--timestamp-format vtt
When --timestamp-format is omitted, CLI infers srt or vtt from the --timestamp-out extension and falls back to json.
When tempo or silence duration is adjusted, timestamps already align with the processed audio. Do not subtract removed silence or adjust subtitle times again. See Configuration for the option.
Save raw timestamp JSON
cast "Hello, world. This is a test." \
--out hello.wav \
--timestamp-out hello.timestamps.json
JSON is useful when another tool will create captions, animate text, or align visuals manually.
Choose the caption unit
Cast v1.0.12 or later supports --caption-unit sentence|word|char for SRT and WebVTT. The default is sentence; choose word for one word per cue or char for one character per cue.
cast "Hello, world. This is a test." \
--out hello.wav \
--timestamp-out hello.srt \
--caption-unit word
--timestamp-granularity chooses the alignment data returned by the API; --caption-unit chooses how that data becomes subtitle cues. Word and character caption units request the matching alignment granularity automatically when --timestamp-granularity is omitted. An explicit granularity must match the caption unit or be both. Raw JSON keeps its original structure; do not use --caption-unit word or char with JSON output. For Japanese (jpn) or Chinese (zho), use --caption-unit char.
Choose granularity
cast "Hello, world." \
--out hello.wav \
--timestamp-out hello.srt \
--timestamp-granularity both
For languages without whitespace between words, such as Japanese (jpn) or Chinese (zho), use character-level timestamps for usable subtitle timing:
cast "こんにちは。世界。" \
--language jpn \
--out hello.wav \
--timestamp-out hello.srt
Caption workflow for agents
Create narration audio and captions from script.txt.
Use the CLI.
Write audio to ./video/voiceover.wav.
Write subtitles to ./video/voiceover.srt.
Keep the subtitle file next to the audio file.
Output choices
| Output | Use when |
|---|---|
.srt | Video editors, Shorts/Reels/TikTok caption import |
.vtt | Web video players and browser-based previews |
.json | Custom rendering, karaoke highlights, lip-sync, downstream automation |