CLI

Timestamps & captions

Generate timestamp alignment data, SRT subtitles, or WebVTT captions with the CLI.

The CLI can call Typecast Timestamp TTS and save alignment data alongside the generated audio. Use this when an agent needs subtitles for Shorts, caption timing for social video, karaoke-style highlights, or lip-sync metadata.

Generate subtitles

# Save audio and SRT subtitles
cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.srt

# Save audio and WebVTT subtitles
cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.vtt \
  --timestamp-format vtt

When --timestamp-format is omitted, CLI infers srt or vtt from the --timestamp-out extension and falls back to json.

When tempo or silence duration is adjusted, timestamps already align with the processed audio. Do not subtract removed silence or adjust subtitle times again. See Configuration for the option.

Save raw timestamp JSON

cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.timestamps.json

JSON is useful when another tool will create captions, animate text, or align visuals manually.

Choose the caption unit

Cast v1.0.12 or later supports --caption-unit sentence|word|char for SRT and WebVTT. The default is sentence; choose word for one word per cue or char for one character per cue.

cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.srt \
  --caption-unit word

--timestamp-granularity chooses the alignment data returned by the API; --caption-unit chooses how that data becomes subtitle cues. Word and character caption units request the matching alignment granularity automatically when --timestamp-granularity is omitted. An explicit granularity must match the caption unit or be both. Raw JSON keeps its original structure; do not use --caption-unit word or char with JSON output. For Japanese (jpn) or Chinese (zho), use --caption-unit char.

Choose granularity

cast "Hello, world." \
  --out hello.wav \
  --timestamp-out hello.srt \
  --timestamp-granularity both

For languages without whitespace between words, such as Japanese (jpn) or Chinese (zho), use character-level timestamps for usable subtitle timing:

cast "こんにちは。世界。" \
  --language jpn \
  --out hello.wav \
  --timestamp-out hello.srt

Caption workflow for agents

Create narration audio and captions from script.txt.
Use the CLI.
Write audio to ./video/voiceover.wav.
Write subtitles to ./video/voiceover.srt.
Keep the subtitle file next to the audio file.

Output choices

OutputUse when
.srtVideo editors, Shorts/Reels/TikTok caption import
.vttWeb video players and browser-based previews
.jsonCustom rendering, karaoke highlights, lip-sync, downstream automation
⌘I