> ## Documentation Index
> Fetch the complete documentation index at: https://typecast.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Timestamps & captions

The CLI can call Typecast Timestamp TTS and save alignment data alongside the generated audio. Use this when an agent needs subtitles for Shorts, caption timing for social video, karaoke-style highlights, or lip-sync metadata.

## Generate subtitles

```bash
# Save audio and SRT subtitles
cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.srt

# Save audio and WebVTT subtitles
cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.vtt \
  --timestamp-format vtt
```

When `--timestamp-format` is omitted, CLI infers `srt` or `vtt` from the `--timestamp-out` extension and falls back to `json`.

When tempo or silence duration is adjusted, timestamps already align with the processed audio. Do not subtract removed silence or adjust subtitle times again. See [Configuration](/docs/cli-reference/configuration) for the option.

## Save raw timestamp JSON

```bash
cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.timestamps.json
```

JSON is useful when another tool will create captions, animate text, or align visuals manually.

## Choose the caption unit

Cast v1.0.12 or later supports `--caption-unit sentence|word|char` for SRT and WebVTT. The default is `sentence`; choose `word` for one word per cue or `char` for one character per cue.

```bash
cast "Hello, world. This is a test." \
  --out hello.wav \
  --timestamp-out hello.srt \
  --caption-unit word
```

`--timestamp-granularity` chooses the alignment data returned by the API; `--caption-unit` chooses how that data becomes subtitle cues. Word and character caption units request the matching alignment granularity automatically when `--timestamp-granularity` is omitted. An explicit granularity must match the caption unit or be `both`. Raw JSON keeps its original structure; do not use `--caption-unit word` or `char` with JSON output. For Japanese (`jpn`) or Chinese (`zho`), use `--caption-unit char`.

## Choose granularity

```bash
cast "Hello, world." \
  --out hello.wav \
  --timestamp-out hello.srt \
  --timestamp-granularity both
```

For languages without whitespace between words, such as Japanese (`jpn`) or Chinese (`zho`), use character-level timestamps for usable subtitle timing:

```bash
cast "こんにちは。世界。" \
  --language jpn \
  --out hello.wav \
  --timestamp-out hello.srt
```

## Caption workflow for agents

```text
Create narration audio and captions from script.txt.
Use the CLI.
Write audio to ./video/voiceover.wav.
Write subtitles to ./video/voiceover.srt.
Keep the subtitle file next to the audio file.
```

## Output choices

| Output | Use when |
|--------|----------|
| `.srt` | Video editors, Shorts/Reels/TikTok caption import |
| `.vtt` | Web video players and browser-based previews |
| `.json` | Custom rendering, karaoke highlights, lip-sync, downstream automation |

<Tip>
  For social video, generate captions in the same step as audio. It keeps the final narration and subtitle timing tied to the exact same synthesis result.
</Tip>
