A CLI tool to download YouTube subtitles with intelligent fallbacks and optional translation.
- Smart fallback chain: Official transcripts → Auto-generated captions → Whisper transcription
- Whisper control: Force Whisper for quality, or disable it to avoid API costs
- Multi-language support: Request subtitles in multiple languages simultaneously
- Automatic translation: Translate transcripts to target languages via LLM (OpenAI-compatible)
- Flexible output: SRT, VTT, plain text, or article formats
- Pipe mode: Output to stdout for piping to other tools (like Fabric)
- Batch processing: Process multiple URLs from command line or file
- Cookie support: Bypass YouTube restrictions using browser cookies
- Organized storage: Auto-naming with video title and upload date
- Smart caching: Whisper results cached to disk for retry resilience
# Install globally with uv
uv tool install .
# Or with pip
pip install .
# For development (editable mode)
uv pip install -e .After updating the code, reinstall to pick up changes:
uv pip install -e .This tool relies on:
- Python 3.11+
- yt-dlp – Downloads YouTube audio and extracts metadata (installed automatically; current releases are recommended)
- ffmpeg – Required by yt-dlp for audio extraction
ffmpeg needs to be installed separately via your system package manager:
# macOS
brew install ffmpeg
# Ubuntu/Debian
sudo apt install ffmpeg
# Windows (with winget)
winget install ffmpegInitialize a config file:
yt config initView current configuration (or defaults):
yt config showThe config file is located at ~/.config/yt/config.yaml:
# Target languages for transcripts
languages:
- en
- ja
# Output settings
output:
format: srt # srt, vtt, txt, or article
filename_date: upload # upload, request (today), or none
pipe_mode: false # When true, output transcript to stdout for piping
article:
length: original # original, long, medium, short
metadata: frontmatter # frontmatter, header, footer, none
# Optional: per-language length profiles
# Precedence: CLI --length > exact language key > global length
# length_by_language:
# en: long # English articles use 'long' profile
# ja: medium # Japanese articles use 'medium' profile
# ko: short # Korean articles use 'short' profile
# Deduplication: skip article generation if identical video already processed
# Requires metadata: frontmatter (uses video_id from frontmatter)
# Use --force to bypass deduplication and regenerate
dedup:
enabled: false # Skip generating article if duplicate video_id found
recursive: false # When true, scan subdirectories for duplicates
# Output directories (~ is expanded)
storage:
audio_dir: "~/YouTube Subtitles/Audio"
transcript_dir: "~/YouTube Subtitles/Transcripts"
article_dir: "~/YouTube Subtitles/Articles"
discard_audio: false # Delete audio after Whisper transcription
# Logging
logging:
file: "~/YouTube Subtitles/yt.log" # Log file path (omit to disable)
# YouTube/yt-dlp settings (optional)
youtube:
cookies_from_browser: chrome # or firefox, safari, edge, brave
# cookies_file: ~/cookies.txt # alternative: path to cookies.txt
# player_client: web # force client: web, android, ios, tv
# OpenAI-compatible Whisper API (used when YouTube captions unavailable)
transcription:
base_url: https://api.openai.com/v1/audio/transcriptions
model: whisper-1
api_key_env: OPENAI_API_KEY
use_whisper: auto # auto (fallback), force (always), never (YouTube only)
# OpenAI-compatible Chat API for translation
llm:
base_url: https://api.openai.com/v1/chat/completions
model: gpt-4o
api_key_env: OPENAI_API_KEYExport your API keys before running:
export OPENAI_API_KEY="your-api-key"# Single video
yt "https://www.youtube.com/watch?v=VIDEO_ID"
# Multiple videos
yt "URL1" "URL2" "URL3"
# From a file (one URL per line)
yt --input urls.txt# From Chrome
yt "URL" --cookies-from-browser chrome
# From specific Chrome profile
yt "URL" --cookies-from-browser "chrome:Profile 1"
# From Safari or Firefox
yt "URL" --cookies-from-browser safari
yt "URL" --cookies-from-browser firefox
# Or use exported cookies.txt (Netscape format)
yt "URL" --cookies ~/Downloads/cookies.txtNote: If cookie-authenticated requests fail (e.g., due to YouTube's JavaScript challenges), the tool automatically retries without cookies. This provides the best of both worlds—cookies work when needed, anonymous fallback when they don't.
YouTube client fallback: When YouTube rejects a media URL with HTTP 403,
ytretries the audio download through the Android player client. This is automatic;--player-clientcan still be used to force a specific client.
# SRT format (default)
yt "URL" --format srt
# WebVTT format
yt "URL" --format vtt
# Plain text (no timestamps)
yt "URL" --format txt
# Article format (LLM-generated article from transcript)
yt "URL" --format article
# Article with length control
yt "URL" --format article --length short # Concise summary
yt "URL" --format article --length medium # Standard article
yt "URL" --format article --length long # Comprehensive article
yt "URL" --format article --length original # Full rewrite (default)Articles are saved to ~/YouTube Subtitles/Articles/ as .md files.
You can configure different length profiles for different languages using length_by_language in your config file:
output:
article:
length: original # Global default
length_by_language:
en: long # English articles use 'long' profile
ja: medium # Japanese articles use 'medium' profile
ko: short # Korean articles use 'short' profileLength precedence (highest to lowest):
- CLI
--lengthflag - always takes precedence - Exact language match in
length_by_language - Global
lengthsetting
Example: If config has length: original and length_by_language: {en: long}, then:
yt "URL" --languages en→ useslong(exact match)yt "URL" --languages ja→ usesoriginal(global default)yt "URL" --languages en --length short→ usesshort(CLI override)
Prevent regenerating articles for videos you've already processed:
output:
article:
metadata: frontmatter # Required for dedup
dedup:
enabled: true # Skip if duplicate video_id found
recursive: false # Search subdirectories tooHow it works:
- Requires
metadata: frontmatterto embedvideo_idin generated articles - Before generating an article, scans the article directory for existing files with matching
video_id - If
recursive: true, also searches nested subdirectories - Skips generation if a duplicate is found
- Use
--forceto bypass deduplication and regenerate anyway
Example:
# First run: generates article
yt "https://www.youtube.com/watch?v=abc123" --format article
# Second run: skips (duplicate detected)
yt "https://www.youtube.com/watch?v=abc123" --format article
# Force regeneration
yt "https://www.youtube.com/watch?v=abc123" --format article --forceDeduplication works across all languages - if you've processed a video in English, requesting Japanese will also be skipped (same video_id).
Articles can include video metadata. Configure with output.article_metadata:
| Value | Description |
|---|---|
frontmatter |
YAML frontmatter (default) - works with Obsidian, Jekyll, Hugo |
header |
Visible header block with title, author, links |
footer |
Source citation at the end |
none |
No metadata |
Example with frontmatter:
---
schema_version: 1
title: "Video Title"
author: "Channel Name"
video_id: "dQw4w9WgXcQ"
url: https://www.youtube.com/watch?v=dQw4w9WgXcQ
language: en
upload_date: 2024-12-13
request_date: 2026-01-21
profile: long
---
Article content...The frontmatter includes:
schema_version: Format version (currently 1)video_id: YouTube video ID (used for deduplication)language: Target language codeprofile: Article length profile used (original,long,medium,short)
# Override config languages
yt "URL" --languages en,ja,ko
# Disable translation (save source language only)
yt "URL" --no-translateOutput transcript to stdout for piping to other tools (like Fabric):
# Pipe transcript to another tool
yt "URL" --pipe | fabric --pattern summarize
# Pipe mode without saving files
yt "URL" --pipe --no-save | other-tool
# Get plain text for easier processing
yt "URL" --pipe --format txt | fabric --pattern extract_wisdomIn pipe mode:
- Only the first language is output to stdout (for clean piping)
- Status messages go to stderr
- Transcript content goes to stdout
- Files are still saved for all languages (use
--no-saveto disable)
Control when Whisper transcription is used:
# Auto mode (default): Use YouTube captions when available, fall back to Whisper
yt "URL" --use-whisper auto
# Force mode: Always use Whisper, skip YouTube captions
yt "URL" --use-whisper force
# Never mode: Only use YouTube captions, skip video if none available
yt "URL" --use-whisper never| Mode | Behavior |
|---|---|
auto |
Smart fallback: YouTube captions → Whisper (default) |
force |
Always transcribe via Whisper API (requires API key) |
never |
YouTube captions only, no API costs |
Note: Whisper transcription results are cached to disk. If translation fails afterward, retrying won't re-transcribe the audio.
# Force overwrite existing files
yt "URL" --force
# Delete audio after transcription
yt "URL" --discard-audio
# Use specific config file
yt "URL" --config /path/to/config.yaml
# Force specific YouTube client (for compatibility)
yt "URL" --player-client android
# Enable verbose output
yt "URL" --verbose| Command | Description |
|---|---|
yt <urls> |
Process YouTube URLs and download transcripts |
yt config show |
Display current configuration (or defaults) |
yt config init |
Create a new configuration file |
| Option | Description |
|---|---|
urls |
YouTube URLs (positional arguments) |
--input, -i |
File containing URLs, one per line |
--config |
Path to config.yaml (default: ~/.config/yt/config.yaml) |
--format, -f |
Output format: srt, vtt, txt, or article |
--length |
Article length: original, long, medium, short (for --format article) |
--languages, -l |
Comma-separated target languages (e.g., en,ja,ko) |
--pipe, -p |
Pipe mode: output transcript to stdout |
--no-save |
Don't save files (only useful with --pipe) |
--force |
Overwrite existing output files |
--no-translate |
Skip translation; save only source language |
--discard-audio |
Delete audio file after Whisper transcription |
--use-whisper |
Whisper usage: auto (fallback), force (always), never (skip) |
--cookies |
Path to cookies.txt (Netscape format) |
--cookies-from-browser |
Browser to extract cookies from |
--player-client |
Force YouTube client: web, android, ios, tv |
--verbose, -v |
Enable verbose output for debugging |
Files are saved with the naming pattern:
{YYYY-MM-DD} - {Video Title} [{language}].{ext}
The date prefix is controlled by the output.filename_date config option:
| Value | Description |
|---|---|
upload |
Video upload date (default) - good for archiving |
request |
Today's date - good for personal knowledge management |
none |
No date prefix - just {Video Title} [{language}].{ext} |
Examples with filename_date: upload:
2024-10-26 - How AI Works [en].srt2024-10-26 - How AI Works [ja].srt
Examples with filename_date: request (assuming today is 2026-01-21):
2026-01-21 - How AI Works [en].srt2026-01-21 - How AI Works [ja].srt
Examples with filename_date: none:
How AI Works [en].srtHow AI Works [ja].srt
This project was inspired by the yt helper function in Fabric, an open-source framework for augmenting humans using AI. While Fabric's yt extracts YouTube transcripts for use with AI patterns, this tool extends that concept with:
- Multi-language support with automatic translation
- Whisper API fallback when captions are unavailable
- Article generation from transcripts
- Organized file storage with configurable naming
- Flexible output formats (SRT, VTT, TXT, article)
Works great with Fabric! Use --pipe mode to feed transcripts directly into Fabric patterns:
yt "URL" --pipe --format txt | fabric --pattern extract_wisdom
yt "URL" --pipe | fabric --pattern summarizeMIT