Podcast Transcript Fetcher

SkillMedia

Use when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast. Tier 2 (RSS+Groq Whisper) is the recommended approach -- fast, free, and most reliable. Also use when asked to "get transcript", "find episode", "summarize podcast", or "search podcast content". Do not use for general web scraping or non-podcast audio transcription.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Podcast Transcript Fetcher skill

What this skill tells your AI

The instructions your AI receives, as published by varnan-tech/opendirectory in skills/podcast-transcript-fetcher/SKILL.md and read by ahel’s review.

Fetch transcripts from 5 supported podcasts. Tier 2 (RSS+Groq Whisper) is the recommended approach -- fast, free, and the most reliable across all podcasts. Tier 1 free sources are best-effort (limited availability). Tier 3 Taddy API is the premium/commercial option.

Quick Reference

# Get latest episode transcript (auto-detects best method)
python scripts/get_transcript.py "Lenny's Podcast" --latest

# Search by episode title or number
python scripts/get_transcript.py 20vc --episode "Marc Andreessen"
python scripts/get_transcript.py dwarkesh --episode 15

# Force specific method
python scripts/get_transcript.py "cheeky pint" --latest --method whisper
python scripts/get_transcript.py a16z --latest --method taddy

# Save to file
python scripts/get_transcript.py lennys --latest --output transcript.md

# List all supported podcasts
python scripts/get_transcript.py --list-podcasts

Supported Podcasts

PodcastTier 1 (best-effort)Tier 2 RSS+Whisper [RECOMMENDED]Tier 3 Taddy (premium)
Lenny's PodcastGitHub archive (269 transcripts)✅ Substack RSS✅ Covered
Dwarkesh PodcastWebsite scrape + Substack PDF✅ Substack RSS✅ Covered
Cheeky Pint(none)✅ Transistor.fm RSS✅ Covered
20VCSubstack PDF✅ Libsyn RSS✅ Covered
A16z PodcastWebsite scrape✅ Simplecast RSS✅ Covered

Implementation

1. Install Dependencies

# Core (always required)
pip install requests

# Cloud transcription (recommended — fast, free tier)
pip install groq
export GROQ_API_KEY="your-key"  # Get at https://console.groq.com

# Local transcription (free, needs ~5GB RAM)
pip install faster-whisper

# Audio compression (for Groq's 25 MB limit — Windows: winget/scoop)
#   winget install ffmpeg  or  scoop install ffmpeg

# Taddy API (commercial, optional)
export TADDY_API_KEY="your-key"  # Get at https://taddy.org

2. Get a Transcript

The script auto-selects the best method. Tier 2 is the default recommendation:

Tier 1 → Tier 2 (RECOMMENDED) → Tier 3
(best-effort)  (Whisper)  (Taddy API premium)

Tier 1: Free direct sources (best-effort, limited availability)

  • Lenny's: Clones ChatPRD/lennys-podcast-transcripts and searches by title
  • Dwarkesh: Substack PDF scrape
  • 20VC: Substack PDF scrape
  • A16z: Website scrape
  • Cheeky Pint: No Tier 1 sources available
  • Note: Tier 1 sources are best-effort and limited. Tier 2 (RSS+Whisper) is the recommended approach.

Tier 2: RSS + Whisper transcription [RECOMMENDED]

  • Downloads MP3 from podcast RSS feed
  • Compresses if >25 MB (ffmpeg)
  • Transcribes via Groq Whisper API (free tier, ~10s per hour of audio)
  • Fast, free, and works for every podcast in the registry
  • Default recommendation for all use cases

Tier 3: Taddy API (commercial/premium)

  • Requires TADDY_API_KEY ($75/mo+)
  • Use for large-scale or production transcript needs
  • Covers all 5 podcasts with auto-transcription

3. Analyze with AI

Once you have a transcript, pipe it to the agent for analysis:

I have this transcript from [podcast]. Can you:
1. Summarize the key arguments
2. Extract 3 actionable insights
3. Identify any controversial claims
4. Compare with [other podcast] on the same topic

Supported Workflows

Single Episode

ScenarioCommand
Latest episodeget_transcript.py "Lenny's Podcast" --latest
Specific episode by titleget_transcript.py 20vc --episode "Sam Altman"
Episode by numberget_transcript.py dwarkesh --episode 42
Force Whisper transcription (Tier 2, recommended)get_transcript.py a16z --latest --method whisper
Force Taddy API (premium)get_transcript.py lennys --latest --method taddy
Save to Markdownget_transcript.py cheeky-pint --latest --output episode.md
JSON outputget_transcript.py dwarkesh --latest --json

Cross-Podcast Search & Batch

ScenarioCommand
Search all podcasts by keywordget_transcript.py --search "Marc Andreessen"
Search by guest nameget_transcript.py --guest "Sam Altman"
Search within one podcastget_transcript.py "Lenny's Podcast" --search "vibe coding"
Batch-transcribe last N episodesget_transcript.py "Dwarkesh Podcast" --last 5
Search + transcribe top matchesget_transcript.py --search "AI safety" --transcribe
Pipeline with custom countget_transcript.py --search "scaling laws" --transcribe --transcribe-count 5
Filtered search pipelineget_transcript.py "A16z Podcast" --search "crypto" --transcribe

Output Structure

Batch transcription saves to output/ with per-podcast subdirectories:

output/dwarkesh-podcast/Dwarkesh Podcast_2024-01-15_agi-is-still-30-years-away.md
output/20vc/20 Minutes VC (20VC)_2024-03-10_funding-round-analysis.md

Each file includes a YAML frontmatter header:

---
podcast: Dwarkesh Podcast
episode: AGI is still 30 years away
date: 2024-01-15
url: https://...
source: whisper
---

Podcast Registry

The registry at scripts/podcasts.json maps each podcast to its RSS feeds, transcript sources, and API endpoints. To add new podcasts:

{
  "id": "new-podcast",
  "name": "New Podcast",
  "rss": "https://example.com/feed.xml",
  "transcript_sources": {
    "primary": {"type": "website_scrape", "url": "https://example.com"}
  }
}

Troubleshooting

ProblemSolution
"No transcript found"Tier 2 (RSS+Whisper) is the recommended approach. If auto mode fails, try --method whisper to force it.
RSS fetch failsRSS feeds may change; check scripts/podcasts.json for current URLs
Audio download slowLarge MP3s can take minutes on slow connections
Groq rate limitedWait or switch to local faster-whisper
Taddy not returning transcriptsSome episodes lack transcripts; try --method whisper
Podcast not in registryAdd it to scripts/podcasts.json
Unicode error on WindowsFixed: script auto-reconfigures stdout to UTF-8; saved files use UTF-8 encoding
Audio > 25 MB for GroqInstall ffmpeg: winget install ffmpeg (Windows) or brew install ffmpeg (macOS)

RSS Feed Status (as of 2026-06)

PodcastOld Feed (broken)Current Feed
Cheeky Pintfeeds.transistor.fm/the-cheeky-pint (404)feeds.transistor.fm/cheeky-pint-with-john-collison
20VCfeeds.simplecast.com/3GxrMqOd (404)feeds.libsyn.com/61840/rss
A16zfeeds.simplecast.com/0cJfpoz2 (404)feeds.simplecast.com/JGE3yC0V

Common Mistakes

  • Forgetting API keys: Set GROQ_API_KEY in your env or .env file
  • Relying on Tier 1 free sources: Tier 1 is best-effort and limited. Always fall back to Tier 2 (RSS+Whisper) which is the recommended method.
  • Not cloning the Lenny's repo first: The GitHub archive must be cloned locally for Tier 1 to work
  • Using --method taddy without TADDY_API_KEY: Falls through silently; set the key or use auto mode

Signals

GitHub stars
642
Forks
68
Last commit
Aug 2026

ahel review

  • K1binfo
    installs-packages
  • K1binfo
    installs-packages (in scripts/get_transcript.py)
  • K1binfo
    installs-packages (in README.md)
  • K1binfo
    installs-packages (in references/podcasts.md)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Catalog kind
skill
Gateway key
podcast-transcript-fetcher
Source
github.com/varnan-tech/opendirectory