Rawkode Academy Video Importer
SkillMediaImport YouTube videos into Rawkode Academy: run the video-importer, upload R2 assets, verify transcoding, trigger and verify transcription, create content/videos entries, and avoid claiming triggered jobs are complete before artifacts exist.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Rawkode Academy Video Importer skill
What this skill tells your AI
The instructions your AI receives, as published by rawkode-academy/rawkode-academy in .agents/skills/rawkode-academy-video-importer/SKILL.md and read by ahel’s review.
Use this skill when importing a YouTube video into rawkode-academy content, especially when the task mentions projects/rawkode.academy/tasks/video-importer, R2 uploads, transcoding, transcription, captions, or episode descriptions.
Critical Rule
Separate evidence for each stage:
- Import/upload complete means the importer state has
completedand the original R2 assets return200. - Transcoding complete means
https://content.rawkode.academy/videos/<cuid>/stream.m3u8returns200, or Cloud Run shows the execution completed successfully. - Transcription complete means
https://content.rawkode.academy/videos/<cuid>/captions/en.vttreturns200.
Do not say transcoding or transcription completed just because the job/workflow was triggered.
Import
- Start at the repo root and check for unrelated work:
jj status
Preserve unrelated user changes.
- Before running dependency installs, Bun workspace commands, tests, or generated CI checks, synchronize the workspace from the repository root:
cuenv sync -A
-
Extract the video ID from the YouTube URL. For
https://www.youtube.com/watch?v=tKc8Rna3vDQ, usetKc8Rna3vDQ. -
Check importer state:
cd projects/rawkode.academy/tasks/video-importer
cuenv exec -e production uv run python youtube_to_r2.py <youtube-id> --show-state
- Run the importer:
cuenv exec -e production uv run python youtube_to_r2.py <youtube-id>
If the user already downloaded the video, find the file and pass it explicitly:
find /Users/rawkode/Downloads -maxdepth 1 -type f \( -iname '*.mp4' -o -iname '*.mkv' -o -iname '*.mov' -o -iname '*.webm' \) -print
cuenv exec -e production uv run python youtube_to_r2.py <youtube-id> --local-video /absolute/path/to/video.mp4
The importer downloads or prepares the video, extracts MP3 audio, uploads original.mkv, original.mp3, and thumbnail.webp to rawkode-academy-content, triggers Cloud Run transcoding-job, and prints the CUID plus SQL metadata. The downloaded JPEG thumbnail is only a temporary conversion source; thumbnail.webp is the canonical runtime asset. It keeps state in ~/.youtube_to_r2_state/<youtube-id>.json.
If a turn is interrupted, poll the running session if available before starting a new import. Avoid duplicate uploads for the same YouTube ID.
- Verify import/upload:
jq '{cuid, completed_steps, artifacts, title: .video_info.title, duration: .video_info.duration, upload_date: .video_info.upload_date}' /Users/rawkode/.youtube_to_r2_state/<youtube-id>.json
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/original.mkv
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/original.mp3
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/thumbnail.webp
If an older import has both thumbnail.jpg and thumbnail.webp, remove the JPEG after verifying WebP:
cd projects/rawkode.academy/tasks/video-importer
cuenv exec -e production uv run python migrate_thumbnails_to_webp.py --video-id <cuid> --workers 1
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/thumbnail.webp
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/thumbnail.jpg # should return 404
Content
Create the content file from importer output and nearby entries:
- Path:
content/videos/shows/<show>/<year>/<slug>.md id: importer CUIDslug: clean show suffixes like-rawkode-liveunless nearby entries keep thempublishedAt: YouTube upload date as ISO UTC if no better timestamp existsduration: importer duration in secondsshow: matching show ID, commonlyrawkode-liveorcloud-native-compasstype,category,technologies,guests,resources: match nearby entries
If a guest is identified by GitHub handle, use the handle for both filename and id, for example content/people/davidmdm.mdx with id: davidmdm.
If the video references a missing technology, add content/technologies/<id>/index.mdx before referencing it. Keep it concise and use primary project sources.
Transcoding
The importer triggers transcoding but does not prove it completed.
Check the public HLS playlist first:
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/stream.m3u8
If it is 404, check Cloud Run executions:
gcloud run jobs executions list \
--job=transcoding-job \
--region=europe-west2 \
--project=rawkode-academy-production \
--limit=10 \
--format=json
Find the execution whose container env has VIDEO_ID=<cuid>.
status.conditions[]withtype: Completed,status: Truemeans transcoding completed.status.runningCount: 1andCompletedstatusUnknownmeans it is still running.CompletedstatusFalseor failed task counts mean it failed; inspect the execution logs before retrying.
Transcription
The YouTube importer does not trigger transcription. The transcription service is projects/rawkode.academy/platform/transcriptions, a Cloudflare Worker with a Cloudflare Workflow named transcribe.
Before scheduling transcription, make sure the live API can resolve the video, because the workflow calls https://api.rawkode.academy:
curl -fsS https://api.rawkode.academy \
-H 'Content-Type: application/json' \
--data '{"query":"query($id:String!){ videoByID(id:$id){ id title streamUrl thumbnailUrl } }","variables":{"id":"<cuid>"}}'
Trigger transcription:
cd projects/rawkode.academy/platform/transcriptions
cuenv exec -e production bun scripts/schedule_one.ts <cuid> en
Do not pass manual keyterms on the CLI. Transcription keyterms are derived from live content metadata: video title/terms, episode code/terms, show name/terms, guest name/terms, and technology name/terms. If those GraphQL fields or content metadata changed locally, deploy the website/API before deploying or running the transcription Worker, because the workflow queries https://api.rawkode.academy.
Check workflow state:
cuenv exec -e production bunx wrangler workflows instances list transcribe
cuenv exec -e production bunx wrangler workflows instances describe transcribe <workflow-id>
Check artifacts:
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/transcription/deepgram.json
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/captions/en.deepgram.vtt
curl -fsS -I https://content.rawkode.academy/videos/<cuid>/captions/en.vtt
captions/en.vtt is the final caption file used by watch pages.
Known transcription states:
Completedpluscaptions/en.vttreturning200means transcription is done.Waitingcan mean the workflow is between retries; rundescribefor the exact step error.- Transcription uses Cloudflare Workers AI
@cf/deepgram/nova-3, not a direct Deepgram API key. IfASR_PAYMENT_REQUIREDappears, the old direct-Deepgram worker is still running; deploy the current worker and start a fresh workflow.
Episode Description
Only review the transcript after captions/en.vtt exists. Fetch the VTT, skim the major topic flow, then update the video description to be concise and production-facing. Prefer direct summaries over YouTube boilerplate. Keep resources grounded in links mentioned in the episode or obvious primary project links.
Treat generated captions as evidence for the description, not proof that the transcript was manually copyedited. If the VTT still has ASR rough edges, say so plainly unless you actually cleaned and re-uploaded the transcript.
Validation
Run the website content/type check:
cd projects/rawkode.academy/website
bun run astro check
If sandboxed execution fails with listen EPERM for the Cloudflare/Vite inspection port, rerun with approval. Existing Vite dependency-scan noise can appear; use the final exit code and Astro diagnostic summary.
Always report blocked or in-flight work plainly: uploaded, transcode running, transcription waiting, captions missing, etc.
Signals
- GitHub stars
- 32
- Forks
- 12
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
rawkode-academy-video-importer- Source
- github.com/rawkode-academy/rawkode-academy