intake-prefetch
SkillFiles & storageRun the first part of an online intake (listing + audio acquire) locally when the HF side can't, YouTube bot-checks from HF IPs, flaky Drive/archive.org fetches, wrong-content source files, then hand the run back to the Inspector for GPU align. Includes cleanup.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the intake-prefetch skill
What this skill tells your AI
The instructions your AI receives, as published by qud-technologies/quranic-universal-audio in .claude/skills/intake-prefetch/SKILL.md and read by ahel’s review.
The align pipeline is: plan (list the source) → mint → acquire (HF Job) → align → split → sidecars → assemble. Ref: docs/reference/align-pipeline.md. This skill covers doing plan listing and acquire on this machine, then resuming online.
Env for every snippet: HF_TOKEN from the main checkout's .env; prod bucket hetchyy/quranic-inspector-bucket; prod API https://hetchyy-quranic-universal-audio.hf.space. In Git Bash prefix curl calls to /api/... with MSYS_NO_PATHCONV=1. Work in the session scratchpad, never the repo.
When to use
| Symptom | Cause | Do |
|---|---|---|
Build plan fails could not list … / Sign in to confirm you're not a bot | YouTube blocks HF/AWS IPs (listing and download), all clients; cookies rot within hours | §1 + §2 |
Acquire fails BotCheckError for YouTube sources | same | §2 |
Acquire fails on Drive (quota exceeded, 403) or archive.org 5xx | host throttling; HTTP 5xx is already retried 4× in-job | wait + Retry once; still failing → §2 |
| Coverage report lists a surah as missing / a file "repeats surah N" | a source file holds the wrong surah (e.g. two files with the same audio) | §4 |
The durable fix for YouTube is a residential proxy in Space secret INSPECTOR_YTDLP_PROXY — if set, just Retry online instead.
1. List locally (plan)
Home IPs list fine without cookies:
import dataclasses, json, requests
from qua_shared.schemas import IntakeSource
from services.admin.intake_plan.enumerate import enumerate_source # sys.path: repo root + inspector/
listing = enumerate_source(IntakeSource(method="playlist", playlist_url=URL))
body = {"listing": dataclasses.asdict(listing)}
r = requests.post(f"{API}/api/admin/intake/{RID}/plan", json=body,
headers={"Authorization": f"Bearer {HF_TOKEN}"}) # owner bearer works on intake/align routes
Then in Requests → the plan panel: set includes + identity, press Align. The run starts and acquire fails on the blocked sources — expected; it has now written staging/<slug>/<run_id>/groups.json.
2. Acquire locally
- Download
staging/<slug>/<run_id>/groups.jsoninto a local mount dirM/staging/<slug>/<run_id>/. - Run the real job against that dir (it fetches, encodes to canonical mp3, bakes peaks):
A video that dies mid-stream (SLUG=<slug> RUN_ID=<run_id> INSPECTOR_BUCKET_MOUNT=M PYTHONPATH=<repo root> ACQUIRE_WORKERS=8 python qua_jobs/acquire_audio.pybytes read, more expected) usually works withyt-dlp -f 140; transient 403s just need a re-run (the job skips finished files). - Upload
M/reciters/<slug>/audio/*.mp3(+peaks/*.json.gzfor single-chapter sources) to the same paths in the prod bucket withhuggingface_hub.batch_bucket_files(add=[(local_path, dest)])— pass file paths, not bytes (bytes overwrite is ~25× slower). - Inspector → Retry the run. Acquire sees every file present and skips; align proceeds on GPU.
Slots: a combined source lands in audio/<200 + position among INCLUDED plan entries>.mp3. Changing includes after prefetch shifts slots — re-run §2.
3. Watch
GET /api/admin/reciter/<slug>/align/status (bearer ok) until done/succeeded. A failed stage → read staging/<slug>/<run_id>/{acquire,split}.json, fix, Retry (resumes from the failed stage).
4. Wrong-content source file
Upstream mirrors can share a bad file (e.g. 046 actually holding surah 47). Don't cross-source to YouTube offsets in the manifest — releases need a real per-chapter link:
- Find the right recording (another upload / juz video of the same session); verify by cross-correlation and aligner probe.
- Upload the corrected file to a QUD archive.org item (the built-in browser pane is signed in; the user publishes/approves).
- Align that one chapter and splice it into the bucket files (
detailed,segments,low_confidence_v2,auto_split_v1,chapter_sources,coverage_report,pipeline_meta, edit histories,catalog/audio_manifest/<slug>.jsonwithsource_url+ recomputed checksum). Keep each file's exact JSON format; one bucket write per process; verify by re-download. - DB delivery
chapter_count/total_duration_sec: the PATCH route rejects bearer — use an inspector-admin_bootstrap.run(..., safe_write=True)one-off with--prod --yes-prodandINSPECTOR_DEV_OWNER_HF_ID/_LOGINset (pauses + resumes the Space itself). - Restart the prod Space if you wrote the bucket out-of-band, then check
/api/seg/chapters/<slug>,/api/seg/validate/<slug>(missing_verses0) and the audio proxy.
Cleanup
- Delete the local mount dir and any downloaded source audio / probes from the scratchpad.
- Never leave a cookies file on disk; never extract browser cookies.
- Stop any local helper servers or background downloads you started.
- Don't delete
staging/oraudio/2xx.mp3slots yourself — split removes the slots once every cut succeeds; a failed split keeps them for Retry.
Signals
- GitHub stars
- 42
- Forks
- 5
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
intake-prefetch- Source
- github.com/qud-technologies/quranic-universal-audio