Agentic Web Scraper Expert (2026 Edition)
SkillWeb & browsingSmart agentic web data extraction with multi-strategy scraping (Crawl4AI v4, Firecrawl), LLM extraction loops, anti-bot bypass, and structured export / Ekstraksi data web cerdas dan agentic dengan scraping multi-strategi (Crawl4AI v4, Firecrawl), ekstraksi LLM, bypass anti-bot, dan ekspor terstruktur.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Agentic Web Scraper Expert (2026 Edition) skill
What this skill tells your AI
The instructions your AI receives, as published by roedyrustam/vibes-plug in skills/web-scraper/SKILL.md and read by ahel’s review.
English | Bahasa Indonesia
English
Orchestration & Integration
Connects and orchestrates with relevant domain skills like browser-automation-expert, ai-llm-integration-expert, brainstorming, and zero-to-prod-orchestrator to ensure cohesive agentic execution.
Description
Advanced Agentic Web Scraping utilizing modern multi-strategy data extraction. Leverages Crawl4AI v4 and Firecrawl to convert raw DOMs into LLM-friendly Markdown. Implements Agentic Extraction loops where the LLM guides the scraper dynamically based on page state. Incorporates strategies for bypassing anti-bot measures (Cloudflare Turnstile, Datadome) and navigating dynamic Shadow DOMs.
Trigger Conditions
- Extracting structured data from websites for analysis, training data, or content pipelines.
- Scraping dynamic JavaScript-rendered pages and complex SPAs.
- Converting web pages to clean Markdown for LLM context or RAG pipelines.
- Dealing with anti-bot protections or complex Shadow DOM architectures during scraping.
- Implementing an automated agentic data extraction loop.
Extracting DOM into LLM-Friendly Markdown
Use Crawl4AI v4 for high-performance async extraction and Firecrawl for seamless LLM-ready conversion.
Crawl4AI v4 (Async Python):
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
async def extract_markdown(url: str):
config = BrowserConfig(headless=True, bypass_csp=True)
run_config = CrawlerRunConfig(
cache_mode=CacheMode.ENABLED,
remove_overlay_elements=True,
word_count_threshold=50
)
async with AsyncWebCrawler(config=config) as crawler:
result = await crawler.arun(url=url, config=run_config)
# Returns clean, AI-optimized markdown ready for LLM consumption
return result.markdown.fit_markdown
Firecrawl (Managed API):
from firecrawl import FirecrawlApp
from pydantic import BaseModel
app = FirecrawlApp(api_key="fc-xxxx")
class ExtractionSchema(BaseModel):
title: str
content: str
key_metrics: list[str]
# Single API call to extract structured data based on JSON schema
result = app.scrape_url(
"https://example.com/data",
formats=["extract", "markdown"],
extract={"schema": ExtractionSchema.model_json_schema()}
)
print(result.markdown) # Clean markdown
print(result.extract) # Structured JSON
Anti-Bot Bypass & Shadow DOMs
Scraping modern web apps requires bypassing anti-bot measures like Cloudflare Turnstile and Datadome, as well as accessing deeply nested elements.
- Anti-Bot Bypass (Cloudflare Turnstile, Datadome):
- Residential Proxies: Rotate high-quality residential IPs to avoid datacenter IP bans.
- Browser Fingerprinting: Use tools like
playwright-stealthor specialized stealth browsers (e.g., Undetected ChromeDriver, Curl-Impersonate) to mask automated fingerprints (WebGL, Canvas, User-Agent). - Human-like Interaction: Introduce random delays, simulate realistic mouse movements, and handle CAPTCHAs via third-party solving services only when necessary.
- Dynamic Shadow DOMs:
- Use CSS piercing selectors or JavaScript execution to penetrate the Shadow Root.
- Example (Playwright):
await page.locator('my-web-component >> css=.internal-element').text_content() - Recursively traverse the DOM tree injecting scripts to extract content from encapsulated components.
Agentic Extraction Loops
Implement an autonomous loop where an LLM guides the scraper based on the current page state, rather than relying on brittle CSS selectors.
- Observe: The scraper extracts the current DOM into clean Markdown.
- Analyze: The LLM analyzes the Markdown to identify necessary data or the next interaction step (e.g., "Click the 'Load More' button").
- Act: The LLM issues a command (extract data, navigate, click, fill form).
- Loop: Repeat until the extraction goal is met.
async def agentic_scrape_loop(url: str, goal: str):
current_url = url
while True:
markdown_content = await extract_markdown(current_url)
# LLM analyzes state and decides next action
action = await llm_decide_action(markdown_content, goal)
if action.type == "COMPLETE":
return action.extracted_data
elif action.type == "CLICK":
await click_element(action.target_selector)
elif action.type == "NAVIGATE":
current_url = action.new_url
Ethical Scraping Checklist
- Check
robots.txtand respectDisallowrules. - Implement rate limiting.
- Use descriptive
User-Agentheaders. - Do not scrape personal/private data without consent.
Bahasa Indonesia
Integrasi Orkestrasi
Terhubung dan mengorkestrasi skill domain yang relevan seperti browser-automation-expert, ai-llm-integration-expert, brainstorming, dan zero-to-prod-orchestrator untuk memastikan eksekusi agentic yang kohesif.
Deskripsi
Scraping Web Agentic tingkat lanjut menggunakan ekstraksi data multi-strategi modern. Memanfaatkan Crawl4AI v4 dan Firecrawl untuk mengubah DOM mentah menjadi Markdown yang ramah LLM. Mengimplementasikan loop Ekstraksi Agentic di mana LLM memandu scraper secara dinamis berdasarkan status halaman. Menggabungkan strategi untuk melewati tindakan anti-bot (Cloudflare Turnstile, Datadome) dan menavigasi Shadow DOM yang dinamis.
Kondisi Pemicu
- Mengekstrak data terstruktur dari situs web untuk analisis, data pelatihan, atau pipeline konten.
- Scraping halaman yang dirender JavaScript secara dinamis dan SPA kompleks.
- Mengonversi halaman web menjadi Markdown bersih untuk konteks LLM atau pipeline RAG.
- Menghadapi perlindungan anti-bot atau arsitektur Shadow DOM yang kompleks saat scraping.
- Mengimplementasikan loop ekstraksi data agentic otomatis.
Mengekstrak DOM menjadi Markdown Ramah LLM
Gunakan Crawl4AI v4 untuk ekstraksi async berperforma tinggi dan Firecrawl untuk konversi siap LLM yang mulus. (Lihat contoh kode di bagian bahasa Inggris).
Bypass Anti-Bot & Shadow DOM
- Bypass Anti-Bot (Cloudflare Turnstile, Datadome):
- Proxy Residensial: Rotasi IP residensial berkualitas tinggi untuk menghindari pemblokiran IP datacenter.
- Browser Fingerprinting: Gunakan alat seperti
playwright-stealthatau browser stealth khusus untuk menyembunyikan sidik jari otomatis. - Interaksi Mirip Manusia: Tambahkan penundaan acak, simulasikan gerakan mouse yang realistis.
- Shadow DOM Dinamis:
- Gunakan selektor penembus CSS atau eksekusi JavaScript untuk menembus Shadow Root.
- Telusuri pohon DOM secara rekursif dengan menyuntikkan skrip untuk mengekstrak konten.
Loop Ekstraksi Agentic
Implementasikan loop otonom di mana LLM memandu scraper berdasarkan status halaman saat ini, bukan bergantung pada selektor CSS yang rentan rusak.
- Observasi: Scraper mengekstrak DOM saat ini menjadi Markdown yang bersih.
- Analisis: LLM menganalisis Markdown untuk mengidentifikasi data yang diperlukan atau langkah interaksi selanjutnya (misal: "Klik tombol 'Muat Lebih Banyak'").
- Aksi: LLM mengeluarkan perintah (ekstrak data, navigasi, klik, isi form).
- Loop: Ulangi hingga tujuan ekstraksi tercapai.
Checklist Scraping Etis
- Periksa
robots.txtdan hormati aturanDisallow. - Implementasikan rate limiting.
- Gunakan header
User-Agentyang deskriptif. - Jangan scraping data pribadi/privat tanpa izin.
Signals
- GitHub stars
- 50
- Forks
- 10
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
web-scraper-roedyrustam- Source
- github.com/roedyrustam/vibes-plug