Agent Browser Skill

SkillWeb & browsing

Lets your agent open web pages, click, fill forms, take screenshots, and scrape content via a browser CLI.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Agent Browser Skill skill

About this capability

Browser automation through the repo agent-browser CLI. Explicit helper for navigation, forms, screenshots, scraping, and web-app checks. Prefer Browser Use or Playwright when available. Do NOT load for: sharing URLs, embedding links, or editing screenshot files.

What this skill tells your AI

The instructions your AI receives, as published by chachamaru127/claude-code-harness in skills/agent-browser/SKILL.md and read by ahel’s review.

ブラウザ自動化を行うスキル。agent-browser CLI を使用して、UI デバッグ・検証・自動操作を実行します。


トリガーフレーズ

このスキルは以下のフレーズで自動起動します:

  • 「ページを開いて」「URLを確認して」
  • 「クリックして」「入力して」「フォームに」
  • 「スクリーンショットを撮って」
  • 「UIを確認して」「画面をテストして」
  • "open this page", "click on", "fill the form", "screenshot"

機能詳細

機能詳細
ブラウザ自動化See references/browser-automation.md
AI スナップショットワークフローSee references/ai-snapshot-workflow.md

実行手順

Step 0: agent-browser の確認

# インストール確認
which agent-browser

# 未インストールの場合
npm install -g agent-browser
agent-browser install

Step 1: ユーザーのリクエストを分類

リクエストタイプ対応アクション
URL を開くagent-browser open <url>
要素をクリックスナップショット → agent-browser click @ref
フォーム入力スナップショット → agent-browser fill @ref "text"
状態確認agent-browser snapshot -i -c
スクリーンショットagent-browser screenshot <path>
デバッグagent-browser --headed open <url>

Step 2: AI スナップショットワークフロー(推奨)

ほとんどの操作で、まずスナップショットを取得してから要素参照で操作します:

# 1. ページを開く
agent-browser open https://example.com

# 2. スナップショット取得(AI 向け、インタラクティブ要素のみ)
agent-browser snapshot -i -c

# 出力例:
# - link "Home" [ref=e1]
# - button "Login" [ref=e2]
# - input "Email" [ref=e3]
# - input "Password" [ref=e4]
# - button "Submit" [ref=e5]

# 3. 要素参照で操作
agent-browser click @e2           # Login ボタンをクリック
agent-browser fill @e3 "user@example.com"
agent-browser fill @e4 "password123"
agent-browser click @e5           # Submit

Step 3: 結果の確認

# 現在の状態をスナップショットで確認
agent-browser snapshot -i -c

# または URL を確認
agent-browser get url

# スクリーンショットを取得
agent-browser screenshot result.png

クイックリファレンス

基本操作

コマンド説明
open <url>URL を開く
snapshot -i -cAI 向けスナップショット
click @e1要素をクリック
fill @e1 "text"フォームに入力
type @e1 "text"テキストを入力
press Enterキーを押す
screenshot [path]スクリーンショット
closeブラウザを閉じる

ナビゲーション

コマンド説明
back戻る
forward進む
reloadリロード

情報取得

コマンド説明
get text @e1テキスト取得
get html @e1HTML 取得
get url現在の URL
get titleページタイトル

待機

コマンド説明
wait @e1要素を待機
wait 10001秒待機

デバッグ

コマンド説明
--headedブラウザを表示
consoleコンソールログ
errorsページエラー
highlight @e1要素をハイライト

セッション管理

複数のタブ/セッションを並列管理:

# セッションを指定
agent-browser --session admin open https://admin.example.com
agent-browser --session user open https://example.com

# セッション一覧
agent-browser session list

# 特定セッションで操作
agent-browser --session admin snapshot -i -c

MCP ブラウザツールとの使い分け

ツール推奨度用途
agent-browser★★★第一選択。AI 向けスナップショットが強力
chrome-devtools MCP★★☆Chrome が既に開いている場合
playwright MCP★★☆複雑な E2E テスト

原則: まず agent-browser を試し、うまくいかない場合のみ MCP ツールを使用。


注意事項

  • agent-browser はヘッドレスモードがデフォルト
  • --headed オプションでブラウザを表示可能
  • セッションは明示的に close するまで維持される
  • 認証が必要なサイトはセッションを活用

Signals

GitHub stars
3k
Forks
299
Last commit
Sep 2026
Hacker News mentions
8
Advanced
Catalog kind
skill
Gateway key
agent-browser-chachamaru127
Source
github.com/chachamaru127/claude-code-harness