llms.txt Standard Analysis and Generation Skill

SkillFiles & storage

Analyzes and generates llms.txt files -- a lightweight anti-hallucination facts hedge that tells AI systems your canonical business facts and key pages. NOT an AI-visibility or citation lever (debunked by Google, Zyppy, and SE Ranking). Can validate existing llms.txt files or generate new ones from scratch by crawling the site.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the llms.txt Standard Analysis and Generation Skill skill

What this skill tells your AI

The instructions your AI receives, as published by thesmokedev/geo-skills in skills/geo-llmstxt/SKILL.md and read by ahel’s review.

Purpose

This skill handles everything related to the llms.txt convention (proposed by Jeremy Howard in September 2024) -- a single Markdown file at the domain root that states your canonical business facts and points at your most important pages.

Positioning as of August 2026: llms.txt is a 30-minute anti-hallucination facts hedge, NOT an AI-visibility lever. Do not sell it to clients as a ranking or citation play -- the evidence says it does not move citations:

  • Google Search Central, 2026 (official): "You don't need to create new machine readable files... to appear in generative AI search." Google has stated plainly that no AI-specific file is required or rewarded.
  • Zyppy meta-analysis of 54 studies (via digitalapplied.com, Jun 2026): llms.txt scores 2.0 out of 10 -- the lowest of all 23 measured GEO factors.
  • SE Ranking (seranking.com/blog/llms-txt, Nov 2025): found zero correlation between having an llms.txt file and AI citation/visibility.

What llms.txt Is Actually For

What remains is real but modest. A well-crafted llms.txt is worth roughly 30 minutes of effort because it:

  1. Reduces misrepresentation: Key facts (pricing, features, locations, founding date) are stated explicitly in one place, reducing AI hallucination about your business when a system does consult the file.
  2. Controls the narrative: You choose which pages and facts are presented as canonical, rather than leaving inference to a crawler.
  3. Costs almost nothing: It is a single static Markdown file. There is no maintenance burden beyond updating it when facts change.

What it does not do: improve rankings, increase citation rates, or substitute for indexation, crawlability, or content quality. Treat it as hygiene -- like a favicon or a humans.txt with a purpose -- and never present it in an audit as a visibility win.


The llms.txt Specification

File Location

The file MUST be located at the root of the domain:

https://example.com/llms.txt

Format Specification

The file uses Markdown formatting with specific conventions:

# [Site Name]

> [One-sentence description of what the site/business does. Keep under 200 characters.]

## Docs

- [Page Title](https://example.com/page-url): Concise description of what this page covers and why it matters.
- [Another Page](https://example.com/another-page): Description of content.

## Optional

- [Less Critical Page](https://example.com/optional-page): Description.

Detailed Format Rules

1. Title (Required)

# Site Name
  • Must be the first line of the file.
  • Should be the official business/site name.
  • Use the H1 heading format (single #).

2. Description (Required)

> Brief description of the site/business
  • Must appear immediately after the title.
  • Use Markdown blockquote format (>).
  • Keep under 200 characters.
  • Should clearly state what the business does and who it serves.
  • Avoid marketing fluff -- be factual and specific.

3. Main Sections (Required -- at least one)

Use H2 headings (##) to organize pages by category. Common section names:

Section NamePurposeExample Content
## DocsPrimary documentation or key pagesProduct pages, service descriptions, core content
## OptionalSecondary pages worth knowing aboutBlog posts, supplementary resources
## APIAPI documentationAPI reference, authentication guides
## BlogBlog or news contentRecent/popular articles
## ProductsProduct catalogProduct pages, pricing
## ServicesService offeringsService descriptions, process pages
## AboutCompany informationAbout page, team, mission
## ResourcesEducational/reference contentGuides, tutorials, whitepapers
## LegalLegal documentsTerms of service, privacy policy
## ContactContact informationContact page, support channels

4. Page Entries (Required)

Each entry follows the format:

- [Page Title](URL): Description of page content

Rules for page entries:

  • Title: Use the actual page title or a clear descriptive title.
  • URL: Must be a full, absolute URL (not relative paths).
  • Description: 10-30 words describing what the page covers. Be specific about the information available.
  • Order: List pages in order of importance within each section.
  • Limit: Include 10-30 page entries total. Prioritize your most authoritative and useful pages.

5. Key Facts Section (Recommended)

## Key Facts
- Founded in [year] by [founder(s)]
- Headquarters: [City, Country]
- [X] customers/users in [Y] countries
- Key products: [Product A], [Product B], [Product C]
- Industry: [Industry classification]

This section provides quick reference data that AI systems frequently need to answer user queries about your business.

6. Contact Section (Recommended)

## Contact
- Website: https://example.com
- Email: hello@example.com
- Support: support@example.com
- Phone: +1-555-123-4567
- Address: 123 Main St, City, State, ZIP, Country

llms-full.txt (Extended Version)

In addition to llms.txt, sites can provide /llms-full.txt -- an extended version with more detail.

Differences from llms.txt:

Featurellms.txtllms-full.txt
LengthConcise (50-150 lines)Comprehensive (150-500+ lines)
Page entries10-30 key pages30-100+ pages
Descriptions10-30 words per entry30-100 words per entry, may include key facts from each page
AudienceQuick AI comprehensionDeep AI analysis
Sections3-6 sections8-15 sections
Key factsBusiness-level factsPage-level facts and data points

Both files can coexist. The convention is that AI systems that choose to consult llms.txt may optionally follow the link to llms-full.txt for deeper detail -- but note there is no evidence (as of Aug 2026) that major AI platforms fetch either file systematically, which is why this skill frames llms.txt as a facts hedge rather than a visibility lever.


Analysis Mode

When checking an existing llms.txt file:

Step 1: Fetch the File

  1. Use WebFetch to retrieve [domain]/llms.txt.
  2. Also check for [domain]/llms-full.txt.
  3. Record HTTP status code:
    • 200: File exists -- proceed to validation.
    • 404: File does not exist -- recommend generation.
    • 403: File exists but is blocked -- flag as misconfiguration.
    • 301/302: Redirect -- follow and note the redirect.

Step 2: Validate Format

Check each structural element:

ElementCheckSeverity if Missing
H1 TitlePresent, matches business nameCritical
Blockquote descriptionPresent, under 200 chars, factualHigh
At least one H2 sectionPresentCritical
Page entries with URLsAt least 5 entries presentHigh
URLs are absoluteAll URLs use full https:// pathsHigh
URLs are validAll URLs return 200 statusMedium
Descriptions presentEvery entry has a description after the colonMedium
Key Facts sectionPresent with business informationMedium
Contact sectionPresent with at least emailLow
Reasonable length30-200 linesLow
No broken MarkdownProper formatting throughoutMedium

Step 3: Assess Content Quality

Rate the llms.txt on these dimensions:

Completeness (0-100):

  • Does it cover all major site sections visible in the navigation?
  • Are the most important/highest-traffic pages included?
  • Is the Key Facts section present with accurate business data?
  • Does it include recent/updated content?

Accuracy (0-100):

  • Do descriptions accurately reflect page content?
  • Are URLs valid and pointing to the correct pages?
  • Are Key Facts verifiable and current?
  • Is the business description accurate?

Usefulness (0-100):

  • Would an AI system understand the site's purpose from this file alone?
  • Are descriptions specific enough to differentiate pages?
  • Are the most citation-worthy pages highlighted?
  • Is the organization logical and intuitive?

Overall llms.txt Score = (Completeness * 0.40) + (Accuracy * 0.35) + (Usefulness * 0.25)

Step 4: Compare Against Site Content

  1. Crawl the site's main navigation and sitemap.
  2. Identify important pages NOT listed in llms.txt.
  3. Check if any listed URLs are broken or redirected.
  4. Verify that the business description matches current homepage messaging.
  5. Flag stale entries (pages that have been significantly updated since the llms.txt was written).

Generation Mode

When creating a new llms.txt file from scratch:

Generator script: skills/geo/scripts/llmstxt_generator.py automates the crawl-and-assemble flow below. Use it as the fast path for the facts hedge -- its output is a canonical-facts file to keep AI systems from hallucinating your basics, not a ranking play. Budget ~30 minutes total including review.

Step 1: Site Discovery

  1. Fetch the homepage and extract:
    • Site name (from <title>, <meta property="og:site_name">, or H1)
    • Business description (from meta description or hero section)
    • Main navigation links
    • Footer links
  2. Fetch /sitemap.xml to discover all public pages.
  3. Identify the site's primary business type (SaaS, E-commerce, Local, Publisher, Agency).

Step 2: Page Prioritization

Categorize all discovered pages and select the most important ones:

Always Include:

  • Homepage
  • About / Company page
  • Pricing page (if exists)
  • Primary product/service pages (top 3-5)
  • Contact page
  • Documentation landing page (if exists)

Include if High Quality:

  • Top blog posts (by apparent importance, recency, or comprehensiveness)
  • Case studies or customer stories
  • Key resource/guide pages
  • FAQ page
  • Careers page (for large companies)

Skip:

  • Thin category/tag pages
  • Pagination pages
  • Login/signup pages
  • Legal boilerplate (unless specifically relevant)
  • Duplicate or near-duplicate content
  • Pages with minimal unique content

Step 3: Write Descriptions

For each selected page:

  1. Fetch the page content using WebFetch.
  2. Read the H1, meta description, and first 2-3 paragraphs.
  3. Write a description that:
    • Is 10-30 words long
    • States what information is on the page
    • Mentions specific topics, data, or features covered
    • Avoids marketing language ("best," "leading," "revolutionary")
    • Uses factual, informative language

Good description examples:

  • Explains the three pricing tiers (Free, Pro, Enterprise) with feature comparison and annual/monthly costs.
  • Details the company's founding in 2018, team of 45 employees, and office locations in Austin and London.
  • Covers integration setup for Slack, Salesforce, and HubSpot with step-by-step guides and API endpoints.

Bad description examples:

  • Our amazing pricing page! (marketing language, no specifics)
  • Learn more about our company. (too vague)
  • Click here for details. (not descriptive)

Step 4: Compile Key Facts

Gather key business facts from the site:

  • Year founded
  • Founder name(s)
  • Headquarters location
  • Number of employees (if public)
  • Number of customers/users (if public)
  • Key products or services (list top 3-5)
  • Industry classification
  • Notable clients or partnerships (if public)
  • Key differentiators (what makes this business unique)
  • Recent milestones or achievements (last 12 months)

Step 5: Assemble the File

Construct the llms.txt following this template:

# [Site Name]

> [One clear sentence: what the business does, who it serves, and its primary value proposition. Under 200 characters.]

## Docs

- [Most Important Page](https://example.com/page): Description covering the key content on this page.
- [Second Page](https://example.com/page-2): Description of this page's content and value.
- [Third Page](https://example.com/page-3): What users and AI systems will find here.

## Products

- [Product A](https://example.com/product-a): Core features, target users, and pricing model for Product A.
- [Product B](https://example.com/product-b): What Product B does and how it differs from Product A.

## Resources

- [Guide Title](https://example.com/guide): Comprehensive guide covering [topic] with [X] sections and practical examples.
- [Blog Post](https://example.com/blog/post): Analysis of [topic] with original data from [source].

## Key Facts

- Founded in [year] by [name(s)]
- Headquartered in [City, Country]
- [Specific metric: e.g., "Serves 10,000+ businesses in 40 countries"]
- [Key differentiator: e.g., "Only platform offering real-time X and Y integration"]
- Industry: [Classification]

## Contact

- Website: https://example.com
- Email: [primary contact email]
- Support: [support URL or email]

Step 6: Validate the Generated File

Before outputting:

  1. Verify all URLs are reachable (200 status).
  2. Confirm total entry count is between 10-30.
  3. Check that no description exceeds 50 words.
  4. Verify the overall file length is 50-150 lines.
  5. Ensure Markdown formatting is clean and consistent.

Output Format

For Analysis Mode

Generate GEO-LLMSTXT-ANALYSIS.md:

# llms.txt Analysis: [Domain]

**Analysis Date:** [Date]
**llms.txt Status:** [Found at URL / Not Found / Error]
**llms-full.txt Status:** [Found / Not Found]

---

## Overall llms.txt Score: [X]/100

| Dimension | Score |
|---|---|
| Completeness | [X]/100 |
| Accuracy | [X]/100 |
| Usefulness | [X]/100 |

---

## Format Validation

| Element | Status | Notes |
|---|---|---|
| H1 Title | [Pass/Fail] | [Notes] |
| Description blockquote | [Pass/Fail] | [Notes] |
| H2 Sections | [Pass/Fail] | [X sections found] |
| Page entries | [Pass/Fail] | [X entries found] |
| URL validity | [Pass/Fail] | [X broken URLs] |
| Entry descriptions | [Pass/Fail] | [X missing descriptions] |
| Key Facts | [Pass/Fail] | [Notes] |
| Contact section | [Pass/Fail] | [Notes] |

---

## Missing Pages

These important pages were found on the site but not in llms.txt:

1. [Page Title](URL) -- [Why it should be included]
2. [Page Title](URL) -- [Why it should be included]

## Improvement Recommendations

1. [Specific recommendation]
2. [Specific recommendation]
3. [Specific recommendation]

## Suggested Updated llms.txt

[Complete rewritten llms.txt file if significant improvements are needed]

For Generation Mode

Output the complete llms.txt file content, ready to be saved to the site's root directory. Also output a brief GEO-LLMSTXT-GENERATION.md report explaining:

  • How many pages were discovered and how many were selected
  • The prioritization rationale
  • Any pages that were borderline (might add later)
  • Recommended update frequency (e.g., monthly for active blogs, quarterly for stable sites)

Best Practices Reference

  1. Update regularly. If your site publishes weekly blog posts, update llms.txt monthly. If your product changes quarterly, update after each release.
  2. Lead with your strongest content. The first entries in each section should be your most authoritative, comprehensive pages.
  3. Be specific in descriptions. "Comprehensive 3,000-word guide to React Server Components with code examples" is far more useful than "React guide."
  4. Include your differentiators. If your site has unique data, original research, or exclusive features, highlight these in descriptions and Key Facts.
  5. Keep it concise. The llms.txt should be scannable in under 60 seconds. Save detail for llms-full.txt.
  6. Use absolute URLs. Always include the full https:// URL, never relative paths.
  7. Test after deployment. After uploading, verify the file is accessible at https://yourdomain.com/llms.txt with no redirects.
  8. Coordinate with robots.txt. Ensure pages listed in llms.txt are not blocked in robots.txt for AI crawlers.
  9. Mirror your site structure. Section names in llms.txt should roughly correspond to your main navigation categories.
  10. Avoid sensitive pages. Do not include internal tools, admin panels, or pages with sensitive information.

Signals

GitHub stars
22
Forks
6
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
geo-llmstxt
Source
github.com/thesmokedev/geo-skills