Generate llms.txt from your website
Enter a domain. The generator finds its sitemap, reads each page's title and description, and gives you an llms.txt draft to edit before you publish it.
Checks robots.txt and XML sitemaps, then reads up to 100 public HTML pages. A scan can take several seconds.
Review before you publish
A sitemap also lists pages an agent has no use for. Remove duplicates and low-value routes such as account or tag pages, and keep the pages that explain the product or help a reader finish a task.
- Review the scan. Edit titles and descriptions, choose a section, and exclude pages that add little context.
- Download the file. Keep the filename exactly
llms.txt. - Publish it at the website root. The public URL should be
https://example.com/llms.txt. - Maintain it with the site. Regenerate it when important pages or canonical URLs change.
Frequently asked questions
How does the generator find pages?
It checks sitemap declarations in robots.txt, falls back to /sitemap.xml, follows sitemap indexes, and reads metadata from each public HTML page.
What happens when a site has more than 100 pages?
The scan stops at 100 pages to keep the tool responsive. A curated llms.txt needs fewer pages than that, so a large site should list only the pages that define its product and documentation.
Does llms.txt control AI crawlers?
No. Use robots.txt and access controls for crawler permissions. llms.txt is a discovery and context file.