Train on your website

Crawling, page limits and keeping answers fresh

How crawling works

Chatixy starts at the URL you provide and follows internal links up to your plan’s page limit, indexing the visible text content of each page. Pages behind logins, paywalls or robots.txt blocks are skipped.

Two limits apply. The per-site page limit caps a single crawl (Starter 150, Pro 2,500, Business 10,000 pages). On top of that, a daily crawl allowance caps the pages indexed across *all* your websites in one day (Starter 150, Pro 12,500, Business 50,000 pages) and resets at 00:00 UTC. If a crawl is refused because the daily allowance is used up, the knowledge tab tells you so and you can run it again after the reset.

The crawl usually finishes in a few minutes for typical marketing sites; very large sites are processed in the background and your support agent improves as pages are added.

Excluding pages

You can exclude paths you never want the support agent to learn from - careers pages, legal archives, internal tools.

  • Add path patterns like /blog/archive/* in the project’s knowledge settings
  • Excluded pages are removed from the index on the next crawl
  • robots.txt disallow rules are always respected

Keeping answers fresh

You can trigger a manual re-crawl at any time from the knowledge tab - useful right after you publish new docs or change pricing. A re-crawl is available again 14 days after the last one finished.

You can also switch on scheduled re-crawls per website and pick the cadence yourself - every 14, 30 or 90 days. The same options are available on every plan.

Tips

  • Trigger a manual re-crawl after big site updates so the support agent never quotes stale content.

Important

  • Removing a page from your site does not instantly remove it from the index - it disappears on the next crawl, or you can delete the source manually.