Website

Crawl and index any website as a knowledge base

Adding a Website

Select Web from the available sources and proceed.

website

Fetch Document

Enter your website URL and click Fetch document. This lists all links found on the website. You can navigate away and complete setup later after all pages are fetched.

Load All Paths

When enabled, Marseil AI recursively crawls all links on the website until all pages are fetched. Recommended to avoid missing pages. Disable to only extract text from the specified URL.

Paths to Skip

To skip certain paths like /blogs or /private, enter comma-separated paths in this field.

Fetched Pages

Review and deselect pages you don’t want to include in training. Once configured, click Create document and your agent will be ready.

You can also preview how much storage the website will consume. Consider upgrading if you have a large site.

Sync Pages

When the website content is updated, sync the document to pull the latest changes. Let us know in Feature Requests if you want automatic syncing.