What it is
Crawl4AI turns any website into clean, LLM-ready Markdown for RAG, AI agents and data pipelines. You can run the open-source crawler yourself as a Python library, a Docker server or the crwl CLI. Alternatively, the hosted Crawl4AI Cloud offers scrape, search and extract through one API, with MCP for agents.
Who it's for
- Developers building RAG pipelines or data pipelines that need clean Markdown from web pages
- Teams building AI agents that need web scraping, search and extraction, including via MCP
- Engineers who want to self-host a crawler (library or Docker) with full control over browsers, proxies and sessions
- Developers who prefer a hosted API and don't want to run browsers or proxies
Requirements
Requirements
- Python and pip (for the library install)
- A browser installed via crawl4ai-setup (or playwright install as a fallback)
- Docker, for the self-hosted server
- An LLM key or provider for LLM extraction and the CLI question mode (not needed for CSS/XPath extraction)
- A Crawl4AI Cloud key, for the hosted API
Setup
Install the library
Install with pip, then run the one-time browser setup.
bashpip install -U crawl4ai crawl4ai-setup # installs the browser, onceCheck the installation
Run the doctor command to verify the install.
bashcrawl4ai-doctor # checks the installationManual browser install (if setup fails)
If the browser setup fails, install Chromium by hand.
bashpython -m playwright install --with-deps chromiumRun the Docker server
The server needs a token; without one it answers only inside its container. Allow about 10 seconds for it to start.
bashexport CRAWL4AI_API_TOKEN="$(openssl rand -hex 32)" docker run -d -p 11235:11235 --name crawl4ai --shm-size=1g \ -e CRAWL4AI_API_TOKEN="$CRAWL4AI_API_TOKEN" \ unclecode/crawl4ai:latest
Examples
Crawl a page to Markdown
pythonimport asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(url="https://news.ycombinator.com")
print(result.markdown)
asyncio.run(main())What it does: Minimal library usage: fetch a page and print its Markdown.
Fit Markdown with a content filter
pythonimport asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
from crawl4ai.content_filter_strategy import PruningContentFilterLXML
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator
async def main():
run_config = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
markdown_generator=DefaultMarkdownGenerator(
content_filter=PruningContentFilterLXML(threshold=0.48, threshold_type="fixed", min_word_threshold=0)
),
)
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
result = await crawler.arun(url="https://en.wikipedia.org/wiki/Web_crawler", config=run_config)
print(len(result.markdown.raw_markdown), "characters of raw Markdown")
print(len(result.markdown.fit_markdown), "characters after the filter")
asyncio.run(main())What it does: Uses a pruning filter to strip boilerplate and compares raw and fit Markdown lengths.
Structured extraction without an LLM
pythonimport asyncio, json
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, JsonCssExtractionStrategy
schema = {
"name": "Quotes",
"baseSelector": "div.quote",
"fields": [
{"name": "text", "selector": "span.text", "type": "text"},
{"name": "author", "selector": "small.author", "type": "text"},
{"name": "tags", "selector": "a.tag", "type": "list", "fields": [{"name": "tag", "type": "text"}]},
],
}
async def main():
run_config = CrawlerRunConfig(
extraction_strategy=JsonCssExtractionStrategy(schema),
scan_full_page=True, # scroll to the end, so the page loads every quote
scroll_delay=0.5,
cache_mode=CacheMode.BYPASS,
)
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
result = await crawler.arun(url="https://quotes.toscrape.com/scroll", config=run_config)
quotes = json.loads(result.extracted_content)
print(f"Extracted {len(quotes)} quotes")
print(json.dumps(quotes[0], indent=2))
asyncio.run(main())What it does: A CSS schema extracts quotes from an infinite-scroll page, scrolling the full page first.
Command line usage
bash# A page as Markdown
crwl https://news.ycombinator.com -o markdown
# Deep crawl, breadth first, at most 10 pages
crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
# Ask a question about a page (needs an LLM key: crwl config)
crwl https://www.example.com/products -q "Extract all product prices"What it does: The crwl CLI covers Markdown output, deep crawling and LLM questions about a page.
Scrape via Crawl4AI Cloud
bashcurl -s https://api.crawl4ai.com/scrape \
-H "Authorization: Bearer $CRAWL4AI_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://news.ycombinator.com"}' | jq -r .markdownWhat it does: Gets a page as Markdown from the hosted API with no local browser.
Pros & cons
Pros
- Pro:Free, Apache 2.0 open source library with Docker server and CLI options
- Pro:Produces LLM-ready Markdown, with fit-Markdown filters and citation lists
- Pro:Offers extraction with or without an LLM (CSS/XPath/regex schemas, or LLM providers)
- Pro:Rich browser control: persistent profiles, proxies, sessions, stealth mode, deep and adaptive crawling
Cons
- Con:Self-hosting means you run the browsers, and handle proxies and JS-heavy pages or bot walls yourself
- Con:Web search (/search, /answer) is only in the paid cloud; /answer is marked experimental
- Con:Cloud is a soft launch where prices can change, and it requires a key
Images
