unclecode/crawl4ai

Open-source Python crawler that turns websites into LLM-ready Markdown. Use it as a library, Docker server or CLI, or via the hosted Crawl4AI Cloud API.

  • 85.1k GitHub stars
  • Python
  • ⚖️ Apache-2.0
  • 🎯 Beginner
unclecode/crawl4ai preview image

What it is

Crawl4AI turns any website into clean, LLM-ready Markdown for RAG, AI agents and data pipelines. You can run the open-source crawler yourself as a Python library, a Docker server or the crwl CLI. Alternatively, the hosted Crawl4AI Cloud offers scrape, search and extract through one API, with MCP for agents.

Who it's for

  • Developers building RAG pipelines or data pipelines that need clean Markdown from web pages
  • Teams building AI agents that need web scraping, search and extraction, including via MCP
  • Engineers who want to self-host a crawler (library or Docker) with full control over browsers, proxies and sessions
  • Developers who prefer a hosted API and don't want to run browsers or proxies

Requirements

Requirements

  • Python and pip (for the library install)
  • A browser installed via crawl4ai-setup (or playwright install as a fallback)
  • Docker, for the self-hosted server
  • An LLM key or provider for LLM extraction and the CLI question mode (not needed for CSS/XPath extraction)
  • A Crawl4AI Cloud key, for the hosted API

Setup

  1. Install the library

    Install with pip, then run the one-time browser setup.

    bash
    pip install -U crawl4ai
    crawl4ai-setup        # installs the browser, once
  2. Check the installation

    Run the doctor command to verify the install.

    bash
    crawl4ai-doctor     # checks the installation
  3. Manual browser install (if setup fails)

    If the browser setup fails, install Chromium by hand.

    bash
    python -m playwright install --with-deps chromium
  4. Run the Docker server

    The server needs a token; without one it answers only inside its container. Allow about 10 seconds for it to start.

    bash
    export CRAWL4AI_API_TOKEN="$(openssl rand -hex 32)"
    docker run -d -p 11235:11235 --name crawl4ai --shm-size=1g \
      -e CRAWL4AI_API_TOKEN="$CRAWL4AI_API_TOKEN" \
      unclecode/crawl4ai:latest

Examples

Crawl a page to Markdown

python
python
import asyncio
from crawl4ai import AsyncWebCrawler

async def main():
    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun(url="https://news.ycombinator.com")
        print(result.markdown)

asyncio.run(main())

What it does: Minimal library usage: fetch a page and print its Markdown.

Fit Markdown with a content filter

python
python
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
from crawl4ai.content_filter_strategy import PruningContentFilterLXML
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator

async def main():
    run_config = CrawlerRunConfig(
        cache_mode=CacheMode.BYPASS,
        markdown_generator=DefaultMarkdownGenerator(
            content_filter=PruningContentFilterLXML(threshold=0.48, threshold_type="fixed", min_word_threshold=0)
        ),
    )
    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        result = await crawler.arun(url="https://en.wikipedia.org/wiki/Web_crawler", config=run_config)
        print(len(result.markdown.raw_markdown), "characters of raw Markdown")
        print(len(result.markdown.fit_markdown), "characters after the filter")

asyncio.run(main())

What it does: Uses a pruning filter to strip boilerplate and compares raw and fit Markdown lengths.

Structured extraction without an LLM

python
python
import asyncio, json
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, JsonCssExtractionStrategy

schema = {
    "name": "Quotes",
    "baseSelector": "div.quote",
    "fields": [
        {"name": "text", "selector": "span.text", "type": "text"},
        {"name": "author", "selector": "small.author", "type": "text"},
        {"name": "tags", "selector": "a.tag", "type": "list", "fields": [{"name": "tag", "type": "text"}]},
    ],
}

async def main():
    run_config = CrawlerRunConfig(
        extraction_strategy=JsonCssExtractionStrategy(schema),
        scan_full_page=True,   # scroll to the end, so the page loads every quote
        scroll_delay=0.5,
        cache_mode=CacheMode.BYPASS,
    )
    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        result = await crawler.arun(url="https://quotes.toscrape.com/scroll", config=run_config)
        quotes = json.loads(result.extracted_content)
        print(f"Extracted {len(quotes)} quotes")
        print(json.dumps(quotes[0], indent=2))

asyncio.run(main())

What it does: A CSS schema extracts quotes from an infinite-scroll page, scrolling the full page first.

Command line usage

bash
bash
# A page as Markdown
crwl https://news.ycombinator.com -o markdown

# Deep crawl, breadth first, at most 10 pages
crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10

# Ask a question about a page (needs an LLM key: crwl config)
crwl https://www.example.com/products -q "Extract all product prices"

What it does: The crwl CLI covers Markdown output, deep crawling and LLM questions about a page.

Scrape via Crawl4AI Cloud

bash
bash
curl -s https://api.crawl4ai.com/scrape \
  -H "Authorization: Bearer $CRAWL4AI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://news.ycombinator.com"}' | jq -r .markdown

What it does: Gets a page as Markdown from the hosted API with no local browser.

Pros & cons

Pros

  • Pro:Free, Apache 2.0 open source library with Docker server and CLI options
  • Pro:Produces LLM-ready Markdown, with fit-Markdown filters and citation lists
  • Pro:Offers extraction with or without an LLM (CSS/XPath/regex schemas, or LLM providers)
  • Pro:Rich browser control: persistent profiles, proxies, sessions, stealth mode, deep and adaptive crawling

Cons

  • Con:Self-hosting means you run the browsers, and handle proxies and JS-heavy pages or bot walls yourself
  • Con:Web search (/search, /answer) is only in the paid cloud; /answer is marked experimental
  • Con:Cloud is a soft launch where prices can change, and it requires a key

Images