MCP Server Community MCP

D4Vinci/Scrapling

Adaptive Python web scraping framework with anti-bot fetchers, spiders for full crawls, and an MCP server so AI agents can scrape pages.

  • 86.5k GitHub stars
  • Python
D4Vinci/Scrapling preview image

What it is

Scrapling is an adaptive web scraping framework that covers everything from a single request to a full-scale crawl. Its parser can relocate elements after a website changes, its fetchers handle anti-bot systems like Cloudflare Turnstile, and its spider framework supports concurrent crawls with pause/resume and proxy rotation. It also ships an MCP server that lets AI chatbots and agents such as Claude and Cursor scrape through it.

Who it's for

  • Web scrapers who want scrapes that survive website design changes
  • Developers building concurrent, multi-session crawls in Python
  • Users of AI agents (Claude, Cursor, etc.) who want scraping tools exposed through MCP
  • Teams building RAG pipelines who need pages or whole sites converted to Markdown
  • Scrapy users who want to parse responses with Scrapling's parser

Requirements

Requirements

  • Python (the package is distributed on PyPI as Scrapling)
  • Playwright's Chromium or Google Chrome for browser-based fetching with DynamicFetcher

Examples

Adaptive stealth fetch and scrape

python
python
from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)  # Fetch website under the radar!
products = p.css('.product', auto_save=True)                                        # Scrape data that survives website design changes!
products = p.css('.product', adaptive=True)                                         # Later, if the website structure changes, pass `adaptive=True` to find them!

What it does: Fetches a page with the stealth fetcher, saves the selected elements, then re-finds them with adaptive=True if the site's structure changes.

Full crawl with a Spider

python
python
from scrapling.spiders import Spider, Response

class MySpider(Spider):
  name = "demo"
  start_urls = ["https://example.com/"]

  async def parse(self, response: Response):
      for item in response.css('.product'):
          yield {"title": item.css('h2::text').get()}

MySpider().start()

What it does: Defines a Scrapy-like spider with start_urls and an async parse callback that yields items, then starts the crawl.

Convert a page to Markdown

python
python
page.markdown()

What it does: The README describes this one-liner as turning any page into clean, sanitized, LLM-ready Markdown without an LLM in the loop.

Pros & cons

Pros

  • Pro:Parser can relocate elements after a site changes, using auto_save and adaptive=True
  • Pro:Fetchers include stealth and fingerprint spoofing, and can bypass Cloudflare Turnstile/Interstitial
  • Pro:Spider framework offers pause/resume, AutoThrottle, proxy rotation, robots.txt compliance and built-in JSON/JSONL/CSV/XML export
  • Pro:MCP server narrows pages with CSS selectors and strips prompt-injection content before the AI sees them

Cons

  • Con:Bypass is documented for Cloudflare; for Akamai, DataDome, Kasada and Incapsula the README points to a separate third-party service
  • Con:The provided README excerpt gives no install or MCP configuration commands; those are in the external docs

Images