PyScrappy
View on GitHubAdaptive Python web scraping toolkit + MCP server for AI agents. Self-healing selectors that survive site changes, TLS-fingerprint stealth to bypass anti-bot filters, CSS/XPath parsing, and 24 built-in scrapers, clean, structured, LLM-ready data from any URL.
Python web-scraping toolkit exposed as an MCP server, giving AI agents tools to pull structured, LLM-ready Markdown/JSON from any URL. Adds self-healing CSS/XPath selectors, TLS-fingerprint stealth, Playwright JS rendering, and 24 built-in site scrapers.
Use Cases
Give AI agents live web data via MCP toolsConvert any URL into clean LLM-ready Markdown/JSONScrape JS-heavy sites with Playwright renderingCrawl whole sites from sitemap.xmlExtract news headlines and RSS feedsFetch stock/crypto/currency quotesE-commerce price and product scrapingRun scrapers through local Ollama models with tool callingBuild custom scrapers as distributable pluginsExport scraped data to CSV/Parquet/ExcelBypass anti-bot filters with TLS impersonationProvide retrieval data for LLM pipelinesParallel batch scraping of many URLsSelf-healing selectors after site markup changes
Built With
- Language
- Python
- Frameworks
- FastMCP · MCP SDK · httpx · BeautifulSoup · lxml · Playwright · pandas · curl_cffi · PyYAML · pyarrow · openpyxl · anyio
Tags
web-scraping · mcp-server · self-healing-selectors · stealth · tls-fingerprint · css-selectors · xpath · llm-ready · structured-data · fastmcp · cli · concurrent-scraping · proxy-support · sitemap-crawling · plugins · markdown-export