Crawl4AI
Crawl4AI is an open-source web crawler and scraper for LLMs and AI agents. It turns any website into clean, LLM-ready Markdown for RAG, agents, and data pipelines, and can also be used hosted with one key.
Visit unclecode/crawl4aiOverview
Crawl4AI is an open-source web crawler and scraper for LLMs and AI agents. It converts websites into clean, LLM-ready Markdown for RAG, AI agents, and data pipelines. It can be self-hosted for free with Python, CLI, Docker, and MCP, or used through Crawl4AI Cloud with one API key.
Key Features
- Clean LLM-ready Markdown with headings, lists, tables, code blocks, and citation hints.
- Structured data extraction with CSS, XPath, regex, or LLM strategies into typed JSON.
- Browser control with persistent profiles, remote CDP connections, sessions, proxies, stealth mode, and Chromium, Firefox, and WebKit.
- Deep crawling with BFS, DFS, best-first strategies, crash recovery, adaptive crawling, and URL discovery.
- Self-hosting through the Python library, CLI, Docker REST API, MCP server, dashboard, and playground.
- Crawl4AI Cloud adds hosted scrape, batch, search, answer, and extract APIs plus MCP for agents.
Use Cases
- Feed clean Markdown into RAG pipelines and AI agents.
- Extract typed JSON records from websites with CSS, XPath, regex, or LLM extraction.
- Run web search and direct answers through the hosted API.
- Crawl many pages at scale for data pipelines and research.
Getting Started
- Install the open-source library with pip install -U crawl4ai and run crawl4ai-setup.
- Use the Python AsyncWebCrawler to fetch a URL and print result.markdown.
- For cloud use, get an API key and call the hosted scrape, search, answer, or extract endpoints with a Bearer token.
- Add the cloud MCP server to agents such as Claude Code, Codex, Cursor, or OpenCode using one configuration line.
Deployment & Requirements
- The library supports Python 3.10 through 3.13.
- crawl4ai-setup installs and sets up the browser; if it fails, install Playwright Chromium manually.
- The Docker server requires CRAWL4AI_API_TOKEN; without a token it answers only inside its container.
- Crawl4AI Cloud requires an API key and uses the hosted service instead of running browsers or proxies yourself.
Before You Adopt
- License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
- The hosted answer endpoint is marked experimental.
- Launch pricing is discounted for the first year and may change, while credit already held keeps its value.
- The self-hosted Docker server requires an API token; without one it only responds inside its container.
- Enterprise compliance maturity is in progress: SOC 2 Type I is complete and Type II is under way.