← Back to all spotlights

Firecrawl: Turn Entire Websites into Clean LLM-Ready Markdown

Mendable's Firecrawl turns messy, JavaScript-bloated websites into clean Markdown and structured JSON for LLMs. Here is how to run it locally.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

Feed raw HTML into a Large Language Model and watch your context window implode. A modest documentation page routinely packs forty kilobytes of tracking pixels, nested <div> wrappers, SVG icons, and legal disclaimers about cookies before you even reach a single paragraph of actual content.

If you have spent any time building Retrieval-Augmented Generation (RAG) pipelines or AI agents, you know the quiet despair of writing brittle Playwright scripts that crash the moment a single CSS class changes. That is where Firecrawl comes in.


[ Unruly Web: SPAs, Cookie Modals, Nav Menus ]
                      │
                      ▼
            [ Firecrawl Engine ]
   ├── Headless Orchestration (Playwright/Chrome)
   ├── Anti-Bot Bypass & JS Hydration
   └── Semantic HTML-to-Markdown Parser
                      │
                      ▼
[ Clean, Structured Markdown / JSON -> Vector DB / LLM ]

Direct Definition: What Is Firecrawl?

Firecrawl is an open-source crawling and scraping engine created by Mendable AI. It takes any URL, navigates complex client-side JavaScript, handles sitemaps and sub-pages recursively, strips away navigational and stylistic cruft, and outputs LLM-ready Markdown or structured JSON schemas.

  • Repository: mendableai/firecrawl
  • Primary Interfaces: REST API, Python SDK, Node.js SDK, Self-hosted Docker
  • Core Output Formats: Markdown, Cleaned HTML, Structured JSON (via LLM extraction)

Why Standard Scrapers Fail Modern RAG

Developers on Reddit and technical YouTube channels have spent the past year debating whether web scraping is essentially a solved problem. The consensus? Scraping static text is solved; scraping for generative AI is an absolute minefield.

Traditional libraries like Beautiful Soup or Scrapy expect predictable document trees. Modern documentation portals, however, are client-rendered React or Next.js applications wrapped in Cloudflare challenges, shadow DOMs, and dynamic lazy-loading triggers.

If you use raw Playwright, you end up maintaining a bespoke mini-browser farm. You must handle proxy rotation, page timeouts, infinite scrolls, and then write custom heuristics to separate the documentation article from the sidebar navigation.

Firecrawl handles the entire pipeline in one shot. It does not just scrape a single page; it discovers routes, traverses subdomains, bypasses bot detection, waits for client-side hydration, and transforms the resulting DOM into crisp GitHub-flavoured Markdown.


Technical Architecture

Under the bonnet, Firecrawl operates as a distributed system designed for resilience:

1. Queue Management: Built on BullMQ backed by Redis, orchestrating concurrent crawl jobs across multiple worker processes.

2. Headless Execution: Uses isolated Chromium instances via Playwright to ensure client-side single-page applications (SPAs) render completely before extraction.

3. Markdown Conversion Engine: Rather than naively stripping tags, Firecrawl parses structural landmarks (<main>, <article>, heading hierarchies) and strips non-content elements (<nav>, <footer>, modal dialogs, scripts) before mapping semantic HTML to equivalent Markdown tokens.

4. LLM Schema Extraction: A native /extract endpoint allows developers to pass an arbitrary Pydantic or Zod schema, transforming messy pages directly into typed JSON without a separate prompting stage.

CapabilityTraditional Scraping (Scrapy/BS4)Raw Headless (Playwright/Puppeteer)Firecrawl
JS-Rendered SPAsFails (requires SSR)SupportedSupported natively
Markdown OutputManual regex/parsers neededManual parsing neededNative & optimised for LLMs
Recursive CrawlingManual crawler logicManual state managementBuilt-in /crawl engine
Structured JSONComplex post-processingComplex post-processingBuilt-in /extract endpoint
Bot MitigationFragile headersManual stealth pluginsIntegrated anti-detection

Local Setup: Self-Hosting Firecrawl

While Mendable runs a managed cloud tier, Firecrawl is open-source under an AGPL-3.0 licence and straightforward to run on your own hardware via Docker Compose.

1. Clone and Configure


git clone https://github.com/mendableai/firecrawl.git
cd firecrawl
cp .env.example .env

Ensure your .env file contains your basic environment defaults. If you plan to use the structured LLM extraction endpoint locally, plug in your OpenAI or local Ollama endpoint keys accordingly.

2. Launch the Stack

Run Docker Compose to spin up Redis, the API server, and Playwright worker nodes:


docker compose up -d

Verify that the API is alive:


curl http://localhost:3002/v1/scrape \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'

Practical Usage Examples

1. Scrape a Single Dynamic Page to Markdown

Using Python, grabbing a page and preparing it for an embedding model requires only a few lines:


from firecrawl import FirecrawlApp

app = FirecrawlApp(api_url="http://localhost:3002")

# Scrape dynamic site into clean markdown
scrape_result = app.scrape_url(
    url="https://docs.docker.com/engine/install/",
    params={"formats": ["markdown", "html"]}
)

print(scrape_result["markdown"][:500])
# Output: Clean H1, H2, and shell commands without cookie banners or top nav menus

2. Crawling an Entire Domain with Sub-page Control

If you are indexing an entire knowledge base for a vector store, use the asynchronous crawl endpoint:


crawl_job = app.crawl_url(
    url="https://docs.python.org/3/tutorial/",
    params={
        "limit": 50,
        "scrapeOptions": {
            "formats": ["markdown"]
        }
    },
    poll_interval=5
)

for page in crawl_job["data"]:
    print(f"Indexed: {page['metadata']['title']} -> {len(page['markdown'])} chars")

3. Extracting Structured Data Directly

Instead of scraping text and piping it into a separate structured output chain, Firecrawl handles semantic extraction directly:


schema = {
    "type": "object",
    "properties": {
        "release_version": {"type": "string"},
        "features": {"type": "array", "items": {"type": "string"}},
        "breaking_changes": {"type": "boolean"}
    },
    "required": ["release_version", "features"]
}

extracted = app.scrape_url(
    url="https://github.com/astral-sh/uv/releases",
    params={
        "formats": ["extract"],
        "extract": {"schema": schema}
    }
)

print(extracted["extract"])

Key Takeaways for AI Engineers

  • Context Window Efficiency: Converting messy DOM trees directly to Markdown reduces raw token usage by up to 70% compared to raw HTML ingestion.
  • Deterministic RAG Ingestion: Provides structural guarantees (headings, code fences, clean lists) that preserve document hierarchy during vector chunking.
  • Open Source Control: Complete self-hosting avoids per-page SaaS scraping bills while running securely inside your own private cloud or local infrastructure.

🛡️ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.