← Back to all sparks
F

Firecrawl

AI-ASSISTANTS
Velocity7.5

Web scraping and crawling API that turns websites into clean, LLM-ready markdown and structured data.

Firecrawl now sells corpora, not crawls - and just aimed one at coding agents.

retrievalindexescoding agentsbenchmarksagent infrastructure
Current state
Firecrawl has converted a scraping API into a retrieval layer built on three owned indexes: Research (arXiv), Life Sciences (41M papers), and now Developer (70M+ READMEs, issues, PRs and OpenAPI specs). Each launches with a recall number attached, including 0.63 recall@10 on its own DevDex set for the newest one. The scrape side has settled into token-efficiency formats - Question, Highlights, and an excerpt-scored /search - that return passages rather than pages. Research Index is now free outright.
Where it's heading
The centre of gravity has moved from selling fetches to selling corpora, with published benchmarks as the competitive axis. Developer Index is the first index pointed at a market far larger than research: every coding agent that needs to answer questions about library behaviour from primary sources. Giving away research retrieval while charging two credits per developer search shows where Firecrawl expects revenue to sit.
Prediction
Expect more verticals on the /search/<domain> pattern, each launched with a self-published recall benchmark. The entries do not indicate whether the free research tier is permanent or an acquisition period.

Recent moves

  1. 27d ago

    Firecrawl Developer Index

    ⚡ SPARK

    The third index in four months, and the first pointed away from papers at the artifacts coding agents actually need. It confirms the shift from crawling on demand to owning corpora, and moves Firecrawl into a market much larger than research retrieval.

  2. 1mo ago

    Life Sciences in Firecrawl Research Index

    ⚡ SPARK

    A second domain for Research Index, but the weight of the release sits in the pricing line rather than the corpus. Making the whole index free reframes it from a paid product into a distribution channel for the indexes that do charge.

  3. 1mo ago

    Introducing our most accurate /search yet

    ⚡ SPARK

    The excerpt-scoring rebuild made /search return the passages that answer a query instead of whole pages, folding the Highlights idea into the default endpoint. This is the point where the scrape side stopped being a fetch tool and became an answer layer.

  4. 2mo ago

    Web-scale /monitor

    Web-scale /monitor extends change-watching from named URLs to open-ended queries across the web, with a plain-English goal deciding which matches justify an alert. It broadens an existing endpoint rather than opening a new surface.

  5. 2mo ago

    v2.11.0: Research Index, keyless access, PII redaction

    The release that bundled Research Index with keyless endpoint access, automatic PII redaction and a cached deterministic JSON extractor. The packaging is the story: index access and cheaper structured scraping arrived in the same version.

  6. 3mo ago

    Firecrawl Research Index

    ⚡ SPARK

    The origin of the index strategy: the first time Firecrawl sold a pre-built corpus rather than a fetch, and the first time it attached a published recall figure to the claim. Every index release since follows this template.