Diffbot API review
Best for entity data

Diffbot is a Menlo Park company that crawls the public web on its own hardware and sells the results in two forms. Extract reads a page you name, classifies it as an article, product, discussion or another type, and returns structured JSON without per-site rules. The Knowledge Graph is the pre-built version: a graph Diffbot says holds over 10 billion entities, from organisations and people to articles.
The credit table carries the argument. Extracting a page costs 1 credit, and so does a Natural Language document. Exporting one entity from the Knowledge Graph costs 25, a facet query record 100, and an Enhance lookup 25 when it finds a match. On the $299 Startup plan, that is the difference between 250,000 pages and 10,000 entities from one allowance.
For SEO that makes Diffbot a supplier of facts rather than rankings. It has no search volumes, no backlinks and no Google SERPs. Its Web Search API queries Diffbot's own English-language index, which leaves out robots.txt opt-outs and generally excludes platform-gated sites such as Reddit, so it cannot say where a page ranks. It can describe an entity, or turn a competitor's site into structured records.
Our scorecard
Our editorial read across six dimensions, scored 1 to 5. This is our own assessment, not an aggregate of user reviews, and it is deliberately kept out of the page's structured data.
- Data coverage4/5
Any page, a graph Diffbot sizes at over 10 billion entities, and its own web index. No rankings, volumes or links whatsoever.
- Pricing2/5
Nothing between free and $299, Crawl starts at $899, and an entity costs 25 times a page.
- Documentation4/5
Complete reference with public credit and rate-limit tables. Some pages still name plans the pricing page no longer sells.
- Ease of integration4/5
A token and a URL for Extract. DQL is a query language to learn, eased by Python and TypeScript libraries, skills and MCP.
- Reliability4/5
APIs Diffbot says it has maintained for over 15 years, and a graph rebuilt every four to five days. No SLA on the self-serve plan cards.
- Support3/5
Dedicated support is an Enterprise line. Staff do answer in public threads, which is helpful but not a service level.
An editorial read, not an aggregate of user reviews, and deliberately kept out of this page's structured data.
Diffbot API in their own words
From Diffbot API's own YouTube channel. Vendor marketing rather than an independent review, included because it is the quickest way to see the product before you sign up.
What you can pull
Analyze
1 credit, 2 via proxy
Classifies any URL into a page type and extracts it with that type's schema; with fields=allLinks it also lists every URL in an XML sitemap
Article, Product and other page types
1 credit a page
Dedicated extractors for articles, products, images, video and discussions, with events, lists and jobs in beta; mode=llm returns markdown
Crawl and Bulk Extract
0 to spider, 1 per page; Crawl on Plus
Crawl, once called Crawlbot, spiders a site from seed URLs and extracts each page into a collection searchable with DQL; Bulk takes a URL list
Knowledge Graph search (DQL)
25 an entity, 100 a facet record
Structured queries over organisations, people, articles, places and products, exported as JSON, JSONL, CSV or Excel
Enhance
25 on a match, 100 with refresh
Matches a name, URL or other identifier to a Knowledge Graph organisation or person and returns the full record
Natural Language
1 credit per 10,000-character document
Entities, facts, salience and sentiment from raw text, with each entity linked to its Knowledge Graph ID
Web Search
Pricing page and docs disagree
Ranked results with a content chunk each from Diffbot's own English index, filterable with after, before, site and url operators
MCP server and agent skills
Standard credit costs
A remote server at mcp.diffbot.com exposing extract, web search, enhance, entity resolution, crawl and DQL, plus installable agent skills
Who Diffbot API is for
The fit is SEO work about entities and content rather than positions. Agencies assemble organisation markup and sameAs links from a structured record; content teams audit what competitors publish at article level; product teams pull clean article or product fields from pages they do not control. Extract at 1 credit a page does most of the work, and the Knowledge Graph answers what a crawl cannot.
Look elsewhere for rankings, keyword volumes or links, none of which it provides, and if your budget sits well under $299 a month once the free plan runs out. Crawling an entire competitor site also requires the $899 Plus plan, which changes the sums for a small agency.
Strengths and limits
What it does well
- Extraction without per-site rules: Analyze classifies a page and applies that type's schema, so a new source needs no parser.
- A Knowledge Graph Diffbot sizes at over 10 billion entities, queryable with DQL or matched one record at a time through Enhance.
- Empty answers are free: a graph search returning nothing costs no credits, and Enhance charges only when it matches.
- Crawl charges nothing for spidering, only for pages it extracts, and follows robots.txt disallow rules by default.
- Ten thousand free credits renew every month, rather than arriving once as a trial.
- Its own crawl and index, independent of Google and Bing, so Web Search is not another reseller of scraped SERPs.
- Diffbot's GitHub organisation hosts an MCP server, remote or self-hosted, alongside agent skills and Python and TypeScript libraries.
Where it falls short
- Nothing sits between the free plan and $299 a month, and Crawl needs the $899 Plus plan.
- Graph data costs 25 credits an entity, so the $299 allowance covers 10,000 records, about $30 per thousand.
- Rankings, search volumes, keyword difficulty and backlinks are all absent.
- Web Search is not Google. The index is English only and generally leaves out platform-gated sites such as Reddit, so it cannot stand in for a SERP API.
- Unusual layouts are the weak spot, and a Diffbot employee has said so publicly: non-standard page types are where it falls short.
- Diffbot can switch proxies on for a whole blocked domain, adding a credit to every call there with no change on your side.
- The documentation lags the pricing page, naming a Professional plan and a Startup crawl policy that no longer match, and disagreeing on how web searches are counted.
Diffbot API pricing
Four plans. Free gives 10,000 credits a month at 5 requests a minute, with no card. Startup costs $299 a month for 250,000 credits at 5 a second, then $0.001 a credit. Plus costs $899 for a million at 25 a second, then $0.0009, and is the first plan with Crawl. Enterprise is quoted. Students and academic researchers get Startup access free.
Price the job in entities before pages. A page is 1 credit and an entity 25, so Startup covers 250,000 extracted pages or 10,000 Knowledge Graph records: roughly $1.20 per thousand pages against $30 per thousand entities. Facet records and refreshed Enhance lookups cost 100. Graph searches returning nothing and Enhance calls with no match are not charged.
Two costs hide in the proxy rules. A request through Diffbot's default proxy is 2 credits, and Diffbot may enable proxies across a whole domain it finds blocked, which quietly doubles the cost of every call to that site. Paid-plan overage is billed automatically at the plan's per-credit rate on the next invoice, with no separate fee.
Web Search is where Diffbot's own pages disagree. The pricing page counts a search as 1 credit inside the plan allowance, while the Web Search documentation says a free token includes 100,000 queries a month at 60 a minute and quotes sales pricing down to $0.02 per thousand. Confirm with Diffbot which applies before building on it.
Free
$0
10,000 credits a month
- No card, and every API except Crawl and Bulk
- 5 requests a minute across all APIs
- Caps of 400 graph entities and 500 NLP calls
- A 429 at the quota, never an overage bill
Startup
$299 / mo
250,000 credits, then $0.001 each
- 5 requests a second
- 250,000 pages or 10,000 entities
- No Crawl or Bulk: both begin on Plus
- Free for students and academic researchers
Plus / Enterprise
$899 / mo or custom
1,000,000 credits, then $0.0009 each
- 25 requests a second and 25 active crawls
- Three user seats on Plus
- Enterprise adds 100+ crawls and dedicated support
- Roughly $22.50 per 1,000 entities on Plus
Request example
A minimal call against the live endpoint, with credentials read from the environment rather than pasted inline.
# A page, classified and extracted: 1 credit. curl -s -G "https://api.diffbot.com/v3/analyze" \ --data-urlencode "token=$DIFFBOT_TOKEN" \ --data-urlencode "url=https://www.example.com/blog/some-post" # An organisation from the Knowledge Graph: 25 credits, # and nothing at all if Enhance finds no match. curl -s -G "https://kg.diffbot.com/kg/v3/enhance" \ --data-urlencode "token=$DIFFBOT_TOKEN" \ --data-urlencode "type=Organization" \ --data-urlencode "name=Diffbot" # Both calls draw on the same allowance, so budget in entities: # 250,000 credits on Startup is 250,000 pages or 10,000 records.
Rate limits and quotas
- Free: 10,000 credits a month at 5 calls a minute, capped at 10,000 Extract calls, 500 Natural Language calls, 400 DQL entities and 400 Enhance entities.
- Startup, $299 a month: 250,000 credits, 5 calls a second, overage $0.001 per credit.
- Plus, $899 a month: 1,000,000 credits, 25 calls a second, overage $0.0009 per credit, 25 active crawls, 3 seats.
- Credit costs: 1 per extracted page, 2 through a datacenter proxy, 0 for spidering, 1 per Natural Language document up to 10,000 characters.
- Knowledge Graph: 25 credits per exported or enhanced entity, 100 per facet record or refreshed Enhance.
- Crawl and Bulk share a limit of 1,000 jobs per token, with up to 30 seeds running at once in a job.
- Non-enterprise credits refresh monthly and do not roll over; the free plan returns a 429 at its quota.
- Web Search: English-only index, 1 to 5 query values per request, with after, before, site and url operators.
- Crawl obeys robots.txt disallow and crawl-delay by default but ignores Allow directives.
- Monthly subscriptions, cancellable anytime; custom invoicing requires a prepaid annual Startup or Plus plan.
What practitioners say on Reddit
Too little practitioner discussion exists to report a sentiment. No thread we found describes running Diffbot in production. The Bright Data, Oxylabs and ScrapingBee pages blame r/webscraping's moderation for thin discussion, but that is not the cause here, since Diffbot threads in that subreddit have survived. It surfaces in passing instead, as a name in a list or as the price a builder is trying to undercut.
The three comments below are context, not a verdict. The signal repeated across 2022, 2023 and 2026 is the $299 floor, raised by people looking for something cheaper. Entity data is the strength that gets named. And one site owner complained that a Diffbot user agent ignored their robots.txt, which nobody in that thread confirmed.
Diffbot staff also reply in public, and one reply is worth knowing before a trial: an employee wrote that non-standard page types are where Extract falls short. That is the vendor's own statement rather than a user's, and it is reported on this page as such.
“The direct competition is Diffbot but is damn expensive ($299/month min plan)”
r/webscraping on Reddit“Diffbot - knowledge-graph style, strong for entity data.”
r/gtmengineering on Reddit“I have diffbot disallowed in my robots.txt I see the bot crawling my site anyways”
r/selfhosted on Reddit
Verdict
Diffbot belongs in an SEO stack as a source of structured facts, not as an SEO API. Extract turns arbitrary pages into clean fields at 1 credit without rules, and the Knowledge Graph answers entity questions that fetching pages on your own cannot. Organisation markup, competitor content audits and entity research are the jobs it is built for.
Budget in entities. The same 250,000 credits on the $299 plan are a quarter of a million pages or ten thousand graph records, and crawling a whole site needs the $899 plan. Start with the free plan's 10,000 credits, which buy 400 entities, and count how many records your question genuinely needs before paying for a tier.
Keep it away from rankings. It holds no Google SERPs, volumes or links, and its Web Search runs on Diffbot's own English index rather than Google's, so it complements a SERP provider such as DataForSEO instead of replacing one. Pair the two when an entity question and a ranking question sit in the same report.
Diffbot API FAQ
How is Diffbot API pricing structured?
Diffbot has four plans. Free carries 10,000 credits a month, Startup is $299 for 250,000, Plus is $899 for a million and Enterprise is quoted. Overage costs $0.001 a credit on Startup and $0.0009 on Plus, added to the next invoice. What a credit buys depends on the product: a page costs 1, a Knowledge Graph entity 25.What does Diffbot Knowledge Graph data cost?
Twenty-five credits for each entity exported or matched through Enhance, and 100 for a facet query record or a refreshed Enhance lookup. On Startup that is 10,000 entities a month, about $30 per thousand, or $25 per thousand as overage. Searches returning no entities, Enhance calls without a match and searches in the dashboard all cost nothing.Is Crawlbot still part of Diffbot, and which plan includes it?
Under a newer name. Diffbot's docs now call it the Crawl API, though parts of the documentation still use the old one. Crawl is limited to the $899 Plus plan and above, with 25 active crawls on Plus. Spidering costs nothing; each page it extracts costs 1 credit. It follows robots.txt by default.Which fields does the Diffbot Article API return?
Clean body text and normalised HTML plus the title, author, date, site name, language, sentiment, images and tags linked to Knowledge Graph entities, for 1 credit a page. Analyze does the same for any page type after classifying it, and adding mode=llm returns markdown ready for a language model.How does the Diffbot Natural Language API handle raw text?
Give it raw text and it returns the entities mentioned, the facts connecting them, how central each entity is and the sentiment towards it, with entities linked to Knowledge Graph IDs. Each document of up to 10,000 characters costs 1 credit, and the free plan allows 500 calls a month. The open-facts feature is currently switched off.Can Diffbot Web Search replace a Google SERP API?
Not for SEO. It searches Diffbot's own crawl rather than Google, so it returns relevant pages, not Google positions. The index is English only and generally omits platform-gated sites such as Reddit. It suits grounding an AI agent or finding sources, while rank tracking still needs a provider that collects Google results, such as DataForSEO or Serper.Who is behind Diffbot and how is it funded?
Diffbot Technologies Corp. is based in Menlo Park, California, and runs its own crawler from a data centre in the state. Its press page records $2 million from technology investors in 2012 and a $10 million Series A led by Tencent and Felicis Ventures in 2016, and lists no later round. In January 2025 it launched an open-weight language model.Where do Diffbot competitors beat it for SEO work?
Split the question in two. For fetching pages, ScraperAPI costs less per page on its entry plan but returns HTML rather than structured fields. For search data, DataForSEO or Serper, because Diffbot has none. For entity records, the comparison is with company-data vendors rather than SEO APIs, and many teams pair Diffbot with a SERP provider instead of choosing between them.
This is an independent review. We have no affiliate relationship with Diffbot API, earn nothing if you sign up, and no provider pays for placement in the directory. Prices were read from the vendor's own pages in September 2026 and change without notice.