Data & analytics

Data scraping & pipelines: the tech stack of 38 real products

Scraping, ETL, data integration and enrichment.

What sets them apart

Picks at least twice as common here as among products overall, in decisions at least 10 of them show.

The typical stack

The leading pick where at least 8 of them show the decision and the leader has at least a quarter of it.

DecisionMost common pickShareAt default usage
DatabaseMySQL10 of 27 · 37%—
Frontend FrameworkReact13 of 22 · 59%—
Backend FrameworkFastAPI9 of 17 · 53%—
LLM APIOpenAI API9 of 15 · 60%$10/mo · GPT-6 Luna
AI SDK & Agent FrameworkLlamaIndex5 of 12 · 42%—
Database Access & ORMsSQLAlchemy7 of 10 · 70%—
HostingCloudflare4 of 10 · 40%$12/mo · Workers Paid
Transactional EmailResend3 of 9 · 33%$20/mo · Pro 50k

Monthly bills add up to about $42 at the calculators' default usage, list prices. Set your own usage →

What they chose, decision by decision

Among the data scraping & pipelines that show each choice, from makers' products and open-source code alike.

Small samples, fewer than 8 products: Authentication (Clerk 3, Auth.js 2) · Background Jobs & Cron (Celery 3, RabbitMQ 2) · Payments (Stripe 5, Paddle 1) · Vector Database (LanceDB 4, Qdrant 3) · Image & Video Hosting (sharp 3, Cloudinary 1) · Product Analytics (PostHog 4, Mixpanel 1)

Data scraping & pipelines we track

21 makers' products and 17 open-source projects. Makers' products first.

AirtopBuilds AI agents that log in to websites, browse them and extract data, from a plain-English description of the task.
PulpMinerConverts Any Webpage Into Realtime JSON API 🟢
SaturnTurn Japan's public data into an AI-ready spreadsheet WS
fileAIClassify, extract, enrich, and validate any file
SliqAutomated data cleaning, in minutes, not hours or days.
AgentQLA query language and API that lets developers select and scrape data from web pages with natural-language-like queries instead of XPath or CSS selectors.
Supametas.AIStreamline unstructured data into LLM RAG-ready datasets
jitsuJitsu is an open-source Segment alternative. Fully-scriptable data ingestion engine for modern data teams. Set-up a real-time data pipeline in minutes, not days+9
deepcrawl100% free and full open-source edge Firecrawl alternative with better links extraction for agents - that you can deploy to cloudflare or vercel by yourself.+4
ScrapingdogScrapingdog is your all-in-one Web Scraping API, effortlessly managing proxies and headless browsers, allowing you to extract the data you need with ease.
cocoindexIncremental engine for long horizon agents 🌟 Star if you like it!+7
docetlA system for agentic LLM-powered data processing and ETL+4
unstractLLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows+4
DaftHigh-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale+3
vectara-ingestAn open source framework to crawl data sources and ingest into Vectara+3
felderaThe Feldera Incremental Computation Engine+1
nextflowA DSL for data-driven computational pipelines
xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fift
HeadlessXThe undetected self-hosted browser automation platform. Powered by Camoufox (Firefox) for 0% detection rates. Built for speed, privacy, and scalability.+2
NeMo-RetrieverNeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice. NeMo Retriever Library uses specialized NVIDIA NIM microservices to find, contextualiz+1
eyesPublic Opinion Mining System of Taiwanese Forums
lotusOptimized Agentic and LLM Bulk Processing Over Your Data
web-page-monitorWeb Site Page Changes Monitor. 网站网页页面更新变更监控提醒。
cjworkbenchThe data journalism platform with built in training
data-prep-kitOpen source project for data preparation for GenAI applications
DataFlow[SIGMOD'27] Easy Data Preparation with latest LLMs-based Operators and Pipelines.
doctorDoctor is a tool for discovering, crawl, and indexing web sites to be exposed as an MCP server for LLM agents.
spiderman基于 scrapy-redis 的通用分布式爬虫框架
zatoESB, SOA, REST, APIs and Cloud Integrations in Python
Crawling-InfrastructureDistributed crawling infrastructure running on top of severless computation, cloud storage (such as S3) and sophisticated queues.
PyAirbytePyAirbyte brings the power of Airbyte to every Python developer. Powers the Airbyte Cloud Replication MCP.
sparrowStructured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
weibospider:zap: A distributed crawler for weibo, building with celery and requests.
datacapDataCap is integrated software for data transformation, integration, and visualization. Support a variety of data sources, file types, big data related database, relational database, NoSQL database, e
timotaoshu提莫淘书,小说爬虫,用node爬书,node 小说,vue+express+node爬虫