CatchAll by NewsCatcher
Summary
CatchAll by NewsCatcher is a product for building customized datasets from the web based on user-defined criteria.
Similar Articles
NanmiCoder/MediaCrawler
MediaCrawler是一个开源的多平台自媒体数据采集工具,支持小红书、抖音、快手、B站、微博、贴吧、知乎等主流平台的公开信息抓取,基于Playwright浏览器自动化实现,无需JS逆向。
Self-Hosted AI news aggregator using Cloudflare Workers, Vectorize, and Nostr
A self-hosted AI news aggregator that uses Cloudflare Workers, Vectorize, D1, and Nostr to scrape sources like Hacker News, Lobsters, and Reddit, generate AI embeddings, and provide a personalized feed based on upvote history.
@NFTCPS: Finally found out where those repost accounts on X get their content! It's this tool MediaCrawler, a single tool that covers Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. It can scrape public content, comments, likes, and reposts. The best part is it doesn't need JS reverse engineering—it uses browser login state to get signatures directly, …
MediaCrawler is a multi-platform social media data scraping tool that supports public content crawling from Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. It bypasses JS reverse engineering by leveraging browser login state, lowering the technical barrier.
@grgerwcwetwet: Recommend an open-source project Horizon, an AI-powered information radar focused on overseas tech circles. It automatically aggregates content from Hacker News, Twitter, Reddit, GitHub and other platforms, then uses AI to filter, deduplicate and summarize, turning truly valuable information into a daily digest. A relatively...
Recommend the open-source project Horizon, an AI-driven overseas tech news radar. It automatically aggregates content from Hacker News, Twitter, Reddit, GitHub and other platforms, performs filtering, deduplication and summarization, generates bilingual (Chinese-English) daily reports, and supports pushing to Feishu, email, WeChat and other channels.
@Jolyne_AI: Another high-performance crawler/scraper found on GitHub: AnyCrawl — makes data collection easier and more efficient. It bundles three engines: Cheerio, Playwright, and Puppeteer: lightning-fast static page parsing, complex JavaScript rendering…
AnyCrawl is a high-performance open-source crawler/scraping tool that integrates Cheerio, Playwright, and Puppeteer engines. It supports static parsing and JS rendering, batch SERP scraping, site-level crawling, multi-threaded/multi-process concurrency, proxy support, and optimized output formats for LLM data collection.