Tag
This article provides a technical guide to scraping TikTok's private mobile API, resulting in a massive dataset of 4.5 billion videos that has been uploaded to Hugging Face for public access.
The author scraped 5.94 billion TikTok videos and 3.23 billion profiles using a reverse-engineering method and uploaded the full dataset to Hugging Face, with a tutorial and code provided for a fee.
Crawl4ai is a free, AI-oriented scraping tool that converts web content to Markdown, ideal for RAG and Agent development.
A historical researcher criticizes the practice of destroying antique books by cutting their spines for AI model training data, arguing it causes irreversible loss of valuable historical information and data that cannot be recovered.
AI companies are reportedly bulk-buying rare books, scanning them by cutting spines and shredding originals, facilitated by a service called ISBNdb, raising ethical concerns.
AI companies are bulk-buying rare books, scanning them by cutting spines and shredding the originals, facilitated by ISBNdb, with a federal judge ruling it fair use. This irreversible destruction of cultural artifacts has sparked outrage.
AI companies are purchasing and destroying rare physical books to obtain training data that is free from AI-generated content, exploiting legal doctrines like fair use and first-sale doctrine, while raising ethical and preservation concerns.
Codeberg is updating its Terms of Use to explicitly prohibit the extraction of data for training large language models (LLM-extrusions), aiming to protect the platform's repositories from unauthorized use.
SurfSense is an open-source competitive intelligence platform for AI agents, featuring live scraping connectors to Reddit, YouTube, Instagram, TikTok, Google Maps, Google Search, and more, with a REST API and MCP server.
Recommend a site called TikTokHub that provides social media APIs for all platforms globally at extremely low prices, covering Twitter, Xiaohongshu, TikTok, Douyin, Reddit, etc. It can be used for data scraping, monitoring, topic collection, and more.
The open-source project Agent Reach allows AI agents to access 14 internet platforms without complex configuration, solving the pain point of data acquisition for agents.
A commentary on the shifting attitudes towards web scraping for AI training, questioning the sudden condemnation of data collection without permission.
AgentKey is a beginner-friendly tool that allows AI agents to stably scrape data from platforms like X (Twitter) and Xiaohongshu, without manually configuring cookies or managing APIs — just one key is needed.
Data collected from Pokémon Go scans has reportedly been used to train navigation technology for military drones, raising concerns about unintended military applications of consumer data.
Promotes a GitHub repo that lets users describe a dataset in natural language and have AI agents research the web to build a structured table, exportable to CSV, with automatic refreshes.
BigSet is an open-source tool. You input a sentence describing the data you need, and it deploys multiple AI agents to research the web in parallel, automatically inferring schema, deduplicating, verifying, and generating a structured table. It supports scheduled refreshes.
A data enthusiast scraped Founders Podcast episodes to chart how host David Senra measures his dinner meetings with founders, revealing an accelerating social calendar and a median dinner length of 3 hours.
The author open-sourced a tool for fetching Binance smart money data, installable via npm, providing professional data including average holding cost of whales.
Introduces how to use XCrawl and Hermes Agent to build a no-code automated intelligence collection workflow, covering scenarios such as competitor monitoring, Twitter interaction radar, and Amazon product monitoring.
A real estate wholesaler built an OpenClaw skill called zillow-full to automate property research, increasing deal flow from 2 to 11 per month by using AI to score listings against personal criteria.