Tag
AI companies are purchasing and destroying rare physical books to obtain training data that is free from AI-generated content, exploiting legal doctrines like fair use and first-sale doctrine, while raising ethical and preservation concerns.
Codeberg is updating its Terms of Use to explicitly prohibit the extraction of data for training large language models (LLM-extrusions), aiming to protect the platform's repositories from unauthorized use.
SurfSense is an open-source competitive intelligence platform for AI agents, featuring live scraping connectors to Reddit, YouTube, Instagram, TikTok, Google Maps, Google Search, and more, with a REST API and MCP server.
Recommend a site called TikTokHub that provides social media APIs for all platforms globally at extremely low prices, covering Twitter, Xiaohongshu, TikTok, Douyin, Reddit, etc. It can be used for data scraping, monitoring, topic collection, and more.
The open-source project Agent Reach allows AI agents to access 14 internet platforms without complex configuration, solving the pain point of data acquisition for agents.
A commentary on the shifting attitudes towards web scraping for AI training, questioning the sudden condemnation of data collection without permission.
AgentKey is a beginner-friendly tool that allows AI agents to stably scrape data from platforms like X (Twitter) and Xiaohongshu, without manually configuring cookies or managing APIs — just one key is needed.
Data collected from Pokémon Go scans has reportedly been used to train navigation technology for military drones, raising concerns about unintended military applications of consumer data.
Promotes a GitHub repo that lets users describe a dataset in natural language and have AI agents research the web to build a structured table, exportable to CSV, with automatic refreshes.
BigSet is an open-source tool. You input a sentence describing the data you need, and it deploys multiple AI agents to research the web in parallel, automatically inferring schema, deduplicating, verifying, and generating a structured table. It supports scheduled refreshes.
A data enthusiast scraped Founders Podcast episodes to chart how host David Senra measures his dinner meetings with founders, revealing an accelerating social calendar and a median dinner length of 3 hours.
The author open-sourced a tool for fetching Binance smart money data, installable via npm, providing professional data including average holding cost of whales.
Introduces how to use XCrawl and Hermes Agent to build a no-code automated intelligence collection workflow, covering scenarios such as competitor monitoring, Twitter interaction radar, and Amazon product monitoring.
A real estate wholesaler built an OpenClaw skill called zillow-full to automate property research, increasing deal flow from 2 to 11 per month by using AI to score listings against personal criteria.
This article is a video tutorial demonstrating how to use the CodeX AI assistant and the Twitter CLI tool to scrape data from the X platform and convert it into Markdown format without needing an API Key.
Canadian federal and provincial privacy watchdogs have determined that OpenAI violated privacy laws by scraping vast amounts of personal data to train ChatGPT without proper consent.