data-scraping

Tag

Cards List
#data-scraping

AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain

Reddit r/singularity · 15h ago Cached

AI companies are purchasing and destroying rare physical books to obtain training data that is free from AI-generated content, exploiting legal doctrines like fair use and first-sale doctrine, while raising ethical and preservation concerns.

0 favorites 0 likes
#data-scraping

Codeberg: ToU extension to prohibit LLM-extrusions

Hacker News Top · 5d ago

Codeberg is updating its Terms of Use to explicitly prohibit the extraction of data for training large language models (LLM-extrusions), aiming to protect the platform's repositories from unauthorized use.

0 favorites 0 likes
#data-scraping

FOSS NotebookLM connected to Reddit,YT,Google,IG,Tiktok etc

Reddit r/LocalLLaMA · 2026-07-14 Cached

SurfSense is an open-source competitive intelligence platform for AI agents, featuring live scraping connectors to Reddit, YouTube, Instagram, TikTok, Google Maps, Google Search, and more, with a REST API and MCP server.

0 favorites 0 likes
#data-scraping

@gengdaJ: Stunned, this TikTokHub site is absolutely mind-blowing... It has social media APIs for all platforms at home and abroad. Twitter API costs less than one cent per call, and even Xiaohongshu, which has always been the hardest to deal with, is only 7 cents. TikTok, Reddit, WeChat Official Accounts, Douyin are a piece of cake. This is now on my core tool list...

X AI KOLs Timeline · 2026-07-04 Cached

Recommend a site called TikTokHub that provides social media APIs for all platforms globally at extremely low prices, covering Twitter, Xiaohongshu, TikTok, Douyin, Reddit, etc. It can be used for data scraping, monitoring, topic collection, and more.

0 favorites 0 likes
#data-scraping

@1YES_yes1: Guys, I really struck gold today. I feel like many people still haven't realized this. **The biggest problem with Agents has never been that the model isn't strong enough.** It's: **They simply can't get the data.** You ask an Agent to find information online. But the reality is: Twitter/X API costs money. Web scraping requires a subscription. …

X AI KOLs Timeline · 2026-06-29 Cached

The open-source project Agent Reach allows AI agents to access 14 internet platforms without complex configuration, solving the pain point of data acquisition for agents.

0 favorites 0 likes
#data-scraping

So now scraping data without permission is bad for AI training all of sudden?

Reddit r/artificial · 2026-06-27

A commentary on the shifting attitudes towards web scraping for AI training, questioning the sudden condemnation of data collection without permission.

0 favorites 0 likes
#data-scraping

@MagiccatMila: Finally found a tool that can stably scrape data from X (Twitter) and Xiaohongshu! AgentKey -- a more beginner-friendly, stable, and improved version [Agent Reach] http://agentkey.app/@cirila I've always wanted to scrape external data but didn't know how to configure cookies and manage…

X AI KOLs Timeline · 2026-06-25 Cached

AgentKey is a beginner-friendly tool that allows AI agents to stably scrape data from platforms like X (Twitter) and Xiaohongshu, without manually configuring cookies or managing APIs — just one key is needed.

0 favorites 0 likes
#data-scraping

Pokémon Go Scans Trained the Navigation Tech for Military Drones

Hacker News Top · 2026-06-11

Data collected from Pokémon Go scans has reportedly been used to train navigation technology for military drones, raising concerns about unintended military applications of consumer data.

0 favorites 0 likes
#data-scraping

@DivyanshT91162: I think GitHub repos are quietly replacing half the SaaS industry. This one turns a prompt into a live dataset. Type: "…

X AI KOLs Timeline · 2026-06-03 Cached

Promotes a GitHub repo that lets users describe a dataset in natural language and have AI agents research the web to build a structured table, exportable to CSV, with automatic refreshes.

0 favorites 0 likes
#data-scraping

@gkxspace: Found a crazy open-source tool. You input a sentence describing what data you want, and it deploys a group of AI agents to research on various websites in parallel. After a few minutes, it compiles a structured table for you. In fact, the data is all on the internet, but turning it into a usable table has always been a labor-intensive task. In the past, this was an engineering project: combining searches, writing crawlers...

X AI KOLs Timeline · 2026-06-03 Cached

BigSet is an open-source tool. You input a sentence describing the data you need, and it deploys multiple AI agents to research the web in parallel, automatically inferring schema, deduplicating, verifying, and generating a structured table. It supports scheduled refreshes.

0 favorites 0 likes
#data-scraping

@tylerjrichards: Every time David Senra goes to dinner with a founder, he’ll tell the world via @FoundersPodcast all about it, including…

X AI KOLs Timeline · 2026-05-31 Cached

A data enthusiast scraped Founders Podcast episodes to chart how host David Senra measures his dinner meetings with founders, revealing an accelerating social calendar and a median dinner length of 3 hours.

0 favorites 0 likes
#data-scraping

@0xBenniee: The smart money dashboard developed by Binance @binancezh is very useful. Domestic wild funds often use this data to pump prices. By observing the average cost basis of smart money, it's easier to find second-entry points. Open-sourced a smart-money data set, after npm install you can use it directly to get all smart money data...

X AI KOLs Timeline · 2026-05-25 Cached

The author open-sourced a tool for fetching Binance smart money data, installable via npm, providing professional data including average holding cost of whales.

0 favorites 0 likes
#data-scraping

@wsl8297: https://x.com/wsl8297/status/2054798253955375388

X AI KOLs Timeline · 2026-05-14 Cached

Introduces how to use XCrawl and Hermes Agent to build a no-code automated intelligence collection workflow, covering scenarios such as competitor monitoring, Twitter interaction radar, and Amazon product monitoring.

0 favorites 0 likes
#data-scraping

shipped my first openclaw skill: zillow-full. built it because manual property research was eating my weekends.

Reddit r/openclaw · 2026-05-12

A real estate wholesaler built an OpenClaw skill called zillow-full to automate property research, increasing deal flow from 2 to 11 per month by using AI to score listings against personal criteria.

0 favorites 0 likes
#data-scraping

@kfk_ai: Step-by-step guide to fetching X platform data using tools | Easy for beginners. Browse the recommended feed, scrape posts from top influencers, and download long threads as Markdown. No API Key required—just use the CodeX AI assistant to do it all in one click. 00:00 Intro 00:53 Tool overview: No API, no cost, auto-detects login status…

X AI KOLs Timeline · 2026-05-12

This article is a video tutorial demonstrating how to use the CodeX AI assistant and the Twitter CLI tool to scrape data from the X platform and convert it into Markdown format without needing an API Key.

0 favorites 0 likes
#data-scraping

OpenAI violated Canadian privacy laws, federal and provincial watchdogs say

Reddit r/ArtificialInteligence · 2026-05-07 Cached

Canadian federal and provincial privacy watchdogs have determined that OpenAI violated privacy laws by scraping vast amounts of personal data to train ChatGPT without proper consent.

0 favorites 0 likes
← Back to home

Submit Feedback