scraping

Tag

Cards List
#scraping

@DavidOndrej1: Opus 5 is better at web search and scraping our research shows it's 69% better than GPT 5.6 Sol

X AI KOLs Timeline · 2026-08-03 Cached

A tweet shares DeepAPI benchmark results claiming Opus 5 outperforms GPT-5.6 Sol by 69% at creating web search queries, winning all 53 blind comparisons.

0 favorites 0 likes
#scraping

Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

Hacker News Top · 2026-07-27

A judge rejected Google's defense under the DMCA to avoid liability for being scraped, a significant ruling for web scraping and copyright law.

0 favorites 0 likes
#scraping

So Reddit has decided that plain HTML is unsafe

Lobsters Hottest · 2026-07-22 Cached

Reddit now requires login to use its old design (old.reddit.com), citing safety concerns about abusive scraping and automated traffic, which frustrates users who prefer the simpler interface.

0 favorites 0 likes
#scraping

Olostep

Product Hunt · 2026-07-17 Cached

Olostep is a web data API that enables extraction, crawling, and structuring of web data at scale, designed for AI teams, data pipelines, and automation.

0 favorites 0 likes
#scraping

Patreon stops asking AI bots not to scrape — and starts blocking them

TechCrunch AI · 2026-07-17 Cached

Patreon has partnered with Cloudflare to actively block AI bots from scraping creator content for training, moving beyond the voluntary robots.txt approach. The move aims to give creators more control over how their work is used by AI companies.

0 favorites 0 likes
#scraping

Suno snatched millions of songs from YouTube, Genius, and Deezer

The Verge · 2026-07-15 Cached

A hack reveals that AI music generator Suno trained its models by scraping millions of songs from YouTube, Genius, and Deezer, backing up allegations of copyright infringement and exposing customer data.

0 favorites 0 likes
#scraping

Hack suggests AI music generator Suno scraped YouTube for training data

TechCrunch AI · 2026-07-15 Cached

A hack into AI music generator Suno reveals it allegedly scraped decades of audio from YouTube, Deezer, and other sources for training data, raising copyright and DMCA concerns amid ongoing lawsuits from major record labels.

0 favorites 0 likes
#scraping

@Pluvio9yte: This is amazing. Apart from the Xiaohongshu API often going down and needing updates, it basically meets all your scraping needs. I've been using it for almost a month. See the comments for usage.

X AI KOLs Timeline · 2026-07-13 Cached

A user praises a data scraping tool, saying that aside from the occasional need to update the Xiaohongshu API, it basically meets all scraping needs.

0 favorites 0 likes
#scraping

shot-scraper 1.11

Simon Willison's Blog · 2026-07-12 Cached

shot-scraper 1.11 is released with minor improvements such as a longer wait time for server processes and new command options like --js-file and --timeout for consistency.

0 favorites 0 likes
#scraping

@gengdaJ: Awesome, now you can directly use Codex to automatically batch scrape WeChat official account articles. Original, multiple articles at once can be accurately scraped; text, likes, shares, comments, and read counts can also be scraped. Previously, I had to manually operate on the website, now just let Codex handle it, so comfortable~ For the text, just scan a QR code, the login status lasts 4 days. Other data...

X AI KOLs Timeline · 2026-07-09 Cached

Introduces a new yichen-skills tool wechat-mp-batch-exporter that allows using Codex to automatically batch scrape WeChat official account articles, including text, likes, shares, comments, read counts, and other data, with login status maintained for 4 days.

0 favorites 0 likes
#scraping

Show HN: Fortress – a stealth Chromium so your agents stop getting blocked

Hacker News Top · 2026-07-08 Cached

Fortress is a stealth Chromium engine that modifies browser fingerprints at the C++ level to help scrapers and browser agents avoid detection by bot detectors like Cloudflare Turnstile, CreepJS, and Sannysoft. It operates as a drop-in CDP replacement for Playwright and Puppeteer.

0 favorites 0 likes
#scraping

@gkxspace: Many people are asking what this website is. Let me explain: TikHub, simply put, is an API aggregation station for social platform data. It fully connects 16 platforms: Douyin, Xiaohongshu, TikTok, Instagram, YouTube, Bilibili, Weibo, Kuaishou, Zhihu, LinkedIn, WeChat Official Accounts, …

X AI KOLs Timeline · 2026-07-04 Cached

TikHub is an API aggregation station for social platform data, supporting 16 platforms (such as Douyin, Xiaohongshu, TikTok, etc.), providing over 1000 APIs for retrieving videos, comments, user profiles, e-commerce data, and more.

0 favorites 0 likes
#scraping

Every AI Visibility Tool Is Lying to You

Hacker News Top · 2026-07-03 Cached

This article critically examines the accuracy of AI visibility tools that claim to measure brand presence in generative AI responses, arguing that they provide false precision due to nondeterminism, personalization, and scraping biases. It calls for transparency in methodology and warns against treating opaque dashboards as stable truth.

0 favorites 0 likes
#scraping

@yhslgg: Why did I mark only this one as 'the most special' among 14 scraping tools? Lao Yang now explains clearly. Folks, this tool with completely different scraping logic — ScrapeGraphAI, 27,900 stars on GitHub. In a nutshell: you say 'help me scrape all the product names and prices from this page,' and the LLM automatically generates the scraping…

X AI KOLs Timeline · 2026-07-01 Cached

Introduces ScrapeGraphAI, an LLM-based scraping tool that can take natural language descriptions of requirements and automatically generate the scraping workflow, eliminating the need to write selectors or care about HTML structure, supporting local models and multiple integration platforms.

0 favorites 0 likes
#scraping

Reddit will require you to log in to use old.reddit.com

Ars Technica · 2026-06-30 Cached

Reddit will require users to log in to access old.reddit.com, citing the need to combat abusive scraping and automated traffic, which may upset users who prefer the old interface for its simplicity and privacy.

0 favorites 0 likes
#scraping

The Threat of Residential Proxies

Lobsters Hottest · 2026-06-30 Cached

Residential proxies, used for scraping and hiding nefarious activity, are rising as bots now generate more internet traffic than humans. The article explores the technical and ethical challenges, including payment protocols like Cloudflare and Coinbase's x402 standard.

0 favorites 0 likes
#scraping

@geekbb: A CLI tool written in Go that integrates three search capabilities: Web search (Brave/DDG/SearXNG/Exa), code search (Grep/Sourcegraph/GitHub), and library documentation query (Context7). It also supports web scraping and site crawling. For AI...

X AI KOLs Timeline · 2026-06-30 Cached

A blazing-fast, stateless CLI tool written in Go that integrates Web search, code search, and library documentation query. It supports web scraping and site crawling, designed for AI agents and terminal use.

0 favorites 0 likes
#scraping

Over 20 publishers sue OpenAI, Microsoft for training ChatGPT with their content

Reddit r/artificial · 2026-06-29 Cached

35 newspaper publishers across the US have filed a lawsuit against OpenAI and Microsoft, alleging that the companies scraped their copyrighted and paywalled content without permission to train ChatGPT, harming local journalism.

0 favorites 0 likes
#scraping

@sairahul1: https://x.com/sairahul1/status/2066809666592718879

X AI KOLs Timeline · 2026-06-16 Cached

A detailed guide on building an AI-powered content machine that scrapes viral content from platforms like TikTok and Instagram, uses AI to generate platform-native posts, and automates scheduling with tools like ScrapeCreators, Kie.ai, and Postiz.

0 favorites 0 likes
#scraping

browser sessions start failing at around 20 concurrent. nobody warns you about this

Reddit r/AI_Agents · 2026-06-12

Playwright scrapers in production on Node.js start failing around 20 concurrent browser sessions, causing memory spikes and crashes. The developer notes documentation does not warn about this limit.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback