web-scraping

Tag

Cards List
#web-scraping

Sites that block AI training crawlers mostly ignore the answer time bots

Hacker News Top · 2026-07-07 Cached

A study of robots.txt files from top 10,000 sites reveals that most block AI training crawlers like GPTBot while largely ignoring answer-time bots such as OAI-SearchBot, highlighting a blind spot in current web governance for AI.

0 favorites 0 likes
#web-scraping

If your agent reads a webpage, the page can tell it to lie about the page

Reddit r/AI_Agents · 2026-07-03

A developer built a non-AI-based checker that detects hidden instructions on web pages designed to deceive AI agents, addressing a vulnerability where pages can instruct agents to lie about their safety.

0 favorites 0 likes
#web-scraping

@Jolyne_AI: Another high-performance crawler/scraper found on GitHub: AnyCrawl — makes data collection easier and more efficient. It bundles three engines: Cheerio, Playwright, and Puppeteer: lightning-fast static page parsing, complex JavaScript rendering…

X AI KOLs Timeline · 2026-07-03 Cached

AnyCrawl is a high-performance open-source crawler/scraping tool that integrates Cheerio, Playwright, and Puppeteer engines. It supports static parsing and JS rendering, batch SERP scraping, site-level crawling, multi-threaded/multi-process concurrency, proxy support, and optimized output formats for LLM data collection.

0 favorites 0 likes
#web-scraping

@DanKornas: Copying web pages into LLMs shouldn’t mean dragging along the whole browser. .MD this page is a browser extension that …

X AI KOLs Timeline · 2026-07-02 Cached

.MD this page is an open-source browser extension that converts web pages into clean, LLM-ready Markdown using Mozilla's Readability, with features like one-click capture, preview, and export.

0 favorites 0 likes
#web-scraping

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection

arXiv cs.AI · 2026-07-02 Cached

This paper proposes a constrained, verifiable agent framework for open-web data collection that shifts LLM output from free-form code to typed JSON collector configurations, achieving zero execution-stage LLM tokens and low latency on 80 tasks.

0 favorites 0 likes
#web-scraping

Cloudflare’s new policy pushes AI companies to pay for publishers’ content

TechCrunch AI · 2026-07-01 Cached

Cloudflare announces a new default policy to block mixed-use web crawlers from ad-hosted pages starting September 2026, aiming to force AI companies to pay for publishers' content and distinguish search from AI training.

0 favorites 0 likes
#web-scraping

@yhslgg: https://x.com/yhslgg/status/2072243790044442961

X AI KOLs Timeline · 2026-07-01 Cached

This article categorizes 14 web scraping tools into five groups: AI new paradigms, engineering-grade frameworks, browser automation, China-specific platforms, and modern lightweight tools, accompanied by real-world cases and selection recommendations.

0 favorites 0 likes
#web-scraping

Context.dev

Product Hunt · 2026-07-01

Context.dev provides a single API for scraping, enriching, and extracting data from the internet.

0 favorites 0 likes
#web-scraping

@sharbel: Someone built a single CLI that reads Twitter, Reddit, YouTube, GitHub, Bilibili, and XiaoHongShu. Zero API fees. Zero …

X AI KOLs Timeline · 2026-06-30 Cached

Agent Reach is a single CLI tool that gives AI agents instant access to Twitter, Reddit, YouTube, GitHub, Bilibili, and XiaoHongShu without API keys or fees. It is 100% open source and compatible with various agent frameworks.

0 favorites 0 likes
#web-scraping

@NousResearch: Hermes Agent now reads the web up to 60x faster and 49x cheaper. Scraping backends pass clean content straight to the a…

X AI KOLs Following · 2026-06-30 Cached

NousResearch's Hermes Agent now reads web pages up to 60x faster and 49x cheaper by using optimized scraping backends and paging large pages locally.

0 favorites 0 likes
#web-scraping

@PrajwalTomar_: Give this to your agent ↓ https://apify.it/x402-awal Setup docs if you want the walkthrough ↓ https://docs.apify.com/pl…

X AI KOLs Following · 2026-06-30 Cached

This post describes how to set up a Coinbase Agentic Wallet to use Apify's web-data tools by paying USDC on Base over the x402 protocol, eliminating the need for an Apify account or API key.

0 favorites 0 likes
#web-scraping

@CycleDecoded: Stop paying a fortune for crappy crawler software and paying the IQ tax! This open-source tool is insane — it directly exposes all social media platforms' data. The game for traffic matrix and data monetization players is over! Meet MediaCrawler, a GitHub project with over 54,000 stars. In plain language…

X AI KOLs Timeline · 2026-06-30 Cached

MediaCrawler is an open-source multi-platform social media crawler tool with over 54,000 stars on GitHub. It supports data collection from 7 major platforms including Xiaohongshu, Douyin, Bilibili, etc., and features multi-account, IP proxy, breakpoint resume, and AI integration.

0 favorites 0 likes
#web-scraping

@PrajwalTomar_: Hermes agent pro tip. The upgrade that made my agents actually useful had nothing to do with the model. It was giving t…

X AI KOLs Following · 2026-06-29 Cached

A pro tip for Hermes agents: use the free open-source repo Agent Reach to give agents internet access, enabling them to read YouTube, Reddit, X, LinkedIn, and GitHub, reducing token costs and improving research efficiency.

0 favorites 0 likes
#web-scraping

@Jolyne_AI: When doing web scraping and data collection, proxy servers are almost a must, otherwise your IP will be blocked in no time. The bigger problem is that most of the "free proxies" online are dead. I just dug up an open-source project on GitHub: Proxifly. It automatically fetches, updates, and verifies free proxies every 5 minutes, making the "usable" aspect...

X AI KOLs Timeline · 2026-06-29 Cached

Introduced the open-source project Proxifly on GitHub, which automatically fetches and verifies available proxies every 5 minutes, covers 60+ countries, supports multiple protocols, and provides downloads in multiple formats as well as npm installation.

0 favorites 0 likes
#web-scraping

@ecommartinez: 10 GitHub Repositories for Scraping the Entire Internet Save them all. Each one extracts clean data from any website. T…

X AI KOLs Timeline · 2026-06-28 Cached

Tweet de @ecommartinez que lista 10 repositorios de GitHub para hacer web scraping y extraer datos limpios de cualquier sitio web.

0 favorites 0 likes
#web-scraping

So now scraping data without permission is bad for AI training all of sudden?

Reddit r/artificial · 2026-06-27

A commentary on the shifting attitudes towards web scraping for AI training, questioning the sudden condemnation of data collection without permission.

0 favorites 0 likes
#web-scraping

@DanKornas: Build web-scraping agents without hand-wiring every step. Open Agent Builder is a visual workflow builder for creating …

X AI KOLs Timeline · 2026-06-25 Cached

Open Agent Builder is a visual workflow builder for creating AI agent pipelines powered by Firecrawl, enabling drag-and-drop design of multi-step scraping and automation workflows with real-time execution updates.

0 favorites 0 likes
#web-scraping

@laowangbabababa: This is absolutely incredible. Seriously, absolutely incredible. It can truly crawl anything, and it's blazing fast! Previously, I would ask Claude Code to search for something online, but the Twitter API costs money, web scraping requires subscriptions, and Bilibili and Xiaohongshu directly block you. Every platform has its own barrier — such a hassle. Agent-Reach, a free and open-source repository...

X AI KOLs Timeline · 2026-06-24 Cached

Agent-Reach is a free and open-source repository that, with a single command, gives AI agents access to 14 platforms, including web pages, YouTube, RSS, Bilibili, and Xiaohongshu, without complex configuration, significantly lowering the barrier for agents to go online.

0 favorites 0 likes
#web-scraping

@NFTCPS: Finally found out where those repost accounts on X get their content! It's this tool MediaCrawler, a single tool that covers Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. It can scrape public content, comments, likes, and reposts. The best part is it doesn't need JS reverse engineering—it uses browser login state to get signatures directly, …

X AI KOLs Timeline · 2026-06-23 Cached

MediaCrawler is a multi-platform social media data scraping tool that supports public content crawling from Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. It bypasses JS reverse engineering by leveraging browser login state, lowering the technical barrier.

0 favorites 0 likes
#web-scraping

Show HN: Selector Forge – browser extension for AI-generated resilient selectors

Hacker News Top · 2026-06-22 Cached

Selector Forge is a browser extension that uses AI to generate and verify reliable CSS/XPath selectors for web automation, helping developers build robust selectors for testing, scraping, and page automation.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback