@NFTCPS: Finally found out where those repost accounts on X get their content! It's this tool MediaCrawler, a single tool that covers Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. It can scrape public content, comments, likes, and reposts. The best part is it doesn't need JS reverse engineering—it uses browser login state to get signatures directly, …
Summary
MediaCrawler is a multi-platform social media data scraping tool that supports public content crawling from Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. It bypasses JS reverse engineering by leveraging browser login state, lowering the technical barrier.
View Cached Full Text
Cached at: 06/23/26, 04:13 PM
So that’s where those repost bloggers on X get their content! It’s this tool, MediaCrawler — a one-stop tool that can scrape public content, comments, likes, and reposts from Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. The best part is it doesn’t require JS reverse engineering; it uses browser login state to get signatures directly, lowering the barrier significantly. If you know some Python, you can run it. Of course, the author makes it clear: for learning and research only, don’t use it for illegal activities. https://github.com/NanmiCoder/MediaCrawler… Recently, X might be cracking down on repost-type accounts, so be moderate! — # NanmiCoder/MediaCrawler Source: https://github.com/NanmiCoder/MediaCrawler # 🔥 MediaCrawler - Social Media Platform Crawler 🕷️ GitHub Stars (https://github.com/NanmiCoder/MediaCrawler/stargazers) GitHub Forks (https://github.com/NanmiCoder/MediaCrawler/network/members) GitHub Issues (https://github.com/NanmiCoder/MediaCrawler/issues) GitHub Pull Requests (https://github.com/NanmiCoder/MediaCrawler/pulls) License (https://github.com/NanmiCoder/MediaCrawler/blob/main/LICENSE) 中文 English Español > Disclaimer: > > Please use this repository for learning purposes only ⚠️⚠️⚠️⚠️. For cases of illegal crawling, see: (https://github.com/HiddenStrawberry/Crawler_Illegal_Cases_In_China) > > All content in this repository is for learning and reference only. Commercial use is prohibited. No individual or organization may use the content of this repository for illegal purposes or infringe upon the legitimate rights of others. The crawling techniques covered in this repository are for learning and research only and must not be used for large-scale crawling or other illegal activities on any platform. This repository assumes no responsibility for any legal liabilities arising from the use of its content. By using the content of this repository, you agree to all terms and conditions of this disclaimer. > > Click to view the detailed disclaimer. Click here ## 📖 Project Introduction A powerful multi-platform social media data collection tool that supports scraping public information from major platforms like Xiaohongshu (RED), Douyin (TikTok), Kuaishou, Bilibili, Weibo, Tieba, Zhihu, etc. ### 🔧 Technical Principles - Core Technology: Based on Playwright (https://playwright.dev/) browser automation framework to save login state - No JS Reverse Engineering: Uses a browser context with preserved login state to obtain signature parameters via JS expressions - Advantages: No need to reverse complex encryption algorithms, significantly lowering the technical barrier ## ✨ Feature Overview | Platform | Keyword Search | Scrape by Post ID | Nested Comments | Scrape Creator Home | Login State Cache | IP Proxy Pool | Comment Word Cloud | |———|—————|—————––|––––––––|———————|——————|—————|––––––––––| | Xiaohongshu | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | Douyin | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | Kuaishou | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | Bilibili | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | Weibo | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | Tieba | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | Zhihu | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ## 🚀 MediaCrawlerPro Major Release! Open source is not easy, welcome to subscribe for support > Focus on learning the architecture design of mature projects — not just crawling techniques. The code design ideas of the Pro version are worth studying in depth! MediaCrawlerPro (https://github.com/MediaCrawlerPro) Key advantages over the open-source version: #### 🎯 Core Feature Upgrades - ✅ Social Media Content Decomposition Agent (New Feature) - ✅ Resume Interrupted Crawls (Key Feature) - ✅ Multi-Account + IP Proxy Pool Support (Key Feature) - ✅ Remove Playwright Dependency, easier to use - ✅ Full Linux Environment Support #### 🏗️ Architecture Design Improvements - ✅ Code Refactoring and Optimization, more readable and maintainable (decoupled JS signature logic) - ✅ Enterprise-Level Code Quality, suitable for building large-scale crawler projects - ✅ Perfect Architecture Design, high scalability, greater learning value from source code #### 🎁 Additional Features - ✅ Desktop Social Media Video Downloader (great for learning full-stack development) - ✅ Multi-Platform Home Feed Recommendations (HomeFeed) - ✅ AI Agent Skill Support (OpenClaw (https://openclaw.ai/) 🦞 / Claude Code / Cursor one-click install, let Agent automatically crawl data) - [ ] Comment Analysis AI Agent In Development 🚀🚀 Click to view: MediaCrawlerPro Project Homepage (https://github.com/MediaCrawlerPro) More details ## 🚀 Quick Start > 💡 If this project is helpful to you, please give it a ⭐ Star! ## 📋 Prerequisites ### 🚀 Install uv (Recommended) Before proceeding, ensure uv is installed on your computer: - Installation: uv Official Installation Guide (https://docs.astral.sh/uv/getting-started/installation) - Verify: Run uv --version in terminal. If the version number displays, installation was successful - Why recommended: uv is the fastest Python package manager, with fast speed and accurate dependency resolution ### 🟢 Node.js Installation The project depends on Node.js. Please download and install from the official website: - Download: https://nodejs.org/en/download/ - Version Requirement: >= 16.0.0 ### 📦 Install Python Packages shell # Enter project directory cd MediaCrawler # Use uv sync to ensure Python version and dependency consistency uv sync ### 🌐 Install Browser Driver (Optional) > If using the default CDP mode (connecting to an existing Chrome browser), no browser driver installation is needed. Only required when using standard Playwright mode. shell # Install browser driver only for standard Playwright mode uv run playwright install ### 🌍 Chrome Browser Configuration (Recommended) By default, the project uses CDP mode to connect to your existing Chrome browser, reusing existing login states, cookies, extensions, etc., significantly reducing the risk of platform anti-crawling detection. Before use: 1. Install the latest Chrome browser (version >= 144), download (https://www.google.com/chrome/) 2. Enable remote debugging: Enter chrome://inspect/#remote-debugging in Chrome address bar, check “Allow remote debugging for this browser instance” 3. When the page shows Server running at: 127.0.0.1:9222, it’s ready > 💡 Tip: After running the crawler, a confirmation dialog will appear in Chrome — click “Accept”. The program will wait for user confirmation; complete the operation within 60 seconds. > > If you prefer not to use CDP mode, set ENABLE_CDP_MODE = False in config/base_config.py to switch to standard Playwright mode. ## 🚀 Run the Crawler shell # View configuration options in config/base_config.py (with Chinese comments) # Read keyword search posts from config and scrape post info + comments uv run main.py --platform xhs --lt qrcode --type search # Read a list of specific post IDs from config and scrape post info + comments uv run main.py --platform xhs --lt qrcode --type detail # Open the corresponding app to scan QR code and login # Examples for other platforms — run the following command to view: uv run main.py --help ## 🖥️ WebUI Visual Interface MediaCrawler provides a web-based visual interface, allowing you to use the crawler without command line. #### Start WebUI Service shell # Start API server (default port 8080) uv run uvicorn api.main:app --port 8080 --reload # Or start as a module uv run python -m api.main After successful startup, visit http://localhost:8080 to open the WebUI interface. #### WebUI Features - Visual configuration of crawler parameters (platform, login method, crawl type, etc.) - Real-time view of crawler status and logs - Data preview and export #### Interface Preview 🔗 Using Python native venv (Not Recommended) #### Create and Activate Python Virtual Environment > For crawling Douyin and Zhihu, Node.js (version >= 16) is required. shell # Enter project root cd MediaCrawler # Create virtual environment # My Python version is 3.11 — requirements.txt is based on this version # If using another Python version, dependencies may be incompatible — solve manually python -m venv venv # macOS & Linux: activate source venv/bin/activate # Windows: activate venv\Scripts\activate #### Install Dependencies shell pip install -r requirements.txt #### Install Playwright Browser Driver shell playwright install #### Run Crawler (Native Environment) shell # By default, comment crawling is disabled. To enable, modify ENABLE_GET_COMMENTS in config/base_config.py # Other options can also be found in config/base_config.py (with Chinese comments) # Read keyword search posts from config and scrape post info + comments python main.py --platform xhs --lt qrcode --type search # Read specific post IDs from config and scrape post info + comments python main.py --platform xhs --lt qrcode --type detail # Open the corresponding app to scan QR code and login # Examples for other platforms — run: python main.py --help ## 💾 Data Storage MediaCrawler supports multiple storage formats: CSV, JSON, JSONL, Excel, SQLite, and MySQL databases. 📖 Detailed instructions: Data Storage Guide ## 💬 Community Groups - WeChat Group: Join (https://nanmicoder.github.io/MediaCrawler/%E5%BE%AE%E4%BF%A1%E4%BA%A4%E6%B5%81%E7%BE%A4.html) - Bilibili: Follow me (https://space.bilibili.com/434377496), sharing AI and crawling techniques ## 💰 Sponsor Showcase TikHub.io provides 900+ high-stability data interfaces covering TK, DY, XHS, Y2B, Ins, X and 14+ major domestic and international platforms. Provides public data APIs for users, content, products, comments, etc., along with 40M+ cleaned structured datasets. Use invitation code cfzyejV9 to register and top up, receive an extra $2 bonus. Atlas Cloud is a full-modality AI inference platform, allowing developers to access video generation, image generation, and LLM APIs via a unified AI API — call 300+ curated models without maintaining multiple vendor integrations. Atlas Cloud recently launched a coding plan offer for developers with cost-effective API access. — ## 🤝 Become a Sponsor Become a sponsor and showcase your product here for daily exposure! Contact: - WeChat: relakkes - Email: [email protected] — ## ☕ Buy Me a Coffee If this project helps you, feel free to tip — every bit of support keeps me updating ❤️ WeChat Tip Alipay Buy Me a Coffee — ## 📚 Other - FAQ: MediaCrawler Full Documentation (https://nanmicoder.github.io/MediaCrawler/) - Crawler Tutorial: CrawlerTutorial Free Course (https://github.com/NanmiCoder/CrawlerTutorial) - News Crawler Project: NewsCrawlerCollection (https://github.com/NanmiCoder/NewsCrawlerCollection) ## ⭐ Star History Chart If this project helps you, please give it a ⭐ Star to let more people know about MediaCrawler! Star History Chart (https://star-history.com/#NanmiCoder/MediaCrawler&Date) ## 📚 References - Xiaohongshu Signature Repository: Cloxl’s xhs Signature Repo (https://github.com/Cloxl/xhshow) - Xiaohongshu Client: ReaJason’s xhs Repo (https://github.com/ReaJason/xhs) - SMS Forwarder: SmsForwarder Reference Repo (https://github.com/pppscn/SmsForwarder) - Intranet Penetration Tool: ngrok Official Docs (https://ngrok.com/docs/) # Disclaimer ## 1. Project Purpose and Nature This project (hereinafter referred to as “this project”) is created as a technical research and learning tool, aiming to explore and learn web data collection techniques. This project focuses on researching social media platform data crawling techniques for technical exchange among learners and researchers. ## 2. Legal Compliance Statement The developers of this project (hereinafter “developers”) seriously remind users to strictly comply with relevant laws and regulations of the People’s Republic of China when downloading, installing, and using this project, including but not limited to the Cybersecurity Law of the People’s Republic of China, the Anti-Espionage Law of the People’s Republic of China, and all other applicable national laws and policies. Users shall bear all legal responsibilities arising from the use of this project. ## 3. Use Purpose Restriction This project is strictly prohibited from any illegal purpose or non-learning, non-research commercial behavior. This project must not be used for any form of illegal intrusion into others’ computer systems, nor for any infringement of others’ intellectual property rights or other legitimate rights. Users must ensure their use of this project is purely for personal learning and technical research, not for any illegal activities. ## 4. Disclaimer The developers have made the best efforts to ensure the legitimacy and safety of this project but assume no liability for any direct or indirect losses arising from the user’s use of this project, including but not limited to data loss, device damage, legal proceedings, etc. ## 5. Intellectual Property Statement The intellectual property rights of this project belong to the developers. This project is protected by copyright law, international copyright treaties, and other intellectual property laws and treaties. Users may download and use this project under the premise of complying with this statement and relevant laws and regulations. ## 6. Right of Final Interpretation The right of final interpretation of this project belongs to the developers. The developers reserve the right to change or update this disclaimer at any time without prior notice.
Similar Articles
@WY_mask: MediaCrawler: Open-source web scraping tool for Xiaohongshu, Douyin, Weibo, Bilibili, Kuaishou. Supports scraping videos, images, comments, likes, reposts, etc. https://github.com/NanmiCoder/MediaCrawler…
MediaCrawler is an open-source multi-platform self-media data collection tool that supports scraping public information from Xiaohongshu, Douyin, Weibo, Bilibili, Kuaishou and other platforms. No JS reverse engineering required, based on Playwright browser automation.
NanmiCoder/MediaCrawler
MediaCrawler是一个开源的多平台自媒体数据采集工具,支持小红书、抖音、快手、B站、微博、贴吧、知乎等主流平台的公开信息抓取,基于Playwright浏览器自动化实现,无需JS逆向。
@CycleDecoded: Stop paying a fortune for crappy crawler software and paying the IQ tax! This open-source tool is insane — it directly exposes all social media platforms' data. The game for traffic matrix and data monetization players is over! Meet MediaCrawler, a GitHub project with over 54,000 stars. In plain language…
MediaCrawler is an open-source multi-platform social media crawler tool with over 54,000 stars on GitHub. It supports data collection from 7 major platforms including Xiaohongshu, Douyin, Bilibili, etc., and features multi-account, IP proxy, breakpoint resume, and AI integration.
@AmberTreelet: Tiance Ge shared yt-dlp for scraping Douyin, YouTube, Bilibili, Twitter. I'll add some universal scraping tools. FxTwitter: Recommended by @0xCheshire for scraping X. get笔记 (Dedao Brain): WeChat Official Accounts, Xiaohongshu, Douyin, Bilibili, X, Podcasts. Google Chrome extension obsidian web clipp…
Introduces multiple web scraping tools, including yt-dlp, FxTwitter, get笔记, etc., for scraping content from different platforms.
@GYLQ520: Hey self-media friends, pay attention! If you miss this tool, you'll regret it big time. MultiPost, a free open-source browser plugin on GitHub, lets you push content to over a dozen platforms like Weibo, Zhihu, Xiaohongshu with one click, no more copy-pasting one by one. Supports text, images, and videos, plus scheduled posting, auto web scraping, and more...
MultiPost is a free open-source browser plugin on GitHub. It supports one-click syncing and publishing to over a dozen major platforms including Weibo, Zhihu, Xiaohongshu. It features text/image/video publishing, scheduled posting, web scraping, and AI content generation, significantly reducing multi-platform operational costs.