@yaojingang: With the consent of our friend Ba Dao Liu, we are open-sourcing the dataset recently collected from major domestic AI platforms. The cleaned dataset, preprint paper, and first analysis report have been pushed to the GitHub repository (see comments for the link). This should be the latest and most comprehensive public GEO raw dataset for major domestic AI platforms, including Doubao, Dee...
Summary
Friend Ba Dao Liu open-sourced search result datasets from 8 domestic AI platforms, containing 620 standard questions and 210,000 citation records, along with cleaned data, a preprint paper, and an analysis report.
View Cached Full Text
Cached at: 07/21/26, 10:39 AM
With the consent of my friend Ba Dao Liu (拔刀刘), we are open-sourcing the datasets they recently crawled from major domestic AI platforms.
The cleaned dataset, preprint papers, and the first analysis report have been pushed to a GitHub repository. The link is in the comments.
This should be the latest and most comprehensive public GEO raw dataset for major domestic AI platforms.
It includes AI search results and source characteristic data from 8 AI platforms: Doubao, DeepSeek, Yuanbao, Qwen, Kimi, Wenxin, and Baidu AI.
The dataset includes:
- 620 standard questions, covering 6 research dimensions and 31 effective question types, including industry, query intent, prompt style, time sensitivity, trigger intensity, and extreme real-world scenarios.
- 214,119 AI citation observations, with 189,845 retained after exact deduplication. These statistics cover the cited sources and abstracts in replies; the full original reply text is not included.
- Citation records merged into 9,878 canonical domain sources and 107,659 canonical pages; 211,248 records carry valid HTTP URLs, accounting for 98.66%.
- Sources classified into 9 first-level types and 39 second-level types, with deduplicated citation classification coverage reaching 99.77%.
- Raw data, cleaned tables, and analysis datasets are retained and can be queried directly using JSONL, Parquet, and DuckDB. A field dictionary, cleaning rules, and quality report are also included.
- All 64 original shards pass SHA-256 checksums. The cleaning process preserves original values, anomaly markers, and traceability fields for reproducibility and further processing.
Over the past few days, we performed secondary cleaning, labeling, and processing on the data and output the first visual analysis. Some interesting findings and insights:
- Head effect of sources is very pronounced across major AI platforms. The Top 10 sources contribute 41.47% of deduplicated citations. Douyin, Tencent News, Toutiao, and Baidu Baike rank in the top four.
- Platform and community sources account for only 4.68% of domains but receive 29.63% of deduplicated citations. Content distribution platforms carry significant weight in the AI citation ecosystem.
- Brand and corporate official websites contribute only 4.33% of deduplicated citations; adding government and public institutions, owned sources total 9.06%. There is still substantial visibility room for brand-owned content.
- Cross-platform consensus is not as high as expected. Among 66 platform pairs, the highest “question × page” Jaccard similarity is only 34.94%, and 11 pairs have zero shared questions or pages. A single platform’s experience cannot be directly extrapolated to others.
- Content freshness is evident. Among pages with reliable publication dates, 88.77% come from 2025 or 2026; 29.57% of page titles contain a year, and 23.06% contain a list or ranking. Timeliness signals and decision-oriented content are worth further validation.
Feel free to leave comments with questions you care about. I can take the time to produce another data analysis report based on those topics.
Similar Articles
@yaojingang: I downloaded all papers related to GEO, AEO, and AI search from the past two years — 41 in total. Read through them one by one, and gained many new insights and inspirations. All 41 papers have been pushed to the GitHub repository. Feel free to download. The address is at the end of the article. Some key insights to share: 1. GEO is not a replacement for SEO. SEO addresses whether content can be retrieved, indexed, and entered into the candidate set...
The author collected and read 41 papers related to GEO, AEO, and AI search, pushed the collection to GitHub, and shared 10 key insights, including the relationship between GEO and SEO, AI search citation mechanisms, the importance of content structuring, and practical optimization directions.
@GenhuiP78950: Open-sourced my AI tools from the past six months. Not a big project, just scripts I use daily – transcribing Douyin/Bilibili videos, podcast-to-text, WeChat public account articles, industry news scanning… 11 in total. Used them privately, now unified with install scripts and docs.
Open-sourced a collection of 11 AI tool scripts for collecting and transcribing content from multiple channels like Douyin, Bilibili, and WeChat public accounts, making it easy to build a personal knowledge base. Supports direct installation by agents such as Claude Code, Codex, etc.
@jinchenma_ai: The best compilation of high-quality AI sources on the web, save it now! You can throw this article to Codex + Obsidian, let AI compile an index directory, then later when you want AI to search for quality information, you can let it search according to this directory.
Recommend an article that compiles the best high-quality AI sources on the web, and suggest using Codex and Obsidian to compile an index directory for future AI search of quality information.
@jinglian: AI Spark @AISpark1 has directly open-sourced its knowledge base. Not just testing the waters with a few articles, but fully releasing all accumulated content. 247 articles, 6 major modules, continuously updated. Open-source link in the comments, save and read at your leisure. ━━━━━━━━━━━━━━ ① AI Beginner's Guide (…
AI Spark has fully open-sourced its knowledge base, containing 247 articles across six major modules (from beginner's guide to industry insights), continuously updated, suitable for AI learners.
@yan5xu: Sharing an AI company research library I recently compiled: Oh My AI Company.
Author @yan5xu shares their compiled AI company research library 'Oh My AI Company', which includes 75 AI companies and products, 75 investment firms, etc., and introduces a research method combining Similarweb traffic analysis and multi-source cross-validation, aiming to provide a continuously updated market map.