@yaojingang: With the consent of our friend Ba Dao Liu, we are open-sourcing the dataset recently collected from major domestic AI platforms. The cleaned dataset, preprint paper, and first analysis report have been pushed to the GitHub repository (see comments for the link). This should be the latest and most comprehensive public GEO raw dataset for major domestic AI platforms, including Doubao, Dee...

X AI KOLs Timeline Tools

Summary

Friend Ba Dao Liu open-sourced search result datasets from 8 domestic AI platforms, containing 620 standard questions and 210,000 citation records, along with cleaned data, a preprint paper, and an analysis report.

With the consent of our friend Ba Dao Liu, we are open-sourcing the dataset recently collected from major domestic AI platforms. The cleaned dataset, preprint paper, and first analysis report have been pushed to the GitHub repository (see comments for the link). This should be the latest and most comprehensive public GEO raw dataset for major domestic AI platforms. It includes AI search results and source characteristics from 8 AI platforms: Doubao, DeepSeek, Yuanbao, Qianwen, Kimi, Wenxin, Baidu AI, and others. This dataset includes: 1. 620 standard questions covering 6 research dimensions and 31 valid question types, including industry, query intent, prompt style, time sensitivity, trigger intensity, and extreme real-world scenarios. 2. 214,119 AI citation observations, reduced to 189,845 after precise deduplication. These refer to cited sources and summaries within answers; the data does not include full response text. 3. Citations aggregated into 9,878 canonical domain sources and 107,659 canonical pages; 211,248 records (98.66%) have valid HTTP URLs. 4. Sources classified into 9 primary categories and 39 secondary categories, with 99.77% coverage of deduplicated citations. 5. Raw data, cleaned tables, and analysis repositories are retained and can be queried directly via JSONL, Parquet, and DuckDB. A field dictionary, cleaning rules, and quality report are also included. 6. All 64 original shards pass SHA-256 checksums. The cleaning process preserves original values, anomaly markers, and traceability fields for reproducibility and further processing. Over the past few days, we performed secondary cleaning, annotation, and processing, outputting the first visual analysis. Some interesting findings and insights: 1. The source head effect is very pronounced across major AI platforms: the top 10 sources contribute 41.47% of deduplicated citations. Douyin, Tencent News, Toutiao, and Baidu Baike rank in the top four. 2. Platform and community-type sources account for only 4.68% of domain names but capture 29.63% of deduplicated citations. Content distribution platforms play a significant role in the AI citation ecosystem. 3. Brand and corporate official websites contribute only 4.33% of deduplicated citations; combined with government and public institutions, self-owned sources account for 9.06%. There is still substantial room for brand-owned content visibility. 4. Cross-platform consensus is lower than expected: among 66 platform combinations, the highest "question × page" Jaccard similarity is only 34.94%, and 11 combinations share no questions or pages at all. Experiences from a single platform cannot be directly generalized to others. 5. Content freshness is evident: among pages with reliable publication timestamps, 88.77% come from 2025 or 2026; 29.57% of page titles include a year, and 23.06% include rankings or lists. Timeliness signals and decision-oriented content warrant further validation. Feel free to leave comments with topics of interest, and I can take the time to produce another data analysis report based on those questions.
Original Article
View Cached Full Text

Cached at: 07/21/26, 10:39 AM

With the consent of my friend Ba Dao Liu (拔刀刘), we are open-sourcing the datasets they recently crawled from major domestic AI platforms.

The cleaned dataset, preprint papers, and the first analysis report have been pushed to a GitHub repository. The link is in the comments.

This should be the latest and most comprehensive public GEO raw dataset for major domestic AI platforms.
It includes AI search results and source characteristic data from 8 AI platforms: Doubao, DeepSeek, Yuanbao, Qwen, Kimi, Wenxin, and Baidu AI.

The dataset includes:

  1. 620 standard questions, covering 6 research dimensions and 31 effective question types, including industry, query intent, prompt style, time sensitivity, trigger intensity, and extreme real-world scenarios.
  2. 214,119 AI citation observations, with 189,845 retained after exact deduplication. These statistics cover the cited sources and abstracts in replies; the full original reply text is not included.
  3. Citation records merged into 9,878 canonical domain sources and 107,659 canonical pages; 211,248 records carry valid HTTP URLs, accounting for 98.66%.
  4. Sources classified into 9 first-level types and 39 second-level types, with deduplicated citation classification coverage reaching 99.77%.
  5. Raw data, cleaned tables, and analysis datasets are retained and can be queried directly using JSONL, Parquet, and DuckDB. A field dictionary, cleaning rules, and quality report are also included.
  6. All 64 original shards pass SHA-256 checksums. The cleaning process preserves original values, anomaly markers, and traceability fields for reproducibility and further processing.

Over the past few days, we performed secondary cleaning, labeling, and processing on the data and output the first visual analysis. Some interesting findings and insights:

  1. Head effect of sources is very pronounced across major AI platforms. The Top 10 sources contribute 41.47% of deduplicated citations. Douyin, Tencent News, Toutiao, and Baidu Baike rank in the top four.
  2. Platform and community sources account for only 4.68% of domains but receive 29.63% of deduplicated citations. Content distribution platforms carry significant weight in the AI citation ecosystem.
  3. Brand and corporate official websites contribute only 4.33% of deduplicated citations; adding government and public institutions, owned sources total 9.06%. There is still substantial visibility room for brand-owned content.
  4. Cross-platform consensus is not as high as expected. Among 66 platform pairs, the highest “question × page” Jaccard similarity is only 34.94%, and 11 pairs have zero shared questions or pages. A single platform’s experience cannot be directly extrapolated to others.
  5. Content freshness is evident. Among pages with reliable publication dates, 88.77% come from 2025 or 2026; 29.57% of page titles contain a year, and 23.06% contain a list or ranking. Timeliness signals and decision-oriented content are worth further validation.

Feel free to leave comments with questions you care about. I can take the time to produce another data analysis report based on those topics.

Similar Articles

@yaojingang: I downloaded all papers related to GEO, AEO, and AI search from the past two years — 41 in total. Read through them one by one, and gained many new insights and inspirations. All 41 papers have been pushed to the GitHub repository. Feel free to download. The address is at the end of the article. Some key insights to share: 1. GEO is not a replacement for SEO. SEO addresses whether content can be retrieved, indexed, and entered into the candidate set...

X AI KOLs Timeline

The author collected and read 41 papers related to GEO, AEO, and AI search, pushed the collection to GitHub, and shared 10 key insights, including the relationship between GEO and SEO, AI search citation mechanisms, the importance of content structuring, and practical optimization directions.

@GenhuiP78950: Open-sourced my AI tools from the past six months. Not a big project, just scripts I use daily – transcribing Douyin/Bilibili videos, podcast-to-text, WeChat public account articles, industry news scanning… 11 in total. Used them privately, now unified with install scripts and docs.

X AI KOLs Timeline

Open-sourced a collection of 11 AI tool scripts for collecting and transcribing content from multiple channels like Douyin, Bilibili, and WeChat public accounts, making it easy to build a personal knowledge base. Supports direct installation by agents such as Claude Code, Codex, etc.

@jinglian: AI Spark @AISpark1 has directly open-sourced its knowledge base. Not just testing the waters with a few articles, but fully releasing all accumulated content. 247 articles, 6 major modules, continuously updated. Open-source link in the comments, save and read at your leisure. ━━━━━━━━━━━━━━ ① AI Beginner's Guide (…

X AI KOLs Timeline

AI Spark has fully open-sourced its knowledge base, containing 247 articles across six major modules (from beginner's guide to industry insights), continuously updated, suitable for AI learners.