WikiSTAR: A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles
Summary
WikiStar is a system that uses an LLM classifier to tag scientifically meaningful edits in Wikipedia article revision histories, enabling interactive exploration of how scientific knowledge evolves over time. The system includes a benchmark dataset and was validated through a user study with domain experts.
View Cached Full Text
Cached at: 07/15/26, 04:22 AM
# A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles
Source: [https://arxiv.org/html/2607.12441](https://arxiv.org/html/2607.12441)
Omer Ehrlich1∗Nitzan Barzilay1∗Rona Aviram2Tom Hope1,3 1The Hebrew University of Jerusalem2Ben\-Gurion University of the Negev3Allen Institute for AI \(Ai2\) omer\.ehrlich@mail\.huji\.ac\.il, nitzan\.barzilay@mail\.huji\.ac\.il
###### Abstract
Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scientific knowledge evolves over time\. Yet scientifically meaningful revisions are obscured by the sheer volume of routine edits, leaving each article’s scientific history hidden\. We presentWikiStar\(ScientificTracking ofArticleRevisions\), an interactive system for exploring scientifically meaningful changes across an article’s revision history\. Using an LLM classifier with an expert\-designed multi\-label taxonomy,WikiStarfirst tags edit types such as the addition of technical terms, new research findings, and changes in scientific narrative\. Then, through interactive views, an article’s full revision history can be traced at any granularity—from aggregate trends that reveal when and in which sections scientific content was added or refined, down to individual edits—showing how scientific knowledge develops at a scale previously impossible\. In a user study, experts from three domains found thatWikiStarsurfaced new patterns and research questions and enabled previously impractical analyses\. We release our system, code and a human\-annotated benchmark\.111![[Uncaptioned image]](https://arxiv.org/html/2607.12441v1/img/github-logo.png)[Code](https://github.com/omerehrlich/WikiStar)![[Uncaptioned image]](https://arxiv.org/html/2607.12441v1/img/hf-logo.png)[WikiStar\-Benchdataset](https://huggingface.co/datasets/omerehrlich/WikiStar-Bench)![[Uncaptioned image]](https://arxiv.org/html/2607.12441v1/img/hf-logo.png)[WikiStardemo](https://huggingface.co/spaces/omerehrlich/WikiStar)
Figure 1:TheWikiStarpipeline for tracking the history of scientific edits in a Wikipedia article\. \(1\) Splitting each revision into sections and matching them across revisions to recover the previous version of each edited section\. \(2\) Multi\-label edit\-type classification of each edit pair\. \(3\) Exploring through interactive views, from aggregate patterns down to individual edits\.
WikiStar: A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles
Omer Ehrlich1∗Nitzan Barzilay1∗Rona Aviram2Tom Hope1,31The Hebrew University of Jerusalem2Ben\-Gurion University of the Negev3Allen Institute for AI \(Ai2\)omer\.ehrlich@mail\.huji\.ac\.il, nitzan\.barzilay@mail\.huji\.ac\.il
\*\*footnotetext:Equal contribution\.## 1Introduction
Wikipedia, the world’s largest online encyclopedia, is written and continuously rewritten by volunteers around the world\. As a widely consulted source, it shapes how the public understands scientific and scholarly topics\. Its openly accessible revision history lets scholars trace the development of ideas within and between academic domains, and examine how public knowledge evolves\(Benjakobet al\.,[2023](https://arxiv.org/html/2607.12441#bib.bib6),[2021](https://arxiv.org/html/2607.12441#bib.bib14)\)\. Scientific articles on Wikipedia are especially well suited for studying knowledge evolution\. Prior work has shown that they closely track developments in the scientific literature through extensive use of reliable academic sources, making them more than public\-facing summaries\(Simonset al\.,[2024](https://arxiv.org/html/2607.12441#bib.bib24); Benjakobet al\.,[2023](https://arxiv.org/html/2607.12441#bib.bib6)\)\. Their revision histories therefore potentially constitute a rich longitudinal record of how scientific concepts are introduced, debated, refined, and stabilized\.
However, a single article can accumulate thousands of edits over its lifetime—the “AI” article alone has over 18,000 section edits—and scientifically significant edits make up only a small fraction, buried among routine changes such as rephrasing and formatting\. Isolating meaningful edits manually is highly labor\-intensive, so scientifically significant edits remain hidden and inaccessible\.
Prior work classifies Wikipedia edits\(Yanget al\.,[2017](https://arxiv.org/html/2607.12441#bib.bib1); Rajagopalet al\.,[2022](https://arxiv.org/html/2607.12441#bib.bib9)\)and builds tools for exploring authorship and editorial conflicts\(Viégaset al\.,[2004](https://arxiv.org/html/2607.12441#bib.bib28); Flöck and Acosta,[2014](https://arxiv.org/html/2607.12441#bib.bib29); Guoet al\.,[2023](https://arxiv.org/html/2607.12441#bib.bib30)\), but both lines target general\-purpose editing at the whole\-page or token level\. We instead work at the section level and focus on scientific significance, which localizes each edit, reveals how a section’s scientific content develops over time, and drives the interactive views of our system\.
We presentWikiStar\(Scientific Tracking of Article Revisions\), a system that detects, quantifies, and contextualizes scientific textual changes in Wikipedia articles over time \(FigureLABEL:fig:system\-small\)\. We introduce the task of*scientific edit classification*: labeling section\-level edits with the scientifically meaningful changes they make, using a taxonomy of ten expert\-designed labels\.WikiStarextracts section\-level edits, classifies each edit against the taxonomy using an LLM, and presents the results through interactive views that trace how an article’s scientific content evolves across time and sections \(Figure[2](https://arxiv.org/html/2607.12441#S2.F2)\)\. This makes the scientific edit history of Wikipedia accessible to Wikipedia researchers and more broadly to scientists, journalists, and others interested in how science is represented online\.
In a user study, experts in Wikipedia editing, science journalism, and the history and philosophy of science found thatWikiStarsurfaced patterns they could not easily discover manually, prompted new research questions, and enabled analyses that were previously impractical\.
Finally, we releaseWikiStar\-Bench, a human\-annotated dataset of section\-level Wikipedia edits labeled for scientific significance\.
In summary, our main contributions:
1. \(1\)We createWikiStar, a system for exploring how a Wikipedia article’s scientific content evolves over time\. We validateWikiStarin a user study with domain experts from three domains, who found it highly useful\.
2. \(2\)We introduce the task of scientific edit classification of Wikipedia article sections, grounded in our taxonomy of ten labels for scientifically salient changes, and releaseWikiStar\-Bench, a human\-annotated dataset of 1,387 section edits spanning three domains\.
## 2Scientific Edit Classification
Figure 2:Overview of theWikiStarsystem\.\(1\) Extraction and classification:The system splits each Wikipedia revision into sections, matches each edited section to its previous version, and classifies each edit pair into scientific edit types\.\(2\) Interactive exploration:Interactive views let users explore the classified history: TheOver Timeview shows temporal trends in edit types, and theSection Comparisonview compares them across selected sections\. Selecting a point or cell within those views generates a summary of its edits, with access to the underlying edits\.In this section, we first introduce the task of*scientific edit classification*of Wikipedia edits\. To support this task, we releaseWikiStar\-Bench, a human\-annotated benchmark, and use it to compare seven LLMs as classifiers\.
As discussed in the Introduction, since scientifically significant edits are only a small fraction of an article’s edits, we define*scientific edit classification*as labeling each Wikipedia edit with the types of scientific change it makes\. We cast this as a multi\-label classification task over section\-level edits\. Given a section and its preceding version, which we call a*section edit pair*, the task is to assign every applicable label from our taxonomy of ten scientific edit types \(§[2\.2](https://arxiv.org/html/2607.12441#S2.SS2)\), or theNon\-Scientific Editlabel when no scientific change occurs\.
We choose the section as our unit of granularity for two reasons\. First, it carries the right amount of context for classification: whole revisions inflate prompt length and reduce output consistency \(Shivashankar and Steinmetz \([2024](https://arxiv.org/html/2607.12441#bib.bib17)\); Fieblingeret al\.\([2024](https://arxiv.org/html/2607.12441#bib.bib18)\)\), while individual sentences discard the context needed to judge whether a change is scientifically meaningful\. Second, sections localize each edit to a specific aspect of the article, lettingWikiStartrack how scientific content is distributed across sections and develops over time\.
### 2\.1Section Identification and Linking
Turning an article’s revision log into section\-level edits requires two operations: parsing each revision into its constituent sections, and pairing each edited section with its previous version—or, for a newly created section, recognizing that it has none\.
Acquiring section text: Using the MediaWiki APIWikimedia Foundation \([2025](https://arxiv.org/html/2607.12441#bib.bib35)\), we retrieve every revision of the article as raw wikitext with its metadata \(revision ID, timestamp, editor\)\. Each revision is split into sections by a regex\-based parser that follows Wikipedia’s markup conventions\. As each revision contains the article’s full section list, most of it unchanged from the previous one, we retain only the sections that changed, so each entry corresponds to a single edited section\.
Matching section edit pairs: Thescientific edit classificationtask requires the version of each section before and after an edit, which we call a*section edit pair*\. Matching versions is challenging because sections are not stable: editors rename, split and merge them, yet Wikipedia stores each revision as flat full\-article text with no identifier linking section versions\. To locate a section’s earlier version, we search backward through revisions and compute two similarity measures against candidate sections:title similarity\(normalized Levenshtein distance\) andcontent similarity\(cosine similarity over embeddings of the section content\)\. We match the section to the most recent candidate for which either measure exceeds a threshold\. This simple approach worked well in practice; more advanced matching is left to future work\.
### 2\.2Scientific Edits Taxonomy
A closed set of labels is what letsWikiStaraggregate, filter, and visualize edits by type\. We designed a taxonomy of ten scientifically meaningful edit types—such asNew Scientific InformationandChange in Scientific Narrative—plus aNon\-Scientific Editlabel, built by a scientific\-Wikipedia expert and refined iteratively against real edits\. Full definitions appear in Appendix[A](https://arxiv.org/html/2607.12441#A1), with annotated examples for select types in Appendix[E](https://arxiv.org/html/2607.12441#A5)\.
### 2\.3WikiStar\-Bench
We curateWikiStar\-Bench, the first human\-annotated dataset of scientifically labeled section\-level Wikipedia edits, released as an evaluation resource for scientific edit classification\. It contains 1,387 examples across three scientific domains—Biology \(821\), Computer Science \(288\), and Neuroscience \(278\)—each spanning at least 9 popular Wikipedia pages\. Each example is a single section edit pair with its metadata and human\-annotated gold labels from our taxonomy\.
#### Annotation process:
The WikiStar\-Bench dataset was labeled by four annotators: three STEM graduate students and an assistant professor who holds a PhD in biology and researches scientific Wikipedia\. Annotators labeled each section edit pair independently, assigning every applicable label from our 10\-label taxonomy, or theNon\-Scientific Editlabel if none applied\.
#### Annotator agreement:
To ensure reliability of human annotation, all annotators tagged a shared subset of 104 examples\. Annotator agreement using mean pairwise Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2607.12441#bib.bib34)\)is high \(κ=0\.89\\kappa=0\.89\), indicating strong overall agreement\. Agreement is naturally lower on more subjective and interpretive labels \(such as*Change in Scientific Narrative*withκ=0\.7\\kappa=0\.7and*Scientific Clarification Added*withκ=0\.8\\kappa=0\.8\)\.
### 2\.4Results
#### Experimental settings:
AgainstWikiStar\-Bench’s gold labels, we report per\-label precision and recall, and macro\- and micro\-F1 \(our primary metric\)\. We evaluate seven LLMs as classifiers: GPT\-5\.4, GPT\-5\-mini, GPT\-5\.1\(OpenAI,[2025a](https://arxiv.org/html/2607.12441#bib.bib40)\), GPT\-4o\(Hurstet al\.,[2024](https://arxiv.org/html/2607.12441#bib.bib38)\), o3\-mini\(OpenAI,[2025b](https://arxiv.org/html/2607.12441#bib.bib39)\), LLaMA 3\.3 70B\(Dubeyet al\.,[2024](https://arxiv.org/html/2607.12441#bib.bib37)\), and Qwen3\-Next\-80B\(Yanget al\.,[2025](https://arxiv.org/html/2607.12441#bib.bib36)\)\.
#### Prompt refinements:
Our final prompt opens with a general task explanation, then gives each label a definition and a single minimal example\. The most effective refinements over a naive one\-sentence\-definition prompt were asking the model to first decide whether the edit is scientific at all, and flagging common mistakes in the definitions; these raise GPT\-5\.4’s macro\-F1 from0\.690\.69to0\.820\.82\. Few\-shot demonstrations did not help \(macro\-F1 drops to0\.780\.78\)\. Appendix[B](https://arxiv.org/html/2607.12441#A2)reports per\-model results across prompt versions\.222See prompt at[create\_prompt\.py](https://github.com/omerehrlich/WikiStar/blob/master/wiki_pipeline/prompts/create_prompt.py)in our repository\.
Aggregate results:Of the seven models, GPT\-5\-mini and GPT\-5\.4 are the two best\-performing, with comparable scores \(macro\-F10\.830\.83vs\.0\.820\.82; micro\-F10\.860\.86vs\.0\.830\.83\)\. The deployed demo uses GPT\-5\-mini due to its overall comparable results and lower operational cost\. Results for all models are in Appendix[C](https://arxiv.org/html/2607.12441#A3)\. Across labels, GPT\-5\.4 trails the human agreement ceiling \(macro\-F10\.820\.82vs\.0\.910\.91\), showing that scientific edit classification is not trivial and leaves room for improvement\.
#### Per\-Label results:
Table[1](https://arxiv.org/html/2607.12441#S2.T1)reports per\-label results for GPT\-5\.4\. F1 varies substantially by label: strongest on objective labels with explicit textual signals \(such asWikilink Addedwith0\.890\.89, andAcademic References Addedwith0\.880\.88\), where remaining errors mostly concern the scientific qualifier, e\.g\., counting a book as an academic reference\. It is weakest on the interpretive labels \(e\.g\.,Change in Scientific Narrativeat0\.610\.61\), where the model tends to over\-extend the label beyond its intended scope, often misclassifyingScientific Clarificationsas narrative changes\. This tracks lower human agreement on those labels \(Section[2\.3](https://arxiv.org/html/2607.12441#S2.SS3.SSS0.Px2)\)\.
GPT\-5\.4HumanLabelPRF1F1New Sci\. Information\.88\.74\.80\.93Sci\. Information Removed\.84\.92\.87\.94Sci\. Clarification Added\.73\.72\.72\.86Technical Terms Added\.80\.93\.86\.93Researcher Names Added\.93\.82\.87\.92Change in Sci\. Narrative\.54\.70\.61\.72Academic References Added\.90\.86\.88\.97Academic References Removed\.84\.86\.85\.95Wikilink Added\.95\.83\.89\.97Quantitative Information\.75\.86\.80\.95Non\-Scientific Edit\.83\.87\.85\.91Macro avg\.82\.83\.82\.91Micro avg\.84\.83\.83\.92Table 1:Per\-label precision, recall, and F1 for GPT\-5\.4 onWikiStar\-Bench\(n=1,387n\{=\}1\{,\}387\), and the human agreement ceiling—mean pairwise F1 across annotators on overlapping examples \(n=104n\{=\}104\)\.
#### Domain\-Specific Results:
We report GPT\-5\.4’s per\-label F1 for each domain inWikiStar\-Bench: Biology, Computer Science, and Neuroscience in Appendix[D](https://arxiv.org/html/2607.12441#A4)\. Macro\-F1 is stable across all domains\. As on the full benchmark, objective labels retain high performance throughout, while interpretive labels are weakest in every domain\. This consistency indicates that the taxonomy and prompts transfer across scientific domains rather than overfitting to any single one\.
## 3WikiStar: System Overview
WikiStarextracts section edit pairs from an article’s revision history, classifies each with our multi\-label taxonomy \(see §[2\.2](https://arxiv.org/html/2607.12441#S2.SS2)\), and turns the labeled edits into interactive views that let users explore the article’s scientific content across time and sections \(see Figure[2](https://arxiv.org/html/2607.12441#S2.F2)\)\. Each component ofWikiStaraddresses a concrete question researchers pose about an article’s scientific history: what changed, when, where, and what type of change\. We introduce these components through the following scenario:
Consider a historian of science who wants to know how Wikipedia’s “Artificial Intelligence” article has changed over time: which ideas were treated as established, when new topics entered, where editors’ attention concentrated\. The answer lives in the revision history, but reading it by hand is impractical, as it accumulated over 18,000 section edits across two decades\.
### 3\.1History Overview
TheHistory Overviewview is an LLM\-generated summary of how the article’s scientific content has evolved\. Classification produces a short explanation of what changed in each edit; theOverviewgathers these explanations for every edit labeled*New Scientific Information*or*Change in Scientific Narrative*—the two labels capturing new or revised scientific claims—and composes them chronologically into a single structured account\. A brief summary highlights the most significant changes; clickingshow moreexpands it into a period\-by\-period narrative of the article’s lifetime\.
The historian begins with the History Overview: the AI article opens with broad definitions of the field, grows more technical as deep learning enters, and turns recently to societal impact—two decades of edits, visible at a glance\.
### 3\.2Trends Over Time
TheTrends Over Timeview provides an interactive line plot comparing the frequency of selected edit types across time \(see Figure[2](https://arxiv.org/html/2607.12441#S2.F2),Trends Over Time\)\. By default, the view aggregates yearly across all sections, but supports filtering along several variables to suit each user’s focus\. Users can select the time frame, the temporal resolution, and the article’s sections to include\. Clicking a point on the graph generates a summary of the edits it represents \(those of a given edit type within a specific period\), and thesummarize all time periodsbutton extends this feature to an entire edit\-type series at once; both ground the aggregated trends in the edits’ textual content\. Below the summary, additional options let users inspect the underlying edit text or export several types of information\.
Narrowing the Trends Over Time view to the Applications section, the historian notices a sharp rise around 2023 in new scientific information and technical terms, with few academic references\. Clicking the peak, they find the summary ties the surge partly to the public release of LLMs, but also, unexpectedly, to editors choosing to add decades\-old milestones like Deep Blue and Watson\.
### 3\.3Per\-Section Edit Heatmap
TheSection Comparisonview complements the aggregatedTrends Over Timeview with fine\-grained comparison across sections\. For any selected edit type, it presents a heatmap where each row is a section and color intensity shows how that edit type trends within the section over time \(Figure[2](https://arxiv.org/html/2607.12441#S2.F2),Section Comparison\)\. Like the Over Time view, it supports filtering and generates an exportable summary when a cell is clicked\. Together, these enable both cross\-section comparison and targeted drill\-down: users can single out a section, edit type, and period, then surface the actual edits\.
Switching to the Section Comparison view, the historian sees*where*each change happened, not just when:*Change in Scientific Narrative*concentrates in the early 2000s in*Philosophy*and*History*, while*New Scientific Information*surges after 2020 in*Applications*and*Misinformation*—a shift from debating what AI is to tracking what it does\.
### 3\.4Section Evolution
TheSection Evolutionview shows how the article’s sections evolve over time\. It lays out the sections as rows and time periods as columns, coloring each cell to show what happened to that section in that period: whether it first appeared, was edited in a scientifically meaningful way, or was renamed\.
### 3\.5Edit Table
This view presents a filterable table of all edits across the article’s history, together with their metadata\. Users can locate specific edits by filtering on metadata fields or by free\-text search, and can export the table for further analysis\.
## 4User Study
We conduct a user study with three participants, chosen to span complementary perspectives on how scientific knowledge is produced and evolves on Wikipedia\.P1is deeply involved in the Wikipedia community as a writer and editor, and is a PhD candidate in Physics\.P2is a journalist covering cybersecurity and disinformation, with an MA in the history of science\. In their work, they frequently treat Wikipedia’s edit record as a lens for examining how knowledge is constructed, contested, and manipulated\.P3is a researcher in the philosophy and history of science, focused on how scientific ideas and models move between fields\.
Each participant received an introduction toWikiStar, then had 15 minutes to freely explore a scientific Wikipedia article of their choice\. Participants were asked to think aloud\. Afterwards, they rated nine statements about the system’s utility on a 5\-point Likert scale \(Table[2](https://arxiv.org/html/2607.12441#S4.T2)\)\.
Tracing a subject across fields\.P3examined the “Ising model” article, a physics formalism that later spread to other fields\. Puzzled that theSection Comparisonview framed its neuroscience content as “social sciences,” they found the answer inSection Evolution—the section was renamed “Neuroscience” only in 2022:
“It’s very interesting to see how this comes together and explains the structure of the article\.”
Surfacing invisible structural history\.P1usedWikiStarto examine the “Chaos Theory” article, whose foundational research largely predates Wikipedia\. They wanted to trace how its structure changed across twenty\-five years, and while inspecting theSection Evolutionview, they noted:
“It’s kind of a dream for anyone who discusses and thinks of Wikipedia as a source of statistical data\. This is much richer and more detailed than any other tool we currently have in our toolbox\.”
From trends to edits\.P2usedWikiStarto examine the “Vaccine” article, tracing scientific debate and controversy across its history\. While looking at theSection Comparisonview, one cell sparked their interest, showing markedly more added scientific content than its neighbors\. On discovering that they could generate a summary of the edits composing that cell, they enthused:
“That’s gorgeous\! It adds tons of value\. This is a very good analysis that I would actually want and need for my research\.”
They described the experience as “almost like guided reading\.” Turning to theTrends Over Timeview,P2experimented with the summarization feature across several points on the plot, then choseView source editsto inspect the data behind the summaries and reflected:
“Now I can connect my pipeline of thinking and research\! This option to look at the metadata and the text of the edits is important, because I need to have the data itself and validate it\.”
Table 2:Participant agreement ratings on system\-utility statements, collected via a 5\-point Likert scale \(1 = strongly disagree, 5 = strongly agree\)\.Avg\.denotes the mean rating across participants\. A subset of the 9 rated statements is shown; see full set in[our repository](https://github.com/omerehrlich/WikiStar-UserStudy)\.
## 5Related Work
A growing body of work studies how scientific knowledge is represented and diffuses on Wikipedia\(Teplitskiyet al\.,[2015](https://arxiv.org/html/2607.12441#bib.bib26)\)\. Most relevant to us,Benjakobet al\.\([2023](https://arxiv.org/html/2607.12441#bib.bib6)\)treat the revision history as a historiographical record of science, reading editorial revisions as intentional acts that mark how knowledge is introduced, refined, and restructured over time\. Studies of this kind remain rare because they are manual and labor\-intensive, requiring close reading of long revision histories\. In parallel, prior work classifies Wikipedia edits using general edit\-type and intention taxonomies\(Liu and Ram,[2011](https://arxiv.org/html/2607.12441#bib.bib21); Daxenberger and Gurevych,[2012](https://arxiv.org/html/2607.12441#bib.bib20); Yanget al\.,[2017](https://arxiv.org/html/2607.12441#bib.bib1)\), capturing generic operations rather than scientifically meaningful change, while another line visualizes edit provenance and activity, from token\-level authorship\(Flöck and Acosta,[2014](https://arxiv.org/html/2607.12441#bib.bib29)\)to edit wars and vandalism\(Guoet al\.,[2023](https://arxiv.org/html/2607.12441#bib.bib30)\)\. Neither classifies the*scientific*content of edits, and both operate at the page or token level rather than the section\.WikiStaris the first to do both at the section level: it matches each edited section to its previous version and classifies the change into ten scientific types\. Working at the section level lets users move between one section’s history and whole\-article trends, making scientific edit analysis accessible, at scale, to anyone interested in how science evolves and is represented online\.
## 6Conclusion
We presentedWikiStar, a system for tracking how the scientific content of Wikipedia articles develops over time\.WikiStarextracts section\-level edits, classifies each into an expert\-designed taxonomy of ten scientifically meaningful edit types, and presents the results through interactive views that let users move from high\-level trends down to the individual edits behind them\. We releasedWikiStar\-Bench, the first human\-annotated benchmark for scientific edit classification, and used it to compare seven LLMs\.
A user study with participants from Wikipedia editing, science journalism, and the history and philosophy of science suggestsWikiStaris useful across these varied perspectives, enabling new kinds of research into how the representation of science on Wikipedia changes over time\.
## References
- Citation needed? wikipedia and the covid\-19 pandemic\.BioRxiv\.Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p1.1)\.
- O\. Benjakob, O\. Guley, J\. Sevin, L\. Blondel, A\. Augustoni, M\. Collet, L\. Jouveshomme, R\. Amit, A\. Linder, and R\. Aviram \(2023\)Wikipedia as a tool for contemporary history of science: a case study on crispr\.PLoS One18\(9\),pp\. e0290827\.Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p1.1),[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
- J\. Cohen \(1960\)A coefficient of agreement for nominal scales\.Educational and Psychological Measurement20,pp\. 37 – 46\.External Links:[Link](https://api.semanticscholar.org/CorpusID:15926286)Cited by:[§2\.3](https://arxiv.org/html/2607.12441#S2.SS3.SSS0.Px2.p1.4)\.
- J\. Daxenberger and I\. Gurevych \(2012\)A corpus\-based study of edit categories in featured and non\-featured wikipedia articles\.InInternational Conference on Computational Linguistics,External Links:[Link](https://api.semanticscholar.org/CorpusID:11053125)Cited by:[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
- A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan, A\. Goyal, A\. S\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Rozière, B\. M\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. J\. Song, D\. Pintz, D\. Livshits, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. A\. AlBadawy, E\. I\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Nail, G\. Mialon, G\. Pang, G\. Cu\-curell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. M\. Kloumann, I\. Misra, I\. Evtimov, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. R\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Oldham, M\. Rita, M\. Pavlova, M\. H\. M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. S\. Chatterji, O\. Duchenne, O\. cCelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasić, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. S\. M\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. C\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whit\-man, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. E\. Tan, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. K\. Singh, A\. Grattafiori, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Vaughan, A\. Baevski, A\. Fein\-stein, A\. Kallet, A\. Sangani, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Franco, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. R\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, P\. \(\. Huang, B\. Loyd, B\. de Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, D\. Civin, D\. Beaty, D\. Kreymer, S\. Li, D\. Wyatt, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smoth\-ers, F\. Sun, F\. Kreuk, F\. Tian, F\. Ozgenel, F\. Caggioni, F\. \(\. Guzmán, F\. J\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Thattai, G\. Herman, G\. Sizov, G\. Zhang, G\. Lakshminarayanan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, I\. Molybog, I\. Tufanov, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, U\. KamHou, K\. Saxena, K\. Prasad, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Huang, K\. Chawla, K\. Lakhotia, K\. Huang, L\. Chen, L\. Garg, A\. Lavender, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Tsimpoukelli, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. P\. Laptev, N\. Dong, N\. Zhang, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollár, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Maheswari, R\. Howes, R\. Rinott, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Yu\. Sidorov, S\. Pan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Feng, S\. Lin, S\. Zha, S\. Shankar, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. K\. Gupta, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Kohler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. O\. Ajayi, V\. Montanez, V\. Mohan, V\. Kumar, V\. Mangla, V\. Ionescu, V\. A\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wang, X\. Wu, X\. Wang, X\. Xia, X\. Wu, X\. Gao, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Y\. Wang, Y\. Hao, Y\. Qian, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, and Z\. Zhao \(2024\)The llama 3 herd of models\.External Links:[Link](https://api.semanticscholar.org/CorpusID:271571434)Cited by:[§2\.4](https://arxiv.org/html/2607.12441#S2.SS4.SSS0.Px1.p1.1)\.
- R\. Fieblinger, M\. T\. Alam, and N\. Rastogi \(2024\)Actionable cyber threat intelligence using knowledge graphs and large language models\.In2024 IEEE European Symposium on Security and Privacy Workshops \(EuroS&PW\),pp\. 100–111\.Cited by:[§2](https://arxiv.org/html/2607.12441#S2.p3.1)\.
- F\. Flöck and M\. Acosta \(2014\)WikiWho: precise and efficient attribution of authorship of revisioned content\.InProceedings of the 23rd International Conference on World Wide Web \(WWW\),pp\. 843–854\.External Links:[Document](https://dx.doi.org/10.1145/2566486.2568026)Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p3.1),[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
- Y\. Guo, Q\. Han, Y\. Lou, Y\. Wang, C\. Liu, and X\. Yuan \(2023\)Edit\-history vis: an interactive visual exploration and analysis on wikipedia edit history\.In2023 IEEE 16th Pacific Visualization Symposium \(PacificVis\),pp\. 157–166\.External Links:[Document](https://dx.doi.org/10.1109/PacificVis56936.2023.00025)Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p3.1),[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
- O\. A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford, A\. Mkadry, A\. Baker\-Whitcomb, A\. Beutel, A\. Borzunov, A\. Carney, A\. Chow, A\. Kirillov, A\. Nichol, A\. Paino, A\. Renzin, A\. Passos, A\. Kirillov, A\. Christakis, A\. Conneau, A\. Kamali, A\. Jabri, A\. Moyer, A\. Tam, A\. Crookes, A\. Tootoochian, A\. Tootoonchian, A\. Kumar, A\. Vallone, A\. Karpathy, A\. Braunstein, A\. Cann, A\. Codispoti, A\. Galu, A\. Kondrich, A\. Tulloch, A\. Mishchenko, A\. Baek, A\. Jiang, A\. Pelisse, A\. Woodford, A\. Gosalia, A\. Dhar, A\. Pantuliano, A\. Nayak, A\. Oliver, B\. Zoph, B\. Ghorbani, B\. Leimberger, B\. Rossen, B\. Sokolowsky, B\. Wang, B\. Zweig, B\. Hoover, B\. Samic, B\. McGrew, B\. Spero, B\. Giertler, B\. Cheng, B\. Lightcap, B\. Walkin, B\. Quinn, B\. Guarraci, B\. Hsu, B\. Kellogg, B\. Eastman, C\. Lugaresi, C\. L\. Wainwright, C\. Bassin, C\. Hudson, C\. Chu, C\. Nelson, C\. Li, C\. J\. Shern, C\. Conger, C\. Barette, C\. Voss, C\. Ding, C\. Lu, C\. Zhang, C\. Beaumont, C\. Hallacy, C\. Koch, C\. Gibson, C\. Kim, C\. Choi, C\. McLeavey, C\. Hesse, C\. Fischer, C\. Winter, C\. Czarnecki, C\. Jarvis, C\. Wei, C\. Koumouzelis, D\. Sherburn, D\. Kappler, D\. Levin, D\. Levy, D\. Carr, D\. Farhi, D\. Mély, D\. Robinson, D\. Sasaki, D\. Jin, D\. Valladares, D\. Tsipras, D\. Li, P\. D\. Nguyen, D\. Findlay, E\. Oiwoh, E\. Wong, E\. Asdar, E\. Proehl, E\. M\. Yang, E\. Antonow, E\. Kramer, E\. Peterson, E\. Sigler, E\. Wallace, E\. Brevdo, E\. Mays, F\. Khorasani, F\. P\. Such, F\. Raso, F\. Zhang, F\. von Lohmann, F\. Sulit, G\. Goh, G\. Oden, G\. Salmon, G\. Starace, G\. Brockman, H\. Salman, H\. Bao, H\. Hu, H\. Wong, H\. Wang, H\. Schmidt, H\. Whitney, H\. Jun, H\. Kirchner, H\. P\. de Oliveira Pinto, H\. Ren, H\. Chang, H\. W\. Chung, I\. Kivlichan, I\. R\. O’Connell, I\. Osband, I\. Silber, I\. Sohl, I\. Okuyucu, I\. Lan, I\. Kostrikov, I\. Sutskever, I\. Kanitscheider, I\. Gulrajani, J\. Coxon, J\. Menick, J\. W\. Pachocki, J\. Aung, J\. Betker, J\. Crooks, J\. Lennon, J\. R\. Kiros, J\. Leike, J\. Park, J\. Kwon, J\. Phang, J\. Teplitz, J\. Wei, J\. Wolfe, J\. Chen, J\. Harris, J\. Varavva, J\. G\. Lee, J\. Shieh, J\. Lin, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Jang, J\. Q\. Candela, J\. Beutler, J\. Landers, J\. Parish, J\. Heidecke, J\. Schulman, J\. Lachman, J\. Mckay, J\. Uesato, J\. Ward, J\. W\. Kim, J\. Huizinga, J\. Sitkin, J\. Kraaijeveld, J\. Gross, J\. Kaplan, J\. Snyder, J\. Achiam, J\. Jiao, J\. Lee, J\. Zhuang, J\. Harriman, K\. Fricke, K\. Hayashi, K\. Singhal, K\. Shi, K\. Karthik, K\. Wood, K\. Rimbach, K\. Hsu, K\. Nguyen, K\. Gu\-Lemberg, K\. Button, K\. Liu, K\. Howe, K\. Muthukumar, K\. Luther, L\. Ahmad, L\. Kai, L\. Itow, L\. Workman, L\. Pathak, L\. Chen, L\. Jing, L\. Guy, L\. Fedus, L\. Zhou, L\. Mamitsuka, L\. Weng, L\. McCallum, L\. Held, O\. Long, L\. Feuvrier, L\. Zhang, L\. Kondraciuk, L\. Kaiser, L\. Hewitt, L\. Metz, L\. Doshi, M\. Aflak, M\. Simens, M\. Boyd, M\. Thompson, M\. Dukhan, M\. Chen, M\. Gray, M\. Hudnall, M\. Zhang, M\. Aljubeh, M\. Litwin, M\. Zeng, M\. Johnson, M\. Shetty, M\. M\. Gupta, M\. Shah, M\. A\. Yatbaz, M\. Yang, M\. Zhong, M\. Glaese, M\. Chen, M\. Janner, M\. Lampe, M\. Petrov, M\. Wu, M\. Wang, M\. Fradin, M\. Pokrass, M\. Castro, M\. Castro, M\. Pavlov, M\. Brundage, M\. Wang, M\. Khan, M\. Murati, M\. Bavarian, M\. Lin, M\. Yesildal, N\. Soto, N\. Gimelshein, N\. Cone, N\. M\. Staudacher, N\. Summers, N\. LaFontaine, N\. Chowdhury, N\. Ryder, N\. Stathas, N\. Turley, N\. A\. Tezak, N\. Felix, N\. Kudige, N\. S\. Keskar, N\. Deutsch, N\. Bundick, N\. Puckett, O\. Nachum, O\. E\. Okelola, O\. Boiko, O\. Murk, O\. Jaffe, O\. Watkins, O\. Godement, O\. Campbell\-Moore, P\. Chao, P\. McMillan, P\. Belov, P\. Su, P\. Bak, P\. Bakkum, P\. Deng, P\. Dolan, P\. Hoeschele, P\. Welinder, P\. Tillet, P\. Pronin, P\. Tillet, P\. Dhariwal, Q\. Yuan, R\. Dias, R\. Lim, R\. Arora, R\. Troll, R\. Lin, R\. G\. Lopes, R\. Puri, R\. Miyara, R\. H\. Leike, R\. Gaubert, R\. Zamani, R\. B\. Wang, R\. Donnelly, R\. Honsby, R\. Smith, R\. Sahai, R\. Ramchandani, R\. Huet, R\. Carmichael, R\. Zellers, R\. Chen, R\. Chen, R\. R\. Nigmatullin, R\. Cheu, S\. Jain, S\. Altman, S\. Schoenholz, S\. Toizer, S\. Miserendino, S\. Agarwal, S\. Culver, S\. Ethersmith, S\. Gray, S\. Grove, S\. Metzger, S\. Hermani, S\. Jain, S\. Zhao, S\. Wu, S\. Jomoto, S\. Wu, S\. Xia, S\. Phene, S\. Papay, S\. Narayanan, S\. Coffey, S\. Y\. L\. Lee, S\. Hall, S\. Balaji, T\. Broda, T\. Stramer, T\. Xu, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Cunninghman, T\. Degry, T\. Dimson, T\. Raoux, T\. Shadwell, T\. Zheng, T\. Underwood, T\. Markov, T\. Sherbakov, T\. Rubin, T\. Stasi, T\. Kaftan, T\. Heywood, T\. A\. Peterson, T\. Walters, T\. Eloundou, V\. Qi, V\. Moeller, V\. Monaco, V\. Kuo, V\. Fomenko, W\. H\. Chang, W\. Zheng, W\. Zhou, W\. Manassra, W\. Sheu, W\. Zaremba, Y\. Patil, Y\. Qian, Y\. Kim, Y\. Cheng, Y\. Zhang, Y\. He, Y\. Zhang, Y\. Jin, Y\. Dai, and Y\. Malkov \(2024\)GPT\-4o system card\.External Links:[Link](https://api.semanticscholar.org/CorpusID:273662196)Cited by:[§2\.4](https://arxiv.org/html/2607.12441#S2.SS4.SSS0.Px1.p1.1)\.
- J\. Liu and S\. Ram \(2011\)Who does what: collaboration patterns in the wikipedia and their impact on article quality\.ACM Trans\. Manag\. Inf\. Syst\.2,pp\. 11:1–11:23\.External Links:[Link](https://api.semanticscholar.org/CorpusID:15299619)Cited by:[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
- OpenAI \(2025a\)Introducing gpt\-5\.Note:Accessed: 2026\-06\-02External Links:[Link](https://openai.com/index/introducing-gpt-5/)Cited by:[§2\.4](https://arxiv.org/html/2607.12441#S2.SS4.SSS0.Px1.p1.1)\.
- OpenAI \(2025b\)OpenAI o3\-mini system card\.Note:Accessed: 2026\-02\-18External Links:[Link](https://cdn.openai.com/o3-mini-system-cardfeb10.pdf)Cited by:[§2\.4](https://arxiv.org/html/2607.12441#S2.SS4.SSS0.Px1.p1.1)\.
- D\. Rajagopal, X\. Zhang, M\. Gamon, S\. K\. Jauhar, D\. Yang, and E\. Hovy \(2022\)One document, many revisions: a dataset for classification and description of edit intents\.InProceedings of the Thirteenth Language Resources and Evaluation Conference,pp\. 5517–5524\.Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p3.1)\.
- K\. Shivashankar and N\. Steinmetz \(2024\)Contri \(e\) ve: context\+ retrieve for scholarly question answering\.arXiv preprint arXiv:2409\.09010\.Cited by:[§2](https://arxiv.org/html/2607.12441#S2.p3.1)\.
- A\. Simons, W\. Kircheis, M\. Schmidt, M\. Potthast, and B\. Stein \(2024\)Who are the “heroes of crispr”? public science communication on wikipedia and the challenge of micro\-notability\.Public Understanding of Science \(Bristol, England\)33,pp\. 918 – 934\.External Links:[Link](https://api.semanticscholar.org/CorpusID:268060504)Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p1.1)\.
- M\. Teplitskiy, G\. Lu, and E\. Duede \(2015\)Amplifying the impact of open access: wikipedia and the diffusion of science\.Journal of the Association for Information Science and Technology68\.External Links:[Link](https://api.semanticscholar.org/CorpusID:10220883)Cited by:[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
- F\. B\. Viégas, M\. Wattenberg, and K\. Dave \(2004\)Studying cooperation and conflict between authors with history flow visualizations\.InProceedings of the SIGCHI Conference on Human Factors in Computing Systems \(CHI\),Vienna, Austria,pp\. 575–582\.External Links:[Document](https://dx.doi.org/10.1145/985692.985765)Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p3.1)\.
- Wikimedia Foundation \(2025\)MediaWiki Action API\.Note:[https://www\.mediawiki\.org/wiki/API:Main\_page](https://www.mediawiki.org/wiki/API:Main_page)Accessed: 2026\-07\-08Cited by:[§2\.1](https://arxiv.org/html/2607.12441#S2.SS1.p2.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.External Links:[Link](https://api.semanticscholar.org/CorpusID:278602855)Cited by:[§2\.4](https://arxiv.org/html/2607.12441#S2.SS4.SSS0.Px1.p1.1)\.
- D\. Yang, A\. Halfaker, R\. Kraut, and E\. Hovy \(2017\)Identifying semantic edit intentions from revisions in wikipedia\.InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing,pp\. 2000–2010\.Cited by:[§1](https://arxiv.org/html/2607.12441#S1.p3.1),[§5](https://arxiv.org/html/2607.12441#S5.p1.1)\.
## Appendix AAppendix \- Edit\-Type Taxonomy
Table 3:TheWikiStartaxonomy: ten labels capturing scientifically meaningful edit types and aNon\-Scientific Editlabel\. Definitions provided in our prompt released with our[our code repository](https://github.com/omerehrlich/WikiStar)\.
## Appendix BAppendix \- Result Comparison Between Prompt Versions
GPT\-5\-miniGPT\-5\.4LLaMALabelNaiveRef\.NaiveRef\.NaiveRef\.New Sci\. Information\.72\.93\.77\.81\.73\.84Sci\. Information Removed\.69\.82\.79\.87\.46\.49Sci\. Clarification Added\.58\.75\.62\.73\.63\.62Technical Terms Added\.75\.86\.78\.86\.71\.71Researcher Names Added\.58\.85\.71\.87\.73\.81Change in Sci\. Narrative\.40\.52\.32\.61\.29\.28Academic References Added\.63\.93\.64\.88\.63\.72Academic References Removed\.70\.85\.72\.85\.46\.50Wikilink Added\.86\.91\.91\.89\.81\.74Quantitative Information\.78\.84\.69\.80\.76\.73Non Scientific Edit\.56\.85\.63\.85\.63\.63Macro F1\.66\.83\.69\.82\.62\.64Micro F1\.68\.86\.72\.83\.67\.70Table 4:Per\-label F1 for thenaiveprompt \(a single\-sentence definition per label\) versus ourrefinedprompt \(explicit inclusion/exclusion rules, per\-label examples, and an up\-front scientific/non\-scientific decision\), across three backbone models\.Boldmarks each refined score that improves on its naive counterpart\.
## Appendix CAppendix \- Per\-Model Results onWikiStar\-Bench
Table[5](https://arxiv.org/html/2607.12441#A3.T5)reports per\-label precision, recall, and F1 for every model we evaluated on the full benchmark, extending the single\-model view of Table[1](https://arxiv.org/html/2607.12441#S2.T1)\. Metrics are computed against the gold labels; per\-label support is reported in Tables[1](https://arxiv.org/html/2607.12441#S2.T1)and[6](https://arxiv.org/html/2607.12441#A4.T6)\.
PrecisionRecallF1LabelGPT\-5 mini
GPT\-5\.4
GPT\-5\.1
Llama3\.3
o3\-mini
GPT\-4o
Qwen3
GPT\-5 mini
GPT\-5\.4
GPT\-5\.1
Llama3\.3
o3\-mini
GPT\-4o
Qwen3
GPT\-5 mini
GPT\-5\.4
GPT\-5\.1
Llama3\.3
o3\-mini
GPT\-4o
Qwen3
New Sci\. Information\.93\.88\.86\.83\.84\.83\.87\.93\.74\.92\.86\.86\.85\.83\.93\.80\.89\.84\.85\.84\.85Sci\. Information Removed\.81\.84\.76\.62\.62\.62\.46\.83\.92\.68\.41\.41\.41\.09\.82\.87\.72\.50\.49\.50\.15Sci\. Clarification Added\.69\.73\.54\.65\.65\.65\.53\.81\.72\.75\.58\.58\.59\.66\.75\.72\.63\.62\.62\.62\.58Technical Terms Added\.88\.80\.80\.84\.84\.83\.66\.83\.93\.83\.61\.61\.61\.84\.86\.86\.81\.71\.70\.70\.74Researcher Names Added\.77\.93\.79\.81\.80\.81\.68\.95\.82\.91\.81\.81\.81\.94\.85\.87\.84\.81\.81\.81\.79Change in Sci\. Narrative\.41\.54\.22\.18\.20\.20\.25\.74\.70\.87\.60\.55\.57\.63\.52\.61\.35\.28\.30\.30\.36Academic Refs Added\.93\.90\.69\.59\.59\.59\.48\.93\.86\.93\.95\.95\.95\.98\.93\.88\.79\.72\.72\.72\.64Academic Refs Removed\.89\.84\.75\.52\.52\.52\.38\.82\.86\.59\.48\.48\.48\.45\.85\.85\.66\.50\.50\.50\.41Wikilink Added\.94\.95\.93\.89\.89\.89\.82\.88\.83\.71\.63\.63\.63\.41\.91\.89\.81\.74\.74\.74\.54Quantitative Information\.80\.75\.69\.76\.76\.76\.59\.89\.86\.86\.71\.71\.71\.78\.84\.80\.76\.73\.73\.73\.67Non\-Scientific Edit\.94\.83\.86\.85\.84\.84\.77\.78\.87\.67\.50\.50\.50\.43\.85\.85\.75\.63\.62\.62\.56Macro avg\.82\.82\.72\.69\.69\.69\.59\.85\.83\.79\.65\.64\.65\.64\.83\.82\.73\.64\.64\.64\.57Micro avg\.86\.84\.74\.74\.74\.74\.65\.86\.83\.80\.66\.66\.66\.64\.86\.83\.77\.70\.70\.70\.64Table 5:Per\-label performance onWikiStar\-Bench\.Precision \(P\), Recall \(R\), and F1 for each model; best value per row\-block inbold\.Llama3\.3andQwen3denote LLaMA\-3\.3\-70B and Qwen3\-80B\.
## Appendix DAppendix \- Per\-Domain Results on WikiStar\-Bench
Biology820 examplesComputerScience288 examplesNeuroscience278 examplesEdit TypePRF1\#PRF1\#PRF1\#New Scientific Information\.88\.77\.82307\.93\.73\.81135\.86\.69\.77145Scientific Information Removed\.89\.88\.88108\.80\.98\.8850\.79\.92\.8566Scientific Clarification Added\.72\.76\.74149\.77\.71\.7469\.70\.63\.6651Scientific Technical Terms Added\.77\.91\.83214\.86\.95\.90109\.80\.96\.87101Researcher Names Added\.92\.85\.89811\.00\.64\.7828\.90\.86\.8842Change in Scientific Narrative\.52\.73\.6022\.62\.50\.5516\.54\.87\.6715Academic References Added\.94\.86\.90139\.80\.86\.8314\.82\.89\.8545Academic References Removed\.87\.85\.8647\.83\.83\.836\.79\.89\.8426Wikilink Added\.95\.84\.89234\.96\.86\.91109\.95\.78\.85117Add\./Mod\. of Quantitative Information\.77\.87\.81121\.63\.71\.6742\.81\.96\.8845Non Scientific Edit\.90\.89\.89295\.70\.88\.7867\.73\.79\.7672Macro Avg\.83\.84\.831717\.81\.79\.79645\.79\.84\.81725Micro Avg\.85\.84\.851717\.83\.81\.82645\.81\.82\.82725Table 6:Performance Across Domains onWikiStar\-BenchPer\-Label\.Precision \(P\), Recall \(R\), F1, and support \(\#; number of annotated instances\) for the ten scientific edit types plus theNon\-Scientific Editlabel\.
## Appendix EAppendix \- Edit Examples
Table 7:Before/after excerpts for selectedWikiStaredit types; changed span inbold,\[…\]marks elided text\.Similar Articles
WikiSpy
WikiSpy is a tool for tracking changes on Wikipedia.
@hwchase17: https://x.com/hwchase17/status/2071963622298050997
The article discusses the emerging pattern of 'wiki memory' for AI agents, where raw source data is intelligently compressed into a persistent, structured knowledge layer that agents can use efficiently. It compares this to basic RAG and gives examples like DeepWiki and LLM Wiki.
Small edits, large models: How Wikipedia advocacy shapes LLM values
This paper demonstrates that a small coordinated Wikipedia editing campaign can measurably shape how language models handle topics, using animal welfare as a case study.
@Huanusa: The ceiling of personal knowledge bases has arrived! This GitHub LLM Wiki project has already garnered 2800+ Stars, completely leaving ordinary RAG in the dust! It's not the useless mode of "re-retrieving" every time, but lets AI directly help you incrementally build a truly structured Wiki — compile knowledge once, and it continuously evolves...
LLM Wiki is an open-source desktop application that uses LLM to incrementally build a structured knowledge base, supporting knowledge graphs, community detection, Obsidian integration, and Chrome clipping, aiming to replace traditional RAG approaches.
@omarsar0: LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valua…
The author advocates for LLM Wikis as a valuable application of AI, showcasing their PaperWiki project that uses agents to automatically curate and maintain a knowledge base of research papers, improving signal-to-noise ratio and enabling cutting-edge research.