Position: Every Ground Truth is a Human Construction, not an Objective Truth
Summary
This position paper argues that ground truth datasets in machine learning are not objective truths but human constructions shaped by choices, and advocates for articulating these choices to improve reliability, transparency, and accountability.
View Cached Full Text
Cached at: 07/14/26, 04:12 AM
# Position: Every Ground Truth is a Human Construction, not an Objective Truth Source: [https://arxiv.org/html/2607.09668](https://arxiv.org/html/2607.09668) ###### Abstract Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models\. This position paper argues that ground truths are not neutral objective measurements that are naturally given, but instead that they are constructed by arrangements of humans and technologies\. We argue that the ML community will benefit from articulating and discussing these often invisible or unreported choices and acknowledging that reference data sets are contingent, not universal\. Focusing on the situated and context\-dependent nature of ground truths can improve reliability by enabling a better informed perspective on where, when, and how the datasets, and the models they have shaped, can best be used\. We argue for increasing ‘situated reliability’ which includes articulating the limits and strengths of models and their truth claims\. Finally, paying more attention to the construction of ground truths can support transparency, accountability, and interdisciplinary work\. annotation ## 1Introduction In machine learning \(ML\) research and development, the term “ground truth” commonly refers to datasets viewed as containing the true values of a given concept\(Kang,[2023](https://arxiv.org/html/2607.09668#bib.bib27)\)that are used to train and evaluate ML models\. An often overlooked aspect is that these ground truths are not neutral objective measurements naturally given\. All of our “truths” are constructed\.\\commentTo highlight this contingency, We use ground truth as a countable noun in this paper to emphasize thata ground truthfor an ML project is only one of many potentialground truths\. Constructing a ground truth encompasses work beyond just aligning with domain knowledge, yet these additional choices and their consequences are often not discussed or reported\. Figure 1:Every ground truth isconstructedas the result of multiple decisions, not a single objective truth\. Different pathways yield different ground truths \(GT1, GT2, GT3, …\) for the same task\.Our position is that ground truths are constructed, and that ML researchers should articulate the choices involved in their construction and discuss the impact those choices have on their results\. The concept of ground truth has been defined in various ways\. The term has been used in remote sensing since the 1960s and been adopted across fields such as computer science and engineering, psychology, neuroscience, and economics\(Woodhouse,[2021](https://arxiv.org/html/2607.09668#bib.bib64)\)\. One view regards ground truth as the fundamental, real, or underlying facts collected at the source\. In remote sensing, it has been defined as “information obtained by direct measurement at ground level, rather than by interpretation of remotely obtained data”\(Woodhouse,[2021](https://arxiv.org/html/2607.09668#bib.bib64)\)\. A more inclusive definition is “information obtained by direct observation of a real system, as opposed to a model or simulation; a set of data that is considered to be accurate and reliable, and is used to calibrate a model, algorithm, procedure”\(Woodhouse,[2021](https://arxiv.org/html/2607.09668#bib.bib64)\)\. Kang \([2023](https://arxiv.org/html/2607.09668#bib.bib27)\) argues that the concept basically refers to information assumed to be true in the development of ML models\. We follow Jaton’s \([2017](https://arxiv.org/html/2607.09668#bib.bib26)\) definition of ground truths as the result of ground truthing, which is the practice of defining the problem to be solved and what the input values and desired output targets should be\(Jaton,[2021](https://arxiv.org/html/2607.09668#bib.bib25),[2017](https://arxiv.org/html/2607.09668#bib.bib26)\)\. In this paper, we point out that ground truths are the result of multiple decisions about what sources of knowledge are legitimate, how to interpret observations, what values are recorded, and who makes and validates those attributions \(illustrated in Figure[1](https://arxiv.org/html/2607.09668#S1.F1)\)\. All of these decisions and contextual practices impact which ground truths are accepted\. Therefore, “ground truth” is a black\-box term that comprises many component decisions on inclusions and exclusions of people, technologies, and concepts\. Opening up that black box means recognizing the contingency—the ‘it could have been otherwise’ aspect—of ground truths\. It encourages us to think about the possibility of alternative ground truths\. When we see ground truths as plural and produced, we can consider how the legitimacy \(or at the very least, the broad applicability\) of ML results from ground truths could be undermined due to their nature as contextual, contingent, and situated\. Our starting point is that ground truths are not preexisting entities but are constructed to fit the task addressed by the ML model\(Jaton,[2017](https://arxiv.org/html/2607.09668#bib.bib26)\)\. They are not solely technical, machine\-made, digital objects, or natural facts\. Instead, much like the data produced through professional vision in other domains\(Goodwin,[1994](https://arxiv.org/html/2607.09668#bib.bib16)\), they are human constructions\(Jaton,[2024](https://arxiv.org/html/2607.09668#bib.bib23)\)formed by assumptions, organizations, processes, and the use of technology\. Flattening the real world into ground truth datasets requires a series of choices and practices involving problem formulation, equipment, settings, expertise, conditions, and epistemologies\. It is a version of knowledge making through coding, highlighting, and producing material representation\(Goodwin,[1994](https://arxiv.org/html/2607.09668#bib.bib16)\)\. This also applies when using previously collected “real\-world” data such as medical records, in terms of both their initial creation in the healthcare system and the choices of which data, variables, and labels to include\. By unboxing the black boxes of ground truths, we want to highlight the challenging and necessary work of creating ground truths for machine learning\. We urge the machine learning community to stop framing ground truths as neutral, objective, and pre\-existing values\.\\commentSince ground truth datasets are not neutral repositories of facts, we want to highlight how They are made through diverse practices that commonly require interdisciplinary communication and shared understanding, in joint creation and partnership\. In addition, we stress the need to unpack the notion of the expert\-labeler, and we compare views of them as oracle versus as sensor as well as important aspects of engaging the crowd as a label source\. We encourage acknowledgment of and reflection on how ground truths are \(necessarily\) intentional creations\. These practices will increase ground truth quality and thereby the reliability of ML models by \(1\) Increasing the situatedness and hence the sharpness of a ground truth that is produced in collaborative and conscientious practices and \(2\) Allowing us to question the limits of applicability for a given ground truth for ML model training\.We call for thinking of ground truth as situated, specific, contingent, and contextual—which will make it better, sharper, more useful—but will also narrow its usability to particular contexts\. This perspective can also reduce the uncritical adoption of data sets for inappropriate use cases\.Recognizing the work of ground truthing could also benefit interdisciplinary collaboration and improve evaluation, transparency, and technology adoption\. While some aspects of how benchmark datasets are made and \(mis\)used have been discussed previously, the current level of awareness around ground truth construction remains insufficient\. It has not yet led to agreement on how to characterize and communicate the relevant aspects of ground truth data that are context\-specific, situated, and hence could change our conclusions if different data construction decisions were made\. There are still no standard vocabulary or documentation methods to express those limitations, nor standard practices for comparing the impact when different choices are made, such as about who serves as the source of labels\. Our hope is that this paper, through its recommended actions, dialogue, and researcher training that views “data collection” as “data construction”, will strengthen the community’s understanding of datasets and their \(in\)valid uses\. In this paper, we first discuss why this topic matters in Section[2](https://arxiv.org/html/2607.09668#S2)and review related work \(Section[3](https://arxiv.org/html/2607.09668#S3)\), then describe different approaches to the construction of ground truths \(via the concept of “construction sites”\) in Section[4](https://arxiv.org/html/2607.09668#S4)\. There we analyze the construction of ground truths for different types of reference values, such as direct measurements, expert interpretations, crowd\-sourced consensus, and synthetic data\. In Section[5](https://arxiv.org/html/2607.09668#S5), we address aspects of communication and collaboration in constructing sense in ground truthing\. We conclude with calls to action \(Section[6](https://arxiv.org/html/2607.09668#S6)\) and alternative views \(Section[7](https://arxiv.org/html/2607.09668#S7)\), followed by a summary of our conclusions \(Section[8](https://arxiv.org/html/2607.09668#S8)\)\. ## 2Why Ground Truth Clarity Matters We argue for the need to acknowledge that ground truths are constructed, and that the process is \(by necessity\) creative, interpretative, contingent, and therefore subjective\.\\commentTo summarize our preceding arguments, Ground truthing is not the passive capture and representation of neutral, objective facts, and datasets and labels are not mirrors of the natural world\. The decisions, choices, assumptions, conditions, and other factors that impact the construction of a ground truth need to be made visible\. Ground truth contingency also needs to be openly discussed and reflected upon, because every ground truth could have been constructed differently and resulted in very different outcomes\. This is important due to the impact that ground truths have on the development, evaluation, and functioning of ML models\. Ground truthing is, after all, “a task\-bounding process and a form of intentional biasing that hardlines the limits of the algorithm and the possible range of outcomes for an ML system” \(Kang,[2023](https://arxiv.org/html/2607.09668#bib.bib27), p\. 3\)\. It poses a plethora of limits on both how we understand the world and what ML models do in the world\. Ground truth choices made for benchmark data sets can have enormous impact that echoes for decades\. For example, the digit images in the benchmark MNIST data set\(LeCunet al\.,[1994](https://arxiv.org/html/2607.09668#bib.bib32)\)were reduced from 128x128 pixels to 28x28 pixels in the 1990s due to resource limits of computers at the time\(LeCunet al\.,[1998](https://arxiv.org/html/2607.09668#bib.bib33)\)\. This has led to the training of thousands of digit classifiers that operate only on 28x28 pixel images\. These classifiers are incapable of classifying the original full\-resolution images from 1992, much less images from today produced by higher\-resolution scanners, despite modern computing environments with orders of magnitude more memory and computational resources\. Newly acquired digit images today have to be reduced to this artificially tiny size to suit MNIST\-based classifiers, due to a 1990s choice of how to construct a digit ground truth\. What are the consequences of not acknowledging that ground truths are constructed and influenced by multiple factors?Ignorance about context can lead to inappropriate or “off\-label”\(Shimronet al\.,[2022](https://arxiv.org/html/2607.09668#bib.bib52)\)use of a ground truth with severe negative, and even harmful, impacts\. For example, undocumented data processing steps used to create open\-access MRI ground truths create an artificially easier problem, leading to an unfortunate performance drop of 47% when the model was applied to newly collected MRI data\(Shimronet al\.,[2022](https://arxiv.org/html/2607.09668#bib.bib52)\)\. Ambiguous annotation processes in dermatology could result in overestimating accuracy and putting patients at risk\(Stutzet al\.,[2025](https://arxiv.org/html/2607.09668#bib.bib56)\)\. In addition, blind use of a pre\-existing ground truth may not result in new knowledge, but merely reproduce current practices and domain expertise\(Henriksen and Bechmann,[2020](https://arxiv.org/html/2607.09668#bib.bib19)\)\. There is also the conundrum of using human\-generated ground truths for training and evaluation when aiming to surpass that same human capability\(Henriksen and Bechmann,[2020](https://arxiv.org/html/2607.09668#bib.bib19); Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\)\. Moreover, the presumed universality of ground truths \(that they are valid across different contexts and domains\) often does not hold\(Lee and Ribes,[2025](https://arxiv.org/html/2607.09668#bib.bib35)\)\. For instance, the well\-known UCI Adult dataset was originally collected to demonstrate a dataset visualizer, but has been applied far beyond that, showing the context drift of benchmark datasets\(Kohavi,[2025](https://arxiv.org/html/2607.09668#bib.bib28)\)\. We recognize that ground truthing is a necessary, pragmatic step to enable ML development and evaluation\(Kang,[2023](https://arxiv.org/html/2607.09668#bib.bib27)\), manifesting as the “assumptions we need to make”\(Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\), and that it is often performed with some understanding of its incompleteness\. Yet the ML community needs to openly discuss ground truths as human constructions rather than simply asthe true knowledge\. AsKang \([2023](https://arxiv.org/html/2607.09668#bib.bib27), p\. 1\)argues, “With the increasing complexity of the tasks to which ML has been applied over the past six decades, \[…\] agreement on what constitutes adequately stable ground truths for ML systems has become exponentially more complicated\.” In fact, we should discuss whether we should keep calling it “ground truth” at all\(Woodhouse,[2021](https://arxiv.org/html/2607.09668#bib.bib64)\)\. Finally, the position that we argue for is highly important for reliability\. In line with acknowledging ground truths as constructed, we propose adopting an understanding that they result in a reliability that issituated\(reliable only in a given context\)\. Situated reliability is not a quantitative value, but instead a key concept that helps us move from discussing reliability as a generic property to understanding it as necessarily “situated” in a particular context\. If we continue to employ singular ground truths, we should improve the awareness of the limits and contingencies of truth claims and the ways by which we understand the basis of, for example, measuring accuracy against a certain ground truth\.Rechtet al\.\([2019](https://arxiv.org/html/2607.09668#bib.bib47)\)studied this aspect of the commonly used ImageNet dataset, finding that “changes in the sampling strategy can indeed affect model accuracies by a large amount, even if the data source and other parts of the dataset creation process stay the same\.” Careful evaluation and documentation practices can make the situatedness of reliability more visible\. We need to recognizesituated reliabilityas a way to capture the limits of ground truths, whether created by the context of data collection, the expert judgments, the range of sensors, or all other factors shaping its construction\. ## 3Related Work The construction of ground truths means that their perceived quality and value are subject to negotiations between machine learning practitioners and their collaborators about how to establish the best ground truth for the task\. An examination of AI experts’ work of establishing ground truths for medical AI\(Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\)demonstrated the value of using “ground truthing” as a verb\(Jaton,[2017](https://arxiv.org/html/2607.09668#bib.bib26); Henriksen and Bechmann,[2020](https://arxiv.org/html/2607.09668#bib.bib19)\)to describe the joint work of medical and ML expertise\.Mulleret al\.\([2021](https://arxiv.org/html/2607.09668#bib.bib38)\)described thesocialprocess of generating labeled data as the result of human negotiation\. This view is also supported by early lessons learned from studies of the collaborative practices used to create computational tools for real world applications\(Collins,[1990](https://arxiv.org/html/2607.09668#bib.bib6); Suchmanet al\.,[1999](https://arxiv.org/html/2607.09668#bib.bib57)\)\. Takeaways from the field of Computer Supportive Collaborative Work include the value of using ethnographic and ethnomethodological methods\(Goodwin,[2000](https://arxiv.org/html/2607.09668#bib.bib17); Raeithel,[1996](https://arxiv.org/html/2607.09668#bib.bib46)\)when collaborating around technology development to examine how agreement is achieved\(Hughes,[1988](https://arxiv.org/html/2607.09668#bib.bib20); Lave and Wenger,[1999](https://arxiv.org/html/2607.09668#bib.bib30)\), which knowledge artifacts are included or excluded\(Star,[1991](https://arxiv.org/html/2607.09668#bib.bib55)\), and whose framings of a domain are engaged\(Orr,[1996](https://arxiv.org/html/2607.09668#bib.bib41); Resnicket al\.,[1991](https://arxiv.org/html/2607.09668#bib.bib48)\)\. This work set the stage for later studies of how computer scientists could work together with domain experts to develop technologies with real world applications\(Suchman,[2007](https://arxiv.org/html/2607.09668#bib.bib58); Milliganet al\.,[2011](https://arxiv.org/html/2607.09668#bib.bib37)\)and through innovative design methodologies\(Costanza\-Chock,[2020](https://arxiv.org/html/2607.09668#bib.bib7)\)\. These methodologies can inspire analysis and change in ML practices, such as ground truthing, that are proving integral to shaping the development of AI tools now\. Additional aspects of the contingent nature of ground truths include observations that “raw data” is an oxymoron, because data is always “cooked” by processes of collection, use, and analysis\(Gitelman,[2013](https://arxiv.org/html/2607.09668#bib.bib15)\); that data is situated and local\(Loukissas,[2019](https://arxiv.org/html/2607.09668#bib.bib36)\); and that ground truths are not pre\-existing entities but shaped to fit the task of the ML model, as in the case of digital image processing\(Jaton,[2017](https://arxiv.org/html/2607.09668#bib.bib26)\)\. Activities and negotiations of establishing ground truths for ML in healthcare have been conceptualized as truth practices, described as a multi\-modal and multilevel performance of truth involving engineers and medical experts with different methods and knowledge\(Henriksen and Bechmann,[2020](https://arxiv.org/html/2607.09668#bib.bib19)\)\. In such domains, ground truths are assembled and valued based on not only medical expert knowledge, but also temporal and technical qualities, the ability to support generalization, and humanness, as expert\-based labels instill both trust and the risk of misjudgments\(Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\)\. Empirical studies have found that a multitude of factors play into the annotation of data to establish a ground truth schema for medical AI, including external \(regulations, context of creation and use, commercial and operational demands\) and internal factors \(epistemic differences and limits of the labeling process\)\(Zajacet al\.,[2023](https://arxiv.org/html/2607.09668#bib.bib65)\)\. These factors shape and constrain the design of a ground truth schema, and they impact the possibility of achieving responsible AI\.Lebovitzet al\.\([2021](https://arxiv.org/html/2607.09668#bib.bib31)\)argue for the benefit of deconstructing ground truths to better understand why an given AI implementation is not successful\. In a medical setting, they found a critical limitation that they characterized as the ground truth containing the know\-what, but not the know\-how, of clinical practice\. Previous work also addressed how different types of digital reference objects, simulations, and synthetic data are used as ground truths, and in some cases claimed as the ‘perfect’ or ‘known truth’, since they can be controlled and offer flexibility\(Högberg and Winter,[2026](https://arxiv.org/html/2607.09668#bib.bib22)\)\. Additionally, the reuse of ground truth data sets and algorithms in new cases and domains distinct from where they were generated is common practice, and not unique to ML\(Leeet al\.,[2025](https://arxiv.org/html/2607.09668#bib.bib34)\), but could nonetheless be problematic\(Lee and Ribes,[2025](https://arxiv.org/html/2607.09668#bib.bib35)\)with regard to the assumed universality of the constituent ground truths\. Awareness of the risk of mismatches between the context of data collection and that of model application has led researchers to advocate for improved data documentation practices, informing about provenance, creation, and use of ML datasets to mitigate harmful outcomes\(Gebruet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib14)\)\. For large language models, research has shown that user\-informed model alignment \(reinforcement learning from human feedback or RLHF\) is strongly influenced by choices about whose preferences, or value judgments, inform the alignment\(Ouyanget al\.,[2022](https://arxiv.org/html/2607.09668#bib.bib43)\)\. The authors warned that the 40 contractors they employed to align a GPT\-3 model were “clearly not representative of the full spectrum of people who will use and be affected by our deployed models\.” This suggests a need for increased discussion of how ground truth construction can shape and limit model performance and generalizability\. ## 4Ground Truth Construction Sites Depending on the task that the ML model is intended to perform, different levels of expertise are required\. For specialized or high\-stakes tasks, a high level of expertise becomes necessary to provide reference values with a sufficient level of reliability, such as identifying a malignant tumor on a medical image\. For other tasks, common knowledge or non\-specialist lay knowledge may be enough, providing the opportunity to construct ground truths by means such as crowd sourcing\. We present and discuss a set of common “construction sites” where ML ground truths are created: direct measurement of a natural phenomenon, expert interpretation of cases or observations, consensus from a crowd, and synthetic data or simulations\. ### 4\.1Direct measurement of a natural phenomenon \\comment in the natural sciences Collecting data in the form of direct measurements, such as by sensors or instruments \(telescopes, microscopes, microphones, cameras, etc\.\) is commonly seen as the most direct way of obtaining a ground truth\. This type of ground truth construction is largely seen as a factual registration of natural states, and less subjective than settings in which labels hinge on human interpretation\. Yet even this collection of “true” values is dependent on an array of choices and preconditions\. These include necessary decisions such as: what equipment to use \(resolution, sensitivity, data rate, etc\.\), where to place sensors, when and how often to perform the measurement, how values should be collected and what annotations and metadata to include \(what the set of appropriate labels is\), and how generalizable the measurements and labels are\. These decisions may also be influenced by factors unrelated to the phenomenon studied, such as the availability of tools, computational resources, and/or funding\. Different choices will result in different ground truths for the same problem\. An example demonstrating the contingent nature of ground truth appears in a study that sought to catalog all fresh impact craters on Mars\(Daubaret al\.,[2022](https://arxiv.org/html/2607.09668#bib.bib9)\)\. “Fresh” impact craters were defined as those whose creation time could be bounded, i\.e\., those craters that had both “before” \(no crater\) and “after” \(with crater\) images in the data set\. As a result, it is likely that some genuinely fresh impacts were labeled as “not fresh” due to the temporal limits of the data set, independent of the appearance of the crater in a given image\. It is evident that a different source data set, or a different definition of “fresh” impact craters, would lead to a different ground truth\. In another case, investigators trained an ML model to classify the Mars surface into 14 terrain types\. After finding that this model only achieved 74% accuracy, they created a second ground truth that assigned the same pixels to 5 coarser\-grained classes\. The new trained classifier achieved 92% accuracy\(Barrettet al\.,[2022](https://arxiv.org/html/2607.09668#bib.bib1)\)\. It is not possible to say which ground truth is the “right” one without knowing the context of use\. For the motivating case of Mars rover landing site selection, mission planners would need to decide what level of accuracy they need to trust a classifier’s output, as well as which class distinctions are critical for that task \(are 5 classes sufficient, or were all 14 needed?\)\. When not enough data from distant planets are available, analog sites on Earth are sometimes used as substitutes in data collection\. This poses additional questions about how measurements of such sites can reliably function as “truth\-spots”\(Ostrowska,[2026](https://arxiv.org/html/2607.09668#bib.bib42)\)and training data for ML\. ### 4\.2Expert interpretation of a case or observation Another common way to construct ground truths is by collecting post\-hoc expert interpretations that are treated as true labels, especially for tasks requiring highly\-skilled domain expertise\. However, expert\-labeled data may be employed later by investigators who lack the same expertise, especially for benchmark datasets\. One example is the commonly used multivariate, categorical benchmark dataset of physical characteristics of mushrooms that are classified as edible or poisonous\(Schlimmer,[1981](https://arxiv.org/html/2607.09668#bib.bib39)\)\. Labels were derived from expert knowledge in the Audubon Society Field Guide to North American Mushrooms\. Mushrooms with “unknown edibility” were conservatively assigned to the poisonous class\. The binary labels of “edible” versus “poisonous” have been treated as a ground truth ever since, and knowledge of which ones were of unknown edibility has not been preserved\. Could this choice have introduced errors? The situatedness \(what expertise was available at that time and place\) and representation choices \(number of classes\) that shaped this ground truth matter\. We do not know how many ML papers using this dataset would have reported different conclusions had all three original classes been used\. In the field of medicine, ML is often applied to separate the healthy from the pathological\. For image\-based data, it often involves the task of detecting signs of disease\.\\commentin a comparable manner as a human expert interpreter, yet it also requires achieving a reference of what the healthy, or in medical terms; normal, looks like\. Hence, Medical or clinical expert interpretations are commonly used to establish ground truths\. This is a laborious process when datasets are produced from scratch\. In one case, experienced neurologists were recruited to spend hours painting areas of imaged slices of a brain to construct a ground truth for developing brain segmentation algorithms\(Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\)\. The products of this expert labor \(hand\-segmented images\) were subsequently treated as the true values\. Another example is the task of identifying malignant tumors in medical images, where image data are labeled as healthy versus not healthy by multiple radiologists\. However, this type of collection of accurate labels opens up different types of uncertainty, such as that of inter\-expert variability, where radiologists disagree, or intra\-expert variability, where the same radiologist gives different assessments in repeated interpretations\. In addition, not all expert sources are the same: elite hospitals may be perceived as generating the preferred \(best\) ground truth\(Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\)\. When expert knowledge is required, the human expert labeler is generally assumed to be \(1\) an absolutely reliable source of knowledge \(oracle\) and \(2\) an impersonal neutral sensor that registers, detects, and provides data in an objective manner\. Yet\\comment, several examples of work show that the processes and issues involved in getting reliable labels from human experts are never free from subjectivity\.Mulleret al\.\([2021](https://arxiv.org/html/2607.09668#bib.bib38)\)\) observed that in practice, the names and meanings of labels are not fixed, but “malleable, changeable, and negotiable depending on who is applying them\.” As such, they also become more than simple definitions\. Rather, the definitions and the data they are applied to are subject to pragmatic limitations, persons, and organizations\. This makes visible the work required to describe the world\.Mulleret al\.\([2021](https://arxiv.org/html/2607.09668#bib.bib38)\)argue that ground truth looks less like an objective truth, and more like the output of a social process\. Their analysis shows that the social complexity \(and more or less open negotiations\) is sometimes necessary to come to agreement about what the data conveys\. \\comment Research has also shown that Finally, expert\-based medical ground truths contain human uncertainties that can introduce errors or inconsistencies, such as in radiology\(Lebovitzet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib31)\)\. In a similar vein, Jaton \([2023](https://arxiv.org/html/2607.09668#bib.bib24)\) describes how the making of a benchmark dataset for personalized cancer immunotherapy resulted in a ground truth of somewhat contested accuracy, yet it remained in use due to it being the only functional benchmark for the task\. ML practitioners and researchers using data and labels from existing collections of expert judgments as ground truths, such as open datasets or medical register data, are impacted by these inherent uncertainties and choices about how the data was documented and what to include\. ### 4\.3Consensus from a crowd Ground truths are also constructed by engaging crowds to provide post\-hoc labels and confirm, or vote on the accuracy of, the assigned labels\(Cabitzaet al\.,[2023](https://arxiv.org/html/2607.09668#bib.bib3)\)\. Crowds can be recruited as annotators for tasks requiring lay knowledge or some types of expertise\. There are citizen science initiatives such as BirdNet, which invites birdwatchers to contribute sightings, sounds, and labels\(Sullivanet al\.,[2009](https://arxiv.org/html/2607.09668#bib.bib59)\)\. However, much work is needed for crowd\-sourcing to generate reliable ground truths\(Rhee,[2025](https://arxiv.org/html/2607.09668#bib.bib49); Dentonet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib11)\)\. Some large advances in AI have been achieved by recruiting paid annotators through crowdsourcing marketplaces, such as Amazon Mechanical Turk\. One example is the creation of the ImageNet dataset\(Denget al\.,[2009](https://arxiv.org/html/2607.09668#bib.bib10)\), which relied heavily on outsourcing the annotation of thousands of images to humans who were instructed to confirm the presence or absence of a given concept in an image\. In the final dataset, images were included if sufficient agreement was obtained among the annotators\(Dentonet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib11)\)\. The image annotation effort was distributed across∼\\sim49,000 workers from 167 countries\. This arguably very culturally diverse pool of workers were tasked with making meaning from data\. The assumption was that their annotations would reflect universal understandings and interpretations of the data, despite the wide variety of backgrounds they brought to the task\(Dentonet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib11)\)\. Deriving a single ground truth from a crowd is challenging\. For example, in Natural Language Processing \(NLP\) tasks, the meaning of words, sentences, and statements can be interpreted very differently by different annotators and be highly subjective and cultural\. For sentiment analysis, the judgment of whether a statement is positive, negative, or neutral could vary significantly between labelers\. Plank \([2022](https://arxiv.org/html/2607.09668#bib.bib44)\) highlights this as human label variation with regards to NLP specifically but also stresses it as a general matter for ML and Computer Vision \(CV\) and for all stages of the pipeline: data, modeling, and evaluation\. Human label variation \(HLV\) could be caused by inattention or insufficient expertise leading to errors, but some differences cannot be dismissed as mistakes\. HLV can also arise from differences of opinion, concept ambiguity, subjectivity, multiple correct options, or the fact that cultural understandings and responses to an object or emotion shift across time, space, and context\(Rhee,[2025](https://arxiv.org/html/2607.09668#bib.bib49)\)\. Commonly, variation is “resolved” by aggregation into a majority vote, allowing only for one belief, label, or category, obfuscating the real\-world complexity\(Plank,[2022](https://arxiv.org/html/2607.09668#bib.bib44); Cabitzaet al\.,[2023](https://arxiv.org/html/2607.09668#bib.bib3)\)\. Plank \([2022](https://arxiv.org/html/2607.09668#bib.bib44)\) identifies the limitations of optimizing and evaluating by a single ground truth, suggesting it might hamper progress, and argues for preserving and using the variation as informative instead of seeing it as a problem, noise, or disagreement to be removed\(Plank,[2022](https://arxiv.org/html/2607.09668#bib.bib44); Cabitzaet al\.,[2023](https://arxiv.org/html/2607.09668#bib.bib3)\)\. We currently lack standard ways to document the subtleties and limitations of a ground truth created by crowd\-sourced labels\. This is especially problematic for data sets that become community\-wide benchmarks, employed by thousands of other researchers who had no direct involvement in the ground truth construction process\. For example, the ImageNet\-1k documentation111[https://huggingface\.co/datasets/ILSVRC/imagenet\-1k](https://huggingface.co/datasets/ILSVRC/imagenet-1k)does not explain enough of the ground truth construction choices \(sampling, sources, classes, labelers\) to enable others to determine whether it is an appropriate ground truth for their needs\. While some more details appear in the accompanying paper\(Russakovskyet al\.,[2015](https://arxiv.org/html/2607.09668#bib.bib50)\), there is no discussion of the impact of alternative choices and that ImageNet \(like any dataset\) has reliability that situated in its choice of sources\. It is a convenience sample, comprising voluntarily posted images on Flickr and other websites in the 2000s; it is not representative of all human photographers \(or image classes, or labelers\)\.Rechtet al\.\([2019](https://arxiv.org/html/2607.09668#bib.bib47)\)found that different image sampling strategies led to very different ImageNet classifier evaluations\. We advocate that this kind of sensitivity analysis be standardized\. ### 4\.4Synthetic data, simulations, and “known truths” Data created by generative AI models and simulations are also utilized as ground truths\(e\.g\., Hamarnehet al\.,[2008](https://arxiv.org/html/2607.09668#bib.bib18); Högberg and Winter,[2026](https://arxiv.org/html/2607.09668#bib.bib22)\)\. Awareness of the generation practices for these synthesized ground truths is necessary to understand the limitations of their use\. Synthetic ground truths frequently contain some elements of fabrication, such as data augmentation by generating plausible variations to boost generalizability or data imputation to accommodate missing values\. Other examples include the injection of artificial cancer nodules into real lung scans to evaluate the detection abilities of CNNs and humans\(Schultheisset al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib51)\)and the injection of artificial noise in labels to enable controlled studies of model robustness to such noise\(de Vries and Thierens,[2025](https://arxiv.org/html/2607.09668#bib.bib12)\)\. Synthetic data have also been used to address biased datasets, perhaps most famously for facial identification algorithms\(Casconeet al\.,[2025](https://arxiv.org/html/2607.09668#bib.bib4)\)and other computer vision use cases, like training autonomous vehicles\(Tsirikoglouet al\.,[2017](https://arxiv.org/html/2607.09668#bib.bib61)\)\. One member of our author team is conducting research on scientists’ experiences of using synthetic data\. This work finds that two very different modes of ground truthing emerge\. The first is to generate a synthetic ‘ideal case’ of ground truth\(Daston and Galison,[2010](https://arxiv.org/html/2607.09668#bib.bib8)\)using models of scientific phenomena\. Example domains where this occurs include protein biology and climate research\(Tonget al\.,[2025](https://arxiv.org/html/2607.09668#bib.bib60)\)\. The second approach is to combine synthetic data with original data to obtain a larger, pooled training set\. In this process \(which often occurs in medical diagnosis domains\), the synthetic data are viewed as containing as yet unidentified ground truths in the generated variations\. As in Section[4\.2](https://arxiv.org/html/2607.09668#S4.SS2), domain experts are then engaged in evaluating and identifying the ground truths of the dataset, which is composed of real and synthetic data\. These four construction sites show different ways in which ground truths are constructed\.They are each subject to a multitude of choices, actors, equipment, and contingencies that shape the datasets and thereby the ML models they inform\. They also raise the necessity of communication and shared understanding between different actors to establish any ground truth dataset\. ## 5Communication and Collaboration in Ground Truth Construction To construct reliable ground truths requires effort and coordination\. This is true whether collecting new sensor data, using data from existing infrastructures, or generating new synthetic data\. This work indirectly or directly involves multiple actors, such as the ML practitioners, the data collectors \(if different from the ML practitioners themselves\), and the labelers and annotators \(experts or crowd\-sourced actors\) assigned to identify the truth in the data\.\\commentIn this section, we discuss aspects of this coordination and communication\. Processes like deciding what the relevant data and sources are, and labeling and annotating data, require communication, agreement, and alignment between different actors, to construct sense of both the data and the concept to be modeled\. ### 5\.1Constructing sense While the verbs ‘making’ and ‘constructing’ are often used interchangeably, there is an important difference of nuance\. ‘Making sense’ can indicate that one is creating understandings of an ontologically separate and discrete ground truth: it implies that the ground truth can be seen, identified, and captured by labeling it, requiring the assumption that a ground truth exists independently and the goal is to correctly identify it\. We argue that it is more appropriate to refer to ‘constructing sense’ from the data\. This idea refers to the process of aligning the needs of the ML expert \(to create functional categories and labels\) and the domain expert \(to express an understanding of the world, derived from historical and contextual knowledge of their field\) toconstructa ground truth, not merely to label it\. This involves constructing meaning for the data, the deriving model, and the phenomena being modeled\. Changing the terminology from making to constructing acknowledges the situatedness and—to some degree—the arbitrary aspects of ground truthing\. One process in which sense is constructed is the labeling of data\. When labeling and annotating data to construct a ground truth, all actors must have a shared understanding of what is being labeled, the goal of labeling, and how to operationalize it\. Commonly, this means developing and adopting annotation guidelines and instructions to reach consistency and agreement of which label to use when, what the labels entail, and what annotation should look like\(Engdahl,[2024](https://arxiv.org/html/2607.09668#bib.bib13)\)\. This is a contingent communicative process to navigate when working with expert labelers\(Mulleret al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib38)\)as well as crowds\(Dentonet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib11)\)\. When engaging crowds as labelers, reaching a shared understanding might be even more challenging considering the heterogeneity and distributed nature of the crowd\. Constraints on how labeling should be performed also depend on the cost and time restrictions of the ML project, sometimes leading to the use of platforms like Amazon Mechanical Turk\(Dentonet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib11)\)\. In these highly distributed cases, often with no face\-to\-face interaction, good communication practices are vital for constructing reliable ground truths\. With regards to what is treated as ground truth, we urge the field to develop a more reflective and critical review process for data, sources, labels, and documentation for ML development and evaluation\. This is especially crucial when the ground truth dataset is collected from existing data infrastructures that were created for purposes other than ML development, such as for census information, biodiversity, or medical record keeping\. An extra challenge arises when the original dataset creators are unavailable or uninvolved in the ML project\. A critical review is necessary with regards to what the data can convey and how labels can function as reliable ground truths, before being appropriated for ML\. The dataset documentation often functions as a standalone communication tool to establish an understanding of what the dataset includes and what variables and labels mean\. We may hope that the documentation is correct and complete while also questioning what has gone undocumented or undiagnosed\(Högberg,[2025](https://arxiv.org/html/2607.09668#bib.bib21)\)\. Moreover, even when the documentation was written with ML development in mind, there can be practical, organizational aspects of such data collections and labels that prevent them from functioning as the accurate record of facts they are assumed to be when later employed as ground truths\. For example, in practice, the same label could have been used somewhat differently due to organizational variations between hospitals\. When possible, it is desirable to engage directly with domain experts\. Such collaboration can also inform decisions about the most suitable data, data sources, and labels\. This requires highly functioning communication and a shared vocabulary between domain experts and ML practitioners\. Often, an ML practitioner alone cannot properly interpret the performance of the ML model with respect to the original domain and must rely on expert knowledge\. This requires pairing the goal of the machine learning task with domain expertise about why the task matters\(Wagstaff,[2012](https://arxiv.org/html/2607.09668#bib.bib62)\)\. We argue thatmore discussion, reflection, and consideration is neededwith regard toground truthing as a communicative process of creating shared understandings and constructing sense\. ### 5\.2Joint creation and partnership The construction of ground truths is dependent on and shaped byjointcreation and partnership\. We argue that the machine learning community should to a greater extent acknowledge the array of actors that contribute to the construction of ground truth\. Establishing a reliable ground truth requires interdisciplinary or intersectoral work\. The expertise of machine learning and the problem domain must be combined when formulating the problem to solve, determining what outcomes are useful, choosing data sources, and establishing meaningful categories for algorithmic tasks\. This work requires interdisciplinary competencies, including facilitation across knowledge paradigms and insights into the complexities of creating data from the world and knowledge from data\. For models to make a positive impact, it is necessary to understand what ground truth is meaningful and useful for the context of model application, something that ML practitioners may not be able to decipher on their own\. To facilitate collaborative practices in ground truth construction \(and beyond\), we suggest that ML education programs to a larger extent include courses or modules in Computer Supported Collaborative Work and fields like Science and Technology Studies and Critical Data Studies\. In addition to the challenges already discussed, ground truthing must comply with ethical and legal restrictions when using ground truths constructed by others\(Kroes and Verbeek,[2014](https://arxiv.org/html/2607.09668#bib.bib29); Sobel,[2021](https://arxiv.org/html/2607.09668#bib.bib53)\)\. Here, computer science can learn from other fields that have been more confronted with a history of unethical data collection, such as medicine or population science\. We suggest that the ML community foster more discussion about the ethical and legal considerations of ground truth constructions and uses\. This should be a standard practice and preferably a collaborative task together with domain experts and social science experts\(Prainsack and Steindl,[2022](https://arxiv.org/html/2607.09668#bib.bib45); Baumgartneret al\.,[2023](https://arxiv.org/html/2607.09668#bib.bib2)\)\. Recognizing the elements of joint creation and the collaborative aspect of ground truthing foregrounds a number of insights: It matterswho is in the constellation of actorsconstructing ground truths; theresponsibilityfor creating ground truths isdistributed, but not necessarily shared equally, between different members of the constellation; even if an expert or a crowd is attributed with the authority to create ‘correct’ ground truths, the ML community also shares responsibility for how those truths are made and operationalized\. ## 6Call to Action Given the issues presented in this paper, we advocate for: \(1\) greater acknowledgment of ground truths as constructed, \(2\) formulating multiple ground truths for the same problem and comparing evaluation results on each, \(3\) developing a standard vocabulary on ground truth limitations, and \(4\) a standard for how dataset creators should specify relevant contexts for use\. These steps can inform \(5\) improvements to ML courses and educational programs\. In more detail, we propose that the ML community should: 1. 1\.Acknowledge the decisions and work involved in producing ground truthsand the fact that they are socially constructed and not just observed or recorded\. 2. 2\.Recognize thatground truths are situatedand hence restricted to asituated reliability\. This can improve knowledge on where, when, and how ground truth datasets, and the models they have shaped, can be of best use\. This includes to: - •Consider how ground truths arecontextual in space and time, reflecting the positionality of people and organizationsmaking the data, the decisions about categories, the vocabularies and the practical conditions\. This necessarily leads to considerations about how frequently ground truths might require updates or revisions\. - •Embrace humility about the generalizabilityof our models to prevent misuse and negative impacts, andconsider the situatednessof ground truths inevaluation practices, which means that the model might need to be evaluated against additional ground truths\. This might increase costs in terms of constructing the additional datasets as well as the evaluation runtime\. However, these increased costs can be weighed against the costs of inadvertently over\-estimating performance and the societal and personal costs of uncritical adoption of ill\-fitting ML solutions that could have potentially severe consequences\. 3. 3\.Improve our ability to discuss and reflectupon the construction of ground truths\. Because of the situatedness of ground truths, the models that are trained and evaluated on them are also limited\. Including additional data in the training set is not always going to solve the model’s limitations\. Hence: - •We needa language\(terminology\) to speak about limits of aspects such as the expertise of the model\. This can be encouraged by activities such as special issues and workshops dedicated to this topic\. 4. 4\.Promote a reflective, critical, and open approach toreporting and documentinghow ground truths were constructed, their limitations, and how these choices may impact the ML results\. For example, the Datasheets for datasets’\(Gebruet al\.,[2021](https://arxiv.org/html/2607.09668#bib.bib14)\)questions about the purpose, collection practices, and limitations for future use can be augmented with questions about additional \(necessary\) choices in the ground truth construction process\. Dataset creators could list the conditions necessary for their ground truth to be properly aligned with a new setting\. This can be achieved by: - •Includingprecise informationabout datasets and their construction choices indataset documentationand metadata\. - •Articulatinghow other choices in ground truthing could have resulted in different outcomes, andaddressing uncertainties\. - •Creatingmultiple different ground truths andevaluatingthe impact of those choices on the resulting models, as done byRechtet al\.\([2019](https://arxiv.org/html/2607.09668#bib.bib47)\)\. - •Addingcritical reflectionon the ground truth to thelist of criteriafor ML paper submissions \(checklists\) and reviews\. 5. 5\.Improve ML courses and educational programsto include critical reflections on ground truths, training on how to collaborate on knowledge and data construction across domains and paradigms, and legal and ethical issues related to ground truth construction and use\. ## 7Alternative Views First, one practical/pragmatic objection to our argument is the view thatdata sets utilized as ground truths by machine learning are the best representations of natural facts, and not constructed\. However, careful reflection on all of the steps, choices, and actors involved in creating a ground truth reveals that alternatives would lead to a different ground truth\. Ignoring the contingency of the construction process limits our ability to conceive of and enact different, possibly better, choices\. Second, when dealing with expert\-based ground truths, one can argue for viewingthe expert as an oracle, always able to deliver accurate predictions under all circumstances, orthe expert as a sensor, able to perform neutral recording of phenomena and provide accurate measurements\. Both views would suggest that ground truths are objective rather than constructed\. However, these views do not hold given findings about inter\- and intra\-observer variability of medical experts’ assessments\(Soyer,[2018](https://arxiv.org/html/2607.09668#bib.bib54)\)\. The ground truths obtained by engaging crowds, with diverse levels of expertise, likewise contain label variation\. Hence, thatthe crowd can function as an objective sensoris also an unreliable assumption\. Moreover, human label variation is not something that can be fully abolished, and it may be regarded as an opportunity rather than a problem\(Plank,[2022](https://arxiv.org/html/2607.09668#bib.bib44)\)\. A third objection to our argument is thatdataset contingency could be solved through more data, or better multimodal data\. Even if this might result in models with increased performance across different contexts, we argue that there is no solution that can eliminate the contingency of ground truths that we have presented\. Finally, some might argue that our recommendations of acknowledging and documenting the construction of ground truths requirestoo much workor istoo much of a burden\. Our response is that it indeed requires some work, but that work is necessary if we want ground truths to be meaningful, reliable, and used appropriately\. Recognizing ground truths as constructed, with situated reliability, allows the resulting models to be properly applied in the right contexts\. ## 8Conclusions We argue for the need to acknowledge and reflect on the fact that ground truths for machine learning are constructed artifacts, in contrast to the common view of them as objective, naturally given truths\. An increased focus on the situatedness and context\-dependence of ground truths, and the work behind them, can improve the reliability of ML models through a better informed perspective on where, when, and how ground truth datasets, and the models they have shaped, can safely be used\. This informs about the limits and strengths of models and their truth claims by defining a ‘situated reliability’\. We also raise the question of whether we should think only in terms of singular ground truths\. At the very least, we should discuss the work involved in constructing ground truths, what a good ground truth is, and whether there is a need to consider using multiple truths\. Finally, paying more attention to the construction of ground truths can offer greater possibilities to achieve transparency of ML models and improve interdisciplinary work in ML development through better communication\. ## References - A\. M\. Barrett, M\. R\. Balme, M\. Woods, S\. Karachalios, D\. Petrocelli, L\. Joudrier, and E\. Sefton\-Nash \(2022\)NOAH\-H, a deep\-learning, terrain classification system for Mars: results for the ExoMars Rover candidate landing sites\.Icarus371,pp\. 114701\.External Links:ISSN 0019\-1035,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.icarus.2021.114701),[Link](https://www.sciencedirect.com/science/article/pii/S0019103521003560)Cited by:[§4\.1](https://arxiv.org/html/2607.09668#S4.SS1.p4.1)\. - R\. Baumgartner, P\. Arora, C\. Bath, D\. Burljaev, K\. Ciereszko, B\. Custers, J\. Ding, W\. Ernst, E\. Fosch\-Villaronga, V\. Galanos, T\. Gremsl, T\. Hendl, C\. Kropp, C\. Lenk, P\. Martin, S\. Mbelu, S\. Morais dos Santos Bruss, K\. Napiwodzka, E\. Nowak, T\. Roxanne, S\. Samerski, D\. Schneeberger, K\. Tampe\-Mai, K\. Vlantoni, K\. Wiggert, and R\. Williams \(2023\)Fair and equitable AI in biomedical research and healthcare: social science perspectives\.Artificial Intelligence in Medicine144,pp\. 102658\.External Links:ISSN 0933\-3657,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.artmed.2023.102658),[Link](https://www.sciencedirect.com/science/article/pii/S0933365723001720)Cited by:[§5\.2](https://arxiv.org/html/2607.09668#S5.SS2.p3.1)\. - F\. Cabitza, A\. Campagner, and V\. Basile \(2023\)Toward a perspectivist turn in ground truthing for predictive computing\.Proceedings of the AAAI Conference on Artificial Intelligence37\(6\),pp\. 6860–6868\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/25840),[Document](https://dx.doi.org/10.1609/aaai.v37i6.25840)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p4.1)\. - L\. Cascone, M\. D\. Maio, V\. Loia, M\. Maucioni, M\. Nappi, and C\. Pero \(2025\)Synthetic data for fairness: bias mitigation in facial attribute recognition\.IEEE Open Journal of the Computer Society6\(\),pp\. 1703–1714\.External Links:[Document](https://dx.doi.org/10.1109/OJCS.2025.3622694),[Link](https://ieeexplore.ieee.org/abstract/document/11206468)Cited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p2.1)\. - H\. M\. Collins \(1990\)Artificial experts: social knowledge and intelligent machines\.MIT Press,Cambridge, Mass\.\.External Links:ISBN 026203168XCited by:[§3](https://arxiv.org/html/2607.09668#S3.p1.1)\. - S\. Costanza\-Chock \(2020\)Design justice: community\-led practices to build the worlds we need\.The MIT Press,Cambridge, Massachusetts\.External Links:ISBN 9780262043458Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - L\. Daston and P\. Galison \(2010\)Objectivity\.Zone Books,New York\.External Links:ISBN 189095179XCited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p3.1)\. - I\. J\. Daubar, C\. M\. Dundas, A\. S\. McEwen, A\. Gao, D\. Wexler, S\. Piqueux, G\. S\. Collins, K\. Miljkovic, T\. Neidhart, J\. Eschenfelder, G\. D\. Bart, K\. L\. Wagstaff, G\. Doran, L\. Posiolova, M\. Malin, G\. Speth, D\. Susko, and A\. Werynski \(2022\)New craters on Mars: an updated catalog\.Journal of Geophysical Research: Planets127\(7\),pp\. e2021JE007145\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1029/2021JE007145),[Link](https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/2021JE007145),https://agupubs\.onlinelibrary\.wiley\.com/doi/pdf/10\.1029/2021JE007145Cited by:[§4\.1](https://arxiv.org/html/2607.09668#S4.SS1.p3.1)\. - S\. de Vries and D\. Thierens \(2025\)Generating the ground truth: synthetic data for soft label and label noise research\.International Journal of Data Science and Analytics20,pp\. 5603–5615\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/s41060-025-00786-z)Cited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p2.1)\. - J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-Fei \(2009\)ImageNet: a large\-scale hierarchical image database\.In2009 IEEE Conference on Computer Vision and Pattern Recognition,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2009.5206848)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p2.1)\. - E\. Denton, A\. Hanna, R\. Amironesei, A\. Smart, and H\. Nicole \(2021\)On the genealogy of machine learning datasets: a critical history of ImageNet\.Big Data & Society8\(2\),pp\. 20539517211035955\.External Links:[Document](https://dx.doi.org/10.1177/20539517211035955),[Link](https://doi.org/10.1177/20539517211035955)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p2.1),[§5\.1](https://arxiv.org/html/2607.09668#S5.SS1.p3.1)\. - I\. Engdahl \(2024\)Agreements ‘in the wild’: standards and alignment in machine learning benchmark dataset construction\.Big Data & Society11\(2\),pp\. 20539517241242457\.External Links:[Document](https://dx.doi.org/10.1177/20539517241242457),[Link](https://doi.org/10.1177/20539517241242457)Cited by:[§5\.1](https://arxiv.org/html/2607.09668#S5.SS1.p3.1)\. - T\. Gebru, J\. Morgenstern, B\. Vecchione, J\. W\. Vaughan, H\. Wallach, H\. D\. III, and K\. Crawford \(2021\)Datasheets for datasets\.Communications of the ACM64\(12\),pp\. 86–92\.External Links:ISSN 0001\-0782,[Link](https://dl.acm.org/doi/10.1145/3458723),[Document](https://dx.doi.org/10.1145/3458723)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p5.1),[item 4](https://arxiv.org/html/2607.09668#S6.I1.i4.p1.1)\. - L\. Gitelman \(2013\)”raw data” is an oxymoron\.The MIT Press,Cambridge, Mass\.\.External Links:ISBN 9780262518284Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p3.1)\. - C\. Goodwin \(1994\)Professional vision\.American Anthropologist96\(3\),pp\. 606–633\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1525/aa.1994.96.3.02a00100),[Link](https://anthrosource.onlinelibrary.wiley.com/doi/abs/10.1525/aa.1994.96.3.02a00100),https://anthrosource\.onlinelibrary\.wiley\.com/doi/pdf/10\.1525/aa\.1994\.96\.3\.02a00100Cited by:[§1](https://arxiv.org/html/2607.09668#S1.p6.1)\. - C\. Goodwin \(2000\)Action and embodiment within situated human interaction\.Journal of Pragmatics32\(10\),pp\. 1489–1522\.External Links:ISSN 0378\-2166,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/S0378-2166%2899%2900096-X),[Link](https://www.sciencedirect.com/science/article/pii/S037821669900096X)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - G\. Hamarneh, P\. Jassi, and L\. Tang \(2008\)Simulation of ground\-truth validation data via physically\- and statistically\-based warps\.InMedical Image Computing and Computer\-Assisted Intervention – MICCAI 2008,D\. Metaxas, L\. Axel, G\. Fichtinger, and G\. Székely \(Eds\.\),Berlin, Heidelberg,pp\. 459–467\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/978-3-540-85988-8%5F55),ISBN 978\-3\-540\-85988\-8Cited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p1.1)\. - A\. Henriksen and A\. Bechmann \(2020\)Building truths in AI: making predictive algorithms doable in healthcare\.Information, Communication & Society23\(6\),pp\. 802–816\.External Links:[Document](https://dx.doi.org/10.1080/1369118X.2020.1751866),[Link](https://doi.org/10.1080/1369118X.2020.1751866),https://doi\.org/10\.1080/1369118X\.2020\.1751866Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p4.1),[§3](https://arxiv.org/html/2607.09668#S3.p1.1),[§3](https://arxiv.org/html/2607.09668#S3.p3.1)\. - C\. Högberg and P\. Winter \(2026\)Digital phantoms in medical research: synthetic data and the pursuit of ground truth\.Big Data & Society13\(2\),pp\. 20539517261447840\.External Links:[Document](https://dx.doi.org/10.1177/20539517261447840),[Link](https://doi.org/10.1177/20539517261447840)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p4.1),[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p1.1)\. - C\. Högberg \(2025\)“This ground truth is muddy anyway”: ground truth data assemblages for medical AI development\.Sociologisk Forskning62\(1\-2\),pp\. 85–106\.External Links:[Link](https://sociologiskforskning.se/sf/article/view/27826),[Document](https://dx.doi.org/10.37062/sf.62.27826)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p4.1),[§2](https://arxiv.org/html/2607.09668#S2.p5.1),[§3](https://arxiv.org/html/2607.09668#S3.p1.1),[§3](https://arxiv.org/html/2607.09668#S3.p3.1),[§4\.2](https://arxiv.org/html/2607.09668#S4.SS2.p2.1),[§4\.2](https://arxiv.org/html/2607.09668#S4.SS2.p3.1),[§5\.1](https://arxiv.org/html/2607.09668#S5.SS1.p5.1)\. - D\. Hughes \(1988\)When nurse knows best: some aspects of nurse/doctor interaction in a casualty department\.Sociology of Health & Illness10\(1\),pp\. 1–22\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/1467-9566.ep11340102),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1111/1467-9566.ep11340102),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1111/1467\-9566\.ep11340102Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - F\. Jaton \(2017\)We get the algorithms of our ground truths: designing referential databases in digital image processing\.Social Studies of Science47\(6\),pp\. 811–840\.Note:PMID: 28950802External Links:[Document](https://dx.doi.org/10.1177/0306312717730428),[Link](https://doi.org/10.1177/0306312717730428)Cited by:[§1](https://arxiv.org/html/2607.09668#S1.p4.1),[§1](https://arxiv.org/html/2607.09668#S1.p6.1),[§3](https://arxiv.org/html/2607.09668#S3.p1.1),[§3](https://arxiv.org/html/2607.09668#S3.p3.1)\. - F\. Jaton \(2021\)The constitution of algorithms: ground\-truthing, programming, formulating\.MIT press\.External Links:ISBN 9780262363235Cited by:[§1](https://arxiv.org/html/2607.09668#S1.p4.1)\. - F\. Jaton \(2023\)Groundwork for AI: enforcing a benchmark for neoantigen prediction in personalized cancer immunotherapy\.Social Studies of Science53\(5\),pp\. 787–810\.Note:PMID: 37650579External Links:[Document](https://dx.doi.org/10.1177/03063127231192857),[Link](https://doi.org/10.1177/03063127231192857)Cited by:[§4\.2](https://arxiv.org/html/2607.09668#S4.SS2.p5.2)\. - F\. Jaton \(2024\)Ground truths are human constructions\.Issues in Science and Technology40\(2\),pp\. 85–85\.Cited by:[§1](https://arxiv.org/html/2607.09668#S1.p6.1)\. - E\. B\. Kang \(2023\)Ground truth tracings \(GTT\): on the epistemic limits of machine learning\.Big Data & Society10\(1\),pp\. 20539517221146122\.External Links:[Document](https://dx.doi.org/10.1177/20539517221146122),[Link](https://doi.org/10.1177/20539517221146122)Cited by:[§1](https://arxiv.org/html/2607.09668#S1.p1.1),[§1](https://arxiv.org/html/2607.09668#S1.p3.1),[§2](https://arxiv.org/html/2607.09668#S2.p2.1),[§2](https://arxiv.org/html/2607.09668#S2.p5.1)\. - R\. Kohavi \(2025\)The history of the UCI Adult dataset\.External Links:[Link](https://www.linkedin.com/posts/ronnyk_the-history-of-the-uci-adult-dataset-a-commonly-activity-7276122804540882944-FKSG/)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p4.1)\. - P\. Kroes and P\. Verbeek \(2014\)The moral status of technical artefacts\.Springer,Dordrecht\.External Links:ISBN 9789400779143Cited by:[§5\.2](https://arxiv.org/html/2607.09668#S5.SS2.p3.1)\. - J\. Lave and E\. Wenger \(1999\)Learning and pedagogy in communities\.Learners & pedagogy21,pp\. 15–34\.Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - S\. Lebovitz, N\. Levina, and H\. Lifshitz\-Assaf \(2021\)Is AI ground truth really true?: the dangers of training and evaluating AI tools based on experts’ know\-what\.MIS Quarterly45\(3\),pp\. 1501–1526\.External Links:[Document](https://dx.doi.org/10.25300/MISQ/2021/16564)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p4.1),[§4\.2](https://arxiv.org/html/2607.09668#S4.SS2.p5.2)\. - Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner \(1998\)Gradient based learning applied to document recognition\.Proceedings of IEEE86\(11\),pp\. 2278–2324\.External Links:[Document](https://dx.doi.org/10.1109/5.726791)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p2.1)\. - Y\. LeCun, C\. Cortes, and C\. J\.C\. Burges \(1994\)The MNIST database of handwritten digits\.Note:http://yann\.lecun\.com/exdb/mnist/Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p2.1)\. - F\. Lee, S\. Hajisharif, and E\. Johnson \(2025\)The ontological politics of synthetic data: normalities, outliers, and intersectional hallucinations\.Big Data & Society12\(2\),pp\. 20539517251318289\.External Links:[Document](https://dx.doi.org/10.1177/20539517251318289)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p5.1)\. - F\. Lee and D\. Ribes \(2025\)Computational universalism, or, attending to relationalities at scale\.Social Studies of Science0\(0\),pp\. 03063127251345089\.Note:PMID: 40657790External Links:[Document](https://dx.doi.org/10.1177/03063127251345089),[Link](https://doi.org/10.1177/03063127251345089)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p4.1),[§3](https://arxiv.org/html/2607.09668#S3.p5.1)\. - Y\. A\. Loukissas \(2019\)All data are local: thinking critically in a data\-driven society\.The MIT Press,Cambridge, Massachusetts\.External Links:ISBN 9780262039666Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p3.1)\. - C\. Milligan, C\. Roberts, and M\. Mort \(2011\)Telecare and older people: who cares where?\.Social Science & Medicine72\(3\),pp\. 347–354\.Note:13th International Medical Geography SymposiumExternal Links:ISSN 0277\-9536,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.socscimed.2010.08.014),[Link](https://www.sciencedirect.com/science/article/pii/S0277953610006386)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - M\. Muller, C\. T\. Wolf, J\. Andres, M\. Desmond, N\. N\. Joshi, Z\. Ashktorab, A\. Sharma, K\. Brimijoin, Q\. Pan, E\. Duesterwald, and C\. Dugan \(2021\)Designing ground truth and the social life of labels\.InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems,CHI ’21,New York, NY, USA\.External Links:ISBN 9781450380966,[Link](https://doi.org/10.1145/3411764.3445402),[Document](https://dx.doi.org/10.1145/3411764.3445402)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p1.1),[§4\.2](https://arxiv.org/html/2607.09668#S4.SS2.p4.1),[§5\.1](https://arxiv.org/html/2607.09668#S5.SS1.p3.1)\. - J\. E\. Orr \(1996\)Talking about machines: an ethnography of a modern job\.ILR Press,Ithaca, NY\.External Links:ISBN 0801432979Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - A\. E\. Ostrowska \(2026\)Life and ai at nasa: an ethnography of how scientists and engineers make tools to explore other worlds\.Ph\.D\. Thesis,Department of Technology Management and Economics, Chalmers University of Technology,Gothenburg\.Cited by:[§4\.1](https://arxiv.org/html/2607.09668#S4.SS1.p5.1)\. - L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. Lowe \(2022\)Training language models to follow instructions with human feedback\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NeurIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p6.1)\. - B\. Plank \(2022\)The “problem” of human label variation: on ground truth in data, modeling and evaluation\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 10671–10682\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.731/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.731)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p3.1),[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p4.1),[§7](https://arxiv.org/html/2607.09668#S7.p2.1)\. - B\. Prainsack and E\. Steindl \(2022\)Legal and ethical aspects of machine learning: who owns the data?\.InArtificial Intelligence/Machine Learning in Nuclear Medicine and Hybrid Imaging,P\. Veit\-Haibach and K\. Herrmann \(Eds\.\),pp\. 191–201\.External Links:ISBN 978\-3\-031\-00119\-2,[Document](https://dx.doi.org/10.1007/978-3-031-00119-2%5F14),[Link](https://doi.org/10.1007/978-3-031-00119-2_14)Cited by:[§5\.2](https://arxiv.org/html/2607.09668#S5.SS2.p3.1)\. - A\. Raeithel \(1996\)On the ethnography of cooperative work\.InCognition and Communication at Work,Y\. Engeström and D\. Middleton \(Eds\.\),pp\. 319–340\.Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - B\. Recht, R\. Roelofs, L\. Schmidt, and V\. Shankar \(2019\)Do ImageNet classifiers generalize to ImageNet?\.InProceedings of the 36th International Conference on Machine LearningProceedings of the 29th International Coference on International Conference on Machine LearningProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society,K\. Chaudhuri and R\. Salakhutdinov \(Eds\.\),Proceedings of Machine Learning ResearchAIES ’23, Vol\.97,pp\. 5389–5400\.External Links:[Link](https://proceedings.mlr.press/v97/recht19a.html)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p6.1),[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p5.1),[3rd item](https://arxiv.org/html/2607.09668#S6.I1.i4.I1.i3.p1.1)\. - L\. B\. Resnick, J\. M\. Levine, and S\. D\. Teasley \(1991\)Perspectives on socially shared cognition\.1\. ed\. edition,American Psychological Association,Washington\.External Links:ISBN 1557981213Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - J\. Rhee \(2025\)1: emotion AI: deauthorization and minoritized unfeeling\.InHow That Robot Made Me Feel,E\. Johnson \(Ed\.\),External Links:[Document](https://dx.doi.org/10.7551/mitpress/15314.003.0003)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p4.1)\. - O\. Russakovsky, J\. Deng, H\. Su, J\. Krause, S\. Satheesh, S\. Ma, Z\. Huang, A\. Karpathy, A\. Khosla, M\. Bernstein, A\. C\. Berg, and L\. Fei\-Fei \(2015\)ImageNet large scale visual recognition challenge\.Int\. J\. Comput\. Vision115\(3\),pp\. 211–252\.External Links:ISSN 0920\-5691,[Link](https://link.springer.com/article/10.1007/s11263-015-0816-y),[Document](https://dx.doi.org/10.1007/s11263-015-0816-y)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p5.1)\. - J\. Schlimmer \(1981\)Mushroom \[dataset\]\.Note:UCI Machine Learning RepositoryExternal Links:[Link](https://doi.org/10.24432/C5959T)Cited by:[§4\.2](https://arxiv.org/html/2607.09668#S4.SS2.p1.1)\. - M\. Schultheiss, P\. Schmette, J\. Bodden, J\. Aichele, C\. Müller\-Leisse, F\. G\. Gassert, F\. T\. Gassert, J\. F\. Gawlitza, F\. C\. Hofmann, D\. Sasse, C\. E\. von Schacky, S\. Ziegelmayer, B\. R\. Fabio De Marco, M\. R\. Makowski, F\. Pfeiffer, and D\. Pfeiffer \(2021\)Lung nodule detection in chest x\-rays using synthetic ground\-truth data comparing CNN\-based diagnosis to human performance\.Nature Scientific Reports11\(15857\)\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41598-021-94750-z)Cited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p2.1)\. - E\. Shimron, J\. I\. Tamir, K\. Wang, and M\. Lustig \(2022\)Implicit data crimes: machine learning bias arising from misuse of public data\.PNAS119\(13\),pp\. e2117203119\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1073/pnas.2117203119)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p3.1)\. - B\. Sobel \(2021\)A taxonomy of training data: disentangling the mismatched rights, remedies, and rationales for restricting machine learning\.InArtificial Intelligence and Intellectual Property,J\. Lee, R\. Hilty, and K\. Liu \(Eds\.\),External Links:[Document](https://dx.doi.org/10.1093/oso/9780198870944.003.0011)Cited by:[§5\.2](https://arxiv.org/html/2607.09668#S5.SS2.p3.1)\. - P\. Soyer \(2018\)Agreement and observer variability\.Diagnostic and Interventional Imaging99\(2\),pp\. 53–54\.External Links:ISSN 2211\-5684,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.diii.2018.01.009),[Link](https://www.sciencedirect.com/science/article/pii/S2211568418300172)Cited by:[§7](https://arxiv.org/html/2607.09668#S7.p2.1)\. - S\. L\. Star \(1991\)Power, technology, and the phenomenology of conventions: on being allergic to onions\.InA Sociology of Monsters: Essays on Power, Technology and Domination,J\. Law \(Ed\.\),pp\. 22–56\.Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - D\. Stutz, A\. T\. Cemgil, A\. G\. Roy, T\. Matejovicova, M\. Barsbey, P\. Strachan, M\. Schaekermann, J\. Freyberg, R\. Rikhye, B\. Freeman, J\. P\. Matos, U\. Telang, D\. R\. Webster, Y\. Liu, G\. S\. Corrado, Y\. Matias, P\. Kohli, Y\. Liu, A\. Doucet, and A\. Karthikesalingam \(2025\)Evaluating medical AI systems in dermatology under uncertain ground truth\.Medical Image Analysis103,pp\. 103556\.External Links:ISSN 1361\-8415,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.media.2025.103556),[Link](https://www.sciencedirect.com/science/article/pii/S1361841525001033)Cited by:[§2](https://arxiv.org/html/2607.09668#S2.p3.1)\. - L\. A\. Suchman \(2007\)Human\-machine reconfigurations: plans and situated actions\.2\. ed\. edition,Cambridge University Press,Cambridge\.External Links:ISBN 9780521858915Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p2.1)\. - L\. Suchman, J\. Blomberg, J\. E\. Orr, and R\. Trigg \(1999\)Reconstructing technologies as social practice\.American Behavioral Scientist43\(3\),pp\. 392–408\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1177/00027649921955335)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p1.1)\. - B\. L\. Sullivan, C\. L\. Wood, M\. J\. Iliff, R\. E\. Bonney, D\. Fink, and S\. Kelling \(2009\)eBird: a citizen\-based bird observation network in the biological sciences\.Biological Conservation142\(10\),pp\. 2282–2292\.External Links:ISSN 0006\-3207,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.biocon.2009.05.006),[Link](https://www.sciencedirect.com/science/article/pii/S000632070900216X)Cited by:[§4\.3](https://arxiv.org/html/2607.09668#S4.SS3.p1.1)\. - G\. Tong, J\. Chao, W\. Ma, Z\. Zhong, G\. Gupta, and W\. Zhu \(2025\)Leveraging synthetic data to improve regional sea level predictions\.\.Scientific Reports15\(1\)\.External Links:[Document](https://dx.doi.org/10.1038/s41598-025-88078-1)Cited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p3.1)\. - A\. Tsirikoglou, J\. Kronander, M\. Wrenninge, and J\. Unger \(2017\)Procedural modeling and physically based rendering for synthetic data generation in automotive applications\.External Links:1710\.06270,[Link](https://arxiv.org/abs/1710.06270)Cited by:[§4\.4](https://arxiv.org/html/2607.09668#S4.SS4.p2.1)\. - K\. L\. Wagstaff \(2012\)Machine learning that matters\.Madison, WI, USA,pp\. 1851–1856\.External Links:[Link](https://dl.acm.org/doi/10.5555/3042573.3042809)Cited by:[§5\.1](https://arxiv.org/html/2607.09668#S5.SS1.p6.1)\. - I\. H\. Woodhouse \(2021\)On ‘ground’ truth and why we should abandon the term\.Journal of Applied Remote Sensing15\(4\),pp\. 041501\.External Links:[Document](https://dx.doi.org/10.1117/1.JRS.15.041501),[Link](https://doi.org/10.1117/1.JRS.15.041501)Cited by:[§1](https://arxiv.org/html/2607.09668#S1.p3.1),[§2](https://arxiv.org/html/2607.09668#S2.p5.1)\. - H\. D\. Zajac, N\. R\. Avlona, F\. Kensing, T\. O\. Andersen, and I\. Shklovski \(2023\)Ground truth or dare: factors affecting the creation of medical datasets for training ai\.New York, NY, USA,pp\. 351–362\.External Links:ISBN 9798400702310,[Link](https://doi.org/10.1145/3600211.3604766),[Document](https://dx.doi.org/10.1145/3600211.3604766)Cited by:[§3](https://arxiv.org/html/2607.09668#S3.p4.1)\.
Similar Articles
Position: Ideas Should be the Center of Machine Learning Research
This position paper argues that machine learning research should prioritize ideas over benchmarks and theoretical guarantees, proposing an 'Ideas First' framework that values behavioral signatures and tailored experiments to promote equity and scientific understanding.
Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
This position paper argues that a scientific understanding of AI must go beyond post-hoc analysis and instead study the training dynamics that shape model behavior, with implications for predicting, intervening, and designing training procedures for desired properties like capabilities and safety.
On What The Panopticon Cannot See
The essay critiques over-reliance on computational models, advocating for ethical design principles like single-estimation and transparency to preserve unquantifiable human experiences.
Position: Fairness Failure in Generative Models is an Evaluation Problem
This position paper argues that fairness failures in generative models are primarily due to evaluation problems and proposes Fairness Cards as a standardized reporting artifact to improve reproducibility and accountability.
Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation
This paper critiques the use of single-reference ground truth in ASR evaluation, arguing it causes epistemic injustice for speakers with aphasia. It proposes a new metric, Epistemic Injustice Distance, and advocates for WER-Range to account for diverse transcription conventions.