Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation

arXiv cs.AI Papers

Summary

A mixed-stakeholder workshop study exploring how community representatives, police officers, and academics negotiate risk boundaries for AI policing use cases, focusing on racial bias. Findings show broad openness to AI adoption except for recidivism risk assessment, with deliberations centering on practical effectiveness and equitable benefit.

arXiv:2608.05418v1 Announce Type: new Abstract: AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption. We present results from a mixed-stakeholder deliberation workshop bringing together 30 community representatives, police officers, and academics to assess the risks of 13 AI use cases in policing, with an explicit focus on racial bias. We found that participants were broadly open to AI adoption, rejecting only three use cases outright, most notably recidivism risk assessment, where objections targeted the premise rather than the implementation. Our analysis reveals that foregrounding racial equity did not narrow the deliberation. Instead, discussions gravitated toward a fundamental set of questions: does this tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? This integrated reasoning, reminiscent of the curb-cut effect in inclusive design, highlights the benefit of incorporating the racial bias lens into the risk-benefit analysis of AI use cases from the outset.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:46 AM

# Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation
Source: [https://arxiv.org/html/2608.05418](https://arxiv.org/html/2608.05418)
###### Abstract

AI tools are being increasingly adopted in policing in the UK and worldwide\. Racial bias is a known and well\-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption\. We present results from a mixed\-stakeholder deliberation workshop bringing together 30 community representatives, police officers, and academics to assess the risks of 13 AI use cases in policing, with an explicit focus on racial bias\. We found that participants were broadly open to AI adoption, rejecting only three use cases outright – most notably recidivism risk assessment, where objections targeted the premise rather than the implementation\. Our analysis reveals that foregrounding racial equity did not narrow the deliberation\. Instead, discussions gravitated toward a fundamental set of questions: does this tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? This integrated reasoning—reminiscent of the curb\-cut effect in inclusive design—highlights the benefit of incorporating the racial bias lens into the risk\-benefit analysis of AI use cases from the outset\.

## 1Introduction

As law enforcement agencies become increasingly stretched and resource\-constrained, artificial intelligence \(AI\) is viewed as a key strategy for increasing efficiency in operations\(National Police Chiefs’ Council[2023](https://arxiv.org/html/2608.05418#bib.bib119),[2025](https://arxiv.org/html/2608.05418#bib.bib255)\)\. Agencies have been experimenting with data\-driven applications for over a decade — from predictive policing to facial recognition — and recent years have added significant political pressure towards large\-scale adoption\. Many concerns have arisen alongside this expansion\. Among them, the persistence and further amplification of racial bias\. The historical and ongoing disproportionality in the criminal justice system — where Black individuals are highly overrepresented in arrests, incarceration, and victimisation — makes racial bias a particularly urgent and hard\-to\-resolve challenge for AI in policing\.

While the number of individual applications and tools, even in the UK alone, is numerous\(Takaet al\.[2025](https://arxiv.org/html/2608.05418#bib.bib254)\), use cases of AI in policing fall into similar categories\. Amongst the most popular and well\-known arepredictive policingandrecidivism risk assessment, for both of which there is substantial academic literature documenting how they can perpetuate racial bias \(e\.g\.,\(Ensignet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib12); Zilkaet al\.[2023b](https://arxiv.org/html/2608.05418#bib.bib19)\)\)\. Guidelines for responsible AI in policing have emerged in response — such as the AI Playbook for Policing\(National Police Chiefs’ Council[2025](https://arxiv.org/html/2608.05418#bib.bib255)\), which advises forces to map which groups could be differentially affected and ensure evaluations capture race disproportionalities\. However, they offer little guidance on how evaluations should actually be conducted, and in practice, these are rarely carried out or shared publicly\.

The people best placed to identify these risks*early*are rarely consulted at the design stage\. The teams tasked with addressing racial bias within policing are often siloed from those driving AI adoption, and the communities most affected rarely have a seat at the table\. These decisions are typically driven by management and technical teams, and do not include those with direct knowledge of how racial disproportionality operates in practice\. Earlier intervention, at the ideation stage, offers a more effective opportunity to flag approaches that could embed or exacerbate racial disproportionality before sunk costs make course\-correction difficult\.

We propose and demonstrate a method for negotiating acceptable risk limits within mixed\-stakeholder groups, suitable for the conception stage of AI adoption\. During a one\-day in\-person workshop, we convened 30 stakeholders spanning community representatives, police, and academia to consider 13 AI use cases in policing\. Participants deliberated in mixed groups using a red, amber, and green risk framework \([Table1](https://arxiv.org/html/2608.05418#S2.T1)\)\. Community representatives were not required to have technical knowledge of AI; their contributions drew on expertise and lived experience of how racial disproportionality becomes embedded in the criminal justice system\. Our aims were twofold: to establish which use cases participants found acceptable for deployment and under what conditions, and to understand the process by which they arrived at these judgments\. We found that participants were broadly open to AI adoption, rejecting only three use cases outright, but were sceptical of AI’s ability to drive meaningful improvements in policing\. Importantly, we found that foregrounding racial equity did not narrow the deliberation: instead, discussions focused on fundamental questions about whether a tool works, for whom, and under what conditions\. We argue this is reminiscent of the curb\-cut effect in inclusive design — that designing with marginalised communities in mind surfaces better questions and can produce better outcomes for everyone\.

## 2Background and Related Works

### 2\.1AI Use Cases in Policing

There is substantial interest in AI use in law enforcement, in the UK and worldwide\. What ‘counts’ as AI varies, but below we give a brief background to a number of high\-profile use cases that were considered in the workshop\.

#### Forecasting crime hotspots \(“predictive policing”\)

Efforts to identifywherecrime is likely to occur or stable hotspots to guide resource deployments has been undertaken from as early as the 1930s\(Shaw and McKay[1942](https://arxiv.org/html/2608.05418#bib.bib267); Bragaet al\.[2019](https://arxiv.org/html/2608.05418#bib.bib268)\)\. This is usually based on the assumption that where crime has occurred repeatedly is where it is likely to happen in the future\. However, realistically, accurate prediction of whereandwhencrimes will occur is quite challenging\(Government of the United Kingdom[2025](https://arxiv.org/html/2608.05418#bib.bib199)\)\. The utility of “real\-time” hotspot methods is questionable due to the finer the resolution, the variability, and the low number of recorded crimes\. Indeed, evaluations of these prediction tools, such as PredPolTMhave not shown evidence of crime prevention benefits\(Huntet al\.[2014](https://arxiv.org/html/2608.05418#bib.bib10)\)\. In addition, these tools raise strong concerns about racial inequity\(Lum and Isaac[2016](https://arxiv.org/html/2608.05418#bib.bib14); Sankinet al\.[2021](https://arxiv.org/html/2608.05418#bib.bib168)\), even if they improve forecast accuracy\(Mohleret al\.[2015](https://arxiv.org/html/2608.05418#bib.bib15); Leeet al\.[2024](https://arxiv.org/html/2608.05418#bib.bib196)\)\. However, identifying*stable*hotspots of violence can be done with triangulation of police, hospital, and ambulance data\(Tayloret al\.[2016](https://arxiv.org/html/2608.05418#bib.bib17)\)\.

#### Risk of reoffending prediction

Criminal justice systems have relied on the prediction of an individual’s risk of future crime for more than 200 years\. Tools evolved from phrenology’s pseudo\-science in the 1800s, to actuarial models starting in the 1920\-30s, to statistical \(e\.g\., logistic\) models dominating for many decades\(Berket al\.[2014](https://arxiv.org/html/2608.05418#bib.bib197); Greeneet al\.[2022](https://arxiv.org/html/2608.05418#bib.bib198)\)\. An early example of ‘AI’ in policing was the HART tool in Durham Constabulary to aid decisions about police bail\(Oswaldet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib23)\), assessments of ‘dangerousness’ predicting homicide\(Berket al\.[2009](https://arxiv.org/html/2608.05418#bib.bib24)\)and domestic violence\(Berket al\.[2016](https://arxiv.org/html/2608.05418#bib.bib25)\)\. These tools have consistently raised concerns about racial inequity\(Angwinet al\.[2016](https://arxiv.org/html/2608.05418#bib.bib21); Harcourt[2010](https://arxiv.org/html/2608.05418#bib.bib20); Zilkaet al\.[2023b](https://arxiv.org/html/2608.05418#bib.bib19)\)\.

#### Facial recognition

Facial recognition is now used in the UK for both live operations and retrospective investigations\(Home Office[2024](https://arxiv.org/html/2608.05418#bib.bib40)\)\. Live facial recognition \(LFR\) is arguably the most well\-known application of AI in policing and has been in development for at least the last 15 years\(Daviset al\.[2010](https://arxiv.org/html/2608.05418#bib.bib29)\)\. LFR has been scrutinised worldwide since its first pilot, with concerns raised about wrongful arrest and ethnic disproportionality\(Buolamwini and Gebru[2018](https://arxiv.org/html/2608.05418#bib.bib3)\)\. These concerns have been addressed in some instances – the UK’s National Physical Laboratory assessed LFR equity questions\(Mansfield[2023](https://arxiv.org/html/2608.05418#bib.bib51)\), and the Metropolitan Police have examined false positives, finding that only those on wanted lists were arrested\(BBC News[2025](https://arxiv.org/html/2608.05418#bib.bib52)\)\. However, recent events put these in question \(see Box 2\)\.

Table 1:Workshop use case feedback form structure\.

### 2\.2Racial Bias in Policing AI Systems

Over the last decade, critical technical and socio\-technical research has highlighted bias and related issues in algorithmic tools for policing\. Several studies demonstrated that reliance on arrest records as a proxy for underlying crime can cause an exacerbation of racial bias in different types of predictive tools\. This is due to the likelihood of a crime becoming known to law enforcement varying significantly with race, sex, and age\(Baumer and Lauritsen[2010](https://arxiv.org/html/2608.05418#bib.bib261); Bosicket al\.[2012](https://arxiv.org/html/2608.05418#bib.bib262); Butcheret al\.[2022](https://arxiv.org/html/2608.05418#bib.bib129)\)\.\(Zilkaet al\.[2023b](https://arxiv.org/html/2608.05418#bib.bib19)\)demonstrated this for tools predicting the risk of re\-offending\. Within ‘hotspot mapping’ tools, feedback loops can increase patterns of over\-policing in already over\-policed communities\(Ensignet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib12); Lum and Isaac[2016](https://arxiv.org/html/2608.05418#bib.bib14); Sankinet al\.[2021](https://arxiv.org/html/2608.05418#bib.bib168)\)\. For facial recognition, early work documented significantly higher error rates for darker\-skinned individuals\(Buolamwini and Gebru[2018](https://arxiv.org/html/2608.05418#bib.bib3)\), and while technical improvements have narrowed some gaps, equity in practice remains contested and deployment\-dependent\(Rakova and Dobbe[2023](https://arxiv.org/html/2608.05418#bib.bib64)\)\.

### 2\.3Participatory Research on AI and Policing

Alongside the critical work, a growing body of work examines how practitioners and communities relate to data\-driven policing\.Kearneyet al\.\([2024](https://arxiv.org/html/2608.05418#bib.bib116)\)interviewed 40 Police Scotland practitioners and found that officers were averse to tools like predictive risk assessment and facial recognition, and were concerned with the erosion of community policing through datafication\.Zilkaet al\.\([2023a](https://arxiv.org/html/2608.05418#bib.bib115)\)explored police perspectives on algorithmic transparency and found mixed views, highlighting the importance of transparency for public trust alongside operational concerns\. Few works have directly engaged community representatives on AI in policing; one exception isHaqueet al\.\([2024](https://arxiv.org/html/2608.05418#bib.bib27)\)who brought together community members, technical experts, and law enforcement around a crime\-mapping application, finding that community members were more likely than domain experts to question the core motivation of the tool rather than its technical execution\.Ziosi and Pruss \([2024](https://arxiv.org/html/2608.05418#bib.bib28)\)interviewed community organisations, researchers, and public sector actors about the Chicago crime prediction algorithm, finding that community\-impacted groups used evidence of algorithmic bias to centre liberation and healing, while public sector actors used the same evidence to reaffirm existing power structures\. Scholars have also raised concerns about algorithmic tools deployed without community input, particularly regarding their impact on marginalised groups\(Richardsonet al\.[2019](https://arxiv.org/html/2608.05418#bib.bib265); Brayne[2017](https://arxiv.org/html/2608.05418#bib.bib266)\)\.

### 2\.4Responsible AI in Policing Frameworks

Acknowledging that we have moved beyond interest in AI to live use in policing, several practitioner frameworks for the responsible use of AI in policing have been developed — in the UK, this includes the AI Playbook for Policing\(National Police Chiefs’ Council[2025](https://arxiv.org/html/2608.05418#bib.bib255)\), the Responsible AI Checklist for Policing\(National Police Chiefs’ Council and PROBabLE Futures[2025](https://arxiv.org/html/2608.05418#bib.bib259)\), and the Police Foundation’s review of AI in policing\(Muir and O’Connell[2025](https://arxiv.org/html/2608.05418#bib.bib2)\)— and broader reviews of AI deployment in policing have been undertaken\(Zilkaet al\.[2022](https://arxiv.org/html/2608.05418#bib.bib121); Berk[2021](https://arxiv.org/html/2608.05418#bib.bib11); Leeet al\.[2024](https://arxiv.org/html/2608.05418#bib.bib196)\)\. At the regulatory level, the landscape is shifting rapidly but unevenly across jurisdictions\. The EU AI Act represents the most comprehensive risk\-based AI legislative response to date\. It classifies a range of policing AI applications as either prohibited or high\-risk\. With high\-risk deployments subject to mandatory impact assessments, transparency requirements, and human oversight obligations\. Critics have noted, however, that the Act’s exceptions for law enforcement are broad enough to permit many of the practices it ostensibly restricts\(European Parliament and Council of the European Union[2024](https://arxiv.org/html/2608.05418#bib.bib61)\)\. In the UK, which sits outside the EU regulatory framework post\-Brexit, the approach remains sector\-led and primarily voluntary\. The Algorithmic Transparency Recording Standard \(ATRS\), made mandatory for central government departments in December 2024, requires public sector organisations to publish information about how and why they use algorithmic tools\. However, policing remains operationally independent: only two police force ATRS records had been published as of 2025, and compliance is not currently mandatory for forces\(Government Digital Service[2024](https://arxiv.org/html/2608.05418#bib.bib62)\)\.

### 2\.5The Police Race Action Plan

UK’s Police Race Action Plan \(PRAP\)\(College of Policing and National Police Chiefs’ Council[2022](https://arxiv.org/html/2608.05418#bib.bib1)\)was developed jointly by the National Police Chiefs’ Council \(NPCC\) and the College of Policing to address the race disparities affecting Black people that policing cannot currently fully explain\. Launched in 2022, the plan sets out key areas for improvement across policing structured around four themes:

- •Culture and Workforce: a police service that represents and supports its Black officers, staff and volunteers\.
- •Powers and Procedures: a police service that is fair, respectful, and equitable in its actions towards Black people\.
- •Trust and Reconciliation: a police service that routinely involves Black people in its governance\.
- •Safety and Victimisation: a police service that protects Black people from crime and seeks justice for victims\.

The plan’s commitments include introducing mandatory training on racism, anti\-racism, and Black history; adopting a new “explain or reform” approach to race disparities in the use of police powers such as stop and search, use of Taser, and other types of force; reviewing misconduct and disciplinary processes to reduce racial disparities; and better enabling Black people to have their voices heard in local communities and within policing itself\. The current plan, however, does not include the police’s rapid inclusion of new technologies and the associated risk of increased racial bias\. This work bridges this gap directly and supplements PRAP with a specific AI\-focused mixed\-stakeholder deliberation\.

ThemeCase \#Use CaseRisk ClassificationCulture & Workforce1AI in Recruitment3±03\\pm 0\(n=5n=5\)2AI review of body\-worn footage to identify misconduct2\.8±0\.42\.8\\pm 0\.4\(n=5n=5\)3Bias\-auditing dashboards2±1\.22\\pm 1\.2\(n=4n=4\)Powers & Procedures4Predictive policing \(hotspot mapping\)3±03\\pm 0\(n=3n=3\)5AI\-assisted classification / analysis / summary of reports2\.3±1\.22\.3\\pm 1\.2\(n=3n=3\)6Risk assessment / predictive profiling5±05\\pm 0\(n=4n=4\)7Live Facial recognition3±03\\pm 0\(n=3n=3\)Trust & Reconciliation8Community\-sentiment analysis from public data1\.4±0\.91\.4\\pm 0\.9\(n=5n=5\)9Sentiment Analysis / Summary of Community Feedback1\.8±1\.11\.8\\pm 1\.1\(n=5n=5\)10Public\-facing AI virtual assistant1\.8±1\.11\.8\\pm 1\.1\(n=5n=5\)Safety & Victimisation11Risk Prediction for Victimisation3\.8±1\.13\.8\\pm 1\.1\(n=5n=5\)12AI transcription / translation of calls / statements / interviews3±1\.43\\pm 1\.4\(n=5n=5\)13AI advice to decide if to carry the investigation forward4\.2±1\.14\.2\\pm 1\.1\(n=5n=5\)Table 2:A summary of themes and AI use cases in policing discussed amongst participants\. The average risk classification is reported, alongside the standard deviation\.nnis the number of groups which explicitly classify the risk level for the use case\. Risk levels are enumerated as 1 \(low risk or green\), 3 \(medium risk or amber\), and 5 \(unacceptable risk or red\)\.

## 3Methods

### 3\.1Participants

Participants were invited to register for the workshop by the organisers based on the recommendations of the central Police Race Action Plan \(PRAP\) team\. Four main groups of participants were invited: \(a\) community representatives involved in the reduction of racial bias, discrimination, and disproportionality; \(b\) members of the Police working on reducing disproportionality; \(c\) members of the Police working on innovation and AI adoption; \(d\) academics who work on the responsible deployment of AI in policing\.

Individuals were recruited from the PRAP teams’ networks and rolling recommendations, and—in one case—by a nomination from another invitee\. Community representatives who have been or are currently engaged with PRAP were the main source being drawn upon for community representation\. Several academics were invited by the organisers directly\. Invitations were sent to approximately 45 potential participants, of which 35 registered for the workshops, and 30 attended for the day, excluding 2 organisers\. Of those, 15 were police, 11 were community representatives, and 4 were academics\. In terms of demographics, approximately half of attendees belonged to UK minority ethnic communities; gender balance was roughly even\.

### 3\.2Agenda and Use Cases

After a short introduction, the participants were divided into 6 groups for the morning, and a different set of 6 groups for the afternoon session\. In each session, the groups reviewed a set of AI use cases relevant to policing, divided into four core PRAP themes: \(a\)*Culture & Workforce*\(referring to the recruitment, retention and progression of black officers and staff\), \(b\)*Powers & Procedures*\(relating to fair, respectful and equitable use of police powers\), \(c\)*Trust & Reconciliation*\(relating to community involvement and anti\-racist culture\), and \(d\)*Safety & Victimisation*\(relating to the protection and experiences of black victims of crime\)\. Before the workshops, we asked the PRAP team to highlight any use cases they wanted include in the workshop\. Additional use cases were added by the organisers based on a literature review\. The full list of use cases can be found in[Table2](https://arxiv.org/html/2608.05418#S2.T2)\.

For each use case, one author created an information sheet that included a deployed example, the intended goal, and known risks associated with the use case\. The author attempted to present the information objectively and use non\-technical language\. The provided information was not comprehensive; it was intended as a conversation starter and leveller, since some participants were less familiar with AI and its associated risks than others\. The information sheets for all use cases can be found in the supplementary material\.

### 3\.3Risk framing and use case feedback

To structure the discussion, we asked participants to complete an assessment worksheet as shown in[Table1](https://arxiv.org/html/2608.05418#S2.T1)\. As a group, participants were asked to classify the use cases in terms of risk: \(a\)*Green*– low risk, can be deployed with minimal checks; \(b\)*Amber*– medium risk, deploy only in a specific context and with rigorous checks; \(c\)*Red*– unacceptable risk, do not deploy\. Participants were then asked to justify their selection, indicate allowed use cases and pre\-deployment checks, and highlight what needs to be evaluated and monitored\. Any additional comments were welcomed, and appended to the group’s worksheet\.

### 3\.4Analysis

After the workshop, the 76 completed worksheets were organised by use case, scanned, and typed by one of the authors\. The analysis proceeded in two stages\. In the first stage, one author conducted a thematic analysis of the worksheet responses for each use case, identifying recurring concerns, points of consensus and disagreement, and the reasoning participants used to justify their risk classifications\. Three additional authors, who were also present in the discussions on the day, reviewed and annotated the emerging themes, ensuring they represented not only the worksheets, but the conversations during the workshop \(which were not recorded\)\. In the second stage, two authors examined the questions participants appeared to be implicitly asking when arriving at their risk classifications — what we refer to as back\-engineered questions\. This was cross\-checked by two additional authors to ensure the findings were faithful to the discussions on the day\. Where direct quotes are used in the results section, they are drawn verbatim from the worksheets\. Risk classification was translated to a 5\-point scale, with low being 1, medium 3, and unacceptable risk 5, and summarised for each use case\. When a group marked two risk categories \(e\.g\., low and medium\), the number in\-between the categories \(e\.g\., 2\) was assigned\. Average risk scores reported in[Table2](https://arxiv.org/html/2608.05418#S2.T2)reflect the mean across groups and may therefore take non\-integer values\.

Individual risk assessment and predictive profilingtools, such as COMPAS \(US\), and OGRS and OASys \(UK\), aim to offer an efficient, consistent, and objective risk categorisation with respect to reoffending\. Usually they estimate the risk of re\-arrest within a fixed time period \(e\.g\., 12 months\)\. The best models achieve just under 70% accuracy\. In the UK tools, race is explicitly excluded as an input but validation takes place for different ethnic groups to ensure comparability and minimise inequity\(Moore[2015](https://arxiv.org/html/2608.05418#bib.bib250)\)\. Nonetheless, these often tools rely on arrest and/or prior convictions as proxy for previous offending \(which tends to be higher than official records\)\. But if the probability of being arrested is itself racially biased, or subsequent justice decisions are, tools will reflect and compound that bias\.\(Zilkaet al\.[2023b](https://arxiv.org/html/2608.05418#bib.bib19)\)\.This use case gathered the most objections\. While two groups did not rank it, the other four groups ranked it as an unacceptable risk\. Unlike other use cases where objections are rooted in implementation rather than the idea itself,here the objection is clearly to the idea itself– “bias issues are baked into this”\. Participants also failed to see the practical benefit:“would it be able to prevent reoffending?” This strong objection is worth noting, particularly as policing rebrands itself as more “predictive” or “precise” as part of its“AI Revolution”\. Most tools of this kind focus on re\-arrest rather than harm, with no emphasis on rehabilitation, context, or enabling non\-reoffending\. However, participants acknowledged that the question can be reformulated: rather than predicting who will reoffend, could the question become“what will help this person not to reoffend?” — whether that is housing support, employment, or rehabilitation services\. This re\-framing is much more challenging to implement in practice, but it represents a more constructive and ethically defensible way forward for individualised predictive tools within criminal justice\. We note that in the UK, probation risk assessment tools OASys and Asset, include questions on housing, welfare, and benefits with the explicit aim to form part of intervention plans and risk management strategies, which then go onto inform supervision activities \(Moore, 2015\)\.Box 1: Strong objection to predicting the risk of re\-arrest\.
### 3\.5Limitations

Several study limitations should be acknowledged\. First, all groups reviewed use cases in the same fixed order, which may have introduced order effects: fatigue, anchoring, or carry\-over reasoning from earlier use cases may have shaped how later ones were assessed\. Future workshops like this should counterbalance the order of use cases across groups\. Second, recruitment was conducted primarily through the networks of the PRAP team, which means participants were largely pre\-selected for their existing engagement with issues of racial bias and disproportionality in policing\. This was a conscious choice as racial bias in policing was central to the workshop; however, we acknowledge this sample is unlikely to be representative of the broader police workforce\. This workshop did not aim to present the public’s views, but of stakeholders representing diverse interests\. Third, workshop discussions were not recorded so that participants could feel that their conversations were private and that they could communicate freely\. Additionally, authors present on the day observed discussions across groups and were able to cross\-check that the thematic findings reported were reflective of the conversations as they occurred, and not only of what participants chose to commit to paper\.

Fourth, we acknowledge the possibility of in\-group power dynamics\. When assigning groups for the workshop, we aimed to balance the police and community representatives and assign one academic to each group\. To avoid only dominant voices emerging, we remixed the groups for the afternoon session, such that groups only contained individuals who had not previously been grouped\. The information sheets shared with participants also helped ensure a common base of knowledge around each of the use cases\. Organisers moved between groups throughout morning and afternoon sessions, encouraging groups to try to allocate even amounts of time to each use case and supporting quieter individuals to join in\. Some community representatives appeared to initially feel aware of their lack of AI technical expertise and needed encouragement to share their views\. Interestingly, many police present had similar levels of AI technical expertise, but did not seem to need as much encouragement to share\. Police participants’ confidence in their understanding of how use cases are deployed in practice could translate to confidence talking about the responsible use of AI\. However, the expertise of community representatives regarding ways in which disproportionality is baked into the criminal justice system did not translate as quickly into confidence talking about the responsible use of AI\. By the end of the day, all participants appeared to feel confident fully engaging and sharing views\.

Finally, we reflect on positionality\. The authorship team approaches this work primarily from an academic perspective, with expertise in responsible AI and algorithmic fairness\. One author brings a more practitioner\-oriented experience, and another has a background in public policy\. All authors believe that racial inequality is a critical issue in policing and the criminal justice system, and this may have shaped the analysis\.

## 4Results of Use Case Discussions

In this section, we report on the analysis of 76 completed worksheets for 13 AI in policing use cases\. Not all 6 groups completed all use cases\. Use cases were distributed across morning and afternoon sessions, and groups varied in how many they were able to assess within the time available\. In addition, not all completed worksheets assigned a risk categorisation\. The n values in[Table2](https://arxiv.org/html/2608.05418#S2.T2)reflect the number of groups that explicitly assigned a risk level for each use case\.

### 4\.1Theme A: Culture & Workforce

For this theme, there were 3 use cases, all marked between low and medium risk by all groups\.

#### AI in Recruitment

The use of automation in recruitment is on the rise in UK policing\. Oleeo, a recruiting platform for Police, claims on their website that their platform is used by 70% of UK police forces\(Oleeo[2025](https://arxiv.org/html/2608.05418#bib.bib269)\)\. The first use case therefore focused on AI in recruitment of new officers, particularly in risk screening\. The main perceived benefits were efficiency and cost savings, and a better experience for the candidates\. Another suggested advantage was reduction in bias, by hiding the candidate’s name and gender which is called “blind recruitment”\.

All 6 groups completed this use case, but only 4 ranked it, all assigning Amber \(medium\) risk\. Participants agreed that diversifying the workforce was mission\-critical, but were sceptical of the ability of AI to make headway in this challenge\. The existing, non\-tech recruitment process was described as “a dumpster fire, even before AI”\. Participants found the concept of “blind recruitment” naive and unlikely to lead to meaningful change, a sentiment backed up by the literature\(Simonset al\.[2021](https://arxiv.org/html/2608.05418#bib.bib13)\)\. The resistance to blind fairness was a strong common theme across groups, as was the concern that human\-in\-the\-loop mitigations are proposed without clarity on how they would work effectively in practice\.

Participants highlighted that training data and assessment criteria reflect existing—non\-diverse—norms, which means models can only be trained to identify the types of officers already hired in the past\.111This is a well known issue that does not only affect policing, usually known as the “credit assignment problem”\.They acknowledged the potential benefits of automation—as there may be real gains in time, money, and candidate experience—but highlighted that this might not translate to genuine progress, i\.e\., to getting the best potential officers or shaping the force to be what it needs to be to future\-proof its relationship with a diverse population\. Considering the need for change, there was also a worry that automating the hiring process can make it more rigid and harder to reform\. This use case was an example of an overarching theme in the discussions, where participants felt AI is being used as a “shiny fix” for a process with more fundamental problems, without proper evaluations to see whether it brings any real improvements\.

#### AI Review of Body\-Worn Camera

This use case was relatively well\-received, as it adds accountability and does not replace an existing human process\. However, participants noted that the underlying assumption—that all behaviour is recorded—is undermined in practice: cameras are sometimes off, a minimum 30\-second buffer exists, and officers, particularly bad actors, may use them selectively\. This led to scepticism about the system’s potential added value\. However, most groups agreed there is real potential for improving professional behaviour, accessibility, and trust, especially, if the system should be used not only as a “stick” but also as a “carrot”, with good behaviour highlighted as well, motivating officers acceptance and cooperation\. Fundamental questions were raised about how the act of recording might change the behaviour of everyone involved, and whether it could interfere with helpful, open, and empathetic policing\. The surveillance implications for the public depend entirely on how the system is used, what its goals are, and what commitments are upheld\. For evaluation, monitoring and random sampling were suggested, though a concern was raised that supervisors may lack the resources to review footage effectively, increasing the risk of over\-reliance on AI\.

#### Bias\-Auditing Dashboards

Discussions on this use case were briefer than the previous ones, likely due to time constraints\. Generally, the prevailing tone was scepticism rooted in a lack of specificity—without understanding how the auditing is done, it is very hard to know whether it can be helpful\. Some of that scepticism stemmed from the act of quantification itself and the absence of cultural context: “understanding context is more useful than surface\-level disproportionality”\. Several participants suggested limiting the audit to HR or keeping it project\-specific to make it more actionable\. Groups who rated this use case as low risk did so on the basis that it is not a decision\-making tool but one that offers an opportunity for reflection and transparency—potentially more useful for monitoring AI decision\-making rather than human decision\-making\. Notably, none of the discussions mentioned the potential for ethical whitewashing or other cynical uses of such a tool\.

### 4\.2Theme B: Powers & Procedures

#### Predictive Policing \(Hotspot Mapping\)

Perhaps surprisingly, participants did not object to predictive policing or hotspot mapping per se—the ability to plan and strategise resource deployment was viewed positively\. Instead, the objections concerned implementation and usage\. Participants highlighted two core issues: 1\) the reliance on arrest data as the primary input, and 2\) the goal of optimising for enforcement\. Arrest data does not faithfully represent harm, victimisation, or community sentiment on feeling unsafe, and crime non\-reporting patterns are often systematically biased\. Participants echoed concerns from literature\(Ensignet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib12)\)about feedback loops, i\.e\., predictions informing interventions, which generate new data used for further predictions, compounding existing patterns of over\-policing of racial minority communities\. For almost all groups, discussion of data led naturally to discussion of goals, and a clear consensus that increasing arrests is*not in itself a worthy goal*\. Participants feared that this technology supports what might be called a “whack\-a\-mole model” of crime prevention, noting crime displacement and a lack of connection to wider policing strategy or the social causes of crime\. One comment captured this well: there is a need to “balance between public health and justice approaches”\. On a positive note, participants were optimistic about the potential to use this technology as a resource not just for enforcement but for multi\-agency collaboration, strategy, and policymaking, employing both wider data sources and a broader range of outputs\.

#### AI\-Assisted Processing of Incident Reports

There is a clear appetite for using AI in this context, provided it is deployed safely and responsibly\. The motivations articulated by participants—extracting patterns, informing new officers about cases, and managing large volumes of information—all centre on AI as an assistant supporting human decisions, rather than as an automated decision\-maker\. However, this is often not the guiding principle in the design of current tools\. For AI\-generated summaries and co\-pilot tools, resistance came from several directions: \(a\) a false sense of efficiency \(the time saved generating a report may be offset by the time needed to verify it, or worse, it may not be verified at all\); \(b\) accountability questions around whether officers can be held responsible for AI\-generated reports without undermining the efficiency gains; \(c\) lack of contextual, cultural, and compassionate understanding; and \(d\) potential downstream consequences for prosecution and court procedures\. Evaluation is a major, non\-trivial effort – red\-teaming, stress\-testing, hallucination rates, and information retrieval accuracy are all important, but collectively represent a significant undertaking that may quickly become outdated\. At its core, the discussion about risk\-benefit boiled down to one central question: where, when, and how is AI adoption worth it?

#### Risk Assessment \(Predictive Profiling\) & Live Facial Recognition

Individualised Risk Assessment gathered the most objections of any of the discussed use cases\. In contrast, the use of Live Facial Recognition was generally well accepted under certain conditions\. Both use cases are discussed in detail in Box 1 and Box 2, respectively\.

Live Facial Recognition \(LFR\)—as deployed by the Metropolitan Police \(“the Met”; serving the greater London area\)—functions as a real\-time aid to help officers locate wanted individuals from a watchlist, with safeguarding as an additional declared benefit\. Historically, commercial facial recognition systems were shown to have higher rates of misidentification for Black individuals, and Black women in particular\(Buolamwini and Gebru[2018](https://arxiv.org/html/2608.05418#bib.bib3)\)\. However, LFR systems have since improved, and the Met claims that their system can be operated at settings where there is no statistically significant difference in demographic performance across groups\.Within our workshop, LFR was considered acceptable by the majority of groups, but only under clearly defined conditions: 1\) use only against high\-harm and/or sexual offenders; 2\) do not use for minor crimes \(“most wanted, prolific, and high harm, but not cannabis use”\); and 3\) limit racial bias to a minimum\. Additional conditions included transparency about who is on the list, where and how the technology is deployed, and transparency around outcomes and actual bias levels, alongside human oversight, clear governance, and significant community consultation\. The discussion felt more refined than for other use cases, likely due to considerable effort by UK policing to communicate the conditions under which LFR is deployed and to demonstrate positive outcomes\. The bias discussion extended beyond technical accuracy to include watchlist composition and deployment locations as potential sources of bias\. Evaluation criteria included false and true positives, value of use and outcomes, and bias across multiple dimensions\.After the workshop, two noteworthy developments occurred\. In March 2026, Essex Police paused LFR deployments after a commissioned study found significantly worse performance for Black individuals\(Bland and Verrey[2026](https://arxiv.org/html/2608.05418#bib.bib65)\)\. Specifically that study found that ”Black people were 27 per cent more likely to be correctly identified than any other ethnicity and 31 per cent more likely than White people” \(p\.4\)\. In April 2026, the High Court dismissed a judicial review challenge to the Met’s revised LFR policy, finding it imposed sufficient constraints to meet the ECHR “quality of law” test; the court acknowledged that a policy authorising discriminatory use could undermine its legality, but was not persuaded this policy does so\. These events underscore how acceptance of LFR hinges on specific conditions and assurances about its use and impact remains contested\.Box 2: Live Facial Recognition deemed acceptable for high\-harm offenders\.

### 4\.3Theme C: Trust & Reconciliation

#### Sentiment Analysis & Summary of Community Feedback

These use cases are quite similar, with the former focusing on analysing public data such as social media, and the latter on feedback submitted directly to the police\. Both use cases were generally perceived as low to medium risk\. The key justification was that the scale of monitoring involved cannot be—and currently is not—done effectively manually\. Therefore, AI analysis can function as an informative layer on top of existing processes rather than replacing them\. Described as a useful“temperature check”and tool for emerging issues, sentiment analysis was considered acceptable when used to inform rather than to drive action directly\. The risks identified were consistent across both use cases: the sentiment captured is not a balanced picture and can be skewed by“loud voices and keyboard warriors”; marginalised or underrepresented communities may be excluded or crowded out; and the tools struggle with nuance, sarcasm, tone, and local community context\. Evaluation is also a challenge — outputs need to be compared to human assessment, not just assessed on efficiency grounds\. The recommendation across groups was to use these tools alongside other means of gathering community sentiment, and to remain clearly aware of the limitations of the underlying data\.

#### Public\-Facing AI Virtual Assistant

Classified as low to medium risk by all groups, the virtual assistant was considered acceptable when limited to non\-emergency, low\-risk, and informative purposes\. One group noted that the public already expects these kinds of tools, and already uses external LLMs to ask questions about information on police websites\. A key distinction was drawn between a chatbot that simply directs users to the right part of a website \(low risk but also low benefit\), and one that engages more substantively, which raises greater concerns around misinformation, safeguarding, and public trust\. Essential requirements included monitoring for effectiveness and user feedback, a redirection\-to\-an\-operator mechanism, and clear signalling about appropriate use \(e\.g\., do not use in an emergency; do not report a crime through this channel\)\. The concern that misinformation could harm public trust was raised, alongside a cost\-benefit question about whether the tool genuinely redirects demand away from 101 calls\. Accessibility and the ability to detect vulnerability in user queries—something a human operator might more naturally recognise—were also flagged as important design considerations\.

### 4\.4Theme D: Safety & Victimisation

#### AI Transcription / Translation of Victim Calls / Statements / Interviews

This use case did not achieve consensus, being rated low risk by one group but unacceptable by another, with several medium risk assessments in between\. The group that rated it low focused on compliance with existing regulations \(e\.g\., CPIA222The Criminal Procedure and Investigations Act, 1996\.\), data security, governance, and officer training – a procedural approach\. The group that rated it unacceptable was chiefly concerned with the risk of victims and witnesses being misunderstood, and its potential disparate impact:“negatively affect the Black community due to cultural misunderstandings”\. This concern about disproportionate impact was shared broadly among the participants, who observed that“minority groups are already a step behind, so AI increases the discrimination”\.Concerns were raised about where and how transcription would be used, with non\-evidential recordings considered more acceptable, and higher\-stakes uses—such as evidence presented at court or 999 calls—considered not acceptable\. Most groups were sceptical automated transcription would result in genuine time savings if done properly:“how much time is being saved if everything is being checked over?”\. Requirements included human verification of output, transparency, consent, and appropriate control for the victim or witness, including the ability to“stop, start, and \[go\] back”\. Questions of accountability were also raised:”who takes ownership of ensuring that \[the transcript\] is correct?”

#### AI Advice to Decide Whether to Carry an Investigation Forward

Half of the groups classified this use case as unacceptable risk, with the remainder classifying it as medium risk\. The most accepting groups saw potential value if used aspart of a suite of products to support decision\-making”, with the tool providing assistance rather than acting as a“decision\-maker”in itself\. The groups who found it unacceptable focused on the risk of disparate impact and bias, as well as concerns about gamification and broader ethical concerns:“seems completely unethical”and“bare minimum of what policing should be doing”\.Even the more accepting groups expressed concerns about bias, over\-reliance, the effect on victims, and whether the tool would deliver true efficiency compared to a human baseline\.

#### Risk prediction tools for victimisation

See Box 3\.

Risk prediction tools for victimisation—such as Spain’s VioGén\(Álvarezet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib114); Satariano and Toll Pifarré[2024](https://arxiv.org/html/2608.05418#bib.bib190)\)—aim to reduce harm by predicting risk in domestic violence cases\. However, their accuracy is fundamentally limited, not by technology, but the difficulty of the prediction problem\. A particular blind spot is the lack of recording of interventions, which can make cases where violence was successfully prevented incorrectly classified as low risk\. In the UK, the checklist\-based risk\-assessment tool “DASH” was designed to help frontline practitioners identify individuals at highest risk of serious harm, but a series of tragic failures raised concerns that it was missing high\-risk victim\-survivors, failing to capture dynamic risk, the lived realities of victim\-survivors, the cultural nuance affecting minority communities, and has little use in terms of predictive validity\. In 2022, the College of Policing recommended switching to an alternative “DARA” tool which has better predictive validity; yet as of August 2025, 20 of 39 UK police forces still use DASH\.Our groups’ views were divided: two rated this use case as unacceptable risk, three as medium risk, and one left it unmarked\. Where deemed acceptable, justifications included the tool already being in use in some form, the potential for AI to enable greater scale and consistency, and its value as a“second voice”that may reduce human bias\. It was emphasised that the tool should function as“part of a wider system”supporting officer decision\-making, as it was“high risk if used as a sole tool without giving consideration of other factors\.”The subtlety of this assessment was stressed throughout, with cases“often requiring a nuanced understanding of context to make a holistic decision\.”These concerns link to fundamental questions around agency and accountability\. A key requirement was that officers retain full accountability“to remain accountable for decisions made\.”This was tied directly to bias \(“fear of overreliance, lack of accountability on bias from police”\), and an overall worry that the tool would exacerbate cultural stereotypes:“Black women are seen as less vulnerable, so large bias in training data\.”The importance of participatory development was highlighted:“bring in victims/survivors—validating questions being asked, gather their perception on this,”alongside the need for multi\-agency response\. For evaluation, the need to understand the human baseline was emphasised\. However, this is inherently difficult, as successful interventions prevent harm and may therefore resemble misjudgments of risk, making it hard to demonstrate good judgment precisely when it succeeds\.Box 3: Concerns around Risk prediction tools for victimisation exacerbating racial bias and cultural stereotypes\.

## 5Analysis of Risk\-Bounding Process

In addition to considerations on the specific use cases, the workshop helped us gain insights into the process of establishing the risk category for the use cases\. When participants were asked to classify the risk of each use case, they received no guidance on what to consider, except that it should be in the context of racial bias\. No structured framework was provided for how to arrive at a risk category\. Below, we examine the reasoning process that emerged\.

Our main finding is that although the explicit framing was racial bias, the reasoning process that emerged was considerably broader\. Rather than focusing exclusively on risks to racial minorities, participants consistently reasoned about whether a tool delivers genuine benefit for everyone — including, but not limited to, marginalised communities\. The questions groups asked were therefore not only about who might be harmed, but about whether the tool works, for whom, under what conditions, and with what safeguards\. We examine this reasoning process in detail below, and return to its broader implications in the discussion\.

#### Does it work?

The starting point for most groups was whether the tool is actually capable of doing what it claims\. In some cases, this was a matter of the technological capability or the availability of good\-quality, non\-biased data\. However, for several use cases—including recidivism and victimisation risk prediction—the objections were fundamental and concerned in whether the task itself is tractable: predicting victimisation risk, for instance, is a hard prediction problem regardless of the sophistication of the tool\. The question of whether the tool works was therefore not only technical but conceptual – does the use case rest on a sound premise? A sub\-question was“will it effectively help a human decision maker?”As participants had strong objections to autonomous AI decision\-making, groups consistently questioned whether the tool is enabling the officer to make better judgments and decisions, or effectively deciding for them how to act next\. This distinction was central to how groups assigned risk: tools perceived as adding an additional layer of information were more accepted, while those perceived at risk of displacing human judgment received higher risk ratings\. Participants deemed it crucial that accountability for outcomes and for bias can be meaningfully upheld\.

#### Will it deliver genuine benefit?

Even for uses deemed potentially acceptable, groups pressed hard on whether deployment would translate into meaningful benefit in practice\. One concern was the false sense of efficiency: groups asked whether time savings are real if AI is used responsibly, i\.e\., when appropriate verification and correction are accounted for, and whether the tool will actually save money if it opens the organisation up to legal challenges\. Even more importantly, participants questioned whether the tools can deliver benefits beyond efficiency:“Will this use of AI help alleviate the current structural problems or worsen them?”Some groups had a more pragmatic approach to considering the risk\-benefit analysis, using the flawed status quo as a reference point\. Groups considered, given the current process being problematic or already under\-resourced, whether AI is likely to help\. This was often emphasised in the context of bias: if the officers are biased, can“giving them AI reduce overall bias, or will it make it worse?”

#### Will it deliver benefit for*everyone?*

The third and most consistently applied layer of scrutiny asked whether the benefits of a tool would be distributed equitably, and specifically whether marginalised communities would share in them, or bear a disproportionate share of the risks\. Some of these discussions were rooted in well\-known failure points such as biased data, but much of the discussion focused on how loss of context and nuance around language and culture can lead to disparate impact\. Crucially, these questions were framed not only in terms of harm avoidance but also of benefit: would the tool help the force become more equitable? Would it improve outcomes for victims from minority communities? Would it build or erode trust? This inclusive framing meant that racial bias was not treated as a separate checklist item but was woven into the broader assessment of whether the tool delivers on its promises—for everyone, not just those already well\-served by existing processes\. Questions of community involvement followed naturally: groups asked who defines what*good*means, whether community input has been sought, and whether the change is likely to improve trust and relationships with those most affected by policing\.

#### Measuring success and public benefit

Alongside these questions was a practical one: can we actually measure or estimate whether a tool works, and whether it delivers benefits equitably? This requires, as a starting point, a clear and measurable definition of“what does success look like?”, and an understanding of“what evidence is needed to say that it’s effective?”\. A strong condition for acceptance was that use cases be designed to have demonstrable benefits, beyond a simplistic notion of efficiency, and that these are actively monitored, measured, and communicated transparently\.

## 6Discussion

This is a critical time for policing, which faces strong competing pressures: the desire and public expectation to deliver a better, more protective service; the urgent imperative to address deep and longstanding racial inequity; and the reality of shrinking resources\. AI is frequently positioned as the solution to this tension – a way to do more with less, and a clear win when done responsibly\(National Police Chiefs’ Council[2025](https://arxiv.org/html/2608.05418#bib.bib255); Mooreet al\.[2025](https://arxiv.org/html/2608.05418#bib.bib251)\)\. In this framing, efficiency is the primary promise, and concerns around ethics, including racial bias, are treated as blockers to unlocking enormous benefits\. The reality, however, is more complex: there are very few, if any, clear examples of AI in policing delivering genuine, measurable public benefit, yet numerous examples of algorithmic tools exacerbating already unacceptable racial disparities\(Angwinet al\.[2016](https://arxiv.org/html/2608.05418#bib.bib21); Ensignet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib12); Sankinet al\.[2021](https://arxiv.org/html/2608.05418#bib.bib168); Zilkaet al\.[2023b](https://arxiv.org/html/2608.05418#bib.bib19)\)\. This paper presents evidence that challenges this framing, showing that the consideration of racial bias is not an obstacle but a critical lens that should be applied as early as possible when considering AI adoption\.

### 6\.1The Importance of Community Consultation

Although many concerns raised by participants align with findings from responsible AI literature, the risk categorisations did not always follow expected lines\. For example, hotspot mapping—a use case with well\-documented equity concerns\(Ensignet al\.[2018](https://arxiv.org/html/2608.05418#bib.bib12); Lum and Isaac[2016](https://arxiv.org/html/2608.05418#bib.bib14)\)—was relatively well received; participants viewed the underlying goal of strategic resource deployment as legitimate and the problems as rooted in implementation rather than the idea itself\. Similarly, LFR was generally deemed acceptable, despite considerable attention to racial bias concerns\(Buolamwini and Gebru[2018](https://arxiv.org/html/2608.05418#bib.bib3); Rakova and Dobbe[2023](https://arxiv.org/html/2608.05418#bib.bib64)\)\. In contrast, recidivism risk assessment, deployed for decades in the UK and around the world, gathered the strongest objections of any use case discussed, critiquing the idea itself, not only its technical execution\. Participants questioned if such tools could deliver benefit to the public, or help with offender rehabilitation\. Indeed, these tools were introduced in order to*reduce*racial bias and redirect low\-risk offenders from prison, but evidence showed they have failed on both counts\(Kleinberget al\.[2018](https://arxiv.org/html/2608.05418#bib.bib170); Stevenson and Doleac[2024](https://arxiv.org/html/2608.05418#bib.bib214); Zilkaet al\.[2023b](https://arxiv.org/html/2608.05418#bib.bib19)\)\. This distinction—between objections to implementation vs\. to the premise—is hard to surface by any means other than community engagement and reiterates its importance\. The workshop also surfaced many cultural and contextual concerns rooted in lived experience, such as adultification and the perception of Black women as less vulnerable\.

### 6\.2Racial Equality as a Key to Better Design

Current approaches to responsible AI in policing tend to rely on structured, qualitative, evaluation frameworks – responsible AI checklists, impact assessments, and long\-form documentation that works through considerations of bias, transparency, accountability, and other principles one by one \(e\.g\.,\(National Police Chiefs’ Council and PROBabLE Futures[2025](https://arxiv.org/html/2608.05418#bib.bib259); Government Digital Service[2025](https://arxiv.org/html/2608.05418#bib.bib260)\)\)\. These frameworks serve an important function, but they implicitly treat ethical considerations as a separate line of enquiry rather than a fundamental lens to consider when deciding whether a use case is worth pursuing at all\. The deliberative benefit\-risk process we observed in this workshop was strikingly different from this checklist model\. Participants did not work through a list of considerations; instead, they reasoned in an integrated way, with many of these considerations flowing naturally towards a set of fundamental questions: does it work? Will it deliver genuine benefit and will that benefit extend to everyone? We found this process much more reflective of participatory design\(Jacksonet al\.[2023](https://arxiv.org/html/2608.05418#bib.bib59); Labedzkaet al\.[2026](https://arxiv.org/html/2608.05418#bib.bib31)\), with the strong underlying vision that a tool that can deliver benefits to marginalised communities, will also benefit the police and the public as a whole\. This vision mirrors what is known as the curb\-cut effect\(Shneiderman[2020](https://arxiv.org/html/2608.05418#bib.bib37)\): designing with marginalised users in mind consistently surfaces better questions, and produces better outcomes, for everyone\.

### 6\.3People Before Progress

The findings demonstrate that discussion of AI use cases need not be all\-or\-nothing\(Lawalet al\.[2026](https://arxiv.org/html/2608.05418#bib.bib252)\)\. Participants were clearly not against technological progress nor AI: they ranked only 3 out of 13 use cases \(6, 11, & 13\) with an average risk above medium\. In fact, participants in this mixed\-stakeholder workshop were more accepting of AI use cases than police professionals alone\(Kearneyet al\.[2024](https://arxiv.org/html/2608.05418#bib.bib116)\)\. Despite much discussion of human bias and flawed current processes, participants maintained a clear stance in favour of human judgment, consistently distinguishing between tools that inform human judgment and tools that risk displacing it\(Elish[2019](https://arxiv.org/html/2608.05418#bib.bib234); Gaoet al\.[2021](https://arxiv.org/html/2608.05418#bib.bib35)\)\. Indeed, designing AI to support rather than supplant human judgment may be the most effective way to realise its benefits while managing the associated risks\. Recent work \(Labedzka et al\. 2026\) on AI for missing persons investigations illustrates what this may look like in practice: a hybrid system combining LLM\-based summarisation with rule\-based visualisations and source linking, designed to augment officers’ sensemaking while maintaining autonomy and professional accountability\. Such systems show that AI can be built to genuinely help officers — and, through them, the communities they police — rather than creating a veneer of objectivity while quietly eroding the space for professional judgment\.

## 7Conclusion

This work demonstrates that mixed\-stakeholder deliberation, centred on racial equity, is both feasible and valuable at the earliest stages of AI adoption in policing\. Community representatives brought knowledge that technical review alone cannot replicate — surfacing concerns rooted in lived experience and adding invaluable context\. Foregrounding racial bias did not narrow or politicise the deliberation; it deepened it\. This reinforces the notion that ethical considerations and participatory approaches should be at the centre of AI adoption — particularly when better public service, not just efficiency, is the true goal\.

## References

- J\. L\. G\. Álvarez, J\. J\. L\. Ossorio, C\. Urruela, and M\. R\. Díaz \(2018\)Integral monitoring system in cases of gender violence viogén system\.Behavior & Law Journal4\(1\)\.Cited by:[§4\.4](https://arxiv.org/html/2608.05418#S4.SS4.SSS0.Px3.1.pic1.1.1.1.1.1.1)\.
- J\. Angwin, J\. Larson, S\. Mattu, and L\. Kirchner \(2016\)Machine Bias\.Note:ProPublicaExternal Links:[Link](https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.05418#S6.p1.1)\.
- E\. P\. Baumer and J\. L\. Lauritsen \(2010\)Reporting crime to the police, 1973–2005: a multivariate analysis of long\-term trends in the National Crime Survey \(NCS\) and National Crime Victimization Survey \(NCVS\)\.Criminology48\(1\),pp\. 131–185\.Cited by:[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1)\.
- BBC News \(2025\)Metropolitan police facial recognition finds only wanted people on lists\.Note:https://www\.bbc\.co\.uk/news/articles/c4gp7j55zxvoAccessed: 2025Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px3.p1.1)\.
- R\. A\. Berk, S\. B\. Sorenson, and G\. Barnes \(2016\)Forecasting Domestic Violence: A Machine Learning Approach to Help Inform Arraignment Decisions\.Journal of Empirical Legal Studies13\(1\),pp\. 94–115\.External Links:[Document](https://dx.doi.org/10.1111/jels.12098)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1)\.
- R\. A\. Berk \(2021\)Artificial intelligence, predictive policing, and risk assessment for law enforcement\.Annual Review of Criminology4,pp\. 209–237\.External Links:[Document](https://dx.doi.org/10.1146/annurev-criminol-051520-012342)Cited by:[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1)\.
- R\. Berk, J\. Bleich, A\. Kapelner, J\. Henderson, G\. Barnes, and E\. Kurtz \(2014\)Using regression kernels to forecast a failure to appear in court\.arXiv preprint arXiv:1409\.1798\.Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1)\.
- R\. Berk, L\. Sherman, G\. Barnes, E\. Kurtz, and L\. Ahlman \(2009\)Forecasting Murder within a Population of Probationers and Parolees: A High Stakes Application of Statistical Learning\.Journal of the Royal Statistical Society: Series A \(Statistics in Society\)172\(1\),pp\. 191–211\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-985X.2008.00556.x)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1)\.
- M\. Bland and J\. Verrey \(2026\)Live facial recognition: accuracy, watchlists, and deterrence\.Technical reportUniversity of Cambridge\.Note:Commissioned by Essex Police\. Version 6, 12 March 2026External Links:[Link](https://www.essex.police.uk/SysSiteAssets/media/downloads/essex/%0Aabout-us/live-facial-recognition/%0A2026-03-12-lfr-accuracy-watchlists-deterrence-cambs-uni.pdf)Cited by:[§4\.2](https://arxiv.org/html/2608.05418#S4.SS2.SSS0.Px3.1.pic1.1.1.1.1.1.3)\.
- S\. J\. Bosick, C\. M\. Rennison, A\. R\. Gover, and M\. Dodge \(2012\)Reporting violence to the police: predictors through the life course\.Journal of Criminal Justice40\(6\),pp\. 441–451\.Cited by:[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1)\.
- A\. A\. Braga, B\. S\. Turchan, A\. V\. Papachristos, and D\. M\. Hureau \(2019\)Hot spots policing of small geographic areas effects on crime\.Campbell Systematic Reviews15\(3\),pp\. e1046\.External Links:[Document](https://dx.doi.org/10.1002/cl2.1046)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1)\.
- S\. Brayne \(2017\)Big data surveillance: the case of policing\.American sociological review82\(5\),pp\. 977–1008\.Cited by:[§2\.3](https://arxiv.org/html/2608.05418#S2.SS3.p1.1)\.
- J\. Buolamwini and T\. Gebru \(2018\)Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification\.InProceedings of the 1st Conference on Fairness, Accountability and Transparency,S\. A\. Friedler and C\. Wilson \(Eds\.\),Proceedings of Machine Learning Research, Vol\.81,pp\. 77–91\.External Links:[Link](https://proceedings.mlr.press/v81/buolamwini18a.html)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05418#S4.SS2.SSS0.Px3.1.pic1.1.1.1.1.1.1),[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1)\.
- B\. Butcher, M\. Zilka, and A\. Weller \(2022\)Racial disparities in arrests for drug violations in the us: what can we learn from publicly available data?\.InACM conference on Equity and Access in Algorithms, Mechanisms, and Optimization \(EAAMO\),Cited by:[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1)\.
- College of Policing and National Police Chiefs’ Council \(2022\)Police Race Action Plan: Improving Policing for Black People\.Technical reportCollege of Policing\.External Links:[Link](https://assets.college.police.uk/s3fs-public/Police-Race-Action-Plan.pdf)Cited by:[§2\.5](https://arxiv.org/html/2608.05418#S2.SS5.p1.1)\.
- M\. Davis, S\. Popa, and C\. Surlea \(2010\)Real\-time face recognition from surveillance video\.InIntelligent Video Event Analysis and Understanding,Studies in Computational Intelligence, Vol\.332,pp\. 155–194\.Note:Author’s peer\-reviewed version; original published at www\.springerlink\.comExternal Links:[Document](https://dx.doi.org/10.1007/978-3-642-17554-1%5F8),[Link](https://pureadmin.qub.ac.uk/ws/files/6903843/%0AFaceRecognition_colour.pdf)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px3.p1.1)\.
- M\. C\. Elish \(2019\)Moral crumple zones: cautionary tales in human\-robot interaction \(pre\-print\)\.Engaging Science, Technology, and Society \(pre\-print\)\.Cited by:[§6\.3](https://arxiv.org/html/2608.05418#S6.SS3.p1.1)\.
- D\. Ensign, S\. A\. Friedler, S\. Neville, C\. Scheidegger, and S\. Venkatasubramanian \(2018\)Runaway Feedback Loops in Predictive Policing\.InProceedings of the 1st Conference on Fairness, Accountability and Transparency,Proceedings of Machine Learning Research, Vol\.81,pp\. 160–171\.External Links:[Link](https://proceedings.mlr.press/v81/ensign18a.html)Cited by:[§1](https://arxiv.org/html/2608.05418#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05418#S4.SS2.SSS0.Px1.p1.1),[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1),[§6](https://arxiv.org/html/2608.05418#S6.p1.1)\.
- European Parliament and Council of the European Union \(2024\)Cited by:[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1)\.
- R\. Gao, M\. Saar\-Tsechansky, M\. De\-Arteaga, L\. Han, M\. K\. Lee, and M\. Lease \(2021\)Human\-AI collaboration with bandit feedback\.arXiv preprint arXiv:2105\.10614\.Cited by:[§6\.3](https://arxiv.org/html/2608.05418#S6.SS3.p1.1)\.
- Government Digital Service \(2024\)Algorithmic transparency recording standard \(ATRS\): mandatory scope and exemptions policy\.Note:https://www\.gov\.uk/government/publications/algorithmic\-transparency\-recording\-standard\-mandatory\-scope\-and\-exemptions\-policy/algorithmic\-transparency\-recording\-standard\-atrs\-mandatory\-scope\-and\-exemptions\-policyPublished 17 December 2024Cited by:[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1)\.
- Government Digital Service \(2025\)Data and AI ethics framework\.Technical reportGovernment Digital Service, Department for Science, Innovation and Technology\.Note:First published 2018; last updated 18 December 2025External Links:[Link](https://www.gov.uk/government/publications/data-ethics-framework)Cited by:[§6\.2](https://arxiv.org/html/2608.05418#S6.SS2.p1.1)\.
- Government of the United Kingdom \(2025\)AI to help police catch criminals before they strike\.Note:https://www\.gov\.uk/government/news/ai\-to\-help\-police\-catch\-criminals\-before\-they\-strikeAccessed: 2025Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1)\.
- T\. Greene, G\. Shmueli, J\. Fell, C\. Lin, and H\. Liu \(2022\)Forks over knives: predictive inconsistency in criminal justice algorithmic risk assessment tools\.Journal of the Royal Statistical Society Series A: Statistics in Society185\(Supplement\_2\),pp\. S692–S723\.Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1)\.
- M\. R\. Haque, D\. Saxena, K\. Weathington, J\. Chudzik, and S\. Guha \(2024\)Are we asking the right questions?: designing for community stakeholders’ interactions with ai in policing\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–20\.Cited by:[§2\.3](https://arxiv.org/html/2608.05418#S2.SS3.p1.1)\.
- B\. E\. Harcourt \(2010\)Risk as a Proxy for Race\.Criminology and Public Policy\.Note:Also available as University of Chicago Law & Economics Olin Working Paper No\. 535External Links:[Link](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1677654)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1)\.
- Home Office \(2024\)Summary: accuracy and equitability evaluation of IDEMIA facial recognition algorithm for Home Office strategic facial matching\.Note:https://www\.gov\.uk/government/publications/facial\-recognition\-technology\-tests\-national\-physical\-laboratory/summary\-accuracy\-and\-equitability\-evaluation\-of\-idemia\-facial\-recognition\-algorithm\-for\-home\-office\-strategic\-facial\-matchingAccessed: 2024Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px3.p1.1)\.
- P\. Hunt, J\. Saunders, and J\. S\. Hollywood \(2014\)Evaluation of the Shreveport predictive policing experiment\.Technical reportTechnical ReportRR\-531\-NIJ,RAND Corporation,Santa Monica, CA\.Note:Sponsored by the National Institute of JusticeExternal Links:[Link](https://www.rand.org/pubs/research_reports/RR531.html)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1)\.
- C\. B\. Jackson, K\. Crowston, and C\. Østerlund \(2023\)A participatory modeling framework for algorithmic accountability and transparency\.Government Information Quarterly40\(2\),pp\. 101773\.Cited by:[§6\.2](https://arxiv.org/html/2608.05418#S6.SS2.p1.1)\.
- C\. Kearney, J\. Hron, H\. Kosc, and M\. Zilka \(2024\)Beyond use\-cases: a participatory approach to envisioning data science in law enforcement\.InACM Conference on Fairness, Accountability and Transparency \(FAccT\),External Links:[Link](https://doi.org/10.1145/3630106.3659007)Cited by:[§2\.3](https://arxiv.org/html/2608.05418#S2.SS3.p1.1),[§6\.3](https://arxiv.org/html/2608.05418#S6.SS3.p1.1)\.
- J\. Kleinberg, H\. Lakkaraju, J\. Leskovec, J\. Ludwig, and S\. Mullainathan \(2018\)Human decisions and machine predictions\.The quarterly journal of economics133\(1\),pp\. 237–293\.Cited by:[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1)\.
- P\. Z\. Labedzka, D\. Peters, J\. J\. Dudley, and M\. Zilka \(2026\)Human\-AI interaction for time\-critical sensemaking in missing persons investigations\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,External Links:[Document](https://dx.doi.org/10.1145/3772318.3793148)Cited by:[§6\.2](https://arxiv.org/html/2608.05418#S6.SS2.p1.1)\.
- T\. Lawal, A\. Paul, M\. Oswald, and A\. Taylor \(2026\)Remote ai weapons detection: reversing ‘stop and search’ to ‘search and stop’?\.Legal Studies\(English\)\.Note:Engineering and Physical Sciences Research Council through funding from Responsible AI UK\.External Links:[Document](https://dx.doi.org/10.1017/lst.2026.10127),ISSN 0261\-3875Cited by:[§6\.3](https://arxiv.org/html/2608.05418#S6.SS3.p1.1)\.
- Y\. Lee, B\. Bradford, and K\. Posch \(2024\)The effectiveness of big data\-driven predictive policing: systematic review\.Justice Evaluation Journal7\(2\),pp\. 127–160\.Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1),[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1)\.
- K\. Lum and W\. Isaac \(2016\)To Predict and Serve?\.Significance13\(5\),pp\. 14–19\.External Links:[Document](https://dx.doi.org/10.1111/j.1740-9713.2016.00960.x)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1),[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1)\.
- T\. Mansfield \(2023\)Facial recognition technology equitability study\.Technical reportNational Physical Laboratory\.External Links:[Link](https://science.police.uk/site/assets/files/3396/%0Afrt-equitability-study_mar2023.pdf)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px3.p1.1)\.
- G\. O\. Mohler, M\. B\. Short, S\. Malinowski, M\. Johnson, G\. E\. Tita, A\. L\. Bertozzi, and P\. J\. Brantingham \(2015\)Randomized Controlled Field Trials of Predictive Policing\.Journal of the American Statistical Association110\(512\),pp\. 1399–1411\.External Links:[Document](https://dx.doi.org/10.1080/01621459.2015.1077710)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1)\.
- C\. Moore, C\. Gill, N\. Bliss, K\. Butler, S\. Forrest, D\. Lopresti, M\. L\. Maher, H\. Mentis, S\. Shekhar, A\. Stent,et al\.\(2025\)Concerning the responsible use of ai in the us criminal justice system\.Communications of the ACM68\(8\),pp\. 41–44\.Cited by:[§6](https://arxiv.org/html/2608.05418#S6.p1.1)\.
- R\. Moore \(2015\)A compendium of research and analysis on the Offender Assessment System \(OASys\) 2009–2013\.Technical reportMinistry of Justice,London\.External Links:[Link](https://assets.publishing.service.gov.uk/media/5a7f4c9540f0b62305b8552b/compendium-of-research-and-analysis-on-oasys-2009-2013.pdf)Cited by:[§3\.4](https://arxiv.org/html/2608.05418#S3.SS4.1.pic1.1.1.1.1.1.1)\.
- R\. Muir and F\. O’Connell \(2025\)POLICING and artificial intelligence\.Technical reportThe Police Foundation\.Cited by:[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1)\.
- National Police Chiefs’ Council and PROBabLE Futures \(2025\)Responsible AI checklist for policing\.Technical reportNational Police Chiefs’ Council\.External Links:[Link](https://library.college.police.uk/docs/NPCC/Responsible-AI-checklist-for-policing-2025.pdf)Cited by:[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1),[§6\.2](https://arxiv.org/html/2608.05418#S6.SS2.p1.1)\.
- National Police Chiefs’ Council \(2023\)Covenant for using artificial intelligence\(ai\) in policing\.External Links:[Link](https://science.police.uk/site/assets/files/4682/ai_principles_1_1_1.pdf)Cited by:[§1](https://arxiv.org/html/2608.05418#S1.p1.1)\.
- National Police Chiefs’ Council \(2025\)Artificial intelligence \(ai\) playbook for policing\.Technical reportNational Police Chiefs’ Council\.External Links:[Link](https://www.npcc.police.uk/SysSiteAssets/media/downloads/publications/publications-log/science-and-innovation/2025/npcc-ai-strategy.pdf)Cited by:[§1](https://arxiv.org/html/2608.05418#S1.p1.1),[§1](https://arxiv.org/html/2608.05418#S1.p2.1),[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1),[§6](https://arxiv.org/html/2608.05418#S6.p1.1)\.
- Oleeo \(2025\)Police recruitment: Reduce the time it takes to hire qualified police officers\.Note:https://www\.oleeo\.com/industries/police/Oleeo claims its platform is used by 70% of UK police forces, including all forces in Wales and Scotland\. Accessed: May 2026Cited by:[§4\.1](https://arxiv.org/html/2608.05418#S4.SS1.SSS0.Px1.p1.1)\.
- M\. Oswald, J\. Grace, S\. Urwin, and G\. C\. Barnes \(2018\)Algorithmic Risk Assessment Policing Models: Lessons from the Durham HART Model and ‘Experimental’ Proportionality\.Information & Communications Technology Law27\(2\),pp\. 223–250\.External Links:[Document](https://dx.doi.org/10.1080/13600834.2018.1458455)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1)\.
- B\. Rakova and R\. Dobbe \(2023\)A sociotechnical audit: assessing police use of facial recognition\.Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency,pp\. 1334–1346\.Cited by:[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1),[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1)\.
- R\. Richardson, J\. M\. Schultz, and K\. Crawford \(2019\)Dirty data, bad predictions: How civil rights violations impact police data, predictive policing systems, and justice\.New York University Law Review Online94,pp\. 15\.Cited by:[§2\.3](https://arxiv.org/html/2608.05418#S2.SS3.p1.1)\.
- A\. Sankin, D\. Mehrota, S\. Mattu, and A\. Gilbertson \(2021\)Crime prediction software promised to be free of biases\. new data shows it perpetuates them\.The Markup2\.Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1),[§6](https://arxiv.org/html/2608.05418#S6.p1.1)\.
- A\. Satariano and R\. Toll Pifarré \(2024\)An algorithm told police she was safe\. then her husband killed her\.\.The New York Times\.Note:https://www\.nytimes\.com/interactive/2024/07/18/technology/spain\-domestic\-violence\-viogen\-algorithm\.htmlAccessed: October 2024Cited by:[§4\.4](https://arxiv.org/html/2608.05418#S4.SS4.SSS0.Px3.1.pic1.1.1.1.1.1.1)\.
- C\. R\. Shaw and H\. D\. McKay \(1942\)Juvenile delinquency and urban areas\.University of Chicago Press,Chicago\.Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1)\.
- B\. Shneiderman \(2020\)Bridging the gap between ethics and practice: guidelines for reliable, safe, and trustworthy human\-centered AI systems\.ACM Transactions on Interactive Intelligent Systems \(TiiS\)10\(4\),pp\. 1–31\.Cited by:[§6\.2](https://arxiv.org/html/2608.05418#S6.SS2.p1.1)\.
- J\. Simons, S\. Adams Bhatti, and A\. Weller \(2021\)Machine learning and the meaning of equal treatment\.InProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society,pp\. 956–966\.Cited by:[§4\.1](https://arxiv.org/html/2608.05418#S4.SS1.SSS0.Px1.p2.1)\.
- M\. T\. Stevenson and J\. L\. Doleac \(2024\)Algorithmic risk assessment in the hands of humans\.American Economic Journal: Economic Policy16\(4\),pp\. 382–414\.Cited by:[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1)\.
- E\. Taka, T\. Lawal, M\. Calder, M\. Sevegnani, K\. Kotsoglou, E\. McClory\-Tiarks, and M\. Oswald \(2025\)Mapping the probabilistic ai ecosystem in criminal justice in england and wales\.arXiv preprint arXiv:2512\.04116\.Cited by:[§1](https://arxiv.org/html/2608.05418#S1.p2.1)\.
- A\. Taylor, A\. Boyle, A\. Sutherland, and C\. Giacomantonio \(2016\)Using Ambulance Data to Reduce Community Violence: Critical Literature Review\.European Journal of Emergency Medicine23\(4\),pp\. 248–252\.External Links:[Document](https://dx.doi.org/10.1097/MEJ.0000000000000351)Cited by:[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px1.p1.1)\.
- M\. Zilka, C\. Ashurst, L\. Chambers, E\. P\. Goodmann, P\. Ugwudike, and M\. Oswald \(2023a\)Exploring police perspectives on algorithmic transparency: a qualitative analysis of police interviews in the uk\.InProceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization,pp\. 1–19\.Cited by:[§2\.3](https://arxiv.org/html/2608.05418#S2.SS3.p1.1)\.
- M\. Zilka, R\. Fogliato, J\. Hron, B\. Butcher, C\. Ashurst, and A\. Weller \(2023b\)The Progression of Disparities within the Criminal Justice System: Differential Enforcement and Risk Assessment Instruments\.InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency,New York, NY, USA,pp\. 1553–1569\.External Links:[Document](https://dx.doi.org/10.1145/3593013.3594099)Cited by:[§1](https://arxiv.org/html/2608.05418#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.05418#S2.SS1.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2608.05418#S2.SS2.p1.1),[§3\.4](https://arxiv.org/html/2608.05418#S3.SS4.1.pic1.1.1.1.1.1.1),[§6\.1](https://arxiv.org/html/2608.05418#S6.SS1.p1.1),[§6](https://arxiv.org/html/2608.05418#S6.p1.1)\.
- M\. Zilka, H\. Sargeant, and A\. Weller \(2022\)Transparency, governance and regulation of algorithmic tools deployed in the criminal justice system: a uk case study\.InACM Conference on AI, Ethics and Society \(AIES\),External Links:[Link](https://doi.org/10.1145/3514094.3534200)Cited by:[§2\.4](https://arxiv.org/html/2608.05418#S2.SS4.p1.1)\.
- M\. Ziosi and D\. Pruss \(2024\)Evidence of what, for whom? The socially contested role of algorithmic bias in a predictive policing tool\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency,pp\.\.External Links:[Document](https://dx.doi.org/10.1145/3630106.3658991)Cited by:[§2\.3](https://arxiv.org/html/2608.05418#S2.SS3.p1.1)\.

Similar Articles

Why responsible AI development needs cooperation on safety

OpenAI Blog

OpenAI publishes a policy research paper identifying four strategies to improve industry cooperation on AI safety norms: communicating risks/benefits, technical collaboration, increased transparency, and incentivizing standards. The analysis addresses how competitive pressures could lead to under-investment in safety and proposes mechanisms to align incentives toward safe AI development.

AI Epistemic Risks: Emerging Mechanisms & Evidence [R]

Reddit r/MachineLearning

A new paper co-authored by 30 experts examines epistemic risks from AI—threats to our ability to form accurate beliefs and reason well—including mechanisms like persuasion, cognitive offloading, and feedback loops, and outlines directions to mitigate these risks.

AI safety via debate

OpenAI Blog

OpenAI proposes a novel approach to AI safety where two AI agents debate each other while a human judge evaluates their arguments, allowing humans to supervise AI systems whose behavior is too complex to directly understand. The method leverages debate and adversarial reasoning to align advanced AI with human values and preferences.

Computer cops

The Verge

An investigation into the growing business of selling AI technologies to police departments, including facial recognition, chatbots, and automated report-writing tools, and the potential consequences for civil liberties and policing practices.