Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)
Summary
This paper presents a validated adaptation protocol for label-free drone-based crowd counting to ensure safety at mass events like the 2034 FIFA World Cup in Saudi Arabia, addressing domain shift and deployment challenges.
View Cached Full Text
Cached at: 08/19/26, 10:08 AM
# Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)
Source: [https://arxiv.org/html/2608.17625](https://arxiv.org/html/2608.17625)
AlJawharh AlOtaibi\*Jude AlSubaie\*Affiliation:alanoud@daldata\.ai, aljawharh@daldata\.ai, jude@daldata\.aiAffiliation:Riyadh, Saudi Arabia
###### Abstract
Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale\. Drone\-based counting for such venues must hold accuracy on footage unlike anything in its training corpus, without labels, and must warn of dangerous inflow before a crush forms\. We deliver a validated answer built on525525controlled runs, a full\-resolution corpus study, five falsification ablations, and a five\-condition evaluation of a safety interlock, and we resolve three questions that a deployment decision depends on\.We validate the adaptation stage\.Label\-free adaptation is decisive and holds up as conditions worsen: it recovers3131–49%49\\%of shift\-induced error across four corruptions and five severities, with the strongest single method gaining41\.841\.8maeover the frozen source \(95% CI\[34\.1,49\.6\]\[34\.1,49\.6\],p=7\.5×10−10p=7\.5\{\\times\}10^\{\-10\},d=2\.52d=2\.52\)\. We establish a*severity law*separating methods whose absolute protective margin is constant from the one whose margin grows, and a stability budget that identifies which configuration is safe to fly\. On a full\-resolution corpus carrying a genuine\+48\+48maeaerial gap \(reached after retraining the source model to14\.614\.6validationmae, a34%34\\%improvement\), adaptation repairs the dense\-scene undercounting that would otherwise cause a monitor to under\-report a forming crush, and the flux\-based risk module fires on real congestion episodes in22of66full\-length target clips\.We localise where the recoverable error lives\.Building the regime a physics\-informed conservation prior asks for \(300300\-frame clips at200200ms spacing, five times wider than standard, so genuine motion exists between frames\), we determine that the adaptation signal in this task is normalisation\-driven rather than flow\-driven: the continuity residual is provably invariant to the proportional counting errors that domain shift actually produces, a result confirmed by four on/off ablations correlated atr=0\.999r=0\.999and by a40%40\\%input corruption that moves accuracy by0\.050\.05mae\. This tells practitioners where to spend adaptation capacity and where not to\.We derive the optimal deployment policy\.Evaluating a label\-free shift gate as a decision policy, we show that shift magnitude and accuracy damage are rank\-independent \(Spearmanρ=0\.20\\rho=0\.20;ρ=−0\.60\\rho=\-0\.60among genuine shifts\), quantify the58%58\\%of available headroom a magnitude\-based gate forgoes, and establish unconditional adaptation with tail monitoring as the evidence\-backed policy\. We close with a six\-point protocol and the acceptance criteria for the next build\.
## 1Introduction
Crowd disasters are failures of monitoring before they are failures of crowd control\. In nearly every modern stadium and pilgrimage tragedy, the dangerous build\-up of density was under way for minutes before anyone acted on it\. The 2034 FIFA World Cup in Saudi Arabia, and the Hajj gatherings the country manages each year, will place enormous crowds under exactly the conditions in which such build\-ups form\. A system that could watch these crowds from the air and raise a warning while there is still time to intervene would address a problem that existing, manual monitoring handles poorly\.
Drone\-mounted cameras are the natural sensor, and crowd counting from aerial video is a mature enough technique to estimate density in principle\. In practice it breaks at the first contact with a real event\. A counting model trained on one corpus loses accuracy the moment the footage differs in altitude, illumination, optics, or transmission quality, and event footage always differs\. Worse, no ground\-truth counts exist during a live event to correct the model\. It must adapt to the incoming stream using no labels at all\. Label\-free test\-time adaptation \(TTA\), which updates the model from a self\-supervised objective on the test stream itself, is the only practical response\[[2](https://arxiv.org/html/2608.17625#bib.bib2),[1](https://arxiv.org/html/2608.17625#bib.bib1)\]\.
This setting raises a specific and appealing idea\. Between two consecutive frames the number of people in a region can change only through movement across its boundary: people are conserved\. If a counting network’s density predictions are inconsistent with the motion measured by optical flow, that inconsistency is an error the network can correct, without labels\. The same quantity, the flux of people across a line, is also a natural early\-warning signal for congestion\. A single physical law might therefore supply both the adaptation signal and the safety signal the deployment needs\. This paper asks whether it does\.
We build the pipeline our idea implies: a CSRNet density regressor\[[7](https://arxiv.org/html/2608.17625#bib.bib7)\]adapted at test time under a population\-conservation loss computed from RAFT optical flow\[[8](https://arxiv.org/html/2608.17625#bib.bib8)\], and evaluate it against the requirements a safety deployment actually imposes\. Those requirements are stricter than a single benchmark average\. An integrator must know which parts of the pipeline carry the accuracy, how that benefit changes as conditions worsen toward the tail where danger lives, how the system behaves on its worst runs rather than its average ones, and whether it can decide unaided when adaptation is warranted\. We answer each of these on DroneCrowd\[[9](https://arxiv.org/html/2608.17625#bib.bib9)\]with controlled corruptions and a full\-resolution transfer study, using paired statistics, effect sizes, and Holm correction, and two ablations built to expose a component that contributes nothing\. Because the conservation prior is expected to be weakest when frames are close together, we also grant it the regime it favours: a full\-resolution retrain on the complete corpus \(validationmae22\.3→14\.622\.3\\rightarrow 14\.6\) with frames sampled five times further apart than the default\.
Our findings are as follows\.
1. 1\.Adaptation is effective and its benefit is predictable\. It recovers4040–46%46\\%of shift\-induced error at the reference severity and3030–49%49\\%across a five\-level severity sweep \(Section[5](https://arxiv.org/html/2608.17625#S5)\)\.
2. 2\.The benefit does not degrade as corruption worsens; we report a per\-method severity law and identify a stability cost in the combined method, whose worst runs occur at low severity \(Section[5](https://arxiv.org/html/2608.17625#S5)\)\.
3. 3\.On full\-resolution transfer, adaptation removes the dense\-scene undercounting that dominates source error, and the flux indicator fires on real congestion episodes \(Sections[6](https://arxiv.org/html/2608.17625#S6),[9](https://arxiv.org/html/2608.17625#S9)\)\.
4. 4\.The conservation prior does not improve on entropy minimisation in any condition we tested, including the wide\-spacing regime built to favour it\. We explain this with an invariance argument and localise the recoverable error to normalisation statistics \(Section[7](https://arxiv.org/html/2608.17625#S7)\)\.
5. 5\.A label\-free shift score is a poor basis for gating adaptation, because its magnitude does not track the accuracy damage a shift causes; we therefore recommend unconditional adaptation with monitoring of the worst\-run tail \(Sections[8](https://arxiv.org/html/2608.17625#S8),[10](https://arxiv.org/html/2608.17625#S10)\)\.
## 2Related Work
#### Test\-time adaptation\.
Adapting a model to the test stream without labels has converged on the normalisation layers as the point of intervention\. AdaBN\[[1](https://arxiv.org/html/2608.17625#bib.bib1)\]recomputes batch\-normalisation statistics on the target data and needs no gradient step; TENT\[[2](https://arxiv.org/html/2608.17625#bib.bib2)\]adds a single objective, minimising prediction entropy through the BN affine parameters while every convolutional weight stays frozen\. The robustness\-oriented successors, CoTTA\[[3](https://arxiv.org/html/2608.17625#bib.bib3)\]against error accumulation, EATA\[[4](https://arxiv.org/html/2608.17625#bib.bib4)\]through sample selection and anti\-forgetting, and SAR\[[5](https://arxiv.org/html/2608.17625#bib.bib5)\]through sharpness\-aware updates, as well as the gradient\-free LAME\[[6](https://arxiv.org/html/2608.17625#bib.bib6)\], all inherit this frozen\-backbone, normalisation\-centred design\. That shared design is what makes the family the right setting for our question: with capacity confined to the same small parameter space, any advantage a physics prior offers must show up there or nowhere\. We benchmark against AdaBN and TENT and position the robust variants as the next comparison \(Section[10](https://arxiv.org/html/2608.17625#S10)\)\.
#### Crowd counting\.
Density\-map regression with dilated convolutions, as in CSRNet\[[7](https://arxiv.org/html/2608.17625#bib.bib7)\], remains the standard treatment of congested scenes, and we adopt it unchanged so that our findings concern the adaptation objective rather than a new architecture\. DroneCrowd\[[9](https://arxiv.org/html/2608.17625#bib.bib9)\]is the corpus throughout this study; its scale, altitude range, and dense aerial viewpoints are representative of the mass\-gathering setting we target\. VisDrone\[[10](https://arxiv.org/html/2608.17625#bib.bib10)\]defines the adjacent aerial benchmark, and we are explicit that we report no results on it: we name DroneCrowd→\\rightarrowVisDrone the external\-validity milestone this protocol is built to be carried into \(Section[10](https://arxiv.org/html/2608.17625#S10)\)\.
#### Physics\-informed priors\.
Physics\-informed learning\[[11](https://arxiv.org/html/2608.17625#bib.bib11)\]supervises a network with a law its outputs must satisfy, and succeeds where that law genuinely constrains the solution\. Population conservation is the natural instance for counting: with a displacement field from RAFT\[[8](https://arxiv.org/html/2608.17625#bib.bib8)\], the change in count within a region must equal the flux across its boundary\. Our contribution to this programme is a sharp negative characterisation: the precise conditions under which the conservation residual carries gradient for counting, and the invariance that empties it under the shifts that actually occur \(Section[7](https://arxiv.org/html/2608.17625#S7)\)\. Because the argument is stated at the level of the residual rather than the architecture, it transfers to any density\-regression task tempted by the same prior\.
#### Shift detection\.
Label\-free detection of distribution shift\[[12](https://arxiv.org/html/2608.17625#bib.bib12)\]is well developed, but a safety interlock imposes a stronger requirement than the literature usually asks of it: the score must be monotone not in*whether*a shift occurred but in*how much accuracy it costs*\. We show these are different quantities in this task, and that a magnitude score, however well it detects shift, is the wrong basis for gating adaptation \(Section[8](https://arxiv.org/html/2608.17625#S8)\)\.
## 3Method
### 3\.1Backbone and adaptation family
We build on a CSRNet density regressor\[[7](https://arxiv.org/html/2608.17625#bib.bib7)\], mapping each frame to a density mapDt\(x\)D\_\{t\}\(x\)whose integral over a regionΩ\\Omegais the predicted countCt\(Ω\)=∫ΩDt𝑑xC\_\{t\}\(\\Omega\)=\\int\_\{\\Omega\}D\_\{t\}\\,dx\. Following the fully test\-time protocol of TENT\[[2](https://arxiv.org/html/2608.17625#bib.bib2)\], only batch\-normalisation parameters are updated on the test stream; all convolutional weights stay frozen\. Holding everything fixed except the objective ensures each comparison isolates the loss rather than a difference in model capacity\. We compare five configurations:Source\(frozen, no adaptation\),AdaBN\(test\-stream BN statistics\),TENT\(entropy minimisation\),Ours\(conservation residual alone\), andTENT\+Ours\(both objectives\)\.
### 3\.2Population\-conservation prior
People are neither created nor destroyed between consecutive frames, so the count inside a region can change only through motion across its boundary\. Absent sources or sinks inΩ\\Omega,
∂∂t∫ΩDt𝑑x\+∮∂ΩDt𝐯t⋅𝐧𝑑ℓ=0,\\frac\{\\partial\}\{\\partial t\}\\int\_\{\\Omega\}D\_\{t\}\\,dx\\;\+\\;\\oint\_\{\\partial\\Omega\}D\_\{t\}\\,\\mathbf\{v\}\_\{t\}\\cdot\\mathbf\{n\}\\,d\\ell\\;=\\;0,\(1\)where𝐯t\\mathbf\{v\}\_\{t\}is the pixel\-wise displacement field between framesttandt\+1t\{\+\}1, estimated with a frozen pretrained RAFT network\[[8](https://arxiv.org/html/2608.17625#bib.bib8)\]\. Written in divergence form and discretised on the pixel grid, this yields the per\-pixel continuity residual
rt=Dt\+1−Dt\+∇⋅\(Dt𝐯t\),r\_\{t\}\\;=\\;D\_\{t\+1\}\-D\_\{t\}\+\\nabla\\\!\\cdot\\\!\\left\(D\_\{t\}\\mathbf\{v\}\_\{t\}\\right\),\(2\)whose squared magnitudeℒphys=‖rt‖22\\mathcal\{L\}\_\{\\mathrm\{phys\}\}=\\\|r\_\{t\}\\\|\_\{2\}^\{2\}we minimise either alone or added to the entropy objective,ℒ=ℒent\+λℒphys\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{ent\}\}\+\\lambda\\mathcal\{L\}\_\{\\mathrm\{phys\}\}\. Where predictions obey the law the two terms cancel; any imbalance is a candidate label\-free error signal, and Section[7](https://arxiv.org/html/2608.17625#S7)determines precisely which errors it can and cannot see\.
### 3\.3Flux\-based risk indicator
The boundary integral in Eq\. \([1](https://arxiv.org/html/2608.17625#S3.E1)\) yields as a by\-product an inward\-flux signalΦt\(Ω\)=−∮∂ΩDt𝐯t⋅𝐧dℓ\\Phi\_\{t\}\(\\Omega\)=\-\\oint\_\{\\partial\\Omega\}D\_\{t\}\\mathbf\{v\}\_\{t\}\\cdot\\mathbf\{n\}\\,d\\ell: a region taking in people faster than they leave registers sustained positiveΦt\\Phi\_\{t\}before it becomes dangerously dense\. We useΦt\\Phi\_\{t\}as a*relative*congestion\-onset indicator, which is the form in which it is operationally useful today\. Expressing it as an absolute crush threshold requires density in people/m2/\\mathrm\{m\}^\{2\}and hence a meters\-per\-pixel scale to the ground plane; Section[9](https://arxiv.org/html/2608.17625#S9)specifies that calibration as the acceptance criterion for the absolute mode\.
### 3\.4Shift\-gated safeguard
Adaptation modifies a model at inference time, so a mature system should be able to decide without labels whether to intervene\. We instrument a gate that compares the batch\-normalisation statistics of the incoming stream against those cached from clean data, producing a scalar shift scoress, and adapts only whens\>τs\>\\tau, withτ=2sclean\\tau=2s\_\{\\mathrm\{clean\}\}\. Section[8](https://arxiv.org/html/2608.17625#S8)evaluates it as a decision policy, comparing what it delivered against what each alternative policy would have delivered, which is the form a deployment decision requires\.
## 4Experimental Protocol
Table 1:The three experimental tracks\. All adaptation is label\-free and updates only BN parameters\. Base models are stated explicitly, because Track B uses a stronger retrained source and the two error scales are reported separately throughout\.#### Track A: controlled corruption benchmark\.
We evaluate CSRNet on a drone crowd\-counting stream ofn=750n=750frames\. To isolate robustness from scene variability, we hold scene content fixed and apply four synthetic corruptions that emulate documented failure modes of aerial capture: additive Gaussian noise \(sensor noise\), motion blur \(platform and subject motion\), low light \(dusk and night operation\), and JPEG compression \(bandwidth\-limited transmission\), each measured against a clean reference\. The five\-method benchmark runs at the reference severity for55conditions×\\times55methods×\\times55seeded replicates=125=125runs\. The severity sweep extends this over five severity levels for four methods:4×5×4×5=4004\\times 5\\times 4\\times 5=400runs\. Total:𝟓𝟐𝟓\\mathbf\{525\}runs\.
#### Track B: full\-resolution corpus with real inter\-frame motion\.
Track B is the engineering centrepiece of the study and was purpose\-built to test the conservation prior in its strongest regime while simultaneously providing the realistic transfer setting the deployment case needs\. We ingested the full1111GB release, converted trajectory annotations from the native\.matformat, retrained the source model at full resolution \(validationmae14\.614\.6, improved from22\.322\.3\), and rebuilt pair sampling to draw300300\-frame clips at∼200\{\\sim\}200ms spacing, five times wider than Track A, so genuine displacement exists between paired frames for Eq\. \([2](https://arxiv.org/html/2608.17625#S3.E2)\) to constrain\. Track B additionally carries a real aerial domain gap \(\+48\+48maesource degradation\) rather than a synthetic corruption, and its full\-length clips are what make the risk module measurable \(Section[9](https://arxiv.org/html/2608.17625#S9)\)\.
#### Track C: policy evaluation of the safeguard\.
The gate is evaluated on the five Track\-A conditions against both the frozen source and the adapted model, and scored as one of four candidate policies rather than as a binary classifier\.
#### Metrics and analysis\.
We report mean absolute error \(mae\) and root\-mean\-square error \(RMSE\) of the predicted count\. Because replicates share seeds across methods, comparisons are*paired*: we use pairedtt\-tests with 95% confidence intervals and the paired effect size Cohen’sdzd\_\{z\}, corroborated by Wilcoxon signed\-rank tests, and we control families of per\-shift tests with the Holm–Bonferroni procedure\. The Source model is deterministic across seeds, so its comparisons are one\-sample tests of each adaptive method’s replicates against the Source constant\. Stability is reported as across\-replicate coefficient of variation \(CV\) and worst\-replicate error, because a safety application is governed by its tail\.
#### Reporting discipline on base models\.
Track A results use a source model early\-stopped at validationmae26\.126\.1; Track B uses the full\-resolution retrain at14\.614\.6\. Every comparison in this paper is within\-track on a single fixed base model, and no quantity is pooled across tracks\. The two tracks are designed to converge on conclusions, not on absolute error levels, and they do\.
## 5Validated: Adaptation Efficacy and the Severity Law
Figure 1:Adaptation across corruption severity\.\(a\)Typical accuracy \(median countmae, lower is better\): the frozen Source degrades steeply while every adaptive method holds far below it and the protective gap widens\.\(b\)Stability \(worst replicate, min–max band\): TENT\+Ours carries the extreme tail, peaking near113113maeat motion\-blur severity 1 \(best typical accuracy, widest spread\)\. Panel \(b\) is the basis for the stability budget in Section[10](https://arxiv.org/html/2608.17625#S10)\.#### Result 1: adaptation recovers most of the cost of domain shift\.
Averaged over the four corruptions at the reference severity, the unadapted Source model reaches96\.196\.1mae, while every adaptive method lands in the5252–5858range \(Table[5](https://arxiv.org/html/2608.17625#S12.T5)\), a4040–45%45\\%reduction of shift\-induced error\. Aggregating the strongest single method \(AdaBN\) over its2020shifted replicates, the gain over Source is41\.841\.8mae\(95% CI\[34\.1,49\.6\]\[34\.1,49\.6\],p=7\.5×10−10p=7\.5\\times 10^\{\-10\}, Cohen’sd=2\.52d=2\.52\), several times the conventional threshold for a large effect\. The gain is concentrated in the simplest available mechanism, realigning batch\-norm statistics to the incoming stream, which is an operationally welcome finding: the component doing the work is parameter\-free, cheap, and stable\. RMSE reproduces the same ordering\.
#### Result 2: the severity law\.
Table[6](https://arxiv.org/html/2608.17625#S12.T6)reports the sweep\. Source error climbs monotonically from73\.673\.6to112\.0112\.0maeacross severities11–55and no adaptive method follows it: at severity55, TENT holds76\.776\.7and TENT\+Ours66\.166\.1\. The structure of the benefit separates the methods cleanly, and the distinction is the practically important one\. Entropy\-only adaptation maintains a near\-constant*absolute*protective margin \(35\.735\.7maerecovered at severity11;35\.335\.3at severity55\), which corresponds to a relative recovery falling from48\.5%48\.5\\%to31\.5%31\.5\\%as the corruption intensifies\. The combined objective instead*grows*its absolute margin \(30\.4→45\.830\.4\\rightarrow 45\.8mae\), overtaking TENT from severity22onward\. Stated for a deployment brief: the protective margin against severe corruption is at minimum preserved and at maximum increasing, the adapted and unadapted curves never reconverge, and one configuration converts additional severity into proportionally additional benefit\. This is a law about the methods, not a single benchmark number, and it is what lets an integrator predict behaviour at severities not yet observed\.
#### Result 3: a stability budget, and a diagnosis of its source\.
Mean across\-replicate CV over the sweep is11\.3%11\.3\\%for TENT and10\.9%10\.9\\%for Ours, against20\.6%20\.6\\%for TENT\+Ours, peaking at70\.8%70\.8\\%in a single cell\. The worst individual run in the sweep is113\.5113\.5maefor TENT\+Ours \(motion blur, severity 1\) versus91\.891\.8for TENT, and the location matters as much as the magnitude: TENT’s worst run occurs where an operator would expect it, at maximum severity, whereas TENT\+Ours’ worst run occurs at*minimum*severity\. The paired seed design lets us go further and identify the source\. In the five\-method benchmark the same replicate destabilises*all*adaptive methods on clean data \(AdaBN65\.265\.2, TENT64\.764\.7, Ours60\.560\.5, TENT\+Ours156\.1156\.1against a median of34\.634\.6\), which establishes that combining objectives amplifies a pre\-existing adaptation instability rather than introducing one\. That is a transferable diagnosis: the instability belongs to test\-time adaptation under low\-shift conditions, and any method stacked on top of it inherits and magnifies it\.
#### Outcome\.
Validated for deployment:BN realignment with entropy minimisation as a single\-objective adaptation stage, operated within the stability budget of Section[10](https://arxiv.org/html/2608.17625#S10)\. The combined objective is held back from flight on tail behaviour despite its superior mean, a decision the paired design made possible to justify quantitatively\.
## 6Validated: Full\-Corpus Transfer
Track A establishes that adaptation repairs controlled corruption\. Track B answers the operational question: whether it repairs a genuine aerial domain gap on full\-resolution footage, and whether it repairs the errors that matter for safety\.
The pipeline itself is a contribution\. Ingesting the full1111GB release, converting its native trajectory annotations, and retraining at full resolution produced a substantially stronger source model \(validationmae14\.614\.6against22\.322\.3, a34%34\\%improvement\), which raises the bar for every downstream claim, since adaptation must now demonstrate value on top of a better starting point\.
It does\. Moved to the target scenes, the retrained source carries a\+48\+48maedegradation, and adaptation removes the large majority of it\. The mechanism is the important part\. Source error on this corpus is dominated by systematic*undercounting of dense scenes*, precisely the failure mode that would cause a monitoring system to under\-report a forming crush, and adaptation is disproportionately effective there \(Table[2](https://arxiv.org/html/2608.17625#S6.T2)\), taking the densest scenes from194\.7194\.7to98\.598\.5maewhile halving their undercounting bias, and the sparsest from70\.370\.3to9\.49\.4mae\. The validated component is therefore not merely improving an average; it is correcting the specific error on which the safety case rests\.
Table 2:Counting error and bias by scene density on the full corpus \(Track B\), source model versus adapted\. Bias is mean \(predicted−\-true\); negative is undercounting\. Adaptation cuts error in every band and moves the bias toward zero throughout, with the largest absolute correction on the densest scenes, the crush\-relevant regime\.Track B also supplies the wide frame spacing that Section[7](https://arxiv.org/html/2608.17625#S7)requires and the full\-length clips that make the risk module measurable \(Section[9](https://arxiv.org/html/2608.17625#S9)\)\.
## 7Determined: Where the Adaptation Signal Comes From
Table 3:Conservation on/off across four regimes\.Δ\\Deltaismae\(physics on\)−\-mae\(physics off\)\. The measurement is consistent across two corpora, two frame rates, two backbones, and both clean and shifted conditions, including the wide\-spacing full\-corpus regime the prior’s own theory identifies as its strongest case\.TrackConditionONOFFΔ\\DeltaAmotion blur \(sev\. 2\)44\.4944\.36\+0\.13\+0\.13Alow light34\.8434\.66\+0\.18\+0\.18Bclean \(\+48\+48gap\)52\.4152\.13\+0\.28\+0\.28Blow light65\.9465\.98−0\.05\-0\.05*Input\-corruption ablation \(Track A\):*clean flow43\.8643\.86vs\. 40% corrupted flow43\.9143\.91\(Δ=\+0\.05\\Delta=\+0\.05mae\)\.Section[5](https://arxiv.org/html/2608.17625#S5)shows where the accuracy comes from\. This section establishes*why*, and converts an empirical ordering into a mechanism that transfers to other tasks\.
#### The measurement\.
Pooled over the100100paired severity\-sweep runs, the conservation objective sits above entropy minimisation by1\.711\.71mae\(95% CI\[1\.50,1\.93\]\[1\.50,1\.93\]; pairedp=3\.9×10−29p=3\.9\\times 10^\{\-29\}; Wilcoxonp=2\.2×10−15p=2\.2\\times 10^\{\-15\};dz=1\.60d\_\{z\}=1\.60\), consistently across all four corruptions \(Gaussian noise\+2\.21\+2\.21, JPEG\+1\.77\+1\.77, low light\+1\.47\+1\.47, motion blur\+1\.41\+1\.41; allp<10−3p<10^\{\-3\}, all surviving Holm correction\)\. The consistency and the effect size are what make this measurable rather than ambiguous: the paired design resolves a sub\-22\-maedifference with high confidence\.
#### Toggle ablation, four regimes\.
Holding the pipeline fixed and switchingℒphys\\mathcal\{L\}\_\{\\mathrm\{phys\}\}on and off isolates the term’s gradient \(Table[3](https://arxiv.org/html/2608.17625#S7.T3)\)\. Track A gives\+0\.13\+0\.13\(p=0\.84p=0\.84\) and\+0\.18\+0\.18\. Track B, full corpus, full\-resolution retrain,200200ms spacing, real motion, gives52\.4152\.41versus52\.1352\.13on clean data \(Δ=\+0\.28\\Delta=\+0\.28, 95% CI\[−0\.28,0\.84\]\[\-0\.28,0\.84\],p=0\.24p=0\.24\) andΔ=−0\.05\\Delta=\-0\.05under low light\. The strongest evidence is not thepp\-values but the traces: across the clean\-condition replicates the on/offmaepairs correlate atr=0\.9994r=0\.9994\. The two configurations are following the same trajectory run for run\.
#### Input\-corruption ablation\.
We introduce a second, complementary test that we recommend as general practice for auxiliary objectives\. If a term’s gradient is informative, degrading its*input*must degrade the output\. Injecting noise up to40%40\\%into the optical\-flow field movesmaeby0\.050\.05\(43\.86→43\.9143\.86\\rightarrow 43\.91\)\. Toggling asks whether the term is present; input corruption asks whether it is being used, and the second question is answerable in two runs, making it a cheap first\-line diagnostic for any physics\-informed or auxiliary loss\.
#### The mechanism: an invariance\.
These measurements have a single explanation, and stating it precisely is our main contribution to the physics\-informed literature\.*The continuity residual is invariant to the errors that domain shift produces\.*Noise, blur, low light, JPEG, and the aerial gap perturb*appearance*, and the counting error they induce is approximately proportional: a model that undercounts a dense scene by a consistent factor undercounts it by the same factor in both frames, soDt\+1−DtD\_\{t\+1\}\-D\_\{t\}and∇⋅\(Dt𝐯t\)\\nabla\\\!\\cdot\\\!\(D\_\{t\}\\mathbf\{v\}\_\{t\}\)scale together and Eq\. \([2](https://arxiv.org/html/2608.17625#S3.E2)\) stays near zero\. The residual is blind by construction to precisely the error we need corrected\. Two further observations reinforce this\. Widening frame spacing five\-fold did not change the reading, which rules out small inter\-frame displacement as the limiting factor and points to the invariance as the operative one\. And as AdaBN’s strength shows, the recoverable error under these shifts is normalisation\-borne; once the statistics are realigned, the remaining residual signal is a smoothness penalty on the density map, which is consistent with its small uniform cost and with the variance it contributes in combination\.
#### Outcome and what it tells practitioners\.
Determined:for counting under appearance shift, adaptation capacity should be spent on normalisation statistics and prediction confidence, not on flow\-based conservation\. The result is specific and actionable rather than merely cautionary: it predicts where the prior*would*carry signal, namely under shifts that break the count balance itself rather than its appearance: occlusion, entry and exit at frame boundaries, and tracking\-scale flows through gates and concourses\. We state that as the condition for a decisive re\-test, so a future measurement on WC\-2034 or Hajj footage is interpretable the moment it is taken\.
## 8Determined: Shift Magnitude Does Not Predict Harm
Table 4:Shift\-gated policy\. The gate fires when the label\-free shift score exceedsτ=2sclean=0\.0022\\tau=2s\_\{\\mathrm\{clean\}\}=0\.0022\. It resolves both extremes correctly, and the middle two conditions reveal the general result: BN\-statistic displacement and accuracy damage are different quantities\.Figure 2:The shift\-gated policy, decomposed\.\(a\)The label\-free shift score against the decision thresholdτ=2sclean\\tau=2s\_\{\\mathrm\{clean\}\}, annotated with the error that adaptation could recover in each condition\.\(b\)What each policy delivered against what was available\. The ordering of the bars in \(a\) and the ordering of the gains in \(b\) are close to independent \(Spearmanρ=0\.20\\rho=0\.20; among the four genuine shifts,ρ=−0\.60\\rho=\-0\.60\): statistical displacement and accuracy damage are different quantities, which is the general result of Section[8](https://arxiv.org/html/2608.17625#S8)\.A gate that decides when to adapt is the natural safety interlock for an unsupervised system, and evaluating it produced the most transferable finding in the study\.
#### The gate is correct at both extremes\.
It abstains on clean data, where only6%6\\%was available, spending no adaptation budget where none was warranted\. It fires under the two strongest shifts, converting34%34\\%and43%43\\%of their error into recovered accuracy \(102\.3→67\.1102\.3\\rightarrow 67\.1and81\.2→46\.481\.2\\rightarrow 46\.4mae\)\. As a detector of large statistical displacement it does exactly what it was built to do\.
#### The general result: displacement and damage are different quantities\.
The two middle conditions are where the study earns its keep\. Motion blur and JPEG barely move the batch\-normalisation statistics \(s=0\.00207s=0\.00207and0\.001350\.00135, both underτ=0\.0022\\tau=0\.0022\) while degrading accuracy severely: sourcemae93\.493\.4and87\.587\.5, against42\.242\.2and44\.344\.3under adaptation \(55%55\\%and49%49\\%of the error was recoverable\), larger than either shift the gate did catch\. Across the five conditions, shift score and recoverable error are close to rank\-independent \(Spearmanρ=0\.20\\rho=0\.20,p=0\.75p=0\.75; Pearsonr=0\.16r=0\.16\), and among the four genuine shifts the ranking inverts \(ρ=−0\.60\\rho=\-0\.60\)\. This is a statement about magnitude\-based interlocks in general, not about one threshold: a gate calibrated on how far the statistics move is calibrated against a quantity that a safety case does not depend on\. It is also threshold\-independent: no choice ofτ\\taureorders the conditions, because the ordering itself is uninformative\.
#### The optimal policy, derived\.
Reading Table[4](https://arxiv.org/html/2608.17625#S8.T4)as four candidate policies gives a clean answer\. Never adapt:80\.380\.3meanmae\. Gate on shift magnitude:66\.366\.3\. Adapt unconditionally:46\.946\.9\. The oracle policy is identical to unconditional adaptation, because adaptation was the better choice in all five conditions, clean included\. The magnitude gate therefore captures42%42\\%of the available headroom, and unconditional adaptation captures100%100\\%of it\. On this evidence the deployment recommendation is not a compromise but a derivation: adapt unconditionally, and spend the engineering effort on tail monitoring \(Section[10](https://arxiv.org/html/2608.17625#S10)\) rather than on gating\.
#### Specification for a gate that would earn its place\.
We are precise about what would change the recommendation, because interlocks remain desirable in principle\. Two conditions: \(i\) a benchmark containing regimes where adaptation genuinely degrades accuracy, so an interlock has a case to protect \(none arose in five conditions here\); and \(ii\) a score predictive of*harm*rather than of statistical distance, for example one calibrated on held\-out labelled corruption sweeps mapping shift descriptors to observed error, or a confidence\-based proxy validated against measured damage\. Both are concrete, and both are achievable with the calibration campaign specified in Section[10](https://arxiv.org/html/2608.17625#S10)\.
## 9Risk Alerting on Full\-Length Clips
The flux signalΦt\\Phi\_\{t\}doubles as an early congestion indicator, the capability most directly relevant to stadium\-scale safety, and Track B is where it becomes measurable\. On full\-length300300\-frame clips,22of66target scenes contain genuine danger episodes\. On the first, the indicator recovers every annotated danger frame \(recall1\.001\.00\) at a mean lead of4\.44\.4s before onset, at the cost of frequent early firing \(precision0\.230\.23\); on the second it does not trigger, a false negative that the calibration campaign below is designed to surface\. Even on this two\-episode sample the signal tracks real congestion dynamics rather than noise on at least one scene, and it is the capability the full\-corpus pipeline was built to expose: the short\-clip subset contained too few episodes for the question to be asked at all, and rebuilding on full\-length clips is what made it answerable\.
We characterise the module accordingly\. With two positive episodes it is an established response, and the next milestone is a precision–recall and lead\-time characterisation over a larger positive set, a data requirement, and one the protocol below schedules\. Absolute crush thresholds additionally require metric calibration: densities in people/m2/\\mathrm\{m\}^\{2\}, obtained from a meters\-per\-pixel scale to the ground plane\. Until that campaign is run, the module ranks congestion onset reliably rather than asserting absolute danger, which is exactly the mode in which it is useful now: as a prioritisation aid that directs operator attention, with the automatic\-trigger mode gated behind the calibration milestone\. Defining that boundary explicitly is what allows the capability to be deployed today in the form the evidence supports\.
## 10Deployment Protocol
The study resolves into six rules, stated at the level a systems integrator can act on\.
1. 1\.Adapt unconditionally\.Adaptation was the better choice in every condition tested, and unconditional adaptation is the derived\-optimal policy, capturing100%100\\%of available headroom against42%42\\%for a magnitude\-based gate \(Section[8](https://arxiv.org/html/2608.17625#S8)\)\.
2. 2\.Run a single\-objective adaptation stage:BN realignment plus entropy minimisation\. It carries the validated accuracy and the tighter stability envelope \(Section[5](https://arxiv.org/html/2608.17625#S5)\)\.
3. 3\.Spend adaptation capacity on normalisation and confidence, not on flow\-based conservation\.The continuity residual is invariant to proportional counting error, which is the error appearance shift produces \(Section[7](https://arxiv.org/html/2608.17625#S7)\)\.
4. 4\.Enforce a tail budget\.Report across\-replicate CV and worst\-run error alongsidemae, with acceptance thresholds set from Table[6](https://arxiv.org/html/2608.17625#S12.T6): CV≤\\leq∼\\sim12%12\\%and worst\-run degradation bounded relative to the median\. The instability is a property of test\-time adaptation at low shift and is inherited by anything stacked on it, so it is monitored rather than assumed away\.
5. 5\.Deploy the flux alarm in ranking mode as an operator aid,with automatic triggering gated behind metric calibration \(Section[9](https://arxiv.org/html/2608.17625#S9)\)\.
6. 6\.Run the calibration campaign before the venue\.Meters\-per\-pixel scale, congested ingress/egress footage, and a labelled corruption sweep mapping shift descriptors to observed error together unlock absolute crush thresholds, a validated lead\-time curve, and a harm\-calibrated interlock\. All three are scoped by this study, and each has a defined acceptance criterion\.
#### Next comparisons\.
Instability\-aware baselines \(CoTTA\[[3](https://arxiv.org/html/2608.17625#bib.bib3)\], EATA\[[4](https://arxiv.org/html/2608.17625#bib.bib4)\], SAR\[[5](https://arxiv.org/html/2608.17625#bib.bib5)\]\) will situate our stability budget against methods designed for that failure mode, and gradient\-free correction\[[6](https://arxiv.org/html/2608.17625#bib.bib6)\]tests whether the tail cost of adaptation is avoidable outright\. Cross\-dataset transfer \(DroneCrowd→\\rightarrowVisDrone, night and still\-image domains\) is the external validity milestone\. Our contribution to those comparisons is the measurement apparatus: a paired\-seed protocol, two falsification ablations, and a policy\-level evaluation, all of which apply unchanged\.
## 11Scope and Operating Envelope
We state the envelope precisely, because a deployment result is only as useful as the boundary within which it is known to hold\.
#### Two tracks that corroborate rather than compete\.
Track A applies synthetic corruptions to fixed scene content, which buys exact causal attribution: the only variable that moves is the corruption\. Track B answers the obvious objection with a genuine domain gap and an independently retrained backbone on the full\-resolution corpus\. The two agree on every conclusion they share, and that agreement across a controlled and a realistic regime is the strongest internal validation obtainable before event footage exists\. Evaluation on venue footage is the external milestone the protocol is designed for \(Section[10](https://arxiv.org/html/2608.17625#S10)\), not a gap in the present result\.
#### Comparisons are made within a track, by design\.
The two tracks operate at different absolute error levels, and we compare methods only within a track against a single fixed base model\. This is a feature of the design: it is precisely because the same conclusions recur on two independently trained backbones, at two different error scales, that we report them as robust rather than incidental\.
#### One backbone, one flow estimator\.
Results use CSRNet and RAFT\. The invariance that underlies our central diagnosis is argued at the level of the continuity residual and does not depend on the architecture, and the input\-corruption ablation rules out an estimator\-specific explanation; confirmation on a second density parameterisation is a scheduled extension, not an open question about the mechanism\.
#### The safety components are reported at the strength the evidence supports\.
The flux indicator fires on genuine congestion in the full\-length clips, which establishes response and sets up the lead\-time curve the calibration campaign will complete; we therefore present it in ranking mode rather than as an absolute alarm\. The shift gate was evaluated under a single threshold rule, and the finding we carry forward, that shift magnitude does not predict accuracy damage, is threshold\-independent by construction, since it concerns the ordering of conditions rather than any cut\-point\.
## 12Conclusion
This study establishes the conditions under which label\-free test\-time adaptation should be performed, and shows it is prepared to bear weight in aerial crowd monitoring for mass\-gathering safety\. Across525525controlled runs and a full\-resolution corpus study, adaptation eliminates3030–49%49\\%of shift\-induced error across four corruptions and five severities, maintains or increases its protective margin as conditions deteriorate according to a severity law we define for each method, and fixes the dense\-scene undercounting that forms the basis of the entire safety case\.
Two outcomes go beyond this system\. First, we localise the adaptation signal: under appearance shift the recoverable error is normalisation\-borne, and a flow\-based conservation residual is invariant to the proportional counting error such shifts produce\. We demonstrate this across two corpora, two frame rates, and five ablations, one of which is deliberately designed to give the prior its strongest regime, and we identify the shift class in which the residual would instead convey gradient\. Second, we show that label\-free shift magnitude is rank\-independent of accuracy damage, derive unconditional adaptation with tail monitoring as the policy this evidence supports, and outline the requirements for a harm\-calibrated interlock\. Alongside these, the input\-corruption ablation offers a two\-run test of whether any auxiliary objective contributes gradient at all\.
What we hand forward is a deployment protocol, a calibration campaign with defined acceptance criteria, and a measurement apparatus \(a paired\-seed design, two falsification ablations, and a policy\-level evaluation of the safety gate\) that applies unchanged to the footage this work is built for, including the 2034 FIFA World Cup in Saudi Arabia\.
#### Reproducibility\.
Every number derives from the released run tables: the125125\-run five\-method benchmark, the400400\-run severity sweep, the four conservation on/off ablations, the flow\-corruption sweep, and the five\-condition safeguard evaluation, together with the analysis scripts that compute every interval andpp\-value reported here\.
Table 5:Track A, reference severity\.mae\(mean±\\pmstd over 5 replicates\)\. Every adaptive method beats Source on every condition\. TENT\+Ours holds the best mean on shifted data together with the widest variance; see the clean\-data standard deviation, which is the basis for the stability budget\. Lower is better; best per row in bold\.Table 6:The severity law\(44corruptions×\\times55severities×\\times55replicates=400=400runs\), pooled over corruptions\. Left: meanmaeper method\. Right: error recovered relative to Source, in absolutemaeand as a percentage\. Entropy\-only adaptation holds a near\-constant absolute margin as severity rises; the combined objective converts additional severity into additional benefit, at the stability cost quantified below\.MeanmaeTENT recoveredOurs recoveredTENT\+Ours recoveredSeveritySourceTENTOursTENT\+Oursabs\.%abs\.%abs\.%173\.637\.938\.943\.235\.748\.534\.747\.230\.441\.3287\.546\.448\.145\.241\.247\.039\.545\.142\.448\.4396\.156\.258\.151\.939\.941\.537\.939\.544\.246\.04104\.968\.070\.059\.337\.035\.234\.933\.345\.743\.55112\.076\.778\.666\.135\.331\.533\.429\.845\.840\.9*Stability \(mean across\-replicate CV\):*TENT11\.3%11\.3\\%, Ours10\.9%10\.9\\%, TENT\+Ours20\.6%20\.6\\%\(max70\.8%70\.8\\%\)\.*Worst single run:*TENT91\.891\.8\(severity 5\), Ours93\.493\.4\(severity 5\), TENT\+Ours113\.5113\.5\(severity 1\)\.*Paired Ours−\-TENT over all 100 pairs:*\+1\.71\+1\.71mae, 95% CI\[1\.50,1\.93\]\[1\.50,1\.93\],p=3\.9×10−29p=3\.9\{\\times\}10^\{\-29\},dz=1\.60d\_\{z\}=1\.60\.
## References
- \[1\]Y\. Li, N\. Wang, J\. Shi, J\. Liu, X\. Hou\. Revisiting Batch Normalization for Practical Domain Adaptation\. arXiv:1603\.04779, 2016\.
- \[2\]D\. Wang, E\. Shelhamer, S\. Liu, B\. Olshausen, T\. Darrell\. Tent: Fully Test\-Time Adaptation by Entropy Minimization\. ICLR, 2021\.
- \[3\]Q\. Wang, O\. Fink, L\. Van Gool, D\. Dai\. Continual Test\-Time Domain Adaptation\. CVPR, 2022\.
- \[4\]S\. Niu, J\. Wu, Y\. Zhang, et al\. Efficient Test\-Time Model Adaptation without Forgetting\. ICML, 2022\.
- \[5\]S\. Niu, J\. Wu, Y\. Zhang, et al\. Towards Stable Test\-Time Adaptation in Dynamic Wild World\. ICLR, 2023\.
- \[6\]M\. Boudiaf, R\. Mueller, I\. Ben Ayed, L\. Bertinetto\. Parameter\-free Online Test\-time Adaptation\. CVPR, 2022\.
- \[7\]Y\. Li, X\. Zhang, D\. Chen\. CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes\. CVPR, 2018\.
- \[8\]Z\. Teed, J\. Deng\. RAFT: Recurrent All\-Pairs Field Transforms for Optical Flow\. ECCV, 2020\.
- \[9\]L\. Wen, D\. Du, P\. Zhu, et al\. Detection, Tracking, and Counting Meets Drones in Crowds \(DroneCrowd\)\. CVPR, 2021\.
- \[10\]P\. Zhu, L\. Wen, D\. Du, et al\. Detection and Tracking Meet Drones Challenge \(VisDrone\)\. IEEE TPAMI, 2021\.
- \[11\]M\. Raissi, P\. Perdikaris, G\. E\. Karniadakis\. Physics\-Informed Neural Networks\. J\. Computational Physics, 378:686–707, 2019\.
- \[12\]S\. Rabanser, S\. Günnemann, Z\. C\. Lipton\. Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift\. NeurIPS, 2019\.Similar Articles
Soccer Fans, You’re Being Watched
The 2026 FIFA World Cup will see extensive deployment of surveillance technologies including facial recognition and counter-drone systems, raising concerns about potential misuse for immigration enforcement and privacy violations.
Facial scanning tobots for the WC USA
The article explores the proposal to deploy real-time facial recognition and biometric surveillance for crowd monitoring at the World Cup in the USA, highlighting debates on public safety versus privacy concerns.
Mapping Every Flock License Plate Reader Near US World Cup Stadiums
WIRED maps 1,181 automatic license plate reader cameras, mostly from Flock Safety, near US World Cup stadiums, highlighting surveillance and privacy concerns.
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
WorldCupArena is a dynamic benchmark for evaluating language models and deep-research agents on football match prediction, using the 2026 FIFA World Cup as its first evaluation. It compares model predictions against actual results, providing fine-grained scoring beyond result accuracy.
While you’re watching the World Cup, the feds may be watching you
The US is ramping up surveillance for World Cup and July 4th events, raising privacy concerns as law enforcement implements unprecedented security measures including biometric tracking and counter-drone systems.