H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities
Summary
This paper adapts Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP) frameworks to professional handball, introducing Handball-xT and Handball-VAEP using five seasons of tracking-derived event data from the Handball Bundesliga, and releases the code for clubs.
View Cached Full Text
Cached at: 08/14/26, 09:32 AM
# H-VAEP and H-xT:Valuing Offensive On-the-Ball Actionsin Handball by Estimating Probabilities
Source: [https://arxiv.org/html/2608.12926](https://arxiv.org/html/2608.12926)
Julius BroermannAffiliation:Paderborn University, Paderborn, GermanyE\-mail[\{firstname\.lastname\}@uni\-paderborn\.de](mailto:{firstname.lastname}@uni-paderborn.de)Oliver MuellerAffiliation:Paderborn University, Paderborn, GermanyE\-mail[\{firstname\.lastname\}@uni\-paderborn\.de](mailto:{firstname.lastname}@uni-paderborn.de)Michael DoeringAffiliation:Paderborn University, Paderborn, GermanyE\-mail[\{firstname\.lastname\}@uni\-paderborn\.de](mailto:{firstname.lastname}@uni-paderborn.de)Affiliation:SG Flensburg\-Handewitt, Flensburg, GermanyJochen BaumeisterAffiliation:Paderborn University, Paderborn, GermanyE\-mail[\{firstname\.lastname\}@uni\-paderborn\.de](mailto:{firstname.lastname}@uni-paderborn.de)
###### Abstract
Traditional player evaluation in professional handball relies on basic box\-score metrics or heuristic indices, which fail to credit the multi\-player build\-up chain\. While football \(soccer\) analytics has adopted Expected Threat \(xT\) and Valuing Actions by Estimating Probabilities \(VAEP\), these event\-based action valuation frameworks have not yet been adapted to handball\. In this paper, we present the first comprehensive adaptation and evaluation of xT and VAEP for handball, utilizing five seasons of tracking\-derived event data from the Handball Bundesliga\. We develop Handball\-xT \(H\-xT\) using a handball\-native court zoning layout, demonstrating via simulations that it is systematically more robust than standard rectangular grids\. We optimize Handball\-VAEP \(H\-VAEP\) by tailoring its feature space and selecting the context length to limit team\-identity leakage\. Our evaluation shows that H\-VAEP yields exceptionally stable, discriminative, and intuitive player ratings that highlight build\-up play\. Finally, we release our complete code repository to help professional clubs deploy these models\.
###### Keywords:
handball action valuation expected threat VAEP
## 1Introduction
Optimizing player recruitment and game tactics requires a detailed understanding of how individual player actions contribute to winning matches\. However, traditional player evaluation in handball remains heavily reliant on basic box\-score metrics like goals and assists, or on heuristic, fixed\-weight combinations of them such as the Handball Performance Index \(HPI\)\[[7](https://arxiv.org/html/2608.12926#bib.bib19)\]\. While football analytics has successfully transitioned to valuing every on\-ball action using Expected Threat \(xT\)\[[18](https://arxiv.org/html/2608.12926#bib.bib8)\]and VAEP\[[4](https://arxiv.org/html/2608.12926#bib.bib7),[5](https://arxiv.org/html/2608.12926#bib.bib4)\], professional handball has lacked a comparable framework\. Consequently, players who drive the build\-up and create scoring opportunities well before the final shot are systematically undervalued\. The transfer of football\-style action valuation frameworks to handball has historically been hindered by the limitations of scout\-collected event data\. Traditional box\-score statistics record only shots, goals, and assists, omitting the passes and dribbles that constitute the bulk of play\. However, the deployment of Local Positioning Systems \(LPS\) across the Handball Bundesliga \(HBL\) has resolved this constraint and enables the automatic recognition of event sequences from tracking data, which are made available to the clubs\. Yet, a case study of 13 HBL clubs indicates that converting this data into actionable insights remains a major challenge due to scarce analytical resources and specialized expertise\[[8](https://arxiv.org/html/2608.12926#bib.bib5)\]\.
We address this barrier by developing and evaluating Handball\-xT \(H\-xT\) and Handball\-VAEP \(H\-VAEP\) using five seasons of HBL tracking\-derived event data \(2021/22 to 2025/26\)\. To facilitate practical adoption, we release our code repository, including API connectors to data providers\.111The code is available at[https://github\.com/JuliusBroermann/handballaction](https://github.com/JuliusBroermann/handballaction)\.Specifically, our contributions are threefold\. First, we adapt Expected Threat to handball \(H\-xT\) using a court zoning layout that respects the sport’s geometry, validating its robustness over standard rectangular grids via simulation\[[19](https://arxiv.org/html/2608.12926#bib.bib22)\]\. Second, we adapt VAEP to handball \(H\-VAEP\) by optimizing its features, algorithm, and context length to suit the sport’s rapid, high\-scoring dynamics while limiting team\-identity leakage\. Third, we evaluate the resulting ratings under a comprehensive validation framework\[[3](https://arxiv.org/html/2608.12926#bib.bib10),[6](https://arxiv.org/html/2608.12926#bib.bib1),[12](https://arxiv.org/html/2608.12926#bib.bib12)\], demonstrating their strong face validity, reliability, discrimination, and stability compared to traditional performance indicators\.
## 2Related Work
While Expected Threat \(xT\)\[[18](https://arxiv.org/html/2608.12926#bib.bib8)\]and VAEP\[[4](https://arxiv.org/html/2608.12926#bib.bib7),[5](https://arxiv.org/html/2608.12926#bib.bib4)\]are prominent in football, the broader principle of valuing actions via expected possession value has been applied across various invasion sports\[[20](https://arxiv.org/html/2608.12926#bib.bib3)\], highlighting how valuation models must be customized to sport\-specific geometries and rules\. Methodologically, models diverge between discrete state space representations \(e\.g\., zones or transition matrices in hockey and rugby\[[13](https://arxiv.org/html/2608.12926#bib.bib13),[17](https://arxiv.org/html/2608.12926#bib.bib14)\]\) and continuous court models \(e\.g\., expected possession value in basketball and rugby\[[2](https://arxiv.org/html/2608.12926#bib.bib6),[16](https://arxiv.org/html/2608.12926#bib.bib15)\]\)\. Furthermore, models must incorporate dynamic, rule\-specific constraints, such as power plays in ice hockey\[[9](https://arxiv.org/html/2608.12926#bib.bib21)\], shot clocks in basketball\[[15](https://arxiv.org/html/2608.12926#bib.bib16)\], tackle limits in rugby\[[17](https://arxiv.org/html/2608.12926#bib.bib14)\], touch phases in volleyball\[[21](https://arxiv.org/html/2608.12926#bib.bib17)\], and down\-and\-distance situations in American football\[[22](https://arxiv.org/html/2608.12926#bib.bib18)\]\.
In handball, player valuation is dominated by traditional statistics or box\-score\-derived heuristic indices like the HPI\[[7](https://arxiv.org/html/2608.12926#bib.bib19)\], which lack spatial and sequential context\. Existing advanced models are either limited to terminal shots, such as Expected Goals \(xG\) frameworks\[[1](https://arxiv.org/html/2608.12926#bib.bib11),[10](https://arxiv.org/html/2608.12926#bib.bib20)\], or operate on continuous tracking trajectories, such as the spatiotemporal EPV framework PIVOT\[[11](https://arxiv.org/html/2608.12926#bib.bib2)\]\. Consequently, handball lacks a tailored, event\-based framework that values discrete actions while incorporating sport\-specific geometries and rules \(e\.g\., native court zones and passive play constraints\)\. Furthermore, existing player ratings have not been validated using structured meta\-analytics frameworks\[[3](https://arxiv.org/html/2608.12926#bib.bib10),[6](https://arxiv.org/html/2608.12926#bib.bib1),[12](https://arxiv.org/html/2608.12926#bib.bib12)\]\. We bridge these gaps by adapting and systematically evaluating xT and VAEP for handball, with H\-xT representing the discrete state space models and H\-VAEP the models that additionally condition on the preceding actions and the game context\.
## 3Adapting Expected Threat to Handball
The Expected Threat \(xT\) framework\[[14](https://arxiv.org/html/2608.12926#bib.bib9),[18](https://arxiv.org/html/2608.12926#bib.bib8)\]evaluates ball progression actions using a spatial Markov chain\. The court is discretized into zones, and the expected threat valuexT\(z\)xT\(z\)represents the probability that a possession currently in zonezzleads to a goal\. Transitions between zones are determined by movement and shot probabilities, which can be solved recursively as detailed in\[[18](https://arxiv.org/html/2608.12926#bib.bib8)\]\. Once these zone values are computed, any individual ball progression action starting in zonezstartz\_\{\\mathrm\{start\}\}and ending inzendz\_\{\\mathrm\{end\}\}is valued by the difference in expected threat:ΔxT=xT\(zend\)−xT\(zstart\)\\Delta xT=xT\(z\_\{\\mathrm\{end\}\}\)\-xT\(z\_\{\\mathrm\{start\}\}\)\. How to discretize the playing area is a critical design choice in xT\. While football analytics typically relies on rectangular grids \(e\.g\., a16×1216\\times 12layout\), this approach is ill\-suited for handball court geometry, which is defined by curved 6m goal creases and 9m free\-throw arcs\. Consultations with professional coaches confirmed that rectangular cells are unintuitive because grid lines intersect these arcs arbitrarily, destroying tactical interpretability\. Consequently, we developed a handball\-native court zoning layout \(illustrated in Fig\.[1\(c\)](https://arxiv.org/html/2608.12926#S3.F1.sf3)\) featuring angular boundaries originating from the goals and concentric divisions matching the 6m and 9m lines\. Furthermore, since build\-up play occurs almost exclusively in the opponent’s half, we aggregate the entire defensive half into a single zone\. Under a full grid, defensive\-half zones converge to virtually identical expected threat values, with minor differences representing spurious noise from action scarcity\.
To select the optimal number of zones and compare layouts, we adopt the simulation\-based robustness methodology of\[[19](https://arxiv.org/html/2608.12926#bib.bib22)\], which balances the trade\-off between model flexibility \(capturing fine\-grained tactical movements\) and robustness \(avoiding high parameter variance\)\. We utilize tracking\-derived event sequences across five HBL seasons \(2021/22 and 2022/23 as development set; 2023/24, 2024/25, and 2025/26 as held\-out evaluation set\)\. We apply light data cleaning, including removing the first 30 s of each match to exclude pre\-match ceremonial passing that would skew transition probabilities\. We represent the event sequences using a schema adapted from SPADL\[[4](https://arxiv.org/html/2608.12926#bib.bib7)\]by omitting body part features and defining four action types: passes, dribbles \(ball possession\), field shots, and seven\-meter penalty shots\. Since pass success labels are missing in portions of the tracking data, success is imputed by verifying if the passing team retains possession for the subsequent action\. We fit a ground\-truth model,MfullM\_\{\\mathrm\{full\}\}, using the combined 2021/22 and 2022/23 development seasons\. To quantify the robustness\-flexibility trade\-off, we then simulateB=1000B=1000seasons by sampling 306 matches \(representing one HBL season\) from the development set\. For each simulationbb, we fit a modelMbM\_\{b\}and calculate the maximum zone\-level absolute deviation from the ground truth:Db=maxz\|xTb\(z\)−xTfull\(z\)\|D\_\{b\}=\\max\_\{z\}\|xT\_\{b\}\(z\)\-xT\_\{\\mathrm\{full\}\}\(z\)\|\. Following\[[19](https://arxiv.org/html/2608.12926#bib.bib22)\], the robustness metricR90R\_\{90\}is the 90th percentile ofDbD\_\{b\}across all simulations\. To ensure a fair comparison via the total number of zones as a flexibility measure, we introduce a single backcourt zone in the grid model as well\. Fig\.[1\(a\)](https://arxiv.org/html/2608.12926#S3.F1.sf1)shows the robustness curves for both layouts\. For any number of zones, the grid model’s 90th percentile deviation is larger than that of the handball\-native model, demonstrating that our native layout is systematically more robust\. Applying the decision rule of choosing the most flexible configuration withR90≤0\.03R\_\{90\}\\leq 0\.03\[[19](https://arxiv.org/html/2608.12926#bib.bib22)\], the handball\-native layout supports a capacity of 67 zones, whereas the grid model is restricted to 45 zones\.
\(a\)Robustness comparison\.\(b\)Grid\-based xT\.
\(c\)Handball\-native xT\.
Figure 1:Robustness analysis and Expected Threat surfaces\. \(a\) shows the 90th percentile of maximum deviation \(shaded down to the 10th percentile\) as a function of zone count for grid\-based vs\. handball\-native layouts following\[[19](https://arxiv.org/html/2608.12926#bib.bib22)\]\. \(b\) and \(c\) display the xT surfaces \(in %\) fitted on the 2024/25 season\.The 2024/25 season xT surfaces \(Figs\.[1\(b\)](https://arxiv.org/html/2608.12926#S3.F1.sf2)and[1\(c\)](https://arxiv.org/html/2608.12926#S3.F1.sf3)\) reveal key tactical insights\. First, inside the 6m crease, threat increases with the goal angle, reflecting higher expected shot values from central positions\. Second, a low\-threat trough exists directly outside the 9m line compared to the outer backcourt: as attackers approach the 9m boundary, defensive contact increases turnovers\. Once this line is penetrated, threat rises near the central 6m line due to high\-quality shooting opportunities\. These central zones are likely undervalued as H\-xT excludes 7m penalties due to missing foul coordinates\. Third, wing corners outside the 6m line exhibit the highest threat outside the crease\. Wing passes enable players to jump into the crease where defensive contact is restricted, yielding clean shots with minimal interference\.
## 4Adapting VAEP to Handball
The VAEP framework\[[4](https://arxiv.org/html/2608.12926#bib.bib7),[5](https://arxiv.org/html/2608.12926#bib.bib4)\]rates on\-ball actions by their impact on a team’s short\-term scoring and conceding probabilities\. Specifically, it uses machine learning to estimatePscores\(Si\)P\_\{\\mathrm\{scores\}\}\(S\_\{i\}\)andPconcedes\(Si\)P\_\{\\mathrm\{concedes\}\}\(S\_\{i\}\), which represent the probabilities that the team possessing the ball scores or concedes a goal within the nextN=10N=10actions from a given game stateSiS\_\{i\}, represented by the lastkkactions\. An action’s value is then defined as the change in scoring probability minus the change in conceding probability caused by it:V\(ai\)=ΔPscores−ΔPconcedesV\(a\_\{i\}\)=\\Delta P\_\{\\mathrm\{scores\}\}\-\\Delta P\_\{\\mathrm\{concedes\}\}\. We refer the reader to\[[4](https://arxiv.org/html/2608.12926#bib.bib7)\]for the full mathematical details of the framework\.
When adapting VAEP to team handball, we optimize the feature space, the learning algorithm, and the context lengthkkfor the sport’s geometries and high\-scoring dynamics\. To evaluate different configurations, we train models on the 2021/22 season and evaluate on the 2022/23 season, measuring Brier score for calibration and ROC\-AUC for rank ordering\. Uncertainty is quantified by a game\-clustered bootstrap \(1,000 replicates over the 306 test games, drawn once and shared across configurations so that comparisons are paired\)\. We control the family\-wise error rate with Holm’s correction\. Hyperparameters are tuned via 5\-fold cross\-validation on game\-level splits using a tree\-structured Parzen estimator \(50 trials\)\. The original football\-native model serves as our baseline\[[4](https://arxiv.org/html/2608.12926#bib.bib7)\]\(CatBoost with default parameters, original feature set,k=3k=3\)\.
Table 1:Predictive performance \(Brier score↓\\downarrow, ROC\-AUC↑\\uparrow\) of H\-VAEP across configurations of feature sets, algorithms, and context lengthskk\(trained on 2021/22, evaluated on 2022/23\)\. The ‘No Features’ baseline predicts marginal training set probabilities\. Bold entries highlight the baseline configuration \(Original VAEP\), the best feature combination \(1b\+2\+3\), the chosen model \(XGBoost,k=3k=3\), and the best overall model \(k=6k=6\)\.Table[1](https://arxiv.org/html/2608.12926#S4.T1)shows that applying football\-native features directly is suboptimal\. We adapt the feature space in three key areas\. First, to prevent overfitting from sparse, high\-scoring patterns, we replace exact cumulative scores \(which generalize poorly in handball’s high\-scoring environment\) with a bucketed score difference \(1b: categorized from strongly behind to strongly leading\)\. Second, to represent shooting quality more effectively, we replace the polar angle with the goal angle between the rays connecting the ball to the two posts \(2\)\. This visual angle is more critical in handball than in football, as handball attacks allow shots from almost any position during a set play, whereas football possessions feature only brief windows in potential shooting locations\. Third, we capture tactical pace by adding the elapsed seconds since gaining possession \(3\), which distinguishes rapid fast breaks from slower set attacks and indicates upcoming shot pressure under handball’s passive play rule\. As shown in Table[1](https://arxiv.org/html/2608.12926#S4.T1), combining these complementary features \(1b\+2\+3\) yields a compounding performance improvement\. All individual and the combined feature modifications improve both Brier score and ROC\-AUC significantly \(α=0\.05\\alpha=0\.05, Holm corrected\) over the tuned baseline forPscoresP\_\{\\mathrm\{scores\}\}, with the exception of the location replacement, whose ROC\-AUC gain is not significant\. ForPconcedesP\_\{\\mathrm\{concedes\}\}the effects are weaker and split by metric: no feature modification significantly improves the Brier score, while the added context feature \(3\) and the combined set do improve ROC\-AUC significantly \(both Holm correctedp<0\.05p<0\.05\)\. The adaptations therefore sharpen the ranking of conceding risk without measurably improving its calibration\.
In terms of model selection, XGBoost achieves the best performance\. Notably, the gain from customizing the features \(\+0\.00596\+0\.00596scoring ROC\-AUC over the tuned CatBoost baseline, 95% CI\[0\.00509,0\.00678\]\[0\.00509,0\.00678\]\) is more than double the gain from subsequent model class selection \(\+0\.00244\+0\.00244scoring ROC\-AUC,\[0\.00211,0\.00276\]\[0\.00211,0\.00276\]\)\.
\(a\)Scoring target \(PscoresP\_\{\\mathrm\{scores\}\}\)\.\(b\)Conceding target \(PconcedesP\_\{\\mathrm\{concedes\}\}\)\.
Figure 2:Predictive performance \(Brier score, left axis, solid\) and team\-identity leakage \(Macro\-AUC, right axis, dashed\) across context lengthskk\.While extending possession history \(context lengthkk\) provides more sequential context to capture tactical patterns, it increases the risk of overfitting and team\-identity leakage, which occurs when a model identifies team\-specific playing styles and predicts outcomes based on team identity rather than general action quality\[[12](https://arxiv.org/html/2608.12926#bib.bib12)\]\. To quantify this, we train a classifier to predict the identity of the team executing the action from its sequence context, measuring its Macro\-AUC on a stratified 20% game\-level holdout within the 2021/22 season\. As shown in Fig\.[2](https://arxiv.org/html/2608.12926#S4.F2), team leakage increases monotonically with context length, rising from 0\.608 atk=1k=1to 0\.664 atk=3k=3and 0\.698 atk=6k=6\. Although extending context tok=6k=6improves predictive performance \(Table[1](https://arxiv.org/html/2608.12926#S4.T1)\), it incurs substantially higher team\-identity leakage\. Thus, we selectk=3k=3as a balanced compromise between predictive quality and team\-identity bias\. Compared to the original VAEP baseline, our final model configuration \(tuned XGBoost, 1b\+2\+3,k=3k=3\) reduces the scoring Brier score to 0\.11220 \(−0\.00241\-0\.00241, 95% CI\[−0\.00287,−0\.00205\]\[\-0\.00287,\-0\.00205\]\) and the conceding Brier score to 0\.02735 \(−0\.00045\-0\.00045,\[−0\.00054,−0\.00037\]\[\-0\.00054,\-0\.00037\]\), while improving scoring ROC\-AUC to 0\.79582 \(\+0\.01330\+0\.01330,\[0\.01174,0\.01491\]\[0\.01174,0\.01491\]\) and conceding ROC\-AUC to 0\.80683 \(\+0\.01592\+0\.01592,\[0\.01311,0\.01876\]\[0\.01311,0\.01876\]\)\.
\(a\)Scoring target \(PscoresP\_\{\\mathrm\{scores\}\}\)\.\(b\)Conceding target \(PconcedesP\_\{\\mathrm\{concedes\}\}\)\.
Figure 3:Reliability of the final H\-VAEP model, pooled over the three held\-out seasons 2023/24, 2024/25, and 2025/26\. Red dots mark equal\-mass decile bins, blue squares equal\-width bins, and the gray histogram the distribution of the predicted probabilities \(right axis\)\.We also assess the final model configuration’s performance and calibration on the three held\-out seasons \(2023/24 to 2025/26\) with a rolling training and prediction setup: models are trained on seasonn−1n\-1\(starting with 2022/23\) to compute action valuations for seasonnnso that every reported probability is out of sample\. Pooled over the 2\.83M valued game states, the scoring model attains a Brier score of 0\.11341 at a ROC\-AUC of 0\.79128 and the conceding model a Brier score of 0\.02654 at a ROC\-AUC of 0\.79454, against base rates of 17\.3% and 3\.0%\. In addition, Fig\.[3](https://arxiv.org/html/2608.12926#S4.F3)reports the reliability of both classifiers\. The scoring model follows the diagonal closely over the entire populated range \(expected calibration error 0\.006 over decile bins\), including the rare predictions above 50%, indicating that it is well calibrated\. Its largest systematic deviation is an underestimation of 1\.4 percentage points around 15%, and the apparent overestimation in the 80–90% bin rests on just 1,733 of 2\.83M predictions\. The spike of the histogram at 100%, and the exactly calibrated equal\-width bin above 90%, are the 53k successful shots \(1\.9% of the states\), whose scoring probability we replace deterministically by 100% because a goal fixes the outcome of the state by definition\. The conceding model is calibrated equally well \(expected calibration error 0\.001\) and stays on the diagonal up to 75% despite how seldom such states occur\. Only above 80%, where fewer than 0\.02% of the predictions lie, does it overestimate the risk of conceding\.
## 5Empirical Player Ratings and Value Composition
To evaluate player ratings for the 2024/25 season, we restrict the analysis to players with at least 500 total minutes and 250 minutes in offense, using official, human\-verified box\-score statistics from Sportradar and HPI data from the HBL website for baseline comparison\. The 500\-minute threshold represents a comparable standard to the 900 minutes used in football VAEP\[[4](https://arxiv.org/html/2608.12926#bib.bib7)\], reflecting handball’s higher physical intensity and rotation rates where players rarely exceed 50 minutes per match\. The additional offensive playing time filter is necessary to exclude defensive specialists, as the tracking system’s automatically recognized event data do not yet include defensive events\.
The resulting rankings demonstrate strong face validity: the ten players with the highest H\-VAEP per 10 minutes in offense \(H\-VAEP/10o\) are widely recognized elite performers who, between 2022 and 2025, collectively received three IHF World Handball Player of the Year awards, three HBL MVP titles, two EHF Champions League Final Four MVP awards, and two HBL Best Young Player awards\. At the same time, the ranking is not a mere reproduction of box\-score statistics\. The player ranked first in H\-VAEP/10o ranks only 55th in goals and 62nd in goals/10o, 9th in assists, 8th in assists/10o, and 27th in HPI: traditional statistics do not fully capture this player’s exceptional build\-up play, which our model highlights by valuing the complete action sequence\. Conversely, the player leading the league in both goals/10o and assists/10o ranks sixth in H\-VAEP/10o\. While goals and assists capture the terminal actions of possessions, H\-VAEP incorporates the value of the preceding build\-up play and accounts for actions like turnovers, providing a more comprehensive view of offensive contributions beyond box\-score totals\.
Figure 4:Relative Handball\-VAEP decomposition by action type in the 2024/25 season\. The left panel shows the composition of the average H\-VAEP per 10 minutes in offense for each position group\. The right panel shows it for the back players \(left, right, and centre backs\) of the season’s top five teams\. Averages are weighted by offensive playing time\.Fig\.[4](https://arxiv.org/html/2608.12926#S5.F4)presents the action\-type decomposition of H\-VAEP/10o aggregated by position group and team\. The left panel reflects the distinct offensive roles of the positions\. Back players generate the majority of their value through passing \(55% for left and right backs, 63% for centre backs\), consistent with their playmaking responsibilities\. Wings instead accumulate most of their value through dribbles \(55%\) and shots \(33%\), as they finish fast breaks after carrying the ball over long distances and attack the crease from the narrow angle of their corner position\. Their passing contributes almost nothing \(3%\), since passes from the wing typically return the ball to the backcourt and rarely improve the attacking situation\. Pivots derive their value from shots \(51%\) and dribbles \(53%\), while their passing contribution is slightly negative \(−4\-4%\)\. A pivot who receives the ball at the six\-meter line and cannot turn toward the goal usually plays the ball backward, which keeps the attack going but hardly improves the scoring opportunity, and playing out of this congested space additionally carries a considerable turnover risk\.
The right panel shows that the value composition within the same position group differs markedly between teams\. SC Magdeburg’s tactical system heavily emphasizes isolation plays and breakthrough runs, and this is clearly reflected in the ratings\. Their back players generate 38% of their value through dribbles, by far the largest share among the top five teams\. SG Flensburg\-Handewitt’s backs show the opposite profile with the highest share of value through passing \(65%\), which results from their emphasis on fast breaks and a system that relies less on isolation plays and breakthroughs\.
\(a\)Fast break\.
\(b\)Positional attack\.
Figure 5:H\-VAEP valuation of two action sequences by SG Flensburg\-Handewitt in the 2024/25 season\. Arrow color encodes an action’s H\-VAEP, line style and end marker its type, and the numbers its order in the sequence\.Finally, we illustrate the face validity of the valuations themselves on two sequences by SG Flensburg\-Handewitt from the 2024/25 season \(Fig\.[5](https://arxiv.org/html/2608.12926#S5.F5)\)\. In the fast break \(Fig\.[5\(a\)](https://arxiv.org/html/2608.12926#S5.F5.sf1)\), H\-VAEP rates the left back’s long pass as the most valuable action \(\+0\.365\+0\.365\) and splits the pivot’s contribution between the dribble that improves the shooting position \(\+0\.220\+0\.220\) and the shot itself \(\+0\.292\+0\.292\)\. The pivot therefore keeps the credit for creating the opportunity even if the shot is missed, in which case only the shot is penalized\. The goalkeeper receives\+0\.141\+0\.141for the outlet pass that starts the break before the defense is set, a contribution that box\-score statistics and the HPI ignore entirely\. The positional attack \(Fig\.[5\(b\)](https://arxiv.org/html/2608.12926#S5.F5.sf2)\) shows that the model separates passes that look alike\. The centre back’s slow ball to the right back is rated neutral \(−0\.007\-0\.007\), the faster return pass slightly positive \(\+0\.053\+0\.053\), and the hard pass into the run of the left back clearly positive \(\+0\.163\+0\.163\) as it initiates the attack\. The left back then turns a contested 9m shooting position into an uncontested wing shot with a diagonal ball \(\+0\.328\+0\.328\)\. In total, the centre back and left back receive more credit for creating the chance \(\+0\.481\+0\.481\) than the wing for finishing it \(\+0\.407\+0\.407\), reflecting how H\-VAEP distributes value across a build\-up sequence\.
## 6Evaluation of Action Valuation Metrics
\(a\)Within\-season\.\(b\)Cross\-season\.\(c\)Meta\-metrics\.
Figure 6:Reliability and meta\-metrics for different player ratings across 2023/24, 2024/25, and 2025/26\. Diamonds in \(a\) and \(b\) indicate the mean correlation\. \(c\) displays the stability and discrimination meta\-metrics from\[[6](https://arxiv.org/html/2608.12926#bib.bib1)\]\.In addition to the demonstrated face validity, we evaluate the quality, reliability, and utility of H\-xT and H\-VAEP using a comprehensive validation framework following established sports analytics methodologies\[[3](https://arxiv.org/html/2608.12926#bib.bib10),[6](https://arxiv.org/html/2608.12926#bib.bib1),[12](https://arxiv.org/html/2608.12926#bib.bib12)\]\. We rate the three held\-out seasons 2023/24 to 2025/26 under the player selection criteria of Section[5](https://arxiv.org/html/2608.12926#S5)\. Each season’s actions are valued by a model fitted only on the preceding season, which mirrors how a club would deploy the models and ensures that no model was trained on the season it values\.
Following\[[3](https://arxiv.org/html/2608.12926#bib.bib10)\], we measure within\-season reliability via the Pearson correlation \(rr\) between two sets of randomly distributed game days \(for players with≥3\\geq 3games in each set\) and cross\-season reliability using consecutive season pairs\. As shown in Figs\.[6\(a\)](https://arxiv.org/html/2608.12926#S6.F6.sf1)and[6\(b\)](https://arxiv.org/html/2608.12926#S6.F6.sf2), traditional counting statistics like goals and assists are highly stable within\-season \(r≥0\.92r\\geq 0\.92, cross\-seasonr≥0\.80r\\geq 0\.80\), whereas shooting accuracy is the most unstable \(r=0\.49r=0\.49within\-season, thoughr=0\.74r=0\.74cross\-season\)\. Our proposed H\-VAEP and H\-VAEP/10o ratings exhibit exceptional reliability \(within\-seasonr≥0\.98r\\geq 0\.98, cross\-seasonr≥0\.96r\\geq 0\.96\)\. This high stability is likely driven by handball’s high\-scoring nature, which yields a high volume of actions, especially assists and goals, and reduces statistical noise compared to football\.
To evaluate player separation, we compute the discrimination and stability meta\-metrics from Franks et al\.\[[6](https://arxiv.org/html/2608.12926#bib.bib1)\]\(Fig\.[6\(c\)](https://arxiv.org/html/2608.12926#S6.F6.sf3)\)\. H\-VAEP/10o performs best, leading in both discrimination \(0\.9940\.994\) and stability \(0\.9990\.999\)\. Without time normalization, absolute H\-VAEP retains high discrimination \(0\.9900\.990\) but lower stability \(0\.9100\.910\)\. Notably, while absolute H\-xT has low single\-game stability \(0\.1640\.164\), normalizing for playing time \(H\-xT/10o\) raises it to0\.9760\.976\. This confirms that adjusting for offensive minutes yields more stable player ratings across all metrics\.
To determine if our metrics capture novel insights, we analyze their correlation with traditional indicators and compute the independence meta\-metric from\[[6](https://arxiv.org/html/2608.12926#bib.bib1)\]\. H\-xT correlates strongly with assists \(r=0\.78r=0\.78, andr=0\.73r=0\.73for H\-xT/10o with assists/10o\), indicating that progression into high\-threat zones \(e\.g\., wing crease jumps\) often leads to assists\. H\-VAEP and H\-VAEP/10o correlate strongly with goals \(r=0\.72r=0\.72,r=0\.66r=0\.66\) and goals/10o \(r=0\.75r=0\.75,r=0\.77r=0\.77\), respectively\. Consequently, both H\-VAEP \(0\.1470\.147\) and H\-VAEP/10o \(0\.1630\.163\) score low on the independence metric \(where lower values denote higher correlation with other indicators\), similar to goals \(0\.1780\.178\)\. Conversely, shooting accuracy \(0\.6250\.625\) and HPI \(0\.3770\.377\) exhibit stronger independence\. For HPI, this is a direct design choice: it aggregates defensive and disciplinary actions \(e\.g\., blocks, steals, suspensions\) that are excluded from our offensive, on\-ball schema\.
Qualitatively, coaches confirmed that our handball\-specific zoning layout is markedly more intuitive than rectangular grids, aligning with their tactical terminology\. Furthermore, discussions during a workshop with coaches from the first and second divisions of the Handball Bundesliga showed that they intuitively grasped how the H\-VAEP framework decomposes and attributes value across multi\-player build\-up sequences, validating its practical utility for player assessment\.
## 7Conclusion and Future Work
In this work, we adapted, optimized, and evaluated the Expected Threat \(H\-xT\) and VAEP \(H\-VAEP\) frameworks for professional team handball\. By addressing sport\-specific requirements, we developed a handball\-native court zoning layout that respects the sport’s unique geometry, which was shown to be systematically more robust than standard rectangular grids\. Furthermore, we customized the feature space of the VAEP framework to accommodate handball’s high\-scoring dynamics and rapid pace, and we selected an optimal context length to balance predictive quality against team\-identity leakage\. Our empirical evaluation across three held\-out seasons of the Handball Bundesliga demonstrated that our models generate stable, reliable, and intuitive player ratings that align with expert assessments, while successfully crediting non\-terminal actions and build\-up play\.
Several avenues remain for future research\. First, while professional coaches highlight the importance of defensive contributions, defensive actions are not yet automatically recognized by tracking systems, requiring methods to extract defensive events from raw tracking data\. Second, handball performance valuation and credit attribution should expand to off\-ball behavior\. Valuing spatial denial in defense or gap creation in offense requires processing player trajectories rather than discrete event data\. Third, although possession\-based prediction targets align closely with handball’s tactical structure, implementing them requires addressing event\-detection noise and false positives that currently prevent the robust identification of possession boundaries\.
## References
- \[1\]\(2023\)Expected Goals Prediction in Professional Handball using Synchronized Event and Positional Data\.InProceedings of the 6th International Workshop on Multimedia Content Analysis in Sports,MMSports ’23,New York, NY, USA,pp\. 83–91\.External Links:ISBN 979\-8\-4007\-0269\-3,[Document](https://dx.doi.org/10.1145/3606038.3616152)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p2.1)\.
- \[2\]D\. Cervone, A\. D’Amour, L\. Bornn, and K\. Goldsberry\(2016\)A Multiresolution Stochastic Process Model for Predicting Basketball Possession Outcomes\.Journal of the American Statistical Association111\(514\),pp\. 585–599\.External Links:ISSN 0162\-1459,[Document](https://dx.doi.org/10.1080/01621459.2016.1141685)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[3\]J\. Davis, L\. Bransen, L\. Devos, A\. Jaspers, W\. Meert, P\. Robberechts, J\. Van Haaren, and M\. Van Roy\(2024\)Methodology and evaluation in sports analytics: challenges, approaches, and lessons learned\.Machine Learning113\(9\),pp\. 6977–7010\(en\)\.External Links:ISSN 1573\-0565,[Document](https://dx.doi.org/10.1007/s10994-024-06585-0)Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p2.1),[§2](https://arxiv.org/html/2608.12926#S2.p2.1),[§6](https://arxiv.org/html/2608.12926#S6.p1.1),[§6](https://arxiv.org/html/2608.12926#S6.p2.1)\.
- \[4\]T\. Decroos, L\. Bransen, J\. Van Haaren, and J\. Davis\(2019\)Actions Speak Louder than Goals: Valuing Player Actions in Soccer\.InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining,KDD ’19,New York, NY, USA,pp\. 1851–1861\.External Links:ISBN 978\-1\-4503\-6201\-6,[Document](https://dx.doi.org/10.1145/3292500.3330758)Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p1.1),[§2](https://arxiv.org/html/2608.12926#S2.p1.1),[§3](https://arxiv.org/html/2608.12926#S3.p2.1),[§4](https://arxiv.org/html/2608.12926#S4.p1.1),[§4](https://arxiv.org/html/2608.12926#S4.p2.1),[§5](https://arxiv.org/html/2608.12926#S5.p1.1)\.
- \[5\]T\. Decroos, L\. Bransen, J\. Van Haaren, and J\. Davis\(2020\)VAEP: an objective approach to valuing on\-the\-ball actions in soccer \(extended abstract\)\.InProceedings of the Twenty\-Ninth International Joint Conference on Artificial Intelligence,IJCAI’20,pp\. 4696–4700\.External Links:[Document](https://dx.doi.org/10.24963/ijcai.2020/648),ISBN 978\-0\-9992411\-6\-5Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p1.1),[§2](https://arxiv.org/html/2608.12926#S2.p1.1),[§4](https://arxiv.org/html/2608.12926#S4.p1.1)\.
- \[6\]A\. M\. Franks, A\. D’Amour, D\. Cervone, and L\. Bornn\(2016\)Meta\-analytics: tools for understanding the statistical properties of sports metrics\.Journal of Quantitative Analysis in Sports12\(4\),pp\. 151–165\.External Links:ISSN 1559\-0410, 2194\-6388,[Document](https://dx.doi.org/10.1515/jqas-2016-0098)Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p2.1),[§2](https://arxiv.org/html/2608.12926#S2.p2.1),[Figure 6](https://arxiv.org/html/2608.12926#S6.F6),[Figure 6](https://arxiv.org/html/2608.12926#S6.F6.3),[§6](https://arxiv.org/html/2608.12926#S6.p1.1),[§6](https://arxiv.org/html/2608.12926#S6.p3.1),[§6](https://arxiv.org/html/2608.12926#S6.p4.1)\.
- \[7\]Handball\-Bundesliga GmbHHandball Performance Index: Datenbasiert, transparent, fair\!\.Note:Accessed 5 Jun 2026External Links:[Link](https://urlhttps//www.daikin-hbl.de/de/hbl/statistiken/handball-performance-index-die-saisonwerte-im-/%C3/%BCberblick/berechnung-des-hpi)Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p1.1),[§2](https://arxiv.org/html/2608.12926#S2.p2.1)\.
- \[8\]H\. Konzag and N\. Sølvkær Schütz\(2024\)Sports digitalization – realizing the potential value of tracking technologies in professional sports organizations\.Hawaii International Conference on System Sciences 2024 \(HICSS\-57\)\.Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p1.1)\.
- \[9\]G\. Liu and O\. Schulte\(2018\)Deep reinforcement learning in ice hockey for context\-aware player evaluation\.InProceedings of the 27th International Joint Conference on Artificial Intelligence,IJCAI’18,Stockholm, Sweden,pp\. 3442–3448\.External Links:ISBN 978\-0\-9992411\-2\-7Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[10\]A\. Mortelier, F\. Rioult, and J\. Komar\(2024\)What Data Should Be Collected for a Good Handball Expected Goal model?\.InMachine Learning and Data Mining for Sports Analytics,U\. Brefeld, J\. Davis, J\. Van Haaren, and A\. Zimmermann \(Eds\.\),Cham,pp\. 119–130\(en\)\.External Links:ISBN 978\-3\-031\-53833\-9,[Document](https://dx.doi.org/10.1007/978-3-031-53833-9%5F10)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p2.1)\.
- \[11\]O\. Müller, M\. Caron, M\. Döring, T\. Heuwinkel, and J\. Baumeister\(2022\)PIVOT: A Parsimonious End\-to\-End Learning Framework for Valuing Player Actions in Handball Using Tracking Data\.InMachine Learning and Data Mining for Sports Analytics,U\. Brefeld, J\. Davis, J\. Van Haaren, and A\. Zimmermann \(Eds\.\),Cham,pp\. 116–128\(en\)\.External Links:ISBN 978\-3\-031\-02044\-5,[Document](https://dx.doi.org/10.1007/978-3-031-02044-5%5F10)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p2.1)\.
- \[12\]P\. Robberechts, Y\. Rudolph, M\. Van Roy, and A\. Zimmermann\(2024\)ECML PKDD Tutorial on Team Sports Analytics\.External Links:[Link](https://urlhttps//dtai.cs.kuleuven.be/tutorials/sports/ecml2024/)Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p2.1),[§2](https://arxiv.org/html/2608.12926#S2.p2.1),[§4](https://arxiv.org/html/2608.12926#S4.p5.1),[§6](https://arxiv.org/html/2608.12926#S6.p1.1)\.
- \[13\]K\. Routley and O\. Schulte\(2015\)A Markov Game model for valuing player actions in ice hockey\.InProceedings of the Thirty\-First Conference on Uncertainty in Artificial Intelligence,UAI’15,Arlington, Virginia, USA,pp\. 782–791\.External Links:ISBN 978\-0\-9966431\-0\-8Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[14\]S\. Rudd\(2011\)A Framework for Tactical Analysis and Individual Offensive Production Assessment in Soccer Using Markov Chains\.InNew England Symposium on Statistics in Sports,\(en\)\.Note:[https://nessis\.org/nessis11/rudd\.pdf](https://nessis.org/nessis11/rudd.pdf)Cited by:[§3](https://arxiv.org/html/2608.12926#S3.p1.1)\.
- \[15\]N\. Sandholtz and L\. Bornn\(2020\)Markov decision processes with dynamic transition probabilities: An analysis of shooting strategies in basketball\.The Annals of Applied Statistics14\(3\),pp\. 1122–1145\(en\)\.External Links:ISSN 1932\-6157, 1941\-7330,[Document](https://dx.doi.org/10.1214/20-AOAS1348)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[16\]T\. Sawczuk, A\. Palczewska, B\. Jones, and J\. Palczewski\(2024\)A Bayesian Mixture Model approach to expected possession values in rugby league\.PLOS ONE19\(11\),pp\. e0308222\(en\)\.External Links:ISSN 1932\-6203,[Document](https://dx.doi.org/10.1371/journal.pone.0308222)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[17\]T\. Sawczuk, A\. Palczewska, and B\. Jones\(2021\)Development of an expected possession value model to analyse team attacking performances in rugby league\.PLOS ONE16\(11\),pp\. e0259536\(en\)\.External Links:ISSN 1932\-6203,[Document](https://dx.doi.org/10.1371/journal.pone.0259536)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[18\]K\. Singh\(2019\)Introducing Expected Threat \(xT\)\.Note:Accessed 30 Mar 2026External Links:[Link](https://arxiv.org/html/2608.12926v1//urlhttps://karun.in/blog/expected-threat.html)Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p1.1),[§2](https://arxiv.org/html/2608.12926#S2.p1.1),[§3](https://arxiv.org/html/2608.12926#S3.p1.1)\.
- \[19\]K\.W\. van Arem, J\. Söhl, M\. Bruinsma, and G\. Jongbloed\(2025\)The trade\-off between model flexibility and accuracy of the Expected Threat model in football\.InMathSport International 2025,D\. Goossens \(Ed\.\),pp\. 150–155\.External Links:ISBN 978\-90\-835814\-0\-8Cited by:[§1](https://arxiv.org/html/2608.12926#S1.p2.1),[Figure 1](https://arxiv.org/html/2608.12926#S3.F1),[Figure 1](https://arxiv.org/html/2608.12926#S3.F1.3),[§3](https://arxiv.org/html/2608.12926#S3.p2.1)\.
- \[20\]A\. Xarles, S\. Escalera, T\. B\. Moeslund, and A\. Clapés\(2025\)Action Valuation in Sports: A Survey\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\) Workshops,pp\. 6132–6142\.Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[21\]H\. Yoshihara, N\. Ding, and K\. Fujii\(2025\)Play Evaluation Based on Predicting the Outcome of Back\-Row Attacks in Volleyball\.InProceedings of the 13th International Conference on Sport Sciences Research and Technology Support,Marbella, Spain,pp\. 29–37\.External Links:ISBN 978\-989\-758\-771\-9,[Document](https://dx.doi.org/10.5220/0013666100003988)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.
- \[22\]R\. Yurko, S\. Ventura, and M\. Horowitz\(2019\)nflWAR: a reproducible method for offensive player evaluation in football\.Journal of Quantitative Analysis in Sports15\(3\),pp\. 163–183\(en\)\.External Links:ISSN 1559\-0410,[Document](https://dx.doi.org/10.1515/jqas-2018-0010)Cited by:[§2](https://arxiv.org/html/2608.12926#S2.p1.1)\.Similar Articles
Multimodal Injury Risk Prediction in Tennis
This paper proposes a multimodal framework called PART for predicting injury risk in tennis players using machine learning on data from wearables, questionnaires, and video analysis, showing strong performance in holistic athlete assessment.
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking
This paper introduces an auditable LLM harness for exact-score football forecasting, integrating dynamic Poisson-family models with LLM-based contextual reasoning across four system iterations, with exploratory results on 2025-26 English Premier League matches.
A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
This paper introduces a benchmark for predicting spreadsheet user actions, addressing challenges in edit history availability and complex action spaces through manual curation and online evaluation methodology.
Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football
Presents Sim2Win, a team-agnostic event-based system for pre-match outcome prediction and tactical profiling in football, achieving 55.4% accuracy on unseen teams.
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
Proposes AR-VLA, an autoregressive action expert that generates continuous action sequences with long-term memory for context-aware robotic policy training, improving trajectory smoothness and task success rates over reactive VLA models.