Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations

Hugging Face Daily Papers Papers

Summary

A framework using distribution-valued firm characteristics and language-model embeddings provides portfolio risk bounds without cross-asset covariance estimates, demonstrating low-variance allocations with Qwen3-Embedding-8B representations.

Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels. We show that firm-level distribution-valued characteristics can instead provide one-sided certificates of portfolio risk. Under maintained links from characteristics to systematic exposures and from exposures to returns, multi-firm Wasserstein-2 dispersion yields a sharp upper bound on systematic portfolio variance and a corresponding bound for standardized returns. A weighted pairwise relaxation produces an objective that is convex under a checkable condition and requires marginal volatility scales but no cross-asset return covariances. With zero firm-specific slack, the common-map scale changes the certified variance reduction but not the normalized allocation, which depends only on observed information geometry. In a 52-firm panel from 2018-2022, an allocation constructed from Qwen3-Embedding-8B news representations lies between the 0.69th and 1.33rd in-sample variance percentiles across four prespecified capped portfolio populations; equal risk weighting lies between the 21.1st and 28.6th percentiles. The lower in-sample variance ranking relative to equal risk also appears across the reported frozen language-model representations. The framework therefore distribution-valued firm information into a coherent risk bound and an implementable allocation rule constructed without cross-asset return covariances.
Original Article
View Cached Full Text

Cached at: 09/03/26, 11:51 AM

Paper page - Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations

Source: https://huggingface.co/papers/2608.29692

Abstract

A framework using distribution-valued firm characteristics and embedding-based representations provides computable upper bounds on portfolio variance and yields low-variance allocations without cross-asset covariance estimates.

Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels. We show that firm-leveldistribution-valued characteristicscan instead provide one-sided certificates of portfolio risk. Under maintained links from characteristics to systematic exposures and from exposures to returns, multi-firm Wasserstein-2 dispersion yields a sharp upper bound on systematic portfolio variance and a corresponding bound for standardized returns. A weighted pairwise relaxation produces an objective that is convex under a checkable condition and requires marginal volatility scales but no cross-asset return covariances. With zero firm-specific slack, the common-map scale changes the certified variance reduction but not the normalized allocation, which depends only on observed information geometry. In a 52-firm panel from 2018-2022, an allocation constructed fromQwen3-Embedding-8Bnews representations lies between the 0.69th and 1.33rd in-sample variance percentiles across four prespecified capped portfolio populations; equal risk weighting lies between the 21.1st and 28.6th percentiles. The lower in-sample variance ranking relative to equal risk also appears across the reportedfrozen language-model representations. The framework therefore distribution-valued firm information into a coherent risk bound and an implementable allocation rule constructed without cross-asset return covariances.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2608\.29692

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.29692 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.29692 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.29692 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Out-of-Distribution Generalization of Risk Aversion in Language Models

arXiv cs.LG

This paper introduces RiskAverseOOD, a benchmark for measuring how well risk aversion learned in low-stakes gambles generalizes to astronomically high-stakes gambles in language models. Initial results show that models like Qwen3-8B can generalize risk aversion partially across 98 orders of magnitude, though not yet reliably enough for a safety failsafe.

Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite

arXiv cs.LG

This paper proposes a generalized distribution-free semi-supervised learning framework that constructs unbiased risk estimators via linear combinations of component risks, extending PNU learning to multiclass classification while achieving lower variance and providing generalization bounds.