An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis
Summary
This exploratory paper evaluates LLM-assisted rewriting of moderate-complexity financial sentences for DisCoCat-based sentiment analysis, finding that prompt-based compression can reduce circuit complexity by over 70% and slightly improve accuracy compared to a low-complexity baseline.
View Cached Full Text
Cached at: 08/10/26, 08:05 AM
# An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis
Source: [https://arxiv.org/html/2608.07439](https://arxiv.org/html/2608.07439)
###### Abstract
Quantum Natural Language Processing \(QNLP\) provides a grammar\-aware framework for text modeling, with Distributional Compositional Categorical \(DisCoCat\) offering one of its theoretically grounded formulations\. However, prior work on financial sentiment analysis has highlighted practical limitations of DisCoCat, including parser sensitivity, high simulation cost, and difficulty handling longer, more complex sentences\. In this paper, we explore an LLM\-assisted preprocessing workflow that uses controlled rewriting strategies to compress, simplify, or decompose moderate\-complexity financial sentiment sentences into more parser\-compatible and circuit\-efficient variants while preserving sentiment\-bearing meaning\. We compare multiple prompting strategies, LLMs, and filtering configurations against the low\-complexity\-only DisCoCat baseline of Stein et al\. At the circuit level, prompt\-based rewriting substantially reduces corpus\-level complexity, with the strongest compression\-based variants reducing average qubit count and gate count by more than 70% relative to the raw moderate\-complexity subset\. Across repeated training runs, GPT\-4\.1\-mini \+ Prompt B achieves the highest observed mean accuracy, reaching0\.550±0\.0350\.550\\pm 0\.035compared with0\.521±0\.0500\.521\\pm 0\.050for the baseline\. We further find that larger training splits do not necessarily yield better downstream performance; across the evaluated configurations, training\-split size had a moderately negative association with accuracy \(Pearsonr=−0\.446r=\-0\.446\)\. Taken together, these results provide exploratory evidence that LLM\-assisted rewriting can make some moderate\-complexity inputs usable within the evaluated DisCoCat configuration, while identifying prompt design, filtering, and circuit\-aware preprocessing as important considerations in the broader effort toward more scalable and utility\-oriented QNLP for financial sentiment analysis\.
## IIntroduction
In financial analysis, predicting market behavior remains a central objective\. Traditional quantitative models often rely on historical prices, trading volume, and technical indicators to identify market trends and price movements\[[10](https://arxiv.org/html/2608.07439#bib.bib20)\]\. However, these models may not fully capture contextual and behavioral signals arising from external events, financial news, social media discussions, online comments, investor emotions, and the frequency and sentiment of public information that can shape stock price movements\[[9](https://arxiv.org/html/2608.07439#bib.bib21),[10](https://arxiv.org/html/2608.07439#bib.bib20)\]\. Sentiment analysis offers a complementary approach by extracting the polarity and tone of financial news and social media discourse, helping analysts better understand how information environments influence investor reactions and subsequent stock price movements\[[1](https://arxiv.org/html/2608.07439#bib.bib2),[20](https://arxiv.org/html/2608.07439#bib.bib3)\]\. Within this context, natural language processing \(NLP\) has become a key tool for deriving quantitative representations from financial texts, including investor and market sentiment, emotional intensity, uncertainty, and domain\-specific contextual cues embedded in news, corporate disclosures, earnings calls, and social media\[[32](https://arxiv.org/html/2608.07439#bib.bib5),[29](https://arxiv.org/html/2608.07439#bib.bib4),[13](https://arxiv.org/html/2608.07439#bib.bib8),[38](https://arxiv.org/html/2608.07439#bib.bib7),[28](https://arxiv.org/html/2608.07439#bib.bib6)\]\.
As NLP becomes increasingly important for financial sentiment analysis, modern transformer\-based models provide a powerful reference point\. These models can generate fluent human\-like text and adapt to new tasks with limited examples, making them useful for extracting and transforming information from domain\-specific text\[[31](https://arxiv.org/html/2608.07439#bib.bib9)\]\. However, these capabilities are closely tied to model scale\. As language models grow larger, their training requires substantial memory, distributed GPU infrastructure, and sophisticated parallelization strategies\[[30](https://arxiv.org/html/2608.07439#bib.bib10)\]\. Thus, the challenge is not only whether NLP can extract useful signals from financial text, but also whether advanced language\-processing methods can remain accessible, repeatable, and computationally feasible in domain\-specific research settings\.
This tension motivates the exploration of complementary paradigms for language processing\. Quantum computing offers one such direction\. Quantum machine\-learning models can, in some cases, require less training data to achieve good generalization\[[5](https://arxiv.org/html/2608.07439#bib.bib11)\], while Quantum Natural Language Processing \(QNLP\) provides a framework in which grammatical structure can be incorporated directly into language representations\[[17](https://arxiv.org/html/2608.07439#bib.bib12),[26](https://arxiv.org/html/2608.07439#bib.bib14),[24](https://arxiv.org/html/2608.07439#bib.bib15)\]\. Rather than positioning QNLP as a direct replacement for transformer\-based NLP, this perspective frames it as a complementary approach with different computational and representational trade\-offs for financial sentiment analysis\[[35](https://arxiv.org/html/2608.07439#bib.bib16)\]\.
Within QNLP, two common approaches are Quantum\-enhanced Long Short\-Term Memory \(QLSTM\) neural networks\[[11](https://arxiv.org/html/2608.07439#bib.bib18),[6](https://arxiv.org/html/2608.07439#bib.bib17)\]and Quantum\-native Distributional Compositional Categorical \(DisCoCat\) models\[[25](https://arxiv.org/html/2608.07439#bib.bib19),[24](https://arxiv.org/html/2608.07439#bib.bib15)\]\. Prior work suggests that QLSTM currently shows stronger performance than DisCoCat in financial sentiment analysis\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]\. However, this comparison should be interpreted with caution, since the DisCoCat experiments in that study were restricted to low\-complexity inputs due to practical runtime limitations\. The distinction is also substantial in terms of sentence length: the low\-complexity data averages 4\.9 words per sentence, whereas the moderate\-complexity data averages 18\.4, making the latter a longer and more challenging synthetic test condition for DisCoCat\-based processing\. These data should not be interpreted as a direct sample of naturally occurring financial language\. This leaves open whether moderate\-complexity financial sentences can be incorporated into DisCoCat\-based learning under current computational constraints\.
Recent work on Large Language Models \(LLMs\) has shown that semantic compression can be used as a preprocessing strategy to reduce input complexity and improve computational feasibility without modifying the underlying model\[[14](https://arxiv.org/html/2608.07439#bib.bib23)\]\. Inspired by this idea, we use LLMs not as end\-task predictors, but as rewriting tools to compress and decompose moderate\-complexity financial sentences into simpler, sentiment\-preserving forms that are more suitable for lambeq\-based DisCoCat training\. This motivates the central question of this paper:how do LLM\-assisted rewriting and screening configurations affect parser validity, circuit complexity, retained training data, and observed downstream sentiment\-classification accuracy for moderate\-complexity financial sentences under current computational constraints?
To address this question, we evaluate an LLM\-assisted preprocessing workflow that transforms moderate\-complexity financial sentences into simpler, parser\-compatible, and sentiment\-preserving forms for DisCoCat\-based training\. The main contributions of this paper are as follows:
- •We introduce an LLM\-assisted rewriting framework for finance\-oriented DisCoCat that compresses, simplifies, or decomposes moderate\-complexity financial sentences into forms that are more suitable for Bobcat parsing and lambeq\-based circuit construction\.
- •We quantify the effect of rewriting on both linguistic and circuit\-level complexity across the corpus, measuring token length, qubit count, circuit depth, and gate count for raw and rewritten moderate\-complexity sentences\.
- •We evaluate the relationship between rewritten inputs and downstream DisCoCat performance using semantic and sentiment\-preservation proxies, parser validity, retained training size, accuracy, and runtime\.
Empirically, LLM\-assisted rewriting can reduce corpus\-level circuit complexity among successfully parsed outputs, with the strongest compression\-based variants reducing average qubit count and gate count by more than 70% relative to the raw moderate\-complexity subset\. At the classification level, GPT\-4\.1\-mini with Prompt B achieves the highest observed mean accuracy among the evaluated variants, reaching0\.550±0\.0350\.550\\pm 0\.035\. Although this value is higher than the0\.521±0\.0500\.521\\pm 0\.050reference result reported by Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\], the difference should be interpreted descriptively because the comparison is not a fully matched causal experiment\. Larger rewritten datasets also did not necessarily produce higher accuracy\. Together, these findings provide an exploratory proof of concept and an initial step toward more scalable and utility\-oriented QNLP for financial sentiment analysis, while identifying trade\-offs among rewriting, filtering, parser compatibility, retained data, circuit complexity, and runtime\. They do not establish utility\-scale deployment, a general accuracy improvement, or an optimal preprocessing strategy\.
The remainder of this paper is organized as follows\. Section[II](https://arxiv.org/html/2608.07439#S2)presents the background and related work\. Section[III](https://arxiv.org/html/2608.07439#S3)describes the evaluated workflow\. Section[IV](https://arxiv.org/html/2608.07439#S4)reports the experimental results\. Section[V](https://arxiv.org/html/2608.07439#S5)discusses the findings and their implications\. Finally, Section[VI](https://arxiv.org/html/2608.07439#S6)concludes the paper\.
## IIBackground and Related Work
This section reviews the literature that motivates the present study from three complementary perspectives\. First, it outlines the DisCoCat framework and the sentence\-to\-circuit pipeline that underpins grammar\-aware QNLP\. Second, it situates the problem within financial sentiment analysis, with emphasis on prior QNLP studies and the practical limitations observed for DisCoCat on more complex inputs\. Third, it reviews LLM\-based semantic compression and rewriting as a possible preprocessing strategy for reducing input complexity while preserving sentiment\-bearing meaning\. Together, these strands of work define the gap addressed in this paper: the use of LLM\-guided rewriting to make moderate\-complexity financial sentences more usable for DisCoCat\-based learning under current computational constraints\.
### II\-ADisCoCat and QNLP
DisCoCat is one of the main frameworks used to represent linguistic meaning compositionally in QNLP\. It combines distributional semantics with categorical grammar such that syntactic reductions determine how lexical meanings are composed into sentence\-level meaning\[[8](https://arxiv.org/html/2608.07439#bib.bib13),[26](https://arxiv.org/html/2608.07439#bib.bib14),[17](https://arxiv.org/html/2608.07439#bib.bib12),[24](https://arxiv.org/html/2608.07439#bib.bib15)\]\. In this setting, grammar is not merely a preprocessing step; it is part of the meaning\-construction process itself\.
Following the standard DisCoCat formulation, a sentence is represented through three stages\. First, it is parsed into a grammatical expression, typically using a categorial or pregroup\-based representation of syntax\. Second, that grammatical structure is converted into a DisCoCat diagram, where syntactic reductions determine how word meanings interact compositionally\. Third, the resulting diagram is mapped into a trainable computational object, such as a tensor\-network representation or a parameterized quantum circuit, by assigning vector\-space or circuit\-level realizations to grammatical types and lexical items\[[8](https://arxiv.org/html/2608.07439#bib.bib13),[26](https://arxiv.org/html/2608.07439#bib.bib14),[24](https://arxiv.org/html/2608.07439#bib.bib15)\]\. This formulation preserves grammatical structure while linking text to trainable representations\.
Figures[1](https://arxiv.org/html/2608.07439#S2.F1)and[2](https://arxiv.org/html/2608.07439#S2.F2)illustrate this pipeline for the example sentence“Amazon profits soar”\. Figure[1](https://arxiv.org/html/2608.07439#S2.F1)shows the parser\-derived diagram and its reduced form after cup removal, while Figure[2](https://arxiv.org/html/2608.07439#S2.F2)shows the corresponding parameterized circuit\. Together, these figures make clear that the circuit ultimately used for learning is determined by the grammatical structure assigned to the sentence\.
This dependency is important for the present study because sentence complexity is not only a linguistic issue but also a computational one\. Since parser output determines diagram structure, and diagram structure determines the size and form of the resulting trainable representation, more complex sentences may be harder to process and may lead to greater computational burden\. This issue is especially relevant in finance\-oriented QNLP, where sentiment\-bearing meaning often depends on modifiers, negation, and entity\-rich phrasing\. As a result, the applicability of DisCoCat depends not only on the framework itself, but also on whether sentence forms can be transformed into parser\-compatible and trainable representations under computational constraints\.

\(a\)Parser\-derived diagram

\(b\)Reduced diagram after cup removal
Figure 1:First two stages of the compilation process for“Amazon profits soar”\.Figure 2:Final stage of the compilation process for“Amazon profits soar”: the reduced diagram mapped into a parameterized ansatz circuit\.
### II\-BSentiment Analysis in Finance with QNLP
Financial sentiment analysis is a challenging application domain because sentiment is often expressed through qualified outlook, contextual framing, and firm\-specific language rather than through overtly positive or negative lexical cues\[[1](https://arxiv.org/html/2608.07439#bib.bib2),[10](https://arxiv.org/html/2608.07439#bib.bib20),[9](https://arxiv.org/html/2608.07439#bib.bib21),[28](https://arxiv.org/html/2608.07439#bib.bib6),[12](https://arxiv.org/html/2608.07439#bib.bib29),[38](https://arxiv.org/html/2608.07439#bib.bib7),[13](https://arxiv.org/html/2608.07439#bib.bib8)\]\. As a result, financial sentiment is used not only as a classification problem in its own right, but also as a signal for downstream tasks such as forecasting, portfolio design, and market\-behavior analysis\.
Within QNLP, sentiment analysis has become a benchmark because it is expressive enough to require compositional reasoning while remaining manageable for current quantum and hybrid models\. DisCoCat\-based studies have used sentiment tasks to examine whether syntax\-sensitive sentence representations can be trained end to end\. Martinez and Leroy\-Meline\[[25](https://arxiv.org/html/2608.07439#bib.bib19)\]extend earlier binary settings to a multiclass sentiment task, while Ruskanda et al\.\[[34](https://arxiv.org/html/2608.07439#bib.bib35)\]show that ansatz design can affect efficiency and performance\. These studies show that sentiment classification is useful not only as an application benchmark, but also as a way to study design choices within QNLP pipelines\.
A related line of work evaluates quantum\-enhanced recurrent or sequence\-based models for sentiment\-oriented tasks\. Di Sipio et al\.\[[11](https://arxiv.org/html/2608.07439#bib.bib18)\]present early QNLP experiments with a quantum\-enhanced LSTM and Transformer \. Chen et al\.\[[6](https://arxiv.org/html/2608.07439#bib.bib17)\]introduce QLSTM as a hybrid quantum\-classical sequence model and report cases of faster convergence or improved accuracy \. Chu et al\.\[[7](https://arxiv.org/html/2608.07439#bib.bib28)\]further develop this line through complex\-valued embeddings combined with a quantum\-enhanced LSTM for sentiment analysis \. These models provide a comparison point because they are sequence\-oriented, whereas DisCoCat remains grammar\-oriented\.
The finance\-specific QNLP reference point for the present work is Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\], who compare DisCoCat and QLSTM in a financial sentiment case study built from more than one thousand sentences \. Their results show that QLSTM was faster to train, whereas DisCoCat remained constrained in practice\. They also distinguish between low\-complexity and moderate\-complexity inputs, with the latter being longer on average and therefore closer to realistic financial language\. This leaves open the question that motivates the present study: whether DisCoCat can be extended beyond low\-complexity inputs toward sentence forms that better reflect financial language under current computational constraints\.
### II\-CLLMs for Semantic Compression and Text Rewriting
The LLM literature most relevant to this paper concerns semantic compression, simplification, and rewriting\. Gilbert et al\. frame semantic compression as approximate compression that preserves semantic content needed for later reconstruction or downstream use\[[15](https://arxiv.org/html/2608.07439#bib.bib30)\]\. Fei et al\.\[[14](https://arxiv.org/html/2608.07439#bib.bib23)\]study semantic compression as a means of reducing input redundancy before downstream processing\. Liskavets et al\.\[[23](https://arxiv.org/html/2608.07439#bib.bib33)\]investigate prompt compression with an emphasis on retaining information relevant to a given task or query\. Across these studies, the common premise is that text can be shortened in a meaning\-aware manner rather than reduced through simple truncation\.
Related work on text simplification and sentence compression provides a more controlled view of this idea\. Juseod\-DO et al\.\[[19](https://arxiv.org/html/2608.07439#bib.bib32)\]proposed InstructCMP, which shows that instruction\-based LLMs can perform sentence compression under explicit length constraints\. Qiang et al\.\[[33](https://arxiv.org/html/2608.07439#bib.bib34)\]benchmark LLMs across lexical, syntactic, sentence, and document simplification, concluding that they outperform earlier non\-LLM approaches across multiple settings\. Guidroz et al\.\[[18](https://arxiv.org/html/2608.07439#bib.bib31)\]further show that LLM\-based simplification can improve comprehension and reduce cognitive load across several domains, including finance\. Although these studies are not focused on QNLP, they support the use of LLMs as rewriting tools rather than only as end\-task predictors\.
More specifically, this rewriting behavior is typically achieved through prompting, that is, through carefully designed textual instructions that guide LLM outputs without requiring task\-specific fine\-tuning\. The growing relevance of prompting as a mechanism for adapting LLMs to new tasks was reinforced by the strong in\-context learning capabilities reported for GPT\-3\[[4](https://arxiv.org/html/2608.07439#bib.bib36)\], which shifted attention from parameter updates toward prompt engineering as a practical means of controlling model behavior\[[27](https://arxiv.org/html/2608.07439#bib.bib38)\]\. Common prompting strategies include zero\-shot prompting, few\-shot prompting, and chain\-of\-thought prompting, all of which can influence how models generalize to new tasks\[[21](https://arxiv.org/html/2608.07439#bib.bib39),[39](https://arxiv.org/html/2608.07439#bib.bib40)\]\. In general, effective prompts specify the task, provide relevant context, and constrain the desired output format\[[16](https://arxiv.org/html/2608.07439#bib.bib37)\]\. Additional techniques such as iterative refinement and role\-based prompting can further improve alignment between model outputs and task objectives\[[22](https://arxiv.org/html/2608.07439#bib.bib41)\]\. For the present study, this literature is relevant not because the LLM serves as the final classifier, but because prompt design provides the mechanism for steering the model toward controlled sentence rewriting\.
For the present study, the key implication is that rewriting can be controlled\. The objective here is narrower than generic summarization: rewrites must remain sentiment\-bearing while also improving parser compatibility\. This places the problem closer to sentence compression and simplification than to open\-ended generation\. Meaning\-preservation metrics such as MeaningBERT\[[3](https://arxiv.org/html/2608.07439#bib.bib24)\]are useful in this context because they estimate whether rewritten sentences remain semantically close to the source\. However, in finance, semantic similarity alone is insufficient, since a rewrite may remain broadly similar while still altering the sentiment signal relevant to classification\. For this reason, LLM rewriting is used here as a preprocessing mechanism for transforming moderate\-complexity financial sentences into forms that remain suitable for downstream DisCoCat\-based learning\.
### II\-DResearch Gap
Taken together, the literature suggests three observations\. First, DisCoCat\-based QNLP provides a grammar\-aware route from text to trainable quantum representations, but its feasibility depends strongly on parser and circuit complexity\[[26](https://arxiv.org/html/2608.07439#bib.bib14),[17](https://arxiv.org/html/2608.07439#bib.bib12),[24](https://arxiv.org/html/2608.07439#bib.bib15)\]\. Second, financial sentiment analysis is an important application domain with structural challenges, since finance text remains domain\-specific even for classical NLP systems\[[28](https://arxiv.org/html/2608.07439#bib.bib6),[12](https://arxiv.org/html/2608.07439#bib.bib29),[38](https://arxiv.org/html/2608.07439#bib.bib7),[13](https://arxiv.org/html/2608.07439#bib.bib8)\]\. Third, LLMs are capable of semantic compression and simplification in ways that preserve much of the source meaning\[[15](https://arxiv.org/html/2608.07439#bib.bib30),[14](https://arxiv.org/html/2608.07439#bib.bib23),[19](https://arxiv.org/html/2608.07439#bib.bib32),[33](https://arxiv.org/html/2608.07439#bib.bib34),[18](https://arxiv.org/html/2608.07439#bib.bib31)\]\.
What these literatures do not yet provide is a unified treatment of rewriting as a preprocessing layer for finance\-oriented DisCoCat\. Existing QNLP sentiment studies focus on sentiment classification itself, ansatz design, or comparisons between QLSTM and DisCoCat\[[25](https://arxiv.org/html/2608.07439#bib.bib19),[34](https://arxiv.org/html/2608.07439#bib.bib35),[36](https://arxiv.org/html/2608.07439#bib.bib22)\]\. Existing LLM compression studies focus on semantic preservation, length control, or simplification quality\[[15](https://arxiv.org/html/2608.07439#bib.bib30),[19](https://arxiv.org/html/2608.07439#bib.bib32),[14](https://arxiv.org/html/2608.07439#bib.bib23),[33](https://arxiv.org/html/2608.07439#bib.bib34)\]\. However, the combination of these objectives—rewriting moderate\-complexity financial sentences into parser\-compatible forms for Bobcat\-based DisCoCat while preserving sentiment\-bearing meaning—remains largely unexplored\.
This gap motivates the present study\. Rather than using LLMs as end\-task classifiers, we use them as rewriting tools within a grammar\-constrained QNLP pipeline\. Rather than asking only whether financial sentiment can be classified with QNLP, we ask whether more realistic financial sentences can be transformed into forms that are usable for DisCoCat\-based learning under current computational constraints\. In this sense, the present study is positioned not merely as an application of LLM\-based rewriting, but as a proof of concept toward more scalable and utility\-oriented QNLP for financial sentiment analysis\.
## IIIMethodology
This section describes the LLM\-assisted preprocessing workflow used to incorporate moderate\-complexity financial sentences into DisCoCat\-based training\. Following Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\], we use a low\-complexity subset as the baseline and a moderate\-complexity subset as the source for rewriting\. Our main experimental comparison contrasts the low\-complexity\-only DisCoCat baseline of Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]with the evaluated augmented setting, in which LLM\-rewritten moderate\-complexity sentences are incorporated into downstream training\. Figure[3](https://arxiv.org/html/2608.07439#S3.F3)summarizes the workflow: we generate rewritten candidates from the moderate\-complexity subset, screen them for meaning preservation, sentiment consistency, and parser validity, and merge accepted outputs with the low\-complexity baseline for downstream DisCoCat training\.
Moderate1,025 sentences18\.4 words/sent\.Low Complexity1,052 sentences4\.9 words/sent\.Input DatasetsPrompt ASemantic\(1\-to\-1\)Prompt BParser\-Compatible\(1\-to\-1\)Prompt CDecomposition\(1\-to\-many\)LLM RewritingMeaningMeaningBERTSimilaritySentimentFinBERTMajority VoteParserValidity CheckScreeningStrategy AParser\-validNo Additional FilterStrategy BFilteredParser\-validMeaning≥α\\geq\\alphaLabel = Ground TruthFiltering\+MergeRemove NeutralDataset SplitTrain/Val/TestDataset PreparationDisCoCat TrainingBaseline vs Augmented
Figure 3:Exploratory workflow for generating, screening, filtering, and evaluating LLM\-assisted rewrites of moderate\-complexity synthetic financial sentences before combining them with the low\-complexity reference condition\.### III\-ASynthetic Financial Sentiment Data
The workflow begins with two synthetic financial sentiment subsets generated through the ChatGPT\-based data generation approach of Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]: a low\-complexity subset and a moderate\-complexity subset\. The low\-complexity subset serves as the baseline condition, whereas the moderate\-complexity subset provides the source sentences for LLM\-guided rewriting\.
Table[I](https://arxiv.org/html/2608.07439#S3.T1)and Figure[4](https://arxiv.org/html/2608.07439#S3.F4)summarize the token\-length characteristics of both subsets\. The low\-complexity subset contains shorter inputs, ranging from 3 to 9 tokens, with a mean of 4\.85 and a median of 5\. In contrast, the moderate\-complexity subset ranges from 10 to 32 tokens, with a mean of 18\.43 and a median of 18\. This distinction is important because the moderate\-complexity subset more closely reflects realistic financial language, while also introducing greater challenges for DisCoCat\-based processing\.
TABLE I:Token\-length statistics for the low\- and moderate\-complexity subsets\.Figure 4:Token distribution for the low\- and moderate\-complexity subsets\. Dashed lines mark means; dotted lines mark medians\.This distinction also affects the complexity of the resulting DisCoCat representations\. Figure[5](https://arxiv.org/html/2608.07439#S3.F5)illustrates this effect for one low\-complexity sentence \(“Amazon profits soar”\) and one moderate\-complexity sentence \(“rise of fintech has disrupted traditional banking and financial services”\), showing that the moderate\-complexity example yields a substantially larger ansatz circuit\. This difference helps motivate the rewriting stage of the proposed framework\.
Figure 5:Illustrative comparison of circuit complexity for one low\-complexity sentence \(“Amazon profits soar”\) and one moderate\-complexity sentence \(“rise of fintech has disrupted traditional banking and financial services”\)\. The moderate\-complexity example requires more qubits, greater circuit depth, and more gates than the low\-complexity example\.
### III\-BLLM\-Guided Rewriting Strategies
To address this gap, we use LLM\-guided rewriting strategies to transform moderate\-complexity financial sentences into forms that better support downstream DisCoCat\-based learning\. We rely on prompting to steer LLMs toward controlled sentence rewriting rather than direct sentiment prediction\. Consistent with prior work on prompt\-based LLM adaptation\[[4](https://arxiv.org/html/2608.07439#bib.bib36),[16](https://arxiv.org/html/2608.07439#bib.bib37)\], we design the prompts to provide clear rewriting instructions, preserve the financial context of the input sentence, and constrain the output format so that the resulting text remains compatible with downstream DisCoCat processing\. In this setting, we do not seek open\-ended generation; instead, we transform moderate\-complexity financial sentences into simpler forms that reduce linguistic complexity while preserving sentiment\-bearing meaning\.
We evaluate three prompting strategies\. Prompt A performs semantic compression in a one\-to\-one setting and produces a shorter sentence while preserving the main financial meaning and sentiment polarity\. Prompt B performs parser\-compatible rewriting in a one\-to\-one setting and adds stronger structural constraints to improve downstream parseability\. Prompt C performs decomposition by mapping one original sentence into one or more shorter independent sentences when the input contains multiple separable semantic units\. Table[II](https://arxiv.org/html/2608.07439#S3.T2)illustrates the three rewriting strategies on the same input sentence\.
TABLE II:Illustrative example of the three LLM\-guided rewriting strategies applied to the same moderate\-complexity input sentence\.Original sentencePrompt A outputPrompt B outputPrompt C outputStock market has been performing exceptionally well, driving up investor confidence\.Stock market performing exceptionally wellStock market has been performing well\.The stock market has been performing exceptionally well\. The stock market is driving up investor confidence\.
### III\-CScreening and Filtering
After rewriting, we screen candidate outputs using three criteria: \(1\) semantic preservation, \(2\) sentiment consistency, and \(3\) parser validity\. We measure semantic preservation with MeaningBERT\[[3](https://arxiv.org/html/2608.07439#bib.bib24)\]by comparing each rewritten output with its original sentence and obtaining a similarity scoresis\_\{i\}\. We evaluate sentiment consistency with FinBERT\[[2](https://arxiv.org/html/2608.07439#bib.bib25)\]by comparing the predicted sentiment of the rewritten output with the original label\. For Prompt C, one original sentence may yield multiple outputs\. In that case, we classify each decomposed output individually and assign the final predicted label to the original sentence by majority vote across its components\. For example, ifSSis decomposed intos1s\_\{1\},s2s\_\{2\}, ands3s\_\{3\}with predicted labels11,11, and22, then the aggregated label forSSis11\. We then compare this aggregated label with the ground\-truth label\. We check parser validity by attempting to parse each rewritten output with BobcatParser\. We mark a candidate as parser\-valid if BobcatParser converts it into a diagram without failure\.
These signals define two dataset\-construction strategies\. Strategy A retains all parser\-valid outputs without additional filtering\. Strategy B applies stricter filtering by retaining only outputs that satisfy parser validity, sentiment consistency, and a MeaningBERT threshold, that is,si≥ts\_\{i\}\\geq t, wheret∈\{40,50,60,70,80\}t\\in\\\{40,50,60,70,80\\\}\. For the final reported experiments, we uset≥60t\\geq 60as an exploratory operating point selected to represent a practical balance between screening score and retained training size\. This threshold was not optimized using held\-out data and is not treated as an optimal value\.
### III\-DDataset Construction and DisCoCat Training
After screening, we merge the accepted rewritten outputs with the original low\-complexity baseline and construct the datasets for downstream DisCoCat training\. Depending on the experimental variant, we train either on the low\-complexity baseline alone or on a merged dataset that combines the low\-complexity baseline with rewritten moderate\-complexity sentences\. Let𝒟L\\mathcal\{D\}\_\{L\}denote the original low\-complexity baseline and let𝒜\(v\)\\mathcal\{A\}^\{\(v\)\}denote the accepted rewritten outputs for variantvv\. The resulting dataset for each variant is defined as
𝒟\(v\)=𝒟L∪𝒜\(v\)\.\\mathcal\{D\}^\{\(v\)\}=\\mathcal\{D\}\_\{L\}\\cup\\mathcal\{A\}^\{\(v\)\}\.\(1\)
In filtered variants,𝒜\(v\)\\mathcal\{A\}^\{\(v\)\}contains only rewrites that satisfy the selected screening conditions, whereas in unfiltered variants it contains all available rewrites\.
Since the downstream task is binary, we remove neutral instances\. For Prompt C, we assign a common group identifier to all decomposed outputs derived from the same original sentence, and we split the data at the group level to prevent information leakage across the training, validation, and test sets\.
For downstream training, we parse sentences with BobcatParser and convert them into parameterized circuits using the DisCoCat configuration adopted in this study\.
### III\-EEvaluation Protocol
We evaluate the proposed framework at both the circuit level and the classification level\. To characterize feasibility\-related differences associated with rewriting of downstream DisCoCat processing, we measure sentence\-level circuit complexity using the number of qubits, circuit depth, and number of gates\. We compare these metrics between raw moderate\-complexity sentences and their rewritten counterparts\. For Prompt C, which may decompose one source sentence into multiple outputs, we distinguish per\-output circuit complexity and, when relevant, aggregate complexity per original source sentence\.
For downstream classification, we evaluate performance at the sentence level using accuracy:
Accuracy=1N∑i=1N𝕀\(y^i=yi\),\\mathrm\{Accuracy\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\(\\hat\{y\}\_\{i\}=y\_\{i\}\),\(2\)whereNNis the number of evaluated instances,yiy\_\{i\}is the ground\-truth label,y^i\\hat\{y\}\_\{i\}is the predicted label, and𝕀\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\.
When grouped outputs are present, as in Prompt C, we also report group\-level accuracy by aggregating sentence\-level predictions back to the original source sentence through majority vote, with confidence\-based tie\-breaking when necessary\.
To contextualize differences in the number of valid trainable instances produced by each preprocessing strategy, we report the data expansion ratio relative to the low\-complexity baseline:
Expansion\(v\)=n\(v\)n0,\\mathrm\{Expansion\}^\{\(v\)\}=\\frac\{n^\{\(v\)\}\}\{n\_\{0\}\},\(3\)wheren\(v\)=\|𝒟\(v\)\|n^\{\(v\)\}=\|\\mathcal\{D\}^\{\(v\)\}\|is the number of valid instances for variantvv, andn0n\_\{0\}is the number of instances in the Stein et al\. baseline\.
Finally, we conduct an exploratory analysis of the association between DisCoCat training\-split size and downstream accuracy across repeated runs\. This training\-split size is distinct from the total dataset sizennused in the expansion\-ratio analysis\. We report Pearson correlation to characterize linear association and Spearman correlation to characterize monotonic association\. This analysis is descriptive rather than causal, since training size is associated with prompt, model, filtering, and retained sample composition\.
### III\-FExperimental Setup
We useGPT\-4\.1\-miniandQwen2\.5:7Bfor rewriting under three prompt settings\. We selected GPT\-4\.1\-mini and Qwen2\.5:7B to provide an exploratory comparison between a commercially hosted model and an open\-weight model under the same rewriting and downstream DisCoCat workflow\. The purpose of this comparison is to examine whether model family and deployment setting are associated with differences in semantic preservation, parser compatibility, circuit complexity, and downstream behavior; it is not intended to establish a general ranking of model quality\. Because the Qwen Prompt A and Prompt C configurations did not complete downstream training, cross\-model conclusions are restricted primarily to the Prompt B comparison\.
For downstream DisCoCat experiments, we use thelambeqpipeline with the local Bobcat parser, anIQPAnsatz,TketModel, andAerBackendin statevector simulation mode\. We split the data into 80% training, 10% validation, and 10% test sets\. We train withQuantumTrainerandSPSAOptimizerfor 100 epochs using batch size 32, and we monitor validation accuracy during training\. For comparison across variants, we report accuracy as the main evaluation metric\. We implement all experiments in Python usinglambeq,pytket,qiskit/Aer,PyTorch, andtransformers\.
We run part of the experiments locally on a 14\-inch MacBook Pro with an Apple M2 Pro chip, a 10\-core CPU, and 16 GB unified memory, and we run the remaining experiments on the Wahab high\-performance computing environment using one NVIDIA V100 GPU, 8 CPU cores, and 192 GB RAM per GPU node\. When available, we use MPS or CUDA acceleration for rewriting, parser checks, and screening\. Each reported configuration was evaluated across six repeated training runs\. We report means, standard deviations, and minimum values descriptively; formal significance tests, confidence intervals, and effect\-size analyses were not performed\.111Code and selected experiment files are available at[https://github\.com/bllin001/qnlp\-discocat\-llms\-finance\-rewriting](https://github.com/bllin001/qnlp-discocat-llms-finance-rewriting)\.
## IVResults
### IV\-AEffect of Rewriting on Text and Circuit Complexity
We first examine whether prompt\-based rewriting reduces the complexity of moderate\-complexity financial sentences before downstream DisCoCat training\. Figure[6](https://arxiv.org/html/2608.07439#S4.F6)shows the token\-count distributions after rewriting for both GPT\-4\.1\-mini and Qwen2\.5:7B\. Prompt A and Prompt B substantially compress the raw moderate\-complexity inputs, shifting their distributions toward shorter lengths that are closer to the low\-complexity setting\. In contrast, Prompt C behaves differently: rather than enforcing strict compression, it decomposes a single moderate\-complexity sentence into multiple shorter outputs\. As a result, its token\-count distribution remains broader and closer to the original moderate\-complexity range\.
Figure 6:Token\-count distributions after prompt\-based rewriting of the raw moderate\-complexity subset\. Prompt A and Prompt B strongly compress sentence length, whereas Prompt C primarily decomposes inputs into multiple shorter sentences\. Dashed vertical lines indicate the mean token count for each distribution\.This pattern is consistent with the decomposition behavior of Prompt C\. GPT Prompt C produces, on average, 2\.72 sentences per input, with a median of 3, whereas Qwen Prompt C produces 2\.27 sentences per input, with a median of 2\. More specifically, GPT decomposes most inputs into 2–3 sentences, while Qwen more often produces 2\-sentence decompositions\. These results indicate that Prompt C reduces structural complexity primarily through segmentation rather than through direct token compression\.
Figure[7](https://arxiv.org/html/2608.07439#S4.F7)extends this analysis to the circuit level at the corpus scale\. Across all three metrics—qubits, circuit depth, and gates—the raw moderate\-complexity sentences produce the largest circuit distributions\. Prompt A and Prompt B generally shift these distributions toward smaller circuits, especially in terms of qubit count and total number of gates\. Prompt C also reduces circuit complexity relative to the raw moderate subset, but its distributions remain broader and typically higher than those of Prompt A and Prompt B\. This reflects the fact that decomposition simplifies individual sentence\-level circuits without necessarily minimizing the total amount of generated text\.
Figure 7:Corpus\-level circuit complexity across rewriting variants\. From left to right, the boxplots show the distributions of qubit count, circuit depth, and gate count for raw moderate\-complexity sentences and their rewritten counterparts\. Boxes indicate interquartile ranges, horizontal lines indicate medians, whiskers extend to the1\.51\.5IQR limits, and diamonds indicate means\.Table[III](https://arxiv.org/html/2608.07439#S4.T3)quantifies these corpus\-level differences\. Because circuit quantities are defined only for successfully parsed outputs, the reported metrics are computed over outputs that yielded valid sentence\-level circuits; parser compatibility itself is analyzed separately in the next subsection\. Relative to the raw moderate subset, all prompt\-based variants reduce the average number of qubits, circuit depth, and gates\. The strongest overall reduction is obtained by GPT\-4\.1\-mini \+ Prompt A, which lowers the average number of qubits from 23\.16 to 5\.60, circuit depth from 12\.52 to 8\.40, and gates from 118\.07 to 27\.29\. This corresponds to reductions of 75\.81% in qubits, 32\.96% in depth, and 76\.88% in gates\. Qwen2\.5:7B \+ Prompt A follows a similar pattern, reducing qubits by 73\.75%, depth by 29\.83%, and gates by 74\.93%\.
The reductions are strongest for qubits and gates, while reductions in circuit depth are more moderate\. This suggests that prompt\-based rewriting is particularly effective at reducing circuit width and total gate count, whereas depth remains more sensitive to the syntactic structures produced by the parser\. Prompt C also reduces average circuit requirements relative to the raw moderate subset, but its reductions are smaller than those of Prompt A and Prompt B, which is consistent with its decomposition\-based behavior rather than direct compression\.
TABLE III:Corpus\-level circuit complexity across rewriting variants\. We report mean±\\pmstandard deviation for qubits, circuit depth, and gates, together with average percentage reduction relative to the raw moderate subset\. Parser compatibility is analyzed separately in Table[IV](https://arxiv.org/html/2608.07439#S4.T4)\.Taken together, these results show that prompt\-based rewriting reduces circuit complexity through two different mechanisms\. Prompt A and Prompt B primarily act as compression strategies, producing shorter texts that lead to substantially fewer qubits and gates\. Prompt C, in contrast, acts as a decomposition strategy: it reduces the complexity of individual sentence\-level circuits but may preserve or increase the aggregate workload per original sentence\. This distinction is important for interpreting downstream DisCoCat performance, since lower per\-circuit complexity and lower total training cost are related but not equivalent objectives\.
### IV\-BSemantic Preservation, Sentiment Fidelity and Parser Compatibility
Table[IV](https://arxiv.org/html/2608.07439#S4.T4)and Figs\.[8](https://arxiv.org/html/2608.07439#S4.F8)–[10](https://arxiv.org/html/2608.07439#S4.F10)summarize rewriting quality across the three screening dimensions\. Prompt C achieves the strongest semantic preservation for both LLM families, with mean MeaningBERT scores above 90, whereas Prompt A and Prompt B remain in the low\-to\-high 60s\. In sentiment fidelity, Qwen Prompt C obtains the highest overall FinBERT agreement, although all variants preserve negative and positive sentiment more reliably than neutral sentiment\. Parser compatibility remains high for most variants, with GPT Prompt B achieving the strongest Bobcat parse success\. Taken together, these results indicate that no single prompt dominates all evaluated screening criteria: Prompt C best preserves meaning and sentiment, whereas Prompt B best supports parseability\.
Figure 8:Distribution of MeaningBERT similarity scores for rewritten outputs\. Diamonds indicate mean scores\. MeaningBERT is used as an automated semantic\-similarity proxy rather than a complete evaluation of semantic faithfulness\.Figure 9:Automated FinBERT sentiment\-fidelity diagnostics for rewritten outputs\. The Negative, Neutral, and Positive columns report per\-class F1 scores; Overall reports the aggregate agreement measure\. Neutral sentiment is preserved less consistently than negative and positive sentiment\.Figure 10:Bobcat parser success rates for rewritten outputs at the row and sentence levels\. Values indicate the percentage of evaluation units successfully converted into valid parser diagrams\.TABLE IV:Summary of rewriting quality across prompting variants\. MB denotes mean MeaningBERT score, FB denotes overall FinBERT agreement, and BP denotes Bobcat parse success\.Figure[11](https://arxiv.org/html/2608.07439#S4.F11)separates the effects of progressively stricter screening conditions\. Panel \(a\) applies the MeaningBERT threshold alone, panel \(b\) additionally requires FinBERT agreement, and panel \(c\) further requires successful Bobcat parsing\. Att≥60t\\geq 60, the full filtering condition retains 46\.93% of GPT Prompt A, 61\.27% of GPT Prompt B, and 72\.20% of GPT Prompt C\. The corresponding Qwen configurations retain 37\.85%, 51\.32%, and 68\.39%, respectively\. We therefore uset≥60t\\geq 60as an exploratory operating point representing a practical balance between screening strictness and retained training size; this threshold was not optimized using held\-out data\.

\(a\)MeaningBERT threshold only

\(b\)MeaningBERT \+ FinBERT agreement

\(c\)MeaningBERT \+ FinBERT \+ Bobcat validity
Figure 11:Training\-set retention under progressively stricter screening conditions\. Panel \(a\) applies the MeaningBERT threshold only; panel \(b\) additionally requires agreement between FinBERT predictions and the ground\-truth label; and panel \(c\) further requires successful Bobcat parsing\. Each cell reports the retained sample count and percentage relative to the original 1,025\-sentence moderate\-complexity subset\.
### IV\-CDisCoCat Classification Performance
Table[V](https://arxiv.org/html/2608.07439#S4.T5)reports the end\-to\-end DisCoCat classification results across training\-set variants\. In addition to accuracy, the table reports the total number of valid training instances produced by each preprocessing strategy, the corresponding data expansion ratio relative to the Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]baseline, and runtime statistics\. These quantities contextualize the results because they show how the evaluated preprocessing configurations change the number of usable training instances and runtime\. They should not be interpreted as evidence that rewriting itself improves sentence quality, because the configurations also differ in filtering criteria, retained sample composition, and prompt structure\.
Among the evaluated configurations, GPT\-4\.1\-mini with Prompt B achieves the highest observed mean accuracy, reaching0\.550±0\.0350\.550\\pm 0\.035, while expanding the training set to1,7061\{,\}706instances \(1\.98×1\.98\\timesthe Stein et al\. reference size\)\. GPT\-4\.1\-mini with filtered Prompt A obtains a similar mean accuracy of0\.549±0\.0260\.549\\pm 0\.026with1,3171\{,\}317instances \(1\.53×1\.53\\times\)\. These configurations illustrate a trade\-off between retained data, observed accuracy, and runtime rather than a general improvement attributable to rewriting\. The comparison with Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]is descriptive because the experimental conditions are not fully matched\.
Figure[12](https://arxiv.org/html/2608.07439#S4.F12)complements Table[V](https://arxiv.org/html/2608.07439#S4.T5)by showing the run\-level distribution of final test accuracy and runtime across configurations\. GPT\-4\.1\-mini with Prompt B and GPT\-4\.1\-mini with filtered Prompt A have the highest observed final test accuracies among the evaluated augmentation configurations\. Filtered Prompt A has lower mean runtime than Prompt B, although both augmentation configurations require more runtime than the Stein et al\. reference condition\. By contrast, Prompt C variants retain more data but incur substantially higher runtime without corresponding gains in downstream accuracy\.
Figure 12:Run\-level distribution of final test accuracy and runtime across configurations\. Boxplots summarize repeated runs, and overlaid points indicate individual iterations\. The figure highlights the trade\-off between downstream classification performance and computational cost across prompting strategies\.Figure[13](https://arxiv.org/html/2608.07439#S4.F13)provides a complementary view of the training dynamics across configurations\. The GPT\-4\.1\-mini Prompt B unfiltered configuration shows among the highest median training and validation accuracies, while GPT\-4\.1\-mini Prompt A filtered and Qwen Prompt B filtered also show comparatively favorable validation distributions\. In contrast, the Prompt C configurations exhibit higher training and validation losses and lower or more variable accuracy distributions, consistent with their lower downstream accuracy and higher runtime in Table[V](https://arxiv.org/html/2608.07439#S4.T5)\. The effect of filtering is configuration\-dependent: filtered variants are associated with higher validation medians for Prompt A and Qwen Prompt B, but this pattern is not uniform across all prompts\. Because the boxplots pool epoch\-level values across repeated runs, this figure describes optimization dynamics rather than providing an independent significance test or replacing the final test\-set comparison\.
Figure 13:Epoch\-level distributions of training and validation accuracy and loss across evaluated configurations\. Boxplots summarize values pooled across repeated runs, and faint points show sampled epoch values\. The figure describes optimization dynamics and is not a final test\-accuracy comparison\.However, the results also indicate that data expansion alone is not sufficient to increase downstream accuracy\. The Stein et al\. reference condition contains860860instances, whereas GPT\-4\.1\-mini with Prompt B contains1,7061\{,\}706instances \(1\.98×1\.98\\times\) and obtains an observed mean accuracy of0\.550±0\.0350\.550\\pm 0\.035\. GPT\-4\.1\-mini with filtered Prompt A contains1,3171\{,\}317instances \(1\.53×1\.53\\times\) and obtains a similar mean accuracy of0\.549±0\.0260\.549\\pm 0\.026\. In contrast, the filtered and unfiltered GPT Prompt C configurations contain2,7372\{,\}737\(3\.18×3\.18\\times\) and3,1393\{,\}139\(3\.65×3\.65\\times\) instances, respectively, but obtain lower mean accuracies of0\.493±0\.0250\.493\\pm 0\.025and0\.465±0\.0170\.465\\pm 0\.017\. Their mean runtimes also increase to450\.4450\.4and503\.7503\.7minutes, compared with59\.759\.7minutes for the Stein et al\. reference condition\. This pattern indicates that increasing the number of rewritten training instances does not necessarily increase accuracy and may substantially increase computational cost\. The observed differences depend jointly on training\-set size, prompt structure, filtering, retained sample composition, and parser compatibility\.
Figure[14](https://arxiv.org/html/2608.07439#S4.F14)summarizes accuracy across the observed DisCoCat training\-split sizes\. Gray points represent individual runs, whereas blue markers and error bars represent group means and standard deviations\. The relationship is non\-monotonic: the largest training\-split groups do not produce the highest accuracy\. Across the 53 runs, the association was moderately negative, with Pearson’sr=−0\.446r=\-0\.446and Spearman’sρ=−0\.342\\rho=\-0\.342\. Because training\-split size is confounded with prompt, LLM, filtering, and retained sample composition, this pattern should be interpreted descriptively rather than causally\.
Figure 14:Observed final test accuracy across DisCoCat training\-split\-size groups\. Gray points represent individual runs, while blue markers and error bars represent group means and standard deviations\. Thennlabels inside the figure indicate the number of repeated runs in each group, not the total dataset size\.Not all Qwen2\.5:7B\-based variants completed successfully in downstream DisCoCat training\. In addition to the two Qwen Prompt B configurations reported in Table[V](https://arxiv.org/html/2608.07439#S4.T5), we also attempted Qwen Prompt A and Qwen Prompt C under both filtered and unfiltered settings\. However, these runs failed with a dimensionality\-relatedValueErrorduring training \(“setting an array element with a sequence”\), indicating an inhomogeneous shape in the batched representation\. We therefore exclude these failed variants from the comparative accuracy and runtime analysis\. This means that the Qwen\-based comparison reported here is limited to the Prompt B setting\.
TABLE V:End\-to\-end comparison of DisCoCat training\-set variants across repeated runs\. We report the total data sizennafter preprocessing and filtering, the data expansion ratio relative to the Stein et al\. baseline \(n0=860n\_\{0\}=860\), accuracy as mean±\\pmstandard deviation and minimum, and runtime as mean±\\pmstandard deviation and minimum \(in minutes\)\. Bold values indicate the highest observed value among the evaluated augmentation configurations\.
## VDiscussion
The results provide exploratory evidence that some moderate\-complexity financial sentences can be transformed into forms that are usable for DisCoCat\-based learning under the current parser, simulator, and training configuration\. Rather than treating LLM rewriting as an end\-task prediction mechanism, this study uses it as a preprocessing layer that reshapes the linguistic input before Bobcat parsing and circuit construction\. In that sense, the contribution is not simply that LLMs can compress text, but that this compression can be operationalized in a way that affects the practical feasibility of downstream QNLP\.
A first important finding is that rewriting quality is multi\-dimensional\. Results across semantic preservation, sentiment fidelity, and parser compatibility show that no single prompting strategy dominates all criteria\. Prompt C achieves the strongest meaning preservation and the highest or near\-highest sentiment agreement, but these proxy scores do not establish complete preservation of propositions, entities, negation, or financial context\. Prompt B achieves the strongest Bobcat parseability, suggesting that parser\-oriented prompting better supports the structural constraints of the DisCoCat pipeline\. Downstream utility therefore depends on both preserving meaning and generating text that can be parsed and compiled reliably\.
A second key finding is that reducing textual complexity and improving downstream performance are related, but not identical, objectives\. The rewriting analysis shows that Prompt A and Prompt B primarily reduce sentence length, whereas Prompt C reduces structural complexity through decomposition\. These strategies therefore simplify the input in different ways\. The corpus\-level circuit analysis further shows that prompt\-based rewriting shifts the distributions of qubits, circuit depth, and gates downward relative to the raw moderate\-complexity subset, confirming that these benefits extend beyond isolated examples\. Among the compression\-based strategies, Prompt A yields the strongest average reductions in circuit complexity, whereas Prompt B provides the strongest parser compatibility and the best average downstream classification performance\. In contrast, Prompt C retains the most meaning and training data under strict filtering, yet does not achieve the best DisCoCat accuracy\. Parser compatibility and downstream accuracy co\-occurred in some evaluated configurations, but the present study does not isolate parser compatibility as a causal mechanism or establish that it is more important than semantic preservation\.
These findings also motivate a more specialized form of decomposition\. In particular, they suggest the potential value of a tone/polarity\-aware decomposition strategy\. A future Prompt D could isolate distinct components of a financial sentence only when divergent or complementary sentiment\-bearing units can be cleanly separated\. From a circuit perspective, such a strategy could yield more compact ansatz circuits by isolating sentiment\-bearing units into smaller parseable segments, potentially reducing gate count and qubit requirements while preserving compositional signal\. At the same time, the present results indicate that such a strategy would need to be selective: when tone and polarity are deeply intertwined, aggressive decomposition may distort the nuanced sentiment structure required for downstream financial classification\. This strategy was not evaluated in the present study and is proposed as a future extension motivated by the trade\-offs observed for Prompt C\.
The filtering results reinforce this interpretation\. Under increasingly strict fidelity constraints, the retained training set shrinks for all variants, but the rate of shrinkage differs substantially across prompts\. Prompt C remains much more robust at high thresholds, which makes it attractive when the primary goal is to preserve a large amount of high\-fidelity training data\. Nevertheless, the final experiments indicate that the most useful operating point is not the most permissive one, nor the most restrictive one\. For the final experiments, we uset≥60t\\geq 60as an exploratory operating point that represents a practical balance between screening score and retained training size\. Because this value was not selected through held\-out optimization, the results do not establish it as an optimal threshold\.
The final classification results show that increasing the number of rewritten training instances did not consistently increase accuracy in the evaluated configurations\. GPT\-4\.1\-mini with Prompt B expands the training set from860860to1,7061\{,\}706instances and obtains an observed mean accuracy of0\.550±0\.0350\.550\\pm 0\.035, but its mean runtime also increases from59\.759\.7to172\.1172\.1minutes\. The Prompt C configurations expand the dataset further, to2,7372\{,\}737and3,1393\{,\}139instances, but obtain lower mean accuracies of0\.493±0\.0250\.493\\pm 0\.025and0\.465±0\.0170\.465\\pm 0\.017while requiring substantially longer runtimes\. These results show that training\-set expansion alone does not guarantee higher accuracy or lower cost\. Because prompt type, filtering, data composition, and training\-set size vary simultaneously, the study does not establish that data quality matters more than quantity\. It only shows that the evaluated configurations exhibit different trade\-offs among retained data, screening measures, accuracy, and runtime\.
These findings position the workflow as an exploratory proof of concept rather than a complete solution\. The GPT\-4\.1\-mini \+ Prompt B configuration obtained a higher observed mean accuracy than the Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]reference result, but the comparison is not sufficient to establish superiority because the experimental conditions are not fully matched\. The observed outcomes also depend on prompt design, filtering decisions, parser behavior, and the current software stack\. The workflow therefore illustrates a possible way to examine the feasibility of processing longer synthetic inputs, while leaving the causal contribution of each component unresolved\.
Several limitations should be noted\. First, the study builds on the synthetic ChatGPT\-generated financial sentiment subsets introduced by Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]\. These data are useful for controlled experimentation but they do not fully capture the diversity, noise, and contextual structure of naturally occurring financial text\. Second, the evaluation focuses primarily on accuracy\. Although accuracy is appropriate for comparison with the prior DisCoCat setting, it does not exhaustively characterize model behavior\.
Third, although the present results now include corpus\-level circuit complexity analysis, retained training size, downstream classification performance, and training runtime, the reported runtimes do not provide a stage\-by\-stage breakdown of rewriting, MeaningBERT screening, FinBERT screening, Bobcat parsing, circuit construction and downstream training\. Decomposition\-based strategies may reduce the size of individual sentence\-level circuits while simultaneously increasing the number of generated outputs\. Lower per\-circuit complexity therefore does not automatically imply lower total pipeline cost\. The training\-metric boxplots also pool epoch\-level values across repeated runs and should be interpreted as descriptive summaries of optimization dynamics rather than independent statistical observations\.
Additional methodological limitations should also be considered\. The study uses six repeated runs per configuration and reports descriptive summaries without formal significance tests, confidence intervals, or effect\-size estimates\. The filtering procedure combines parser validity, FinBERT agreement, and MeaningBERT thresholds, so downstream differences cannot be attributed to rewriting alone\. The study does not include a classical NLP baseline, a strong non\-LLM or rule\-based rewriting control, or an ablation that separates the effects of rewriting from the effects of filtering and parser\-based selection\. MeaningBERT and FinBERT provide useful automated screening signals, but they are proxies rather than complete assessments of proposition preservation, negation, financial\-entity preservation, or human\-perceived faithfulness\. The thresholdt≥60t\\geq 60was used as an exploratory operating point and was not optimized using held\-out validation data\. The current workflow is also one\-pass: outputs that fail Bobcat parsing are excluded, and parser diagnostics are not returned to the LLM for iterative repair\.
Fourth, the comparison across LLM families is not fully symmetric\. Although Qwen2\.5:7B variants for Prompt A and Prompt C were attempted under both filtered and unfiltered settings, these runs failed during downstream DisCoCat training because of a dimensionality\-related batching error\. Consequently, the Qwen\-based downstream comparison is restricted to Prompt B, and conclusions about differences between LLM families should be interpreted cautiously\.
Finally, the conclusions remain tied to the software stack, parser behavior, simulator settings, model versions, and decoding configuration\. The experiments use simulator\-based execution and do not evaluate hardware execution\. The study does not include human judgments of semantic faithfulness or examine overlap between language\-model pretraining data and the synthetic source material\. These factors may affect the generalizability and reproducibility of the results\.
Future work should first establish a stronger comparative and statistical foundation for the observed results\. Controlled experiments should include a classical or non\-LLM rewriting baseline, a rule\-based simplification baseline, and ablations that separately evaluate rewriting, filtering, parser selection, and training\-set composition\. Future studies should also use matched data splits, predefined threshold\-selection procedures, confidence intervals, formal hypothesis tests, and effect\-size estimates\. In particular, the MeaningBERT thresholdttshould be selected using held\-out validation data or evaluated as a predefined hyperparameter rather than selected only from observed retention curves\. These steps would help separate the effects of rewriting from those of filtering, parser selection, and the composition of the retained training sets\.
Once this comparative foundation is established, a second priority is to complete and broaden the evaluation of models and semantic faithfulness\. The dimensionality and batching failures affecting Qwen Prompt A and Prompt C should be diagnosed and resolved before drawing comparisons across model families\. Subsequent experiments should evaluate newer LLM versions and expand the set of evaluated models to include additional hosted and open\-weight systems across different sizes and architectures\. Human evaluation of semantic faithfulness, including entity, negation, polarity, and financial\-event preservation, would complement the MeaningBERT and FinBERT proxy measures\. To support reproducibility, each experiment should record the model version or checkpoint, decoding settings, and relevant software\-library versions\.
A third priority is to characterize computational cost and scalability across the full pipeline\. Future work should separately measure LLM inference, semantic screening, sentiment screening, Bobcat parsing, diagram construction, circuit compilation, and DisCoCat training\. For Prompt C, the analysis should report both per\-output circuit complexity and aggregate workload per original sentence, since smaller individual circuits may still increase the total number of generated outputs\. Evaluation with additional simulators, hardware backends, and updated parser versions would clarify whether the observed trade\-offs depend on the current software configuration\. Future work could also examine whether circuit\-knitting or other distributed quantum\-computing approximations can support larger DisCoCat circuits, including the weak\-coupling approach proposed by Stenger et al\.\[[37](https://arxiv.org/html/2608.07439#bib.bib1)\]\.
A fourth direction, motivated by the parser and circuit\-construction failures observed in the current workflow, is to introduce parser\-in\-the\-loop iterative repair\. When parser failure diagnostics are available, they could be returned to the LLM together with targeted grammatical or structural constraints for a bounded retry\. Each repaired output would then require re\-evaluation for semantic similarity, sentiment consistency, parser validity, circuit compatibility, and downstream cost\. Such an agentic repair loop should be evaluated as a separate preprocessing variant with a fixed retry budget and explicit stopping conditions, since repeated retries could increase runtime or introduce semantic drift\. The goal would be to determine whether parser feedback improves successful compilation without producing unacceptable changes in sentiment\-bearing meaning\.
Finally, the workflow should be evaluated on naturally occurring financial corpora and across multiple financial subdomains\. A specialized Prompt D for tone\- and polarity\-aware decomposition remains a possible extension, but it should be evaluated as a controlled comparison rather than replace the existing prompts\. Together, these directions would help determine which improvements arise from rewriting itself and which arise from filtering, parser compatibility, or changes in the composition of the retained training sets\.
## VIConclusion
This paper presented an exploratory evaluation of LLM\-assisted rewriting for incorporating moderate\-complexity synthetic financial sentences into DisCoCat\-based sentiment analysis\. Starting from the low\- and moderate\-complexity datasets introduced by Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\], we evaluated whether prompt\-based rewriting could produce inputs that remain usable under the current Bobcat, lambeq, and simulator configuration\.
The results show that Prompt A and Prompt B primarily reduce sentence length, whereas Prompt C reduces structural complexity through decomposition\. Among successfully parsed outputs, the rewriting variants reduced several circuit\-complexity measures relative to the raw moderate\-complexity subset\. Screening results also show that the prompts exhibit different trade\-offs across semantic similarity, sentiment agreement, parser validity, and retained data\. These measures should be interpreted as automated proxies and not as complete evidence of meaning preservation\.
For downstream classification, GPT\-4\.1\-mini with Prompt B achieved the highest mean accuracy among the evaluated variants, while GPT\-4\.1\-mini with Prompt A and filtering obtained a similar mean with lower runtime\. However, the comparison with the Stein et al\.\[[36](https://arxiv.org/html/2608.07439#bib.bib22)\]result is descriptive rather than a matched causal evaluation, and the incomplete Qwen Prompt A and Prompt C runs further limit cross\-model conclusions\. Overall, the findings suggest that LLM\-assisted rewriting can be a useful subject for further investigation of parser and circuit feasibility, but they do not establish a general accuracy improvement or an optimal preprocessing strategy\.
Overall, this study provides an exploratory proof of concept and an initial step toward utility\-scale QNLP for financial sentiment analysis by examining whether LLM\-assisted rewriting can make some longer synthetic financial inputs usable within a reproducible DisCoCat workflow\. The results do not establish utility\-scale deployment, a general accuracy improvement, an optimal preprocessing strategy, or a complete solution to DisCoCat scalability\. Instead, they identify trade\-offs among rewriting, filtering, parser compatibility, retained data, circuit complexity, and runtime, while highlighting the preprocessing, evaluation, and scalability requirements that should be addressed through controlled future experiments\.
## Acknowledgment
This research was supported in part by the Richard T\. Cheng Endowment at ODU\. The authors thank Min Dong of ODU ITS for assistance with HPC resources\. We also thank John P\. T\. Stenger of the U\.S\. Naval Research Laboratory for his support\. We thank the reviewers for their constructive feedback; due to time and space constraints, only some suggestions are reflected in this version, with additional revisions planned for a journal version\. The authors acknowledge the use of Google’s Gemini AI during the preparation of this manuscript\. In accordance with IEEE policy, we note that this generative AI system was used strictly as an advanced copyediting and formatting assistant\. All scientific concepts, experimental data, algorithms, and core analytical arguments remain original human\-authored work\. This work was performed using computational facilities at ODU enabled by grants from the National Science Foundation \(MRI grant no\. CNS\-1828593\), the Virginia Commonwealth Technology Research Fund, and Google Cloud Platform through ODU’s Monarch Sphere initiative\. Any subjective views or opinions expressed in this paper do not necessarily represent the views of the national laboratories, the NSF, or the United States Government\.
## References
- \[1\]S\. Ao\(2018\)Sentiment Analysis Based on Financial Tweets and Market Information\.In2018 International Conference on Audio, Language and Image Processing \(ICALIP\),pp\. 321–326\.External Links:[Document](https://dx.doi.org/10.1109/ICALIP.2018.8455771)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1)\.
- \[2\]D\. Araci\(2019\)FinBERT: Financial Sentiment Analysis with Pre\-trained Language Models\.arXiv\.Note:arXiv:1908\.10063 \[cs\]External Links:[Document](https://dx.doi.org/10.48550/arXiv.1908.10063)Cited by:[§III\-C](https://arxiv.org/html/2608.07439#S3.SS3.p1.10)\.
- \[3\]D\. Beauchemin, H\. Saggion, and R\. Khoury\(2023\)MeaningBERT: assessing meaning preservation between sentences\.Frontiers in Artificial Intelligence6\.External Links:ISSN 2624\-8212,[Document](https://dx.doi.org/10.3389/frai.2023.1223924)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p4.1),[§III\-C](https://arxiv.org/html/2608.07439#S3.SS3.p1.10)\.
- \[4\]T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei\(2020\)Language Models are Few\-Shot Learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.External Links:ISBN 9781713829546Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p3.1),[§III\-B](https://arxiv.org/html/2608.07439#S3.SS2.p1.1)\.
- \[5\]M\. C\. Caro, H\. Huang, M\. Cerezo, K\. Sharma, A\. Sornborger, L\. Cincio, and P\. J\. Coles\(2022\)Generalization in quantum machine learning from few training data\.Nature Communications13,pp\. 4919\.External Links:ISSN 2041\-1723,[Document](https://dx.doi.org/10.1038/s41467-022-32550-3)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p3.1)\.
- \[6\]S\. Y\. Chen, S\. Yoo, and Y\. L\. Fang\(2022\)Quantum Long Short\-Term Memory\.InICASSP 2022 \- 2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 8622–8626\.External Links:ISSN 2379\-190X,[Document](https://dx.doi.org/10.1109/ICASSP43922.2022.9747369)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p3.1)\.
- \[7\]Z\. Chu, X\. Wang, M\. Jin, N\. Zhang, Q\. Gao, and L\. Shao\(2024\)An Effective Strategy for Sentiment Analysis Based on Complex\-Valued Embedding and Quantum Long Short\-Term Memory Neural Network\.Axioms13,pp\. 207\.External Links:ISSN 2075\-1680,[Document](https://dx.doi.org/10.3390/axioms13030207)Cited by:[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p3.1)\.
- \[8\]B\. Coecke, M\. Sadrzadeh, and S\. Clark\(2010\-03\)Mathematical Foundations for a Compositional Distributional Model of Meaning\.arXiv\.Note:arXiv:1003\.4394 \[cs\]Comment: to appearExternal Links:[Link](http://arxiv.org/abs/1003.4394),[Document](https://dx.doi.org/10.48550/arXiv.1003.4394)Cited by:[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p1.1),[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p2.1)\.
- \[9\]A\. E\. de Oliveira Carosia, G\. P\. Coelho, and A\. E\. A\. da Silva\(2021\)Investment strategies applied to the Brazilian stock market: A methodology based on Sentiment Analysis with deep learning\.Expert Systems with Applications184,pp\. 115470\.External Links:ISSN 0957\-4174,[Document](https://dx.doi.org/10.1016/j.eswa.2021.115470)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1)\.
- \[10\]S\. Deng, T\. Mitsubuchi, K\. Shioda, T\. Shimada, and A\. Sakurai\(2011\)Combining Technical Analysis with Sentiment Analysis for Stock Price Prediction\.In2011 IEEE Ninth International Conference on Dependable, Autonomic and Secure Computing,pp\. 800–807\.External Links:[Document](https://dx.doi.org/10.1109/DASC.2011.138)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1)\.
- \[11\]R\. Di Sipio, J\. Huang, S\. Y\. Chen, S\. Mangini, and M\. Worring\(2022\)The Dawn of Quantum Natural Language Processing\.InICASSP 2022 \- 2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 8612–8616\.Note:ISSN: 2379\-190XExternal Links:ISSN 2379\-190X,[Document](https://dx.doi.org/10.1109/ICASSP43922.2022.9747675)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p3.1)\.
- \[12\]K\. Du, F\. Xing, R\. Mao, and E\. Cambria\(2024\)Financial Sentiment Analysis: Techniques and Applications\.ACM Comput\. Surv\.56,pp\. 220:1–220:42\.External Links:ISSN 0360\-0300,[Document](https://dx.doi.org/10.1145/3649451)Cited by:[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[13\]K\. Du, Y\. Zhao, R\. Mao, F\. Xing, and E\. Cambria\(2025\)Natural language processing in finance: A survey\.Information Fusion115,pp\. 102755\.External Links:ISSN 1566\-2535,[Document](https://dx.doi.org/10.1016/j.inffus.2024.102755)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[14\]W\. Fei, X\. Niu, P\. Zhou, L\. Hou, B\. Bai, L\. Deng, and W\. Han\(2024\)Extending Context Window of Large Language Models via Semantic Compression\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 5169–5181\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.306)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p5.1),[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1)\.
- \[15\]H\. Gilbert, M\. Sandborn, D\. C\. Schmidt, J\. Spencer\-Smith, and J\. White\(2023\)Semantic Compression with Large Language Models\.In2023 Tenth International Conference on Social Networks Analysis, Management and Security \(SNAMS\),pp\. 1–8\.Note:ISSN: 2831\-7343External Links:ISSN 2831\-7343,[Document](https://dx.doi.org/10.1109/SNAMS60348.2023.10375400)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1)\.
- \[16\]L\. Giray\(2023\)Prompt Engineering with ChatGPT: A Guide for Academic Writers\.Annals of Biomedical Engineering51\(12\),pp\. 2629–2633\.External Links:[Document](https://dx.doi.org/10.1007/s10439-023-03272-4)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p3.1),[§III\-B](https://arxiv.org/html/2608.07439#S3.SS2.p1.1)\.
- \[17\]R\. Guarasci, G\. De Pietro, and M\. Esposito\(2022\)Quantum Natural Language Processing: Challenges and Opportunities\.Applied Sciences12,pp\. 5651\.External Links:ISSN 2076\-3417,[Document](https://dx.doi.org/10.3390/app12115651)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p3.1),[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[18\]T\. Guidroz, D\. Ardila, J\. Li, A\. Mansour, P\. Jhun, N\. Gonzalez, X\. Ji, M\. Sanchez, S\. Kakarmath, M\. M\. Bellaiche, M\. Á\. Garrido, F\. Ahmed, D\. Choudhary, J\. Hartford, C\. Xu, H\. J\. S\. Echeverria, Y\. Wang, J\. Shaffer, Eric, Cao, Y\. Matias, A\. Hassidim, D\. R\. Webster, Y\. Liu, S\. Fujiwara, P\. Bui, and Q\. Duong\(2025\)LLM\-based Text Simplification and its Effect on User Comprehension and Cognitive Load\.arXiv\.Note:arXiv:2505\.01980 \[cs\]External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.01980)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[19\]Juseon\-Do, J\. Kwon, H\. Kamigaito, and M\. Okumura\(2024\)InstructCMP: Length Control in Sentence Compression through Instruction\-based Large Language Models\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 8980–8996\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.532)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1)\.
- \[20\]C\. Ko and H\. Chang\(2021\)LSTM\-based sentiment analysis for stock price forecast\.PeerJ Computer Science7,pp\. e408\.External Links:ISSN 2376\-5992,[Document](https://dx.doi.org/10.7717/peerj-cs.408)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1)\.
- \[21\]T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. Iwasawa\(2022\)Large Language Models are Zero\-Shot Reasoners\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 22199–22213\.External Links:ISBN 9781713871088Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p3.1)\.
- \[22\]Z\. Lin\(2024\)How to write effective prompts for large language models\.Nature Human Behaviour8\(4\),pp\. 611–615\.External Links:[Document](https://dx.doi.org/10.1038/s41562-024-01847-2)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p3.1)\.
- \[23\]B\. Liskavets, M\. Ushakov, S\. Roy, M\. Klibanov, A\. Etemad, and S\. K\. Luke\(2025\)Prompt Compression with Context\-Aware Sentence Encoding for Fast and Improved LLM Inference\.Proceedings of the AAAI Conference on Artificial Intelligence39,pp\. 24595–24604\.External Links:ISSN 2374\-3468,[Document](https://dx.doi.org/10.1609/aaai.v39i23.34639)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p1.1)\.
- \[24\]R\. Lorenz, A\. Pearson, K\. Meichanetzidis, D\. Kartsaklis, and B\. Coecke\(2023\)QNLP in Practice: Running Compositional Models of Meaning on a Quantum Computer\.Journal of Artificial Intelligence Research76,pp\. 1305–1342\.External Links:ISSN 1076\-9757,[Document](https://dx.doi.org/10.1613/jair.1.14329)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p3.1),[§I](https://arxiv.org/html/2608.07439#S1.p4.1),[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p1.1),[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[25\]V\. Martinez and G\. Leroy\-Meline\(2022\)A multiclass Q\-NLP sentiment analysis experiment using DisCoCat\.arXiv\.Note:arXiv:2209\.03152 \[cs\]External Links:[Document](https://dx.doi.org/10.48550/arXiv.2209.03152)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p4.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1)\.
- \[26\]K\. Meichanetzidis, S\. Gogioso, G\. d\. Felice, N\. Chiappori, A\. Toumi, and B\. Coecke\(2021\)Quantum Natural Language Processing on Near\-Term Quantum Computers\.Electronic Proceedings in Theoretical Computer Science340,pp\. 213–229\.External Links:ISSN 2075\-2180,[Document](https://dx.doi.org/10.4204/EPTCS.340.11)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p3.1),[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p1.1),[§II\-A](https://arxiv.org/html/2608.07439#S2.SS1.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[27\]B\. Meskó\(2023\)Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial\.Journal of Medical Internet Research25,pp\. e50638\.External Links:[Document](https://dx.doi.org/10.2196/50638)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p3.1)\.
- \[28\]K\. Mishev, A\. Gjorgjevikj, I\. Vodenska, L\. T\. Chitkushev, and D\. Trajanov\(2020\)Evaluation of Sentiment Analysis in Finance: From Lexicons to Transformers\.IEEE Access8\.External Links:ISSN 2169\-3536,[Document](https://dx.doi.org/10.1109/ACCESS.2020.3009626)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[29\]N\. Mishra, M\. Anzar, S\. Pandey, and S\. Mishra\(2025\)A Review on Stock Market Trends and Stocks Price Prediction Using Sentiment Analysis and Market Data\.In2025 3rd International Conference on Communication, Security, and Artificial Intelligence \(ICCSAI\),Vol\.3,pp\. 56–62\.External Links:[Document](https://dx.doi.org/10.1109/ICCSAI64074.2025.11063888)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1)\.
- \[30\]D\. Narayanan, M\. Shoeybi, J\. Casper, P\. LeGresley, M\. Patwary, V\. Korthikanti, D\. Vainbrand, P\. Kashinkunti, J\. Bernauer, B\. Catanzaro, A\. Phanishayee, and M\. Zaharia\(2021\)Efficient large\-scale language model training on GPU clusters using megatron\-LM\.InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis,pp\. 1–15\.External Links:ISBN 978\-1\-4503\-8442\-1,[Document](https://dx.doi.org/10.1145/3458817.3476209)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p2.1)\.
- \[31\]S\. Nath, A\. Marie, S\. Ellershaw, E\. Korot, and P\. A\. Keane\(2022\)New meaning for NLP: the trials and tribulations of natural language processing with GPT\-3 in ophthalmology\.British Journal of Ophthalmology106,pp\. 889–892\.External Links:ISSN 0007\-1161, 1468\-2079,[Document](https://dx.doi.org/10.1136/bjophthalmol-2022-321141)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p2.1)\.
- \[32\]N\. Patwardhan, S\. Marrone, and C\. Sansone\(2023\)Transformers in the Real World: A Survey on NLP Applications\.Information14,pp\. 242\.External Links:ISSN 2078\-2489,[Link](https://www.mdpi.com/2078-2489/14/4/242),[Document](https://dx.doi.org/10.3390/info14040242)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1)\.
- \[33\]J\. Qiang, M\. Huang, Y\. Zhu, Y\. Yuan, C\. Zhang, and K\. Yu\(2025\)Redefining Simplicity: Benchmarking Large Language Models from Lexical to Document Simplification\.arXiv\.Note:arXiv:2502\.08281 \[cs\]External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.08281)Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1)\.
- \[34\]F\. Z\. Ruskanda, M\. R\. Abiwardani, I\. Syafalni, H\. T\. Larasati, and R\. Mulyawan\(2023\)Simple Sentiment Analysis Ansatz for Sentiment Classification in Quantum Natural Language Processing\.IEEE Access11,pp\. 120612–120627\.External Links:ISSN 2169\-3536,[Document](https://dx.doi.org/10.1109/ACCESS.2023.3327873)Cited by:[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p2.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1)\.
- \[35\]N\. Stamatopoulos, G\. Mazzola, S\. Woerner, and W\. J\. Zeng\(2022\)Towards Quantum Advantage in Financial Market Risk using Quantum Gradient Algorithms\.Quantum6,pp\. 770\.External Links:[Document](https://dx.doi.org/10.22331/q-2022-07-20-770)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p3.1)\.
- \[36\]J\. Stein, I\. Christ, N\. Kraus, M\. B\. Mansky, R\. Müller, and C\. Linnhoff\-Popien\(2023\)Applying QNLP to Sentiment Analysis in Finance\.In2023 IEEE International Conference on Quantum Computing and Engineering \(QCE\),Vol\.02,pp\. 20–25\.External Links:[Document](https://dx.doi.org/10.1109/QCE57702.2023.10178)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p4.1),[§I](https://arxiv.org/html/2608.07439#S1.p8.2),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p4.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p2.1),[§III\-A](https://arxiv.org/html/2608.07439#S3.SS1.p1.1),[§III](https://arxiv.org/html/2608.07439#S3.p1.1),[§IV\-C](https://arxiv.org/html/2608.07439#S4.SS3.p1.1),[§IV\-C](https://arxiv.org/html/2608.07439#S4.SS3.p2.6),[TABLE V](https://arxiv.org/html/2608.07439#S4.T5.15.7.4),[§V](https://arxiv.org/html/2608.07439#S5.p7.1),[§V](https://arxiv.org/html/2608.07439#S5.p8.1),[§VI](https://arxiv.org/html/2608.07439#S6.p1.1),[§VI](https://arxiv.org/html/2608.07439#S6.p3.1)\.
- \[37\]J\. P\. T\. Stenger, D\. Gunlycke, and N\. Chrisochoides\(2026\-06\)Scalable quantum circuit knitting using a weak\-coupling approximation\.arXiv\.Note:arXiv:2606\.19035 \[quant\-ph\]External Links:[Link](http://arxiv.org/abs/2606.19035),[Document](https://dx.doi.org/10.48550/arXiv.2606.19035)Cited by:[§V](https://arxiv.org/html/2608.07439#S5.p15.1)\.
- \[38\]A\. Todd, J\. Bowden, and Y\. Moshfeghi\(2024\)Text\-based sentiment analysis in finance: Synthesising the existing literature and exploring future directions\.Intelligent Systems in Accounting, Finance and Management31,pp\. e1549\.External Links:ISSN 2160\-0074,[Document](https://dx.doi.org/10.1002/isaf.1549)Cited by:[§I](https://arxiv.org/html/2608.07439#S1.p1.1),[§II\-B](https://arxiv.org/html/2608.07439#S2.SS2.p1.1),[§II\-D](https://arxiv.org/html/2608.07439#S2.SS4.p1.1)\.
- \[39\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-Thought Prompting Elicits Reasoning in Large Language Models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24824–24837\.External Links:ISBN 9781713871088Cited by:[§II\-C](https://arxiv.org/html/2608.07439#S2.SS3.p3.1)\.Similar Articles
Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing
This paper investigates using LLMs to rewrite fragmentary dialogue utterances for improving frozen discourse parsers, finding that zero-shot clarification is unreliable and that error repair through rewriting has a practical ceiling, suggesting rewritability prediction as a key missing capability.
Enhancing Financial Sentiment Analysis via Retrieval Augmented Large Language Models
This paper introduces a retrieval-augmented LLM framework for financial sentiment analysis, achieving 15-48% improvement in accuracy and F1 score over traditional models and LLMs like ChatGPT and LLaMA.
How Robust Is Multimodal Claim Verification to LLM Rewriting?
This paper studies how multimodal claim verification models respond to stylistic text changes induced by LLM rewriting. Evaluating 11 open-weight VLMs (2B–38B), the authors find accuracy is largely robust to natural rewriting and controlled LLM-word injection, though hedging-oriented modifications cause consistent probability shifts across nearly all models.
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs
Proposes a tree-of-thoughts inspired extractive-abstractive approach for legal case judgement summarization using LLMs, with experiments on DeepSeek and LLama showing improved summaries over extractive or abstractive methods alone.
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
This paper examines how LLM-based compression of financial documents can distort investment decisions by losing contextual qualifiers and introducing model-dependent biases, proposing Agentic Context Compression to audit disagreements against the original source.