BCL: Bayesian In-Context Learning Framework for Information Extraction
Summary
BCL is the first optimization framework that uses particle filtering with Bayesian updates to systematically refine label representations for information extraction tasks, showing consistent improvements over existing methods.
View Cached Full Text
Cached at: 06/18/26, 05:45 AM
# BCL: Bayesian In-Context Learning Framework for Information Extraction
Source: [https://arxiv.org/html/2606.18620](https://arxiv.org/html/2606.18620)
Haoliang Liu1Chengkun Cai211footnotemark:1Xu Zhao311footnotemark:1Han Zhu4Shizhou Huang5 Xinglin Zhang6Tao Chen7Jenq\-Neng Hwang8Zhang Huaping9Lei Li9
###### Abstract
Existing information extraction \(IE\) tasks increasingly adopt in\-context learning \(ICL\) with large language models\. However, current approaches either show inconsistent performance across model scales or lack systematic optimization and generalizability\. Building on this, we propose BCL \(Bayesian In\-Context Learning Framework for Information Extraction\), the first optimization framework that uses particle filtering with Bayesian updates to systematically refine label representations across IE tasks\. Through four steps—initialization, observation, weight update, and resampling, BCL generalizes to both sequence labeling and relation classification paradigms\. Extensive experiments demonstrate substantial and consistent improvements over existing approaches\.
BCL: Bayesian In\-Context Learning Framework for Information Extraction
Haoliang Liu1††thanks:Equal contribution\.Chengkun Cai211footnotemark:1Xu Zhao311footnotemark:1Han Zhu4Shizhou Huang5††thanks:Important contribution\.Xinglin Zhang6Tao Chen7Jenq\-Neng Hwang8Zhang Huaping9Lei Li9††thanks:Corresponding authors:lilei@bit\.edu\.cn
††footnotetext:1HiThink Research2University College London
3University of Edinburgh
4The Hong Kong University of Science and Technology
5East China Normal University
6Shanghai Medical Image Insights
7University of Waterloo8University of Washington
9Beijing Institute of Technology## 1Introduction
Recent IE tasks rely on in\-context learning \(ICL\), where large language models \(LLMs\)\(Brownet al\.,[2020](https://arxiv.org/html/2606.18620#bib.bib43)\)are guided by contextual information\. Recent approaches can be broadly categorized into task transfer approaches that reformulate information extraction \(IE\) as auxiliary tasks \(e\.g\., ChatIE\(Weiet al\.,[2023](https://arxiv.org/html/2606.18620#bib.bib76)\), CodeIE\(Liet al\.,[2023a](https://arxiv.org/html/2606.18620#bib.bib67)\)\) and guideline\-based approaches that provide explicit annotation guidelines \(e\.g\., GuideNER\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\)\)\.
Figure 1:Comparison of previous approaches and our work\.Top:Previous methods convert IE tasks to leverage models’ stronger code or chat capabilities, making performance dependent on model\-specific strengths\.Bottom:Our approach directly improves IE performance via automatically generated semantic patterns, regardless of models’ relative strengths across task types\.However, existing approaches face practical limitations\. As illustrated in Figure[1](https://arxiv.org/html/2606.18620#S1.F1)\(top\), task transfer methods show inconsistent performance across model scales\. While potentially effective on ultra\-large commercial models, they often underperform direct IE prompting on smaller models\. ChatIE underperforms one\-shot prompting by a substantial margin on NER tasks, and CodeIE fails on RE tasks with near\-zero micro\-F1\. This inconsistency makes deployment challenging when using lightweight models, which are common in practical settings due to computational constraints\.
Guideline\-based approaches offer an alternative to task transfer methods, but existing work has critical limitations\. GuideNER\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\), the current state\-of\-the\-art, has significant limitations\. First, it uses simple frequency\-based selection without systematic optimization for guideline quality\. Second, it is designed specifically for NER and does not extend to other IE tasks, as evidenced by the absence of RE results in Figure[2](https://arxiv.org/html/2606.18620#S1.F2)\. These limitations motivate the need for a more general and optimized approach\.
Figure 2:Performance comparison of different methods on Qwen\-2\.5\-7B across NER and RE tasks\. The Y\-axis represents F1 score \(%\)\. BCL demonstrates consistent superiority over baseline methods, while ChatIE and CodeIE show substantial degradation on both task types\. GuideNER is applicable only to NER tasks\.Building on these observations, we introduce automatic subcategory generation \(Figure[1](https://arxiv.org/html/2606.18620#S1.F1), bottom\) that decomposes labels into semantically discrete atomic representations\. The key insight is that IE labels are often coarse\-grained: a "Person" label in NER could mean family roles like "father" or "friend" to the model’s prior understanding, while in a specific dataset it may only refer to public figures such as "athlete" or "politician"\. To bridge this gap between the model’s prior knowledge and the dataset’s annotation schema, we represent each label using multiple subcategories as atomic representations, with these subcategory patterns serving as rules to clarify the label’s specific meaning in context\. For example, in NER, "Person" can be represented by subcategories such as "athlete" and "public figure"; in RE, "\[FRESNO, Located\-In, Calif\]" can be decomposed into "\[city, spatial\-containment, state\]" and "\[subregion, asymmetry, super\-region\]"\. Crucially, by discretizing labels into semantic atomic units, we can treat them as controllable discrete variables\. This enables us to optimize label representations using optimization algorithms such as particle filtering, where each rule is a particle with an associated weight, refined through iterative evaluation and Bayesian updates\.
Figure 3:Overview of the BCL framework\. The framework operates through particle filtering with Bayesian updates, alternating between observation and control to progressively optimize the semantic patterns distribution\.We introduceBCL \(Bayesian In\-Context Learning Framework for Information Extraction\), optimizing subcategory patterns through four steps \(Figure[3](https://arxiv.org/html/2606.18620#S1.F3)\): \(1\) initialization—generate initial patterns and set prior weights, \(2\) observation—evaluate via ICL\-based IE to compute likelihoods, \(3\) weight update—refine weights through Bayesian update, \(4\) resampling—eliminate low\-weight particles and diversify high\-performers via LLM mutation\.
Our contributions are:
- •We introduce the key insight of treating context as controllable discrete variables, achieved by decomposing labels into fine\-grained semantic units, enabling systematic optimization methods to be applied\.
- •We develop the first optimization framework using particle filtering with Bayesian updates that generalizes across IE tasks, achieving systematic quality improvement on both sequence labeling and relation classification paradigms\.
- •Extensive experiments demonstrate substantial improvements over existing approaches \(up to 30%\), achieving strong performance while other methods either fail to generalize or show limited effectiveness\.
## 2Related Work
### 2\.1In\-Context Learning for Information Extraction
In\-context learning \(ICL\) enables large language models to adapt to new tasks through demonstration examples without parameter updates\(Brownet al\.,[2020](https://arxiv.org/html/2606.18620#bib.bib43); Weiet al\.,[2022](https://arxiv.org/html/2606.18620#bib.bib63); Minet al\.,[2022](https://arxiv.org/html/2606.18620#bib.bib64)\)\. Traditional ICL approaches for information extraction rely on example\-based demonstrations, where models learn input\-output mappings through pattern recognition\(Donget al\.,[2022](https://arxiv.org/html/2606.18620#bib.bib28); Liet al\.,[2023b](https://arxiv.org/html/2606.18620#bib.bib52)\)\.
Recent work improves ICL for IE by refining demonstration construction, retrieval, and filtering strategies\. C\-ICL\(Moet al\.,[2024](https://arxiv.org/html/2606.18620#bib.bib73)\)incorporates both positive and hard negative examples into demonstrations, while G&O\(Liet al\.,[2024b](https://arxiv.org/html/2606.18620#bib.bib74)\)decomposes generation into intermediate reasoning and structured outputs to improve stability\. GuideNER\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\)replaces demonstrations with LLM\-generated annotation guidelines, and Dr\.ICL\(Luoet al\.,[2024](https://arxiv.org/html/2606.18620#bib.bib31)\)retrieves task\-relevant examples to enhance reasoning performance\. Similarly, MAPS\(Chenet al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib32)\)introduces anchor\-based sampling for fine\-grained entity linking, while recent LLM\-based feature selection methods\(Wanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib34)\)further highlight the importance of iterative filtering for structured extraction\. Related observations also appear in adjacent multimodal understanding settings: Human Motion Instruction Tuning\(Liet al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib35)\)and Multiple Human Motion Understanding\(Liet al\.,[2026](https://arxiv.org/html/2606.18620#bib.bib33)\)show that carefully designed instruction and structured semantic supervision can improve complex motion understanding, suggesting that input organization and guidance are broadly important for structured prediction\.
Beyond demonstration design, recent studies analyze the intrinsic mechanisms of ICL and context utilization\.Shiet al\.\([2026](https://arxiv.org/html/2606.18620#bib.bib29)\)study entropy in context length scaling, whileCaiet al\.\([2025b](https://arxiv.org/html/2606.18620#bib.bib30)\)examine the roles of deductive and inductive reasoning\.Lanet al\.\([2025](https://arxiv.org/html/2606.18620#bib.bib39)\)further propose attention consistency to estimate token importance, providing insights into how models utilize demonstrations during inference\.
For relation extraction, ICL faces challenges in modeling inter\-entity dependencies and contextual patterns\. GPT\-RE\(Wanet al\.,[2023](https://arxiv.org/html/2606.18620#bib.bib65)\)retrieves task\-aware demonstrations with label\-guided reasoning, whileLiet al\.\([2024a](https://arxiv.org/html/2606.18620#bib.bib75)\)propose a recall–retrieve–reason framework to enhance retrieval and reasoning\.Wadhwaet al\.\([2023](https://arxiv.org/html/2606.18620#bib.bib66)\)highlight performance variance across prompts, and CodeIE\(Liet al\.,[2023a](https://arxiv.org/html/2606.18620#bib.bib67)\)reformulates IE as code generation but remains sensitive to demonstration quality\.
Recent studies show that LLMs can perform structured reasoning in complex settings\. CountLLM\(Yaoet al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib40)\)highlights structured dependency modeling, while other work explores retrieval–reasoning in multi\-hop QA\(Jiet al\.,[2026](https://arxiv.org/html/2606.18620#bib.bib58)\), few\-shot generalization without explicit meta\-learning\(Guanet al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib42); Guan,[2025](https://arxiv.org/html/2606.18620#bib.bib53)\), and structured context in visually grounded retrieval\-augmented generation\(Jiet al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib38)\)\. Together, these findings highlight the importance of context utilization\.
### 2\.2Control\-Theoretic and Probabilistic Optimization
Classical control theory\(Åström and Murray,[2021](https://arxiv.org/html/2606.18620#bib.bib10)\)models complex systems as input\-output mappings governed by feedback mechanisms, where external control variables can systematically steer system behavior without directly observing internal states\. Particle filtering\(Gordonet al\.,[1993](https://arxiv.org/html/2606.18620#bib.bib8)\)and sequential Monte Carlo methods\(Doucetet al\.,[2001](https://arxiv.org/html/2606.18620#bib.bib3)\)estimate latent states in high\-dimensional nonlinear systems via population\-based sampling and importance resampling\. In black\-box optimization, Bayesian optimization\(Frazier,[2018](https://arxiv.org/html/2606.18620#bib.bib7); Xuet al\.,[2026](https://arxiv.org/html/2606.18620#bib.bib21)\)builds probabilistic surrogate models with acquisition functions to guide sampling, while Approximate Bayesian Computation\(Beaumontet al\.,[2002](https://arxiv.org/html/2606.18620#bib.bib6); Liu,[2026](https://arxiv.org/html/2606.18620#bib.bib22)\)enables likelihood\-free inference for complex models\. These approaches share a common principle: optimizing system behavior via input\-output observations without access to internal mechanisms\. Evolutionary prompt optimization\(Qiet al\.,[2024](https://arxiv.org/html/2606.18620#bib.bib54)\)applies population\-based search to LLM behavior, but lacks systematic control\-theoretic grounding and focuses on reasoning tasks rather than structured prediction\. In contrast, our work integrates control\-theoretic principles with sequential Monte Carlo methods to optimize demonstration selection in few\-shot learning\.
Recent work begins to apply such principles to controlling LLM behavior\. For example,Caiet al\.\([2025a](https://arxiv.org/html/2606.18620#bib.bib41)\)leverage Bayesian optimization to steer LLM\-driven image editing processes under black\-box settings, demonstrating the effectiveness of probabilistic search for controllable generation\.
### 2\.3Optimization Approaches for LLM Behavior
Various optimization strategies have been explored for LLM behavior control\(Zhaoet al\.,[2026](https://arxiv.org/html/2606.18620#bib.bib44); Cao and Zhao,[2025](https://arxiv.org/html/2606.18620#bib.bib46)\)\. Fine\-tuning\(Weiet al\.,[2021](https://arxiv.org/html/2606.18620#bib.bib49)\)requires substantial resources and labeled data, limiting few\-shot applicability\. Prompt engineering\(Zhouet al\.,[2022](https://arxiv.org/html/2606.18620#bib.bib13); Pryzantet al\.,[2023](https://arxiv.org/html/2606.18620#bib.bib9)\)relies on manual effort or local search heuristics, while evolutionary algorithms\(Qiet al\.,[2024](https://arxiv.org/html/2606.18620#bib.bib54)\)explore prompt spaces but lack systematic guideline optimization for structured prediction\. Existing methods rely on heuristic strategies without control\-theoretic grounding\(Zhaoet al\.,[2021](https://arxiv.org/html/2606.18620#bib.bib12)\)\. Although RLHF\(Ouyanget al\.,[2022](https://arxiv.org/html/2606.18620#bib.bib11)\)and preference optimization\(Rafailovet al\.,[2023](https://arxiv.org/html/2606.18620#bib.bib2)\)address alignment, they modify model parameters rather than optimizing external control inputs like demonstration selection rules\.
## 3Methodology: BCL
Our BCL framework consists of a comprehensive control\-theoretic approach for rule optimization, as illustrated in Figure[3](https://arxiv.org/html/2606.18620#S1.F3)and Figure[4](https://arxiv.org/html/2606.18620#S3.F4)\. Figure[3](https://arxiv.org/html/2606.18620#S1.F3)shows the algorithmic overview with the four key steps of our adaptive filtering process, while Figure[4](https://arxiv.org/html/2606.18620#S3.F4)presents the particle\-based optimization pipeline with iterative generation, evaluation, selection, and mutation phases\.
Figure 4:Overall framework of BCL showing the particle\-based rule optimization pipeline with iterative generation \(Particle Generator\), evaluation\(Posterior Probability Calculator\), selection \(Retain\), and mutation \(Resampler\) phases guided by LLM performance feedback\.### 3\.1Problem Formulation
Given a pre\-trained large language modelℳ\\mathcal\{M\}and a target dataset𝒟=\{\(xi,yi\)\}i=1N\\mathcal\{D\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}with training split𝒟train\\mathcal\{D\}\_\{train\}, development split𝒟dev\\mathcal\{D\}\_\{dev\}and test split𝒟test\\mathcal\{D\}\_\{test\}, our goal is to find an optimal rule listℛ∗\\mathcal\{R\}^\{\*\}that maximizes the model’s information extraction performance on the development set, and then evaluate its generalization performance on the test set\.
#### 3\.1\.1Dataset\-Specific Optimization
Since different datasets have varying data distributions and annotation conventions, we need to adapt the rule selection to each specific dataset\. Using the training portion𝒟train\\mathcal\{D\}\_\{train\}, we extract candidate rules, and then optimize the rule listℛ∗\\mathcal\{R\}^\{\*\}based on performance on the development set𝒟dev\\mathcal\{D\}\_\{dev\}\. This ensures that the final rule list is tailored to both the dataset characteristics and the specific LLM’s behavior\.
#### 3\.1\.2ICL\-based Extraction
For any inputxx, the extraction process follows:
y^=ℳ\(Prompt\(x,Rt\)\)\\displaystyle\\hat\{y\}=\\mathcal\{M\}\(\\text\{Prompt\}\(x,R\_\{t\}\)\)\(1\)whereℳ\\mathcal\{M\}is the pre\-trained LLM,Prompt\(x,Rt\)\\text\{Prompt\}\(x,R\_\{t\}\)constructs the input prompt by combining textxxwith the optimal rule configurationRtR\_\{t\}, andy^\\hat\{y\}is the predicted extraction output\.
#### 3\.1\.3Optimization Objective
We seek to find the optimal rule configuration that maximizes performance on unseen data\. Following standard machine learning practice, we split the available training data into training and validation sets, and optimize rules based on validation performance:
Rt∗\\displaystyle R\_\{t\}^\{\*\}=argmaxRt1\|𝒟val\|∑\(x,y\)∈𝒟valF\(y^,y\)\\displaystyle=\\operatorname\*\{arg\\,max\}\_\{R\_\{t\}\}\\frac\{1\}\{\|\\mathcal\{D\}\_\{\\text\{val\}\}\|\}\\sum\_\{\(x,y\)\\in\\mathcal\{D\}\_\{\\text\{val\}\}\}F\(\\hat\{y\},y\)\(2\)
whereRt=\{\(pi\(t\),ci,wi\(t\)\)\}i=1NR\_\{t\}=\\\{\(p\_\{i\}^\{\(t\)\},c\_\{i\},w\_\{i\}^\{\(t\)\}\)\\\}\_\{i=1\}^\{N\}represents a rule configuration withNNparticles at iterationtt, wherepi\(t\)p\_\{i\}^\{\(t\)\}is theii\-th subcategory pattern \(particle\),cic\_\{i\}is its corresponding entity label, andwi\(t\)w\_\{i\}^\{\(t\)\}is the confidence weight,F\(⋅,⋅\)F\(\\cdot,\\cdot\)is a performance metric \(e\.g\., F1 score\)\. The rule extraction and initial population are derived from𝒟train\\mathcal\{D\}\_\{\\text\{train\}\}, while the optimization objective is evaluated on the held\-out validation set𝒟val\\mathcal\{D\}\_\{\\text\{val\}\}to prevent overfitting\. Final evaluation is performed on𝒟test\\mathcal\{D\}\_\{\\text\{test\}\}\.
#### 3\.1\.4Challenges with Existing Approaches
While existing rule\-based ICL methods like GuideNER have demonstrated the effectiveness of annotation guidelines over examples, they lack systematic optimization strategies for rule selection and combination\. Current approaches rely on heuristic frequency\-based filtering, which fails to capture the complex interdependencies between rules and their collective impact on IE performance\.
#### 3\.1\.5Control\-Theoretic Reformulation
Traditional LLM optimization faces a fundamental challenge: the massive parameter space \(billions of parameters\) renders direct observation and control intractable\. We address this by treating rules as a low\-dimensional, observable interface to the LLM system\.
Recent findings suggest that in\-context learning operates as a rule\-based inference system, where the quality of rules—not examples—determines performance\. This motivates reformulating the optimization as a control system problem, where rules serve as externally controllable state variables:
Rt\+1\\displaystyle R\_\{t\+1\}=f\(Rt,ut\)\(Rule Evolution\)\\displaystyle=f\(R\_\{t\},u\_\{t\}\)\\quad\\text\{\(Rule Evolution\)\}\(3\)yt\\displaystyle y\_\{t\}=h\(Rt\)\(Performance Observation\)\\displaystyle=h\(R\_\{t\}\)\\quad\\text\{\(Performance Observation\)\}\(4\)whereRtR\_\{t\}represents the rule configuration at iterationtt,utu\_\{t\}denotes control actions \(rule modifications\), andyty\_\{t\}is the observed IE performance\.
### 3\.2Adaptive Rule Filtering Algorithm
The discrete, combinatorial nature of rule spaces and nonlinear performance mappings makes classical control methods inappropriate\. We develop an adaptive filtering approach that iteratively estimates optimal rule configurations through performance feedback\.
As illustrated in Figure[3](https://arxiv.org/html/2606.18620#S1.F3), our approach follows a systematic particle filtering process:
#### 3\.2\.1Particle\-based State Representation
Terminology\.To clarify key concepts used throughout this section:
- •Particle: A subcategory pattern paired with its entity label, e\.g\., \(“athlete”, Person\)
- •Rule: A particle with its associated confidence weight in the population
- •Weight: Normalized probabilitywi∈\[0,1\]w\_\{i\}\\in\[0,1\]indicating particle quality, where∑iwi=1\\sum\_\{i\}w\_\{i\}=1
At time steptt, the complete rule configuration is:
Rt=\{\(pi\(t\),ci,wi\(t\)\)\}i=1N\\displaystyle R\_\{t\}=\\\{\(p\_\{i\}^\{\(t\)\},c\_\{i\},w\_\{i\}^\{\(t\)\}\)\\\}\_\{i=1\}^\{N\}\(5\)wherepi\(t\)p\_\{i\}^\{\(t\)\}is theii\-th subcategory pattern \(particle\),cic\_\{i\}is its corresponding entity label, andwi\(t\)w\_\{i\}^\{\(t\)\}is the confidence weight\.
The filtering process consists of four steps:
0\. Initialization \(Rule Extraction\):We initialize the particle population by extracting initial rules from the training dataset using the GuideNER approach\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\)\. For each input\-label pair\(xj,yj\)\(x\_\{j\},y\_\{j\}\)in the training set, we use the LLM to summarize rule patterns:
pi\(0\)=LLMextract\(xj,yj,promptsummary\)\\displaystyle p\_\{i\}^\{\(0\)\}=\\text\{LLM\}\_\{\\text\{extract\}\}\(x\_\{j\},y\_\{j\},\\text\{prompt\}\_\{\\text\{summary\}\}\)\(6\)
The initial particles are assigned prior weights based on their linguistic naturalness under the LLM’s language model\. We generateNNparticles per entity label \(typicallyN=10N=10\) and compute their prior scores:
si\\displaystyle s\_\{i\}=−PPL\(pi\(0\)\),i=1,…,N\\displaystyle=\-\\text\{PPL\}\(p\_\{i\}^\{\(0\)\}\),\\quad i=1,\\ldots,N\(7\)wi\(0\)\\displaystyle w\_\{i\}^\{\(0\)\}=exp\(si\)∑j=1Nexp\(sj\)\\displaystyle=\\frac\{\\exp\(s\_\{i\}\)\}\{\\sum\_\{j=1\}^\{N\}\\exp\(s\_\{j\}\)\}\(8\)where perplexity is computed as:
PPL\(pi\(0\)\)=exp\(−1\|pi\(0\)\|∑j=1\|pi\(0\)\|logP\(tj\|t<j\)\)\\text\{PPL\}\(p\_\{i\}^\{\(0\)\}\)=\\exp\\left\(\-\\frac\{1\}\{\|p\_\{i\}^\{\(0\)\}\|\}\\sum\_\{j=1\}^\{\|p\_\{i\}^\{\(0\)\}\|\}\\log P\(t\_\{j\}\|t\_\{<j\}\)\\right\)\(9\)Lower perplexity indicates the rule text is more natural and coherent according to the model’s internal knowledge, thus receiving higher prior weight\.
1\. Evaluation \(ICL\-based Observation\):Each particle is evaluated through ICL\-based IE inference on a validation batchℬ\(t\)\\mathcal\{B\}^\{\(t\)\}sampled from𝒟train\\mathcal\{D\}\_\{\\text\{train\}\}\. For each particle, we compute their confidence scores:
si\(t\)\\displaystyle s\_\{i\}^\{\(t\)\}=Lc\(pi\(t\),θ\),i=1,…,N\\displaystyle=\\text\{L\}\_\{\\text\{c\}\}\(p\_\{i\}^\{\(t\)\},\\theta\),\\quad i=1,\\ldots,N\(10\)yi\(t\)\\displaystyle y\_\{i\}^\{\(t\)\}=exp\(si\(t\)\)∑j=1Nexp\(sj\(t\)\)\\displaystyle=\\frac\{\\exp\(s\_\{i\}^\{\(t\)\}\)\}\{\\sum\_\{j=1\}^\{N\}\\exp\(s\_\{j\}^\{\(t\)\}\)\}\(11\)whereLc\(pi\(t\),θ\)\\text\{L\}\_\{\\text\{c\}\}\(p\_\{i\}^\{\(t\)\},\\theta\)computes the average log probability when the LLM uses rulepi\(t\)p\_\{i\}^\{\(t\)\}to generate a label sequenceTiT\_\{i\}on a validation sample:
Lc\(pi,θ\)=1\|Ti\|∑j=1\|Ti\|logP\(tj\|t<j,pi,θ\)\\text\{L\}\_\{\\text\{c\}\}\(p\_\{i\},\\theta\)=\\frac\{1\}\{\|T\_\{i\}\|\}\\sum\_\{j=1\}^\{\|T\_\{i\}\|\}\\log P\(t\_\{j\}\|t\_\{<j\},p\_\{i\},\\theta\)\(12\)Hereθ\\thetarepresents the pretrained LLM parameters\. This provides a length\-normalized confidence scoreyi\(t\)∈\(0,1\]y\_\{i\}^\{\(t\)\}\\in\(0,1\]reflecting how well the rule performs on the actual IE task\.
2\. Weight Update \(Bayesian Posterior\):Particle weights are updated by combining prior knowledge with observed performance:
w~i\(t\)\\displaystyle\\tilde\{w\}\_\{i\}^\{\(t\)\}=wi\(t−1\)⋅exp\(β⋅yi\(t\)\)\\displaystyle=w\_\{i\}^\{\(t\-1\)\}\\cdot\\exp\(\\beta\\cdot y\_\{i\}^\{\(t\)\}\)\(13\)wi\(t\)\\displaystyle w\_\{i\}^\{\(t\)\}=w~i\(t\)∑j=1Nw~j\(t\)\\displaystyle=\\frac\{\\tilde\{w\}\_\{i\}^\{\(t\)\}\}\{\\sum\_\{j=1\}^\{N\}\\tilde\{w\}\_\{j\}^\{\(t\)\}\}\(14\)wherewi\(t−1\)w\_\{i\}^\{\(t\-1\)\}is the prior weight andexp\(β⋅yi\(t\)\)\\exp\(\\beta\\cdot y\_\{i\}^\{\(t\)\}\)rewards higher performance\.
3\. Multi\-level Resampling with Rule Mutation:We employ a two\-tier strategy balancing exploitation and exploration, adapting the approach fromQiet al\.\([2024](https://arxiv.org/html/2606.18620#bib.bib54)\):
Tier 1 \(Performance Selection\):Retain the top 50% of particles by weight, removing low\-performing ones\.
Tier 2 \(Diversity Mutation\):Apply semantic mutations to retained particles through LLM\-guided generation:
pi′=LLMmutate\(pi\(t−1\),x\(t\)\)p\_\{i\}^\{\\prime\}=\\text\{LLM\}\_\{\\text\{mutate\}\}\(p\_\{i\}^\{\(t\-1\)\},x^\{\(t\)\}\)\(15\)wherex\(t\)x^\{\(t\)\}provides context\. Three strategies are employed:refinement\(increase specificity\),generalization\(increase coverage\), andcontextualization\(generate domain\-specific variants\)\.
New particles inherit labels from parents and receive weights based on perplexity as described in Equation[9](https://arxiv.org/html/2606.18620#S3.E9):
si′\\displaystyle s\_\{i\}^\{\\prime\}=−PPL\(pi′\),i=1,…,M\\displaystyle=\-\\text\{PPL\}\(p\_\{i\}^\{\\prime\}\),\\quad i=1,\\ldots,M\(16\)wi′\\displaystyle w\_\{i\}^\{\\prime\}=exp\(si′\)∑j=1Mexp\(sj′\)\\displaystyle=\\frac\{\\exp\(s\_\{i\}^\{\\prime\}\)\}\{\\sum\_\{j=1\}^\{M\}\\exp\(s\_\{j\}^\{\\prime\}\)\}\(17\)Lower perplexity indicates greater consistency with the model’s language patterns \(See Appendix[E\.1](https://arxiv.org/html/2606.18620#A5.SS1)for the rationale behind our prior and likelihood selection in the Bayesian filtering framework\.\)\. The updated configuration combines retained and new particles:R\(t\)=ℛkeep∪ℛnewR^\{\(t\)\}=\\mathcal\{R\}\_\{\\text\{keep\}\}\\cup\\mathcal\{R\}\_\{\\text\{new\}\}\.
The iteration continues until: \(a\) performance saturates with relative improvement\(F1\(t\)−F1\(t−3\)\)/F1\(t−3\)<0\.03\(F\_\{1\}^\{\(t\)\}\-F\_\{1\}^\{\(t\-3\)\}\)/F\_\{1\}^\{\(t\-3\)\}<0\.03over 3 iterations, or \(b\) all training samples are exhausted\.
As illustrated in Table[4](https://arxiv.org/html/2606.18620#S4.T4), convergence typically occurs with only 3–5% of training data when performance plateaus\.
## 4Experiment
We conduct extensive experiments on multiple Information Extraction tasks\. Following previous work\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47); Weiet al\.,[2023](https://arxiv.org/html/2606.18620#bib.bib76); Liet al\.,[2023a](https://arxiv.org/html/2606.18620#bib.bib67)\), we use entity\-level F1 score for evaluation\. While our framework is applicable to various IE paradigms, we focus on two representative tasks: sequence labeling \(NER\) and relation classification \(RE\)\. For NER, this requires correct boundary detection and type classification; for RE, correct identification of both entity arguments and their relation type \(see Appendix[E\.2](https://arxiv.org/html/2606.18620#A5.SS2)for precise definitions\)\. Following GuideNER\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\), we measure token cost as the average number of input and output tokens per sample, reflecting computational overhead and inference latency\.
### 4\.1Datasets
We evaluate on six widely\-used IE benchmarks spanning multiple domains and task formulations\. For NER, we use CoNLL\-2003\(Tjong Kim Sang and De Meulder,[2003](https://arxiv.org/html/2606.18620#bib.bib16)\)\(news, 4 entity types\), ACE 2005\(Walkeret al\.,[2006](https://arxiv.org/html/2606.18620#bib.bib59)\)\(news/conversational, 7 types\), and GENIA\(Kimet al\.,[2003](https://arxiv.org/html/2606.18620#bib.bib60)\)\(biomedical, 5 types\)\. For RE, we use NYT\(Riedelet al\.,[2010](https://arxiv.org/html/2606.18620#bib.bib61)\)\(news, 24 relation types\), CoNLL04\(Roth and Yih,[2004](https://arxiv.org/html/2606.18620#bib.bib62)\)\(general domain, 5 types\), and SciERC\(Zhanget al\.,[2024](https://arxiv.org/html/2606.18620#bib.bib77)\)\(scientific papers, 7 types\)\. Appendix[A](https://arxiv.org/html/2606.18620#A1)provides detailed statistics\.
MethodModelNERRETokenCostCoNLL03ACE05GENIANYTCoNLL04SciERCOne\-shotQwen\-2\.5\-3b60\.5527\.0846\.780\.2919\.285\.01385Qwen\-2\.5\-7b62\.8235\.7352\.910\.2828\.105\.54Llama\-3\.1\-8b65\.7840\.9851\.290\.4422\.577\.18Pixtral\-12B60\.2539\.1549\.730\.3821\.707\.86ChatIEQwen\-2\.5\-3b39\.2615\.4133\.320\.0011\.110\.00942Qwen\-2\.5\-7b25\.6411\.8310\.190\.0012\.480\.88Llama\-3\.1\-8b55\.1226\.9135\.200\.0012\.621\.03Pixtral\-12B60\.8529\.2840\.120\.4014\.504\.40CodeIEQwen\-2\.5\-3b45\.9223\.706\.590\.000\.000\.001172Qwen\-2\.5\-7b60\.0022\.7619\.170\.000\.000\.00Llama\-3\.1\-8b0\.050\.060\.000\.000\.000\.00Pixtral\-12B53\.3918\.5720\.620\.000\.000\.00GuideNERQwen\-2\.5\-3b63\.3227\.4341\.49———506Qwen\-2\.5\-7b65\.1041\.5747\.43———Llama\-3\.1\-8b61\.3844\.9742\.86———Pixtral\-12B64\.7637\.6448\.03———BCL\\cellcolorgray\!20Qwen\-2\.5\-3b\\cellcolorgray\!2065\.12\\cellcolorgray\!2035\.46\\cellcolorgray\!2046\.98\\cellcolorgray\!200\.31\\cellcolorgray\!2035\.32\\cellcolorgray\!208\.28501\\cellcolorgray\!20Qwen\-2\.5\-7b\\cellcolorgray\!2072\.83\\cellcolorgray\!2046\.94\\cellcolorgray\!2051\.36\\cellcolorgray\!200\.43\\cellcolorgray\!2042\.46\\cellcolorgray\!209\.57\\cellcolorgray\!20Llama\-3\.1\-8b\\cellcolorgray\!2069\.14\\cellcolorgray\!2053\.10\\cellcolorgray\!2050\.15\\cellcolorgray\!200\.91\\cellcolorgray\!2038\.93\\cellcolorgray\!2012\.06\\cellcolorgray\!20Pixtral\-12B\\cellcolorgray\!2065\.54\\cellcolorgray\!2042\.87\\cellcolorgray\!2050\.65\\cellcolorgray\!200\.60\\cellcolorgray\!2025\.75\\cellcolorgray\!2010\.34Table 1:Performance comparison of different methods \(One\-shot, ChatIE, CodeIE, GuideNER, and BCL\) across Qwen\-2\.5, Llama\-3\.1, and Pixtral models on NER \(CoNLL03, ACE05, GENIA\) and RE \(NYT, CoNLL04, SciERC\) benchmarks\. All results are statistically significant \(p<0\.05p<0\.05\)\. Gray shading highlights the best\-performing method for each model configuration\.
### 4\.2Experiments Setup
All experiments are conducted on a computing cluster equipped with H100 GPUs for computational acceleration\. They are implemented using PyTorch 2\.6\.0 and transformers 4\.51\.3\. We employ four foundation models with varying scales, languages, and training datasets: Qwen2\.5\-3B, Qwen2\.5\-7B, Llama3\.1\-8B and Pixtral\-12B\. This selection allows us to investigate the impact of different model size on our method’s performance\. All models are tested with temperature set to 0\.0 and random seed fixed at 42 to ensure reproducibility of the experiments\. In addition, during the computation of prior probabilities, all models are evaluated in eval mode to disable dropout and other stochastic behaviors\.
For optimization framework, we set the number of particles to 10 based on empirical experience from preliminary experiments\(\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\)\), balancing computational efficiency and exploration capability\. The number of data used in each observation step is determined through grid search over the set \[1, 3, 5, 7, 9, 11, 13, 15\], selecting the value that yields optimal performance for each dataset\. This configuration ensures consistent experimental conditions and enables direct performance comparison across different models and datasets\.
### 4\.3Baselines
We compare against widely\-used ICL approaches in the IE domain\.
Our baselines include:One\-shot, where models perform IE with a single demonstration example to illustrate the task format and desired output structure\.ChatIE\(Weiet al\.,[2023](https://arxiv.org/html/2606.18620#bib.bib76)\)andCodeIE\(Liet al\.,[2023a](https://arxiv.org/html/2606.18620#bib.bib67)\), two task transfer methods that reformulate IE as dialogue or code generation tasks, respectively\. Both methods employ sophisticated example selection strategies and are applicable to both NER and RE tasks\.GuideNER\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\), the current state\-of\-the\-art guideline\-based method for NER, which uses frequency\-based rule selection to assist in\-context learning\. Since GuideNER is specifically designed for NER tasks, we only evaluate it on NER datasets and mark it as "—" for RE tasks\.
This experimental design ensures all baseline methods operate under the same paradigm of pure in\-context learning, enabling fair comparison of different approaches’ effectiveness\.
### 4\.4Main Results
Table[1](https://arxiv.org/html/2606.18620#S4.T1)shows BCL consistently outperforms all baselines across IE benchmarks and model scales\. Task transfer baselines \(ChatIE, CodeIE\) exhibit severe degradation on smaller models: ChatIE achieves only 25\.64 F1 on CoNLL03 \(Qwen\-2\.5\-7B\) versus BCL’s 72\.83, while CodeIE drops to 0\.05 F1 \(Llama\-3\.1\-8B\) versus BCL’s 69\.14, reflecting their reliance on model\-specific capabilities\. Against the rule\-based GuideNER, BCL shows consistent advantages \(72\.83 vs\. 65\.10 on CoNLL03; 51\.36 vs\. 47\.43 on GENIA\), demonstrating systematic Bayesian optimization’s superiority over frequency heuristics\.
BCL’s advantage amplifies on RE tasks where prior methods fail completely\. ChatIE and CodeIE achieve 0\.00 F1 across most RE configurations, while GuideNER is inapplicable \(marked "—"\)\. In contrast, BCL maintains effective performance \(42\.46 F1 on CoNLL04; 12\.06 F1 on SciERC\), demonstrating successful generalization across both NER and RE through adaptive rule optimization\.
This stability enables cost\-efficient deployment: Qwen\-2\.5\-3B with BCL \(65\.12 F1\) matches Llama\-3\.1\-8B’s one\-shot performance \(65\.78 F1\) on CoNLL03, achieving comparable results with 62% fewer parameters\. BCL thus represents the first ICL optimization framework leveraging Bayesian inference for adaptive optimization, achieving consistent performance across diverse IE paradigms and model scales while overcoming the task\-specific limitations of prior methods\.\(Appendix[B](https://arxiv.org/html/2606.18620#A2)illustrates the evolution process through examples\)\.
### 4\.5Ablation Study
We conduct ablation studies to validate the contribution of each component in BCL’s particle filtering framework\. Table[2](https://arxiv.org/html/2606.18620#S4.T2)presents results on CoNLL03 with Qwen\-2\.5\-7B, where we systematically remove each mechanism while keeping others intact\.
Table 2:Ablation study on different components of the proposed method\.Bayesian Weight Update\.Removing Bayesian weight updates causes the most significant performance drop \(\-11\.05 F1 points\), reducing F1 from 72\.83 to 61\.78\. Without this mechanism, the framework degenerates to random search without principled belief updates, demonstrating that Bayesian inference is the core component enabling BCL’s effectiveness\.
Tier\-2 Resampling \(Diversity Mutation\)\.Disabling diversity mutation leads to the second\-largest degradation \(\-9\.85 F1 points\), with F1 dropping to 62\.98\. This confirms that maintaining particle diversity through controlled mutation is crucial for preventing premature convergence to suboptimal solutions and exploring the prompt space effectively\.
Tier\-1 Resampling \(Performance Filtering\)\.While removing this component causes minimal performance loss \(\-1\.54 F1 points\), particle count explodes from 10 to 10,441, revealing that Tier\-1 primarily ensures efficiency by pruning low\-quality particles with negligible performance cost\.
These results validate BCL’s design: Bayesian weight updates and diversity mutation are essential for performance, while performance filtering ensures efficiency by maintaining a compact particle set without performance degradation\.
Effect of Semantic Decomposition\.To isolate the role of semantically coherent subcategories, we conduct an additional ablation where subcategories are randomly reassigned to different entity labels, breaking semantic alignment while keeping the subcategory pool unchanged\. Results \(Appendix[C](https://arxiv.org/html/2606.18620#A3)\) show that this leads to consistent performance degradation, confirming that semantic coherence is critical for effective optimization\.
### 4\.6Cross\-Model Generalization
While our method is designed to optimize inference on a given model, it is also important to understand whether the learned rules capture transferable patterns that extend beyond a specific backbone\. To this end, we study the cross\-model generalization ability of the proposed approach\.
We consider a setting where rules optimized on smaller open\-source models \(e\.g\., Llama\-3\.1\-8B and Qwen\-2\.5\-7B\) are directly applied to a stronger closed\-source model, GPT\-3\.5\-turbo\. This setup reflects a practical scenario in which optimization is performed on accessible models and then deployed on more capable systems\. The inference pipeline remains unchanged, and no additional adaptation is introduced\.
Table 3:Cross\-model generalization results\. Rules optimized on smaller open\-source models can be directly transferred to GPT\-3\.5\-turbo and consistently improve performance\.As shown in Table[3](https://arxiv.org/html/2606.18620#S4.T3), rules learned on smaller models consistently yield performance gains when transferred to GPT\-3\.5\-turbo\. In particular, the transferred rules outperform the direct GPT\-3\.5\-turbo baseline on both datasets\. This suggests that the proposed method captures generalizable reasoning patterns rather than relying on model\-specific behaviors\.
Overall, these results indicate that the learned optimization strategies are not limited to the source model, but can extend to stronger models without additional tuning\.
### 4\.7Parameter Sensitivity Analysis
We perform extensive parameter analysis on two key factors: context window length and training data size\.
#### 4\.7\.1Impact of Training Data Quantity on BCL Performance
Table 4:F1 scores \(%\) for different training set sizes\.We investigate BCL’s data efficiency by evaluating performance across varying training set sizes \(1% to 50%\)\. Table 3 presents F1 scores on Qwen\-2\.5\-7B across three datasets with varying training set sizes\. The results reveal a striking pattern: performance rapidly saturates with minimal data\. On CoNLL03, BCL reaches 69\.12 F1 with only 1% of training data and peaks at 72\.83 F1 with 40%, showing modest improvement \( 3\.7 points\) when scaling from 1% to the optimal setting\. Similar trends appear on GENIA \(peak at 5\-10%: 51\.20\-51\.36 F1\) and ACE05 \(peak at 5%: 46\.94 F1\)\. This data efficiency stems from the particle filtering mechanism’s ability to rapidly update belief distributions after each sample, achieving convergence without overfitting to large training sets\.
#### 4\.7\.2Impact of Context Window Length on BCL Performance
Table 5:F1 scores \(%\) for different context lengths \(in sentences\) during the observation stepThe context window length \(also called batch size\) determines how much information our method considers when computing posterior probabilities at each observation step\. Table[5](https://arxiv.org/html/2606.18620#S4.T5)shows that on Qwen\-2\.5\-7B, performance improves significantly from single\-sentence to five\-sentence contexts due to smoother joint likelihood functions that provide more stable gradients\. However, performance declines beyond 9 sentences as excessive observations lead to overly averaged particle weights, preventing effective updates\.
## 5Conclusion
In this paper, we propose BCL, a general framework that can efficiently traverse training sets and rapidly extract specific relationships from training data using particle\-based methods\. BCL is the first optimization framework using Bayesian inference for IE tasks\. We have conducted extensive and comprehensive experiments to thoroughly demonstrate the effectiveness of our approach\. Additionally, we acknowledge the limitations of our method, such as the generation of a large number of particles to achieve rapid convergence, which increases inference burden and computational costs in practice\. Therefore, our future research will focus on improving particle utilization efficiency to further enhance the overall system performance\.
## Acknowledgement
This work was supported in part by the National Key R&D Program under grants 2025ZD1502903 and 2024YFC3308101\.
## Limitations
Our work has several limitations\. First, while BCL maintains reasonable inference efficiency, the optimization phase requiresO\(K×M\)O\(K\\times M\)iterations, whereKKdenotes the number of particles andMMdenotes the number of training samples\. Each iteration involves CPU\-intensive numerical computations including weight updates and resampling \(detailed analysis in Appendix[E\.3](https://arxiv.org/html/2606.18620#A5.SS3)\)\. This makes the method most suitable for scenarios where optimization costs can be amortized across multiple deployments, rather than one\-time or frequently\-updated applications\.
Second, the particle filtering mechanism may bias optimization toward frequent patterns in the training data, potentially affecting rare entity types\. However, our empirical analysis \(Appendix[D](https://arxiv.org/html/2606.18620#A4)\) shows that BCL is largely robust to such frequency imbalance, with only marginal performance differences across frequency strata\.
## References
- K\. J\. Åström and R\. Murray \(2021\)Feedback systems: an introduction for scientists and engineers\.Princeton university press\.Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- M\. A\. Beaumont, W\. Zhang, and D\. J\. Balding \(2002\)Approximate bayesian computation in population genetics\.Genetics162\(4\),pp\. 2025–2035\.Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§1](https://arxiv.org/html/2606.18620#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p1.1)\.
- C\. Cai, H\. Liu, X\. Zhao, Z\. Jiang, T\. Zhang, Z\. Wu, J\. Lee, J\. Hwang, and L\. Li \(2025a\)Bayesian optimization for controlled image editing via llms\.InAnnual Meeting of the Association for Computational Linguistics \(ACL\),Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p2.1)\.
- C\. Cai, X\. Zhao, H\. Liu, Z\. Jiang, T\. Zhang, Z\. Wu, J\. Hwang, and L\. Li \(2025b\)The role of deductive and inductive reasoning in large language models\.InAnnual Meeting of the Association for Computational Linguistics \(ACL\),Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p3.1)\.
- L\. Cao and J\. Zhao \(2025\)Pretraining on the test set is no longer all you need: a debate\-driven approach to QA benchmarks\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Pdyh3USc2A)Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- C\. Chen, P\. Luo, C\. Feng, T\. Wu, W\. Jiang, and T\. Xu \(2025\)MAPS: a multi\-task framework with anchor point sampling for zero\-shot entity linking\.DATA INTELLIGENCE7\(4\),pp\. 1085–1107\.External Links:[Document](https://dx.doi.org/10.3724/2096-7004.di.2025.0025)Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- Q\. Dong, L\. Li, D\. Dai, C\. Zheng, J\. Ma, R\. Li, H\. Xia, J\. Xu, Z\. Wu, T\. Liu,et al\.\(2022\)A survey on in\-context learning\.arXiv preprint arXiv:2301\.00234\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p1.1)\.
- A\. Doucet, N\. De Freitas, N\. J\. Gordon,et al\.\(2001\)Sequential monte carlo methods in practice\.Vol\.1,Springer\.Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- P\. I\. Frazier \(2018\)A tutorial on bayesian optimization\.arXiv preprint arXiv:1807\.02811\.Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- N\. J\. Gordon, D\. J\. Salmond, and A\. F\. Smith \(1993\)Novel approach to nonlinear/non\-gaussian bayesian state estimation\.InIEE proceedings F \(radar and signal processing\),Vol\.140,pp\. 107–113\.Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- Y\. e\. al\. Guan \(2025\)Learning an efficient optimizer via hybrid\-policy sub\-trajectory balance\.arXiv\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p5.1)\.
- Y\. Guan, Y\. Liu, K\. Zhou, Z\. Shen, J\. Hwang, S\. Belongie, and L\. Li \(2025\)Is meta\-learning out? rethinking unsupervised few\-shot classification with limited entropy\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 4188–4197\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p5.1)\.
- S\. Huang, B\. Xu, Y\. Yu, C\. Li, and X\. A\. Lin \(2025\)GuideNER: annotation guidelines are better than examples for in\-context named entity recognition\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 24159–24166\.Cited by:[§E\.2](https://arxiv.org/html/2606.18620#A5.SS2.p1.1),[§E\.3](https://arxiv.org/html/2606.18620#A5.SS3.p1.8),[§1](https://arxiv.org/html/2606.18620#S1.p1.1),[§1](https://arxiv.org/html/2606.18620#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1),[§3\.2\.1](https://arxiv.org/html/2606.18620#S3.SS2.SSS1.p4.1),[§4\.2](https://arxiv.org/html/2606.18620#S4.SS2.p2.1),[§4\.3](https://arxiv.org/html/2606.18620#S4.SS3.p2.1),[§4](https://arxiv.org/html/2606.18620#S4.p1.1)\.
- Y\. Ji, Z\. Li, R\. Meng, and D\. He \(2025\)Reason\-to\-rank: distilling direct and comparative reasoning from large language models for document reranking\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’25,New York, NY, USA,pp\. 2320–2329\.External Links:ISBN 9798400715921,[Link](https://doi.org/10.1145/3726302.3730070),[Document](https://dx.doi.org/10.1145/3726302.3730070)Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p5.1)\.
- Y\. Ji, Z\. Li, R\. Meng, and D\. He \(2026\)Retrieval–reasoning processes for multi\-hop question answering: a four\-axis design framework and empirical trends\.arXiv preprint arXiv:2601\.00536\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p5.1)\.
- J\. Kim, T\. Ohta, Y\. Tateisi, and J\. Tsujii \(2003\)GENIA corpus—a semantically annotated corpus for bio\-textmining\.Bioinformatics19\(suppl\_1\),pp\. i180–i182\.Cited by:[§4\.1](https://arxiv.org/html/2606.18620#S4.SS1.p1.1)\.
- T\. Lan, J\. Xu, X\. He, J\. Hwang, and L\. Li \(2025\)Attention consistency for llms explanation\.InFindings of the Association for Computational Linguistics: EMNLP,pp\. 1736–1750\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p3.1)\.
- G\. Li, P\. Wang, W\. Ke, Y\. Guo, K\. Ji, Z\. Shang, J\. Liu, and Z\. Xu \(2024a\)Recall, retrieve and reason: towards better in\-context relation extraction\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,pp\. 6368–6376\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p4.1)\.
- L\. Li, S\. Jia, and J\. Hwang \(2026\)Multiple human motion understanding\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 6297–6305\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- L\. Li, S\. Jia, J\. Wang, Z\. Jiang, F\. Zhou, J\. Dai, T\. Zhang, Z\. Wu, and J\. Hwang \(2025\)Human motion instruction tuning\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- P\. Li, T\. Sun, Q\. Tang, H\. Yan, Y\. Wu, X\. Huang, and X\. Qiu \(2023a\)CodeIE: large code generation models are better few\-shot information extractors\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15339–15353\.Cited by:[§1](https://arxiv.org/html/2606.18620#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p4.1),[§4\.3](https://arxiv.org/html/2606.18620#S4.SS3.p2.1),[§4](https://arxiv.org/html/2606.18620#S4.p1.1)\.
- X\. Li, K\. Lv, H\. Yan, T\. Lin, W\. Zhu, Y\. Ni, G\. Xie, X\. Wang, and X\. Qiu \(2023b\)Unified demonstration retriever for in\-context learning\.arXiv preprint arXiv:2305\.04320\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p1.1)\.
- Y\. Li, R\. Ramprasad, and C\. Zhang \(2024b\)A simple but effective approach to improve structured language model output for information extraction\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 5133–5148\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- J\. Liu \(2026\)Discovering what you can control: interventional boundary discovery for reinforcement learning\.External Links:2603\.18257,[Link](https://arxiv.org/abs/2603.18257)Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- M\. Luo, X\. Xu, Z\. Dai, P\. Pasupat, M\. Kazemi, C\. Baral, V\. Imbrasaite, and V\. Zhao \(2024\)Dr\.icl: demonstration\-retrieved in\-context learning\.DATA INTELLIGENCE6\(4\),pp\. 909–922\.External Links:[Document](https://dx.doi.org/10.3724/2096-7004.di.2024.0012)Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- S\. Min, X\. Lyu, A\. Holtzman, M\. Artetxe, M\. Lewis, H\. Hajishirzi, and L\. Zettlemoyer \(2022\)Rethinking the role of demonstrations: what makes in\-context learning work?\.arXiv preprint arXiv:2202\.12837\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p1.1)\.
- Y\. Mo, J\. Liu, J\. Yang, Q\. Wang, S\. Zhang, J\. Wang, and Z\. Li \(2024\)C\-icl: contrastive in\-context learning for information extraction\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 10099–10114\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.\(2022\)Training language models to follow instructions with human feedback\.Advances in neural information processing systems35,pp\. 27730–27744\.Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- R\. Pryzant, D\. Iter, J\. Li, Y\. T\. Lee, C\. Zhu, and M\. Zeng \(2023\)Automatic prompt optimization with" gradient descent" and beam search\.arXiv preprint arXiv:2305\.03495\.Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- B\. Qi, Z\. Qian, Y\. Luo, J\. Gao, D\. Li, K\. Zhang, and B\. Zhou \(2024\)Evolution of thought: diverse and high\-quality reasoning via multi\-objective optimization\.arXiv preprint arXiv:2412\.07779\.Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1),[§3\.2\.1](https://arxiv.org/html/2606.18620#S3.SS2.SSS1.p8.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.Advances in neural information processing systems36,pp\. 53728–53741\.Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- S\. Riedel, L\. Yao, and A\. McCallum \(2010\)Modeling relations and their mentions without labeled text\.InJoint European conference on machine learning and knowledge discovery in databases,pp\. 148–163\.Cited by:[§F\.5](https://arxiv.org/html/2606.18620#A6.SS5.SSS0.Px2.p2.1),[§4\.1](https://arxiv.org/html/2606.18620#S4.SS1.p1.1)\.
- D\. Roth and W\. Yih \(2004\)A linear programming formulation for global inference in natural language tasks\.Cited by:[§4\.1](https://arxiv.org/html/2606.18620#S4.SS1.p1.1)\.
- J\. Shi, Q\. Ma, H\. Liu, H\. Zhao, J\. Hwang, and L\. Li \(2026\)Intrinsic entropy of context length scaling in llms\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p3.1)\.
- E\. F\. Tjong Kim Sang and F\. De Meulder \(2003\)Introduction to the CoNLL\-2003 shared task: language\-independent named entity recognition\.InProceedings of the Seventh Conference on Natural Language Learning at HLT\-NAACL 2003,pp\. 142–147\.External Links:[Link](https://aclanthology.org/W03-0419/)Cited by:[§4\.1](https://arxiv.org/html/2606.18620#S4.SS1.p1.1)\.
- S\. Wadhwa, S\. Amir, and B\. C\. Wallace \(2023\)Revisiting relation extraction in the era of large language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15566–15589\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p4.1)\.
- C\. Walker, S\. Strassel, J\. Medero, and K\. Maeda \(2006\)Ace 2005 multilingual training corpus\.\(No Title\)\.Cited by:[§4\.1](https://arxiv.org/html/2606.18620#S4.SS1.p1.1)\.
- Z\. Wan, F\. Cheng, Z\. Mao, Q\. Liu, H\. Song, J\. Li, and S\. Kurohashi \(2023\)GPT\-re: in\-context learning for relation extraction using large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 3534–3547\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p4.1)\.
- Z\. Wang, Y\. Liang, W\. Sun, Q\. Lin, C\. Xu, and Y\. Zhang \(2025\)A novel feature selection framework based on large language models\.DATA INTELLIGENCE7\(4\),pp\. 1016–1034\.External Links:[Document](https://dx.doi.org/10.3724/2096-7004.di.2025.0078)Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p2.1)\.
- J\. Wei, M\. Bosma, V\. Y\. Zhao, K\. Guu, A\. W\. Yu, B\. Lester, N\. Du, A\. M\. Dai, and Q\. V\. Le \(2021\)Finetuned language models are zero\-shot learners\.arXiv preprint arXiv:2109\.01652\.Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- J\. Wei, Y\. Tay, R\. Bommasani, C\. Raffel, B\. Zoph, S\. Borgeaud, D\. Yogatama, M\. Bosma, D\. Zhou, D\. Metzler,et al\.\(2022\)Emergent abilities of large language models\.arXiv preprint arXiv:2206\.07682\.Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p1.1)\.
- X\. Wei, X\. Cui, N\. Cheng, X\. Wang, X\. Zhang, S\. Huang, P\. Xie, J\. Xu, Y\. Chen, M\. Zhang,et al\.\(2023\)Chatie: zero\-shot information extraction via chatting with chatgpt\.arXiv preprint arXiv:2302\.10205\.Cited by:[§1](https://arxiv.org/html/2606.18620#S1.p1.1),[§4\.3](https://arxiv.org/html/2606.18620#S4.SS3.p2.1),[§4](https://arxiv.org/html/2606.18620#S4.p1.1)\.
- T\. Xu, J\. Liu, N\. Mattei, and Z\. Zheng \(2026\)Fair algorithms with probing for multi\-agent multi\-armed bandits\.Proceedings of the AAAI Conference on Artificial Intelligence40\(32\),pp\. 27332–27340\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i32.39950)Cited by:[§2\.2](https://arxiv.org/html/2606.18620#S2.SS2.p1.1)\.
- Z\. Yao, X\. Cheng, Z\. Huang, and L\. Li \(2025\)CountLLM: towards generalizable repetitive action counting via large language model\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§2\.1](https://arxiv.org/html/2606.18620#S2.SS1.p5.1)\.
- Q\. Zhang, Z\. Chen, H\. Pan, C\. Caragea, L\. J\. Latecki, and E\. Dragut \(2024\)SciER: an entity and relation extraction dataset for datasets, methods, and tasks in scientific documents\.arXiv preprint arXiv:2410\.21155\.Cited by:[§4\.1](https://arxiv.org/html/2606.18620#S4.SS1.p1.1)\.
- J\. Zhao, E\. Min, H\. Wu, Z\. Li, Z\. Sun, H\. Cai, S\. Wang, X\. Chen, and G\. Penn \(2026\)Beyond step pruning: information theory based step\-level optimization for self\-refining large language models\.Proceedings of the AAAI Conference on Artificial Intelligence40\(41\),pp\. 34941–34949\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40798),[Document](https://dx.doi.org/10.1609/aaai.v40i41.40798)Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- Z\. Zhao, E\. Wallace, S\. Feng, D\. Klein, and S\. Singh \(2021\)Calibrate before use: improving few\-shot performance of language models\.InInternational conference on machine learning,pp\. 12697–12706\.Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
- Y\. Zhou, A\. I\. Muresanu, Z\. Han, K\. Paster, S\. Pitis, H\. Chan, and J\. Ba \(2022\)Large language models are human\-level prompt engineers\.InThe eleventh international conference on learning representations,Cited by:[§2\.3](https://arxiv.org/html/2606.18620#S2.SS3.p1.1)\.
## Appendix ADataset Statistics
We provide detailed statistics for all six datasets used in our experiments in Table[6](https://arxiv.org/html/2606.18620#A1.T6)\.
Table 6:Statistics of the datasets used in our experiments\.\|\|Ents\|\|and\|\|Rels\|\|denote the number of entity types and relation types\. \#Train, \#Val and \#Test denote the sample number in each split\.Dataset\|\|Ents\|\|\|\|Rels\|\|\#Train\#Val\#TestNamed Entity RecognitionCoNLL034\-14,0413,2503,453ACE057\-6,202745812GENIA5\-15,0231,6691,854Relation ExtractionNYT\-2456,1955,0005,000CoNLL0445922231288SciERC671,861275551
## Appendix BCase Study
We present a case study using a CoNLL\-2003 sentence to illustrate our three\-stage workflow: particle generator, posterior probability calculator, and resampler\. Here,xxrepresents label sequences,NNdenotes the number of generation rules, andYYrepresents input text\. It is worth noting that the following case studies present simplified examples to illustrate our framework’s workflow\. All experiments follow the technical specifications detailed in Section[3](https://arxiv.org/html/2606.18620#S3)\.
### B\.1Case Study 1: Particle Generator
Figure 5:Case study showing the prompt template for subcategory generation\. The template guides the model to generate semantic subcategories for named entity labels, demonstrated with organization and person entity types\.Particle generation serves as the initial step in our framework to obtain initial particles and compute their corresponding prior probabilities\. We demonstrate this process using the label "organization" as an example\. As shown in Figure[5](https://arxiv.org/html/2606.18620#A2.F5), our prompt template instructs the model to act as a subcategory generation expert, decomposing broad entity labels into semantically distinct subcategories with associated probability weights\. For the "organization" label, the system generates diverse subcategories such as "sports team", "city", "company", "name", and "financial institution"\. Each subcategory represents a different semantic interpretation of how organizations might appear in text, with the probability weights serving as prior probabilities that reflect their typical occurrence frequency\.
### B\.2Case Study 2: Posterior Probability Calculator
Figure 6:Case study demonstrating posterior probability calculation through likelihood estimation\. The model evaluates generated rules against input text to compute particle posterior probabilities based on classification accuracy\.Following particle generation, we compute posterior probabilities by evaluating how well each generated rule performs on the actual classification task\. Figure[6](https://arxiv.org/html/2606.18620#A2.F6)illustrates this process using the input text "EU rejects German call to boycott British lamb\." The system applies the previously generated subcategory rules to identify entities in the given text\. The likelihood of each particle is calculated based on the model’s ability to correctly identify and classify entities according to the generated rules\. The posterior probability calculation applies Bayes’ theorem to update particle weights based on how well each rule matches the observed entity patterns in the input text\. For instance, the "organization" subcategories show varying performance scores: "sports team: 0\.9277", "city: 0\.7073", "company: 0\.4751", demonstrating how different semantic interpretations receive different posterior weights based on their effectiveness\. This likelihood\-based evaluation ensures that particles with better classification performance receive higher posterior probabilities, enabling the system to focus on the most promising labeling hypotheses for subsequent resampling\.
### B\.3Case Study 3: Resampler
Figure 7:Case study demonstrating the resampling stage with particle mutation\.The final stage employs an LLM\-based resampling mechanism that performs context\-aware particle mutation\. As shown in Figure[7](https://arxiv.org/html/2606.18620#A2.F7), given the input text "He said further scientific study was required and if it was found that action was needed it should be taken by the European Union," the system generates contextually relevant subcategories\. The resampling process works as follows: the LLM analyzes the specific context and mutates the original particles to better fit the observed text patterns\. For the "organization" label, the system generates context\-specific subcategories such as "educational institution," "non\-governmental organization," "research institution," and "financial institution," which are more relevant to the scientific and policy context of the input\. Prior probabilities are computed to indicate higher prior probability\. For instance, "educational institution: 0\.340369" and "non\-governmental organization: 0\.297805" receive relatively high priors due to their contextual relevance\. This perplexity\-based weighting ensures that contextually appropriate mutations are favored during the resampling process\.
## Appendix CAblation on Semantic Decomposition
To further validate the necessity of semantically coherent subcategories, we design a controlled ablation experiment that breaks the semantic alignment between subcategories and their corresponding entity labels\.
We take the optimized subcategory set learned by BCL and randomly reassign subcategories to different entity labels\. For example, subcategories such as “company” or “event” may be assigned to the “Person” label instead of semantically consistent categories like “athlete” or “politician”\. Importantly, the subcategory pool remains unchanged, ensuring that the only difference lies in the loss of semantic alignment\.
Table[7](https://arxiv.org/html/2606.18620#A3.T7)shows the results across representative datasets and models\. Breaking semantic coherence leads to consistent performance degradation across all settings\.
Table 7:Effect of semantic decomposition\.On average, shuffling subcategories results in a 7\.38 F1 point drop, demonstrating that semantic coherence plays a crucial role in enabling effective optimization\. Meanwhile, the framework still maintains reasonable performance, suggesting that BCL exhibits graceful degradation rather than catastrophic failure\.
## Appendix DFrequency\-Stratified Analysis on Long\-Tail Entities
To further investigate the potential long\-tail effect discussed in Section X, we conduct a frequency\-stratified evaluation to quantify model performance across entities with different training frequencies\.
We perform experiments on two models \(Llama\-3\.1\-8B and Qwen2\.5\-7B\) and two datasets \(ACE05 and CoNLL2003\)\. Entities in the test set are grouped into three tiers based on their frequency in the training set: Top 30% \(frequent\), Mid 30% \(medium\-frequency\), and Bottom 40% \(rare\)\. We report F1 scores for each group separately\.
Table[8](https://arxiv.org/html/2606.18620#A4.T8)summarizes the results\. Overall, BCL demonstrates strong robustness to frequency imbalance\. The performance gap between frequent and rare entities is generally small \(within 1–2 F1 points\), and in some cases, rare entities even achieve comparable or slightly better performance\.
Table 8:Frequency\-stratified F1 performance of BCL\.These results suggest that, although the particle filtering mechanism theoretically emphasizes frequent patterns, its impact on rare entity performance is limited in practice\. We hypothesize that the Bayesian aggregation process helps mitigate overfitting to frequent patterns by maintaining diverse hypothesis particles\.
## Appendix EImplementation Details
### E\.1Prior and Likelihood Selection in Bayesian Filtering
In our Bayesian filtering framework, particle weights are updated through the combination of prior beliefs and observed evidence\. We use perplexity\-based scores as the prior and task\-specific confidence as the likelihood, based on the following considerations:
\(1\) Rationality of the Prior:Perplexity reflects the degree of consistency between generated rules and the LLM’s internal language distribution\. Lower perplexity indicates that the rule better aligns with the model’s “linguistic intuition,” and as a prior assumption, such rules are more likely to yield high\-quality extraction results\. This provides a reasonable inductive bias: before observing any task\-specific performance, we favor rules that are linguistically coherent and natural according to the model’s pre\-trained knowledge\.
\(2\) Rationality of the Likelihood:The IE confidence scoreLc\\text\{L\}\_\{\\text\{c\}\}\(Equation[12](https://arxiv.org/html/2606.18620#S3.E12)\) measures how well a rule performs on actual extraction tasks\. Unlike perplexity, which only captures linguistic properties, the confidence score directly evaluates the rule’s effectiveness in guiding the LLM to generate correct entity labels\. This task\-specific measurement serves as the likelihood functionp\(D\|θ\)p\(D\|\\theta\), providing empirical evidence to update our beliefs about rule quality based on observed extraction performance\.
\(3\) Separation of Prior and Likelihood:In standard Bayesian updating, the priorp\(θ\)p\(\\theta\)and likelihoodp\(D\|θ\)p\(D\|\\theta\)serve distinct roles\. Similarly, in our framework:
- •Perplexity \(Prior\):Encodes the inductive bias that “rules should be linguistically fluent and coherent”
- •IE Confidence \(Likelihood\):Provides observational evidence through actual task performance
This separation allows newly generated particles to carry reasonable initial beliefs into the evaluation stage, rather than using uninformative uniform priors\. The Bayesian update mechanism \(Equation[13](https://arxiv.org/html/2606.18620#S3.E13)\) then combines these two sources of information:
wi\(t\)⏟posterior∝wi\(t−1\)⏟prior⋅exp\(β⋅yi\(t\)\)⏟likelihood\\underbrace\{w\_\{i\}^\{\(t\)\}\}\_\{\\text\{posterior\}\}\\propto\\underbrace\{w\_\{i\}^\{\(t\-1\)\}\}\_\{\\text\{prior\}\}\\cdot\\underbrace\{\\exp\(\\beta\\cdot y\_\{i\}^\{\(t\)\}\)\}\_\{\\text\{likelihood\}\}\(18\)
\(4\) Robustness via Filtering Dynamics:Even if perplexity imperfectly approximates rule quality, the particle filtering framework provides inherent robustness through its denoising mechanism\. Particles with misleading priors \(low perplexity but poor IE performance\) receive low likelihood scores, causing their weights to exponentially decay through multiplicative updates and be eliminated during resampling\. This mechanism ensures that the final particle distribution is dominated by task performance rather than initial priors\.
### E\.2Evaluation Metric\.
Following previous work\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\), we adopt entity\-level F1 scoring for both NER and RE tasks\. For NER, a predicted entity is considered correct only when both its boundary \(character span\) and type label exactly match the ground truth annotation\. For RE, both entity mentions and their relation type must match exactly\. Precision, recall, and F1 score are computed at the entity/relation level across the entire test set\.
### E\.3Computational Complexity\.
LetKKdenote the number of particles per label andMMdenote the number of training samples used for optimization\. Our method processes batches sequentially through iterative filtering, requiringO\(K×M\)O\(K\\times M\)LLM inference calls in total\. In contrast, methods like GuideNER\(Huanget al\.,[2025](https://arxiv.org/html/2606.18620#bib.bib47)\)requireO\(N\)O\(N\)calls to traverse the complete training set of sizeNN\. Since we operate onM≪NM\\ll Nsamples \(e\.g\.,M≈0\.03NM\\approx 0\.03Nas shown in Section[4\.7\.1](https://arxiv.org/html/2606.18620#S4.SS7.SSS1)\), our method achieves significant efficiency gains\. The weight update \(Equation[14](https://arxiv.org/html/2606.18620#S3.E14)\) and resampling \(Equation[17](https://arxiv.org/html/2606.18620#S3.E17)\) operations involve onlyO\(K\)O\(K\)numerical computations—probability calculations, softmax normalization, sorting, and random sampling—which are negligible compared to LLM inference time\. Therefore, the computational cost is dominated by LLM inference, and our approach is substantially more efficient than full\-dataset traversal methods\.
### E\.4Hyperparameters\.
Key hyperparameters are set as follows: number of particles per labelN=10N=10\(empirically optimal range: 1–20\), selection pressureβ=2\.0\\beta=2\.0\(selected via grid search over\{1,2,5,10\}\\\{1,2,5,10\\\}across multiple datasets\), observation batch size of 4 sentences \(analyzed in Section[4\.7\.2](https://arxiv.org/html/2606.18620#S4.SS7.SSS2)\), and retention ratio of 50% for resampling \(balancing exploitation and exploration\)\. All LLM inference is performed with temperature=0\.0 and random seed=42 for reproducibility\. Convergence is reached when validation F1 score plateaus for 3 consecutive iterations\.
## Appendix FCheck List
This appendix provides additional details required by the ACL Responsible NLP Research checklist\.
### F\.1Data Licenses and Terms of Use \(B2\)
We provide license information for all datasets used in our experiments:
Table 9:License information for all datasets used in experiments\.##### Model Licenses\.
The language models used in our experiments are distributed under the following licenses:
- •Qwen2\.5\-3B/7B: Apache 2\.0 License \(Qwen Team, Alibaba\)
- •Llama\-3\.1\-8B: Llama 3\.1 Community License \(Meta\)
- •Pixtral\-12B: Apache 2\.0 License \(Mistral AI\)
##### Code Availability\.
Our implementation will be released under the MIT License upon acceptance\. The codebase includes all scripts for data preprocessing, particle filtering optimization, and evaluation\.
### F\.2Intended Use and Consistency \(B3\)
All datasets were used in accordance with their intended purposes:
- •CoNLL\-2003: Originally created for the CoNLL\-2003 shared task on language\-independent named entity recognition\. We use it for NER evaluation as intended\.
- •ACE 2005: Developed for entity, relation, and event extraction research\. We use the entity annotations for NER evaluation\.
- •GENIA: Created for biomedical text mining research\. We use it for biomedical NER evaluation as intended\.
- •NYT: Derived from New York Times articles for relation extraction research\. We use it for RE evaluation\.
- •CoNLL04: Designed for joint entity and relation extraction\. We use it for RE evaluation\.
- •SciERC: Created for information extraction from scientific papers\. We use it for scientific RE evaluation as intended\.
Our derived artifacts are intended solely for research purposes in information extraction and should not be used for commercial applications without appropriate licensing\.
### F\.3Privacy and Offensive Content \(B4\)
We did not perform additional anonymization as the original dataset creators have already addressed privacy considerations in their data collection and release procedures\.
### F\.4Artifact Documentation \(B5\)
##### Dataset Coverage\.
Table[10](https://arxiv.org/html/2606.18620#A6.T10)provides detailed documentation of the datasets used:
Table 10:Domain, language, and source documentation for all datasets\.
##### Entity and Relation Types\.
- •CoNLL\-2003 \(4 types\): PER, LOC, ORG, MISC
- •ACE 2005 \(7 types\): Person, Organization, Location, Facility, Weapon, Vehicle, GPE
- •GENIA \(5 types\): DNA, RNA, Protein, Cell Line, Cell Type
- •NYT \(24 relation types\): Including location\-related, person\-related, and organization\-related relations
- •CoNLL04 \(5 relation types\): Located\-In, Work\-For, OrgBased\-In, Live\-In, Kill
- •SciERC \(7 relation types\): Used\-for, Feature\-of, Part\-of, Compare, Hyponym\-of, Evaluate\-for, Conjunction
### F\.5Package Parameters \(C4\)
##### Software Dependencies\.
- •Python 3\.10\.12
- •PyTorch 2\.6\.0
- •Transformers 4\.51\.3
- •NumPy 1\.24\.3
- •scikit\-learn 1\.3\.0 \(for evaluation metrics\)
##### Evaluation Implementation\.
We use theseqevallibrary \(v1\.2\.2\) for NER evaluation with the following settings:
- •mode=’strict’: Exact boundary and type matching required
- •scheme=IOB2: Using IOB2 tagging scheme
For RE evaluation, we implement custom evaluation following prior workRiedelet al\.\([2010](https://arxiv.org/html/2606.18620#bib.bib61)\):
- •A relation is correct if both entity mentions and relation type match exactly
- •Precision, recall, and F1 are computed at the relation\-tuple level
##### Text Preprocessing\.
- •Tokenization: Model\-specific tokenizers from HuggingFace
- •No additional preprocessing \(lowercasing, stemming\) applied
- •Maximum sequence length: 512 tokens \(truncation applied if exceeded\)
### F\.6Use of AI Assistants \(E1\)
In this paper, we employed Large Language Models \(LLMs\) as core components of our proposed BCL framework in three specific stages\. First, we utilized LLMs for rule extraction and particle generation to create initial subcategory patterns from training data, ensuring systematic rule discovery\. Second, we leveraged LLMs for performance evaluation through in\-context learning inference, where models assess rule effectiveness on validation datasets\. Third, we employed LLMs for rule mutation and resampling to generate semantically diverse rule variants during the optimization process\.
Additionally, we used LLMs to assist with writing and language improvement throughout the manuscript preparation process\. This included grammar checking, sentence structure optimization, and clarity enhancement to improve the overall readability of our work\. However, all core research ideas, methodological contributions, experimental designs, and conclusions remain entirely our own intellectual work\.
### F\.7Broader Impact Statement
##### Positive Impacts\.
BCL advances information extraction technology with several benefits:
- •Accessibility: By improving ICL performance on smaller models \(3B\-12B parameters\), BCL democratizes access to effective IE systems for researchers and practitioners with limited computational resources\.
- •Efficiency: The data\-efficient optimization \(converging with 3\-5% of training data\) reduces the annotation burden and computational costs\.
- •Generalizability: The framework’s applicability to both NER and RE tasks provides a unified approach to diverse IE challenges\.
##### Potential Negative Impacts\.
- •Automation of Information Extraction: While beneficial for legitimate applications, improved IE could potentially be misused for unauthorized data harvesting or surveillance\.
- •Environmental Impact: Although more efficient than full fine\-tuning, LLM\-based IE still requires significant computational resources with associated carbon emissions\.
##### Mitigation Strategies\.
We encourage users to:
1. 1\.Apply BCL only to data they have legal rights to process
2. 2\.Implement human oversight in high\-stakes applications
3. 3\.Consider the environmental impact and use appropriately\-sized models for their needsSimilar Articles
LC-ICL: Label-Guided Contrastive In-Context Learning for Robust Information Extraction
This paper proposes LC-ICL, a novel few-shot technique that uses both correct and incorrect examples with error-cause labels to improve large language models' performance on information extraction tasks like named entity recognition and relation extraction.
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
CLBench-V is a benchmark for evaluating multimodal context learning across grounding, new information application, and new knowledge learning. The best model achieves only 0.2847, showing the task remains challenging.
Bootstrap-Conditioned Action Selection with Tabular Foundation Models
The paper proposes BC-ICL, a bootstrap-conditioned action selection method that leverages pretrained tabular foundation models with in-context learning for contextual bandits, improving exploration and regret performance under strict online protocols.
When Context Misleads: In-context Learning with Jurisdiction in Large Language Models
This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.
AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking
AbICL proposes an in-context learning framework for antigen-specific antibody affinity ranking, combining a pretrained structural encoder with a context ranking head to leverage labeled demonstrations for test-time adaptation without gradient updates.