HintMiner: Automatic Question Hints Mining From Q&A Web Posts with Language Model via Self-Supervised Learning
Summary
HintMiner is a novel tool that automatically mines hints for user questions from Q&A web posts like Stack Overflow using a language model trained via self-supervised learning, achieving effective performance in evaluations.
View Cached Full Text
Cached at: 09/16/26, 08:36 AM
# 1Introduction
Source: [https://arxiv.org/html/2609.16060](https://arxiv.org/html/2609.16060)
HintMiner: Automatic Question Hints Mining From Q&A Web Posts with Language Model via Self\-Supervised Learning
Zhenyu ZhangJiuDong Yang
Independent ResearcherIndependent Researcher
###### Abstract
Users often need ask questions and seek answers online\. The Question \- Answering \(QA\) forums such as Stack Overflow cannot always respond to the questions timely and properly\. In this paper, we propose HintMiner, a novel automatic question hints mining tool for users to help them find answers\. HintMiner leverages the machine comprehension and sequence generation techniques to automatically generate hints for users’ questions\. It firstly retrieve many web Q&A posts and then extract some hints from the posts using MiningNet that is built via a language model\. Using the huge amount of online Q&A posts, we design a self\-supervised objective to train the MiningNet that is a neural encoder\-decoder model based on the transformer and copying mechanisms\. We have evaluated HintMiner on 60,000 Stack Overflow questions\. The experiment results show that the proposed approach is effective\. For example, HintMiner achieves an average BLEU score of 36\.17% and an average ROUGE\-2 score of 36\.29%\. Our tool and experimental data are publicly available\.111https://github\.com/zhangzhenyu13/HintMiner\.
## 1Introduction
It is a common practice to seek answers from online Question and Answering \(Q&A\) forums, such as Stack Overflow, Data Science, etc\.\[[29](https://arxiv.org/html/2609.16060#bib.bib23),[5](https://arxiv.org/html/2609.16060#bib.bib37),[3](https://arxiv.org/html/2609.16060#bib.bib36)\]\. These Q&A forums store abundant question related posts accumulated over years\. However, as the posted questions in Q&A sites rely on community members’ voluntary answers, there is no guarantee to obtain timely and satisfactory answers for everyone question\. As a matter of fact, we have found that a large number of questions lack accepted answers in Stack Exchange\. It also costs users lots of time to search from those webs, where the Q&A resource aggregating and reforming methods are quite a necessity\.
In recent years, some methods have been proposed to help users with Q&A\. Some retrieval based methods such as AnswerBot\[[36](https://arxiv.org/html/2609.16060#bib.bib22)\]or the official Stack\-Overflow website specify the key points of answers from the retrieved relevant posts by selecting the most important paragraphs\. Another kind effective Q&A method is machine reading comprehension \(MRC\), which aims to understand the semantics of question and then select a text span from a given passage\[[23](https://arxiv.org/html/2609.16060#bib.bib14),[31](https://arxiv.org/html/2609.16060#bib.bib16),[30](https://arxiv.org/html/2609.16060#bib.bib33),[6](https://arxiv.org/html/2609.16060#bib.bib12)\]as the answer to the question\. However, the MRC cannot combine several spans to form a more rich and semantic\-complete result\. Enlightened by the Q&A systems based on retrieval and MRC such as DrQA\[[6](https://arxiv.org/html/2609.16060#bib.bib12)\], etc\., we build a dedicated automatic question hints mining system to help users\. We targeted at mining hints from Q&A forums while these methods do not utilize the specific Q&A web resources\. And we also try to merge several selected spans to generate semantic rich and complete results while those previous works can only retrieve passages or select independent text spans\.
In this paper, we aim to reuse the existing resources in online Q&A forums to generate useful hints to the user questions\. To that end, we propose a question hints mining tool called HintMiner, which selects and merges several useful segments of texts that can provide some hints for the question\. We formulate HintMiner as:Find the most useful text spans from the relevant posts in Q&A forums and combine them to generate the hints for the question\.Based on the hints provided, it will be much easier for users to get the final answers or help users to clarify and understand the questions\. HintMiner first leverages Elastic Search \(ES222https://www\.elastic\.co/elasticsearch/\) to find relevant information for the question\. It then selects several text\-spans that can provide some hints for a question to form the answer through machine reading comprehension\[[6](https://arxiv.org/html/2609.16060#bib.bib12)\]\. Finally, HintMiner merge the text\-spans to generate semantic rich and complete hints via sequence generation\[[24](https://arxiv.org/html/2609.16060#bib.bib19)\]\. To achieve this, we designed a self\-supervised learning\(SSL\) objective for MiningNet to capture the semantics of questions and the relevant posts and to generate suitable hints\. We construct ”question” \+ ”relative posts” \+ ”proper hints/answers” triplets from millions of online stackoverflow posts\. Then we train the MiningNet to learn to generate such ”hints/answers” with ”questions” \+ ”relative posts” as input\. MiningNet leverages BERT\[[8](https://arxiv.org/html/2609.16060#bib.bib17)\]to encode the question and its relevant posts so as to capture their deep semantics\. The deep semantic representation is further fed to a transformer decoder\[[27](https://arxiv.org/html/2609.16060#bib.bib18)\]that can capture the importance of each input token through the attention mechanism\. With the learned token importance, we build a CopyNet using the copy mechanism\[[13](https://arxiv.org/html/2609.16060#bib.bib39),[41](https://arxiv.org/html/2609.16060#bib.bib38)\]to select a set of relevant tokens from the input to generate hints\.
We have conducted extensive experiments to evaluate HintMiner\. The results show that HintMiner outperforms several information retrieval based methods\. For example, HintMiner achieves an average of 36\.17% BLEU score and 36\.29% ROUGE\-2 score\. Furthermore, MiningNet outperforms several strong retrieval baselines and generation language model baselines\.
Our contributions can be summarized as follows:
- •We build an automatic question hints mining tool called HintMiner, which can help developers solve questions\. We extracted paragraphs from online Q&A forums to build a useful posts dataset\. We also make our code and data publicly available\.
- •We develop MiningNet, a novel self\-supervised learning based model that can capture the semantics of a question and the relevant posts in Q&A forums, and can generate the semantic rich and complete hints for questions\.
- •We have performed extensive evaluation of the proposed approach\. Our results show that HintMiner is effective and outperforms several strong baseline methods\.
Our work is an important step towards intelligent hints mining for Q&A forums\.
## 2The Q&A Forums and Dataset
### 2\.1Question Answering Web Resources
Web users would always ask questions or search relevant answers online\. To solve their problems, all kinds of online users depend heavily on online Q&A sites\. For example, Stack Overflow has become one of the most popular such Q&A sites for developers, and it has accumulated a large number \(over 16 million\) of Q&A posts\. Figure[1](https://arxiv.org/html/2609.16060#S2.F1)shows an example of the posts in Stack Overflow website\. There are mainly six parts in a post: 1\) the title of the question, showing the general concise description of the question, 2\) the detailed description of the question, 3\) the tags assigned by the user who posts the question, indicating the categories of the question, 4\) the list of the answers to the question, including the accepted answer if available, 5\) the linked posts that are marked by Stack Overflow community which are relevant to the current post, and 6\) the related posts that are retrieved by the Stack Overflow system\. It has been found that the responding time can be quite long and many questions may never be answered\[[29](https://arxiv.org/html/2609.16060#bib.bib23)\]\. It is desirable to improve question solution effectiveness by automating the question answering process\. Therefore, it is quite necessity to find proper hints for users’ questions\.
Figure 1:An example of the posts in Stack Overflow
### 2\.2The Construction of the Q&A Dataset
We collected about 17 million posts from four Stack Exchange websites viaarchive\.org, including Stack Overflow333https://stackoverflow\.com/, Artificial Intelligence444https://ai\.stackexchange\.com/, Data Science555https://datascience\.stackexchange\.com/and Cross Validated666https://stats\.stackexchange\.com/\. As illustrated in Figure[1](https://arxiv.org/html/2609.16060#S2.F1), the linked posts are marked by the community and are useful to the question\. There are about 19% of the posts connected with over5M5Mlinks\. We build a Post\-Link Graph where the nodes are posts and the edges are weighted links\. There are two types of links marked by the community\. We set the weights to 0 for links that mark duplicate posts and 1 for the others\. Then we apply Dijkstra algorithm to compute the shortest link distance between each pair of nodes\. Finally, we obtain 4 link distances \(“0”, “1”, “2” and “≥\\geq3”\) because previous researches\[[35](https://arxiv.org/html/2609.16060#bib.bib8),[38](https://arxiv.org/html/2609.16060#bib.bib35)\]show that two posts with a link distanced≥3d\\geq 3is not relevant to each other\. For example, in Figure[2](https://arxiv.org/html/2609.16060#S2.F2), there are 6 marked links betweenA,B,C,D,EA,B,C,D,Eand we complete the rest links \(dashed lines\) except forB,DB,DasdB,D≥3d\_\{B,D\}\\geq 3\. The link distance indicates how useful the post content is to the question of the other post\. The shorter the link distance is, the more useful the post is to the question\.
#### 2\.2\.1Selecting Relevant Posts For Training
In order to train the MiningNet \(Section[3\.2](https://arxiv.org/html/2609.16060#S3.SS2)\), we build a ”question\-passage\-hints” triplets dataset\. For a question of the node \(i\.e\. Post\) in the post\-link graph, we select top 2, 1, 1, 1 paragraphs for posts with distance as 0,1,2,and≥3\\geq 3respectively to construct the relevant passage of current question\. The paragraphs of the passage are randomly shuffled so that the model cannot simply remember order of sentences\. We select the first passage of accepted answer in the post with more than 10 words as thegold hintsfor the question, which is considered to be meaningful\. We removed those questions without neighbors whose distance is 1\. Finally we constructed about 3\.6 million ”question\-passage\-hints” triplets\. Note that we add some less relevant paragraphs whose distance is larger than 1 so that noise and negative content are added to improve the robustness and difficulty of the dataset\.
Figure 2:An Example of Post\-Link Graph
#### 2\.2\.2Selecting Relevant Posts For Inference
We first dumped the posts into the Elastic Search Engine \(ES\), and then retrieve relevant posts from a variety of Q&A forums\. We then perform pre\-processing of the selected posts\. In this work, we aim at generating hints rather than generating code or numerical expressions which usually exists in those scientific forums\. Therefore, we replace a code snippet with\[CODE\], and a mathematical expression with\[NUM\]\. We do not consider hyperlinks either\. To reduce the vocabulary size, we use the BPE algorithm\[[34](https://arxiv.org/html/2609.16060#bib.bib32)\]to perform tokenization, which can transform a compound word into a few tokens\. For each question, we retrieve 5 posts in total\.
The retrieved posts often contain many non\-essential sentences that are useless and can make it difficult for a deep neural network to handle extremely long input\[[16](https://arxiv.org/html/2609.16060#bib.bib41)\]\. Examples of such sentences are ”Maybe my answer can help you”, ”Thank you for your suggestion”, etc\. Therefore, we leverage an ensemble method to filter those sentence, which combines the results of three base algorithms that can identify the important sentences\. 1\)*Lexrank*\[[10](https://arxiv.org/html/2609.16060#bib.bib26)\], a graph based method inspired by Pagerank algorithm\[[32](https://arxiv.org/html/2609.16060#bib.bib31)\], which uses the eigenvector centrality of sentences to select the important sentences\. 2\)*KL greedy search*\[[14](https://arxiv.org/html/2609.16060#bib.bib28)\], an information entropy maximization based method, which uses the KL divergence to compute the relative information gain to greedily select sentences so as to maximize the information entropy of selected sentences\. 3\)*Latent Semantic Analysis \(LSA\)*\[[26](https://arxiv.org/html/2609.16060#bib.bib25)\], which decomposes the sentence\-term matrix using SVD and selects the sentences with the most significant topics via the right singular vectors\. The three base algorithms focus on different aspects of sentence importance\. Therefore, we merge their results and eliminate the sentence repetition\. The resulting set of sentences forms the context passage for MiningNet\.
## 3HintMiner: Generating Hints to User’s Questions
### 3\.1System Overview
In our work, we formulate the problem as follows: given a question and a set of relevant posts, the core problem is to select a set of useful text spans from existing posts and generate the hints to the question\. We process the posts to form the context passage for training \(Section[2\.2\.1](https://arxiv.org/html/2609.16060#S2.SS2.SSS1)\) and inference \(Section[2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2)\)\. The hints are then generated by the MiningNet\.
Figure 3:An Overview of HintMinerFor that purpose, we build HintMiner, which utilizes the techniques of machine reading comprehension\[[6](https://arxiv.org/html/2609.16060#bib.bib12)\]and sequence generation\[[24](https://arxiv.org/html/2609.16060#bib.bib19)\]\. Figure[3](https://arxiv.org/html/2609.16060#S3.F3)shows the overview of HintMiner\. Given a question, we first select the relevant posts from the Q&A forums to form the context passage \(Section[2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2)\)\. Then we feed the context and question to MiningNet, which is an effective deep neural network that can generate the hints to the question by copying and generating tokens from the context\.
### 3\.2The MiningNet Model
Figure 4:An Overview of MiningNet#### 3\.2\.1The Structure of the Model
Figure[4](https://arxiv.org/html/2609.16060#S3.F4)shows the structure of MiningNet, which consists of four parts: a BERT Encoder, a transformer decoder, and a CopyNet\. There are three embeddings in the Input Representation that takes Q&A data as input and outputs the embeddings of input\. These embeddings are extracted from BERT\[[8](https://arxiv.org/html/2609.16060#bib.bib17)\]\. It is worth mentioning that the segment id for tokens in questions is 0 and for tokens in context is 1\. When generating thettht^\{th\}answer token \(AtA\_\{t\}\), the tokens of the generated answer before steptt\(A1,…,At−1A\_\{1\},\.\.\.,A\_\{t\-1\}\) are embedded using position embedding and token embedding only, and are then fed to the transformer decoder\. The⊕\\oplusin Figure[4](https://arxiv.org/html/2609.16060#S3.F4)refers to the use of BERT embedding layer to embed the tokens in text\. Equation[1](https://arxiv.org/html/2609.16060#S3.E1)presents how BERT is used to encode the sequence in our model\. Each token inqq\(a question\) andcc\(the context\) is encoded as adimdimdimension dense vectorTiqT\_\{i\}^\{q\}/TjcT\_\{j\}^\{c\}, and theTclsT\_\{cls\}represents the pooling vector\.
T=\[TCLS,T1q,,…,Tmq,T1c,…,Tnc\]=BERT\(\[q,c\]\]\),whereTiq,Tjc∈Rdim,1≤i≤m,1≤j≤n\\begin\{split\}T=\[T\_\{CLS\},T\_\{1\}^\{q\},,\.\.\.,T\_\{m\}^\{q\},T\_\{1\}^\{c\},\.\.\.,T\_\{n\}^\{c\}\]=BERT\(\[q,c\]\]\),\\\\ where\\,T\_\{i\}^\{q\},T\_\{j\}^\{c\}\\in R^\{dim\},\\,1\\leq i\\leq m,\\,1\\leq j\\leq n\\end\{split\}\(1\)
The output of the BERT encoder is a sequence of vectors representing the semantics of the question and context passage\. The transformer decoder\[[27](https://arxiv.org/html/2609.16060#bib.bib18)\]reads the output of the BERT encoder and then computes the output hidden state of target \(i\.e\. generated answer\) vectors\. The transformer decoder also computes the encoder\-decoder attention score vectors, which uses the multi\-head attention mechanism to pay attention to the “question\+context passage”\. We also build a CopyNet\[[13](https://arxiv.org/html/2609.16060#bib.bib39),[41](https://arxiv.org/html/2609.16060#bib.bib38)\], which takes the encoder\-decoder attention vectors as input and outputs an answer text\. Using CopyNet, certain text spans in the input sequence are selected to be present in the output sequence with the Copy Probability\[[13](https://arxiv.org/html/2609.16060#bib.bib39)\]\. Thus, MiningNet can generate the hints to a question through the copying mechanism by selectively replicating the input text spans\. In this way, we transform the hints generation problem to a MRC problem, where the hints is composed of several selected text spans and combined via generation\. The hints tokensA1A\_\{1\},A2A\_\{2\}, …AnA\_\{n\}are generated one by one through the copying mechanism iteratively untilAnA\_\{n\}is the end token ornnexceeds the hints length limit\. \( The generation length \(nn\) is usually set to a fixed length for satisfactory model performance\[[16](https://arxiv.org/html/2609.16060#bib.bib41)\]\. \)
#### 3\.2\.2Transformer Decoder and CopyNet
In this subsection, we describe the Transformer Decoder and CopyNet models in detail and show how to adapt them to MiningNet\. The encoder vectors \(TT\) are the semantic representation of “question\+context”\. We apply the transformer decoder to them and compute the encoder and decoder attention via Equation[2](https://arxiv.org/html/2609.16060#S3.E2), which defines the compatibility function of the query with the corresponding key in the multi\-head attention\.QQdenotes the decoder hidden vectors at time steptt, which is given as shifted maskedt−1t\-1true answer encoding vectors during training andt−1t\-1predicted answer encoding vectors during testing\.
To select text spans from the context, we apply the copying mechanism \(CopyCopy\) on the encoder\-decoder attention vectors, which can select tokens from input directly\. According to the copy mechanism and the attention in Equation[2](https://arxiv.org/html/2609.16060#S3.E2), the output of CopyNet \(a three layer MLP with same model dimension and activation function used in BERT\) is denoted asp\(At\|A1,…,At−1,T\)=Copy\(attnt\)=softmax\(attnt\)p\(A\_\{t\}\|A\_\{1\},\.\.\.,A\_\{t\-1\},T\)=Copy\(attn^\{t\}\)=softmax\(attn^\{t\}\), whereattnjtattn\_\{j\}^\{t\}is the attention value at stepttforjthj^\{th\}context vector\. Therefore, the probability distribution for generated answer sequencep\(A1\),p\(A2\),…,p\(An\)p\(A\_\{1\}\),p\(A\_\{2\}\),\.\.\.,p\(A\_\{n\}\)can be denoted as Equation[3](https://arxiv.org/html/2609.16060#S3.E3)\.
Attention\(Q,K,V\)=softmax\(Q∗KTdim\)∗V,whereK=V=BERT\(\[q,c\]\)\\begin\{split\}Attention\(Q,K,V\)=softmax\(\\frac\{Q\*K^\{T\}\}\{\\sqrt\{dim\}\}\)\*V,\\\\ where\\,K=V=BERT\(\[q,c\]\)\\end\{split\}\(2\)
p\(A\|T\)=∏tp\(At\|A1,…,At−1,T\)p\(A\|T\)=\\prod\_\{t\}p\(A\_\{t\}\|A\_\{1\},\.\.\.,A\_\{t\-1\},T\)\(3\)
It is worth mentioning that the hints generation process is based on captured semantics of question and context \(passage formed from relevant posts\)\. Essentially, MiningNet simulates a function that maps the semantic representation of thecontext to hintstext spans according to therequirement of the questionvia the attention mechanism\. The context contains the knowledge that can provide some hints for the question, which is represented as encoded vectors\. The BERT encoder leverages its well designed and pre\-trained network to represent the semantics of the question and context, then the decoder computes an attention score for each token of hints that is to be copied from the context based on the semantics\. In the end, the selected text\-spans could represent the most proper hints that can help to clarify and understand the question\.
### 3\.3Self\-supervised Learning Objective
We aim totrain the MiningNet to learn to specify the most useful text\-spans from a given context paragraph for the question\.Therefore, we propose our SSL objective: let the model learn to distinguish important sentences, phrases or words as hints from a noisy context input given a question\. For each post, we extract the top rated K\(=3\) answers and shuffle them randomly to prevent the model remembering to copy the best one always\. We concatenate the K answers and feed the resulting text to the unimportant sentences filtering component to form the context\. We use the best answer as the \(the most useful\) answer to the question of the post\. Finally we form over 1,900,000<question,context,hints\><question,context,hints\>triplets for training the model\.
### 3\.4The Implementation and Training Details
HintMiner utilizes the MiningNet to understand the semantics of questions and posts, and then generate suitable answers to the questions\. To implement MiningNet, we leverage the transformers library\[[33](https://arxiv.org/html/2609.16060#bib.bib2)\]\. We use the BERT\-base as backbone\. For encoder\-decoder attention we use 12 attention heads and the hidden size is the same as the BERT\. We set hyper\-parameters based on previous research\[[2](https://arxiv.org/html/2609.16060#bib.bib40)\]and the pre\-trained BERT encoder structure\[[8](https://arxiv.org/html/2609.16060#bib.bib17)\]\. Therefore, the hyper\-parameters are well fine\-tuned\.
As the encoder\-decoder model suffers from the exposure bias issue\[[24](https://arxiv.org/html/2609.16060#bib.bib19),[40](https://arxiv.org/html/2609.16060#bib.bib20),[15](https://arxiv.org/html/2609.16060#bib.bib21)\], we adopt a hybrid training strategy which firstly uses teacher forcing training\[[24](https://arxiv.org/html/2609.16060#bib.bib19)\]in a supervised way and then uses the policy gradient reinforcement learning with BLEU4\[[20](https://arxiv.org/html/2609.16060#bib.bib29)\]as reward to fine\-tune the model\. We leverage the Adam optimizer\[[12](https://arxiv.org/html/2609.16060#bib.bib34)\]to maximize the probability denoted in Equation[3](https://arxiv.org/html/2609.16060#S3.E3)\. We train the model with initial learning rate as1e−51e\-5for 200,000 training steps with batch size as 32\.
## 4Experiments
Table 1:Evaluation of HintMiner and Compared Methods\. “ROU\-2” denotes ROUGE\-2\.Table 2:Comparison of HintMiner with Different Relevant Post Retrieval MethodsTable 3:Examples of Hints Generated by HintMiner### 4\.1Experimental Design
We conducted experiments to evaluate the effectiveness of HintMiner\. Our evaluation focuses on the following four research questions:
RQ1: How effective is HintMiner in generating hints?
We compare our model with two kinds of representative methods, i\.e\. the text retrieval and text generation techniques\. We list the compared methods as follows:
- •Text retrieval based methods\.AnswerBot\[[36](https://arxiv.org/html/2609.16060#bib.bib22)\]applies the MMR algorithom\[[4](https://arxiv.org/html/2609.16060#bib.bib24)\]to the relevant posts given a question to extract proper paragraphs as answers\.SimCSEis well trained via contrastive learning and can behave rather well in specifying semantic relevant sentences\[[11](https://arxiv.org/html/2609.16060#bib.bib6),[37](https://arxiv.org/html/2609.16060#bib.bib7)\]\. We retrieve the the sentences with highest cosine similarity for a question from the posts\.PageRank\+\+extends the traditional graph\-based important sentences selection approaches\[[10](https://arxiv.org/html/2609.16060#bib.bib26),[18](https://arxiv.org/html/2609.16060#bib.bib27)\]by replacing the tf\-idf features of PageRank\[[32](https://arxiv.org/html/2609.16060#bib.bib31)\]with semantic vectors of SimCSE\.
- •Text generation based methods\.GPT2\[[21](https://arxiv.org/html/2609.16060#bib.bib3)\]an auto\-regressive language model that is pre\-trained to predict the next token in text, which is good at many text generation tasks such as summarization, answer generation, etc\.BART\[[22](https://arxiv.org/html/2609.16060#bib.bib4)\]leverages the advantages of both BERT\[[8](https://arxiv.org/html/2609.16060#bib.bib17)\]and GPT models\[[21](https://arxiv.org/html/2609.16060#bib.bib3)\]to pretrain a encoder\-decoder based language models by applying several language mask strategies in tokens, sentences and whole documents\.UniLM\[[1](https://arxiv.org/html/2609.16060#bib.bib5)\]is a pre\-trained unified language model for both auto\-encoding and partially auto\-regressive language modeling tasks using a a pseudo\-masked language model, which is good at language understanding and generation\. For each method in baselines, we apply same data processing methods as our HintMiner to prompt fair comparable results\. In our implementation, we leveraged the huggingface transformers\[[33](https://arxiv.org/html/2609.16060#bib.bib2)\]to build the neural networks\.
RQ2: How effective is HintMiner when different relevant post retrieval methods are used?
The relevant post retrieval is an important part of HintMiner\. In HintMiner, we use Elastic Search Engine to search for relevant posts\. In this RQ, we evaluate the influence of different retrieval methods\. In Q&A forums such as Stack Overflow, community members often manually mark the linked posts for some questions\. We experimented with the linked posts as the relevant posts \(i\.e\. we directly select posts based on the post\-link graph as described for training in Section[2\.2\.1](https://arxiv.org/html/2609.16060#S2.SS2.SSS1)\)\. Refer to Section[2\.2](https://arxiv.org/html/2609.16060#S2.SS2)for more details\. We also leverage the open online methods such as Stack Exchange Search Engine API777https://api\.stackexchange\.comand Google Search Engine API888https://developers\.google\.com/custom\-searchto retrieve three related posts for each test post\.
Experimental settings:To evaluate the effectiveness of HintMiner, we randomly sampled 60,000 posts accepted answers that do not appear in our training data\. The first paragraph of accepted answer with more than 10 words aregold hintsof the question\. Then, for each question in the posts, we used HintMiner to generate the hints\. Finally we used 4\-gram BLEU score and 2\-gram ROUGE score \(ROUGE\-2\) to evaluate the quality of the generated hints\. For RQ1, RQ2 and RQ4, we set the context length to 500 words in our experiments\. In the generation decoding process, we use the BEAM\-Search algorithm with beam size as 5 and select the best generated texts as the final hint for a question\.
### 4\.2Evaluation Metrics
To evaluate the generated hints, we use two n\-gram language model evaluation metrics, i\.e\. ROUGE and BLEU\[[17](https://arxiv.org/html/2609.16060#bib.bib30),[20](https://arxiv.org/html/2609.16060#bib.bib29)\], which are widely used in machine translation, summarization, and text generation tasks, etc\., to measure the similarity between two sentences\. In our research, we measure whether the generated hints are similar to the gold hints\. The BLEU score uses the common presence of n\-gram count of generated text and reference text to measure the similarity from the precision\-like perspective\. The ROUGE score measures the similarity from the recall\-like perspective\. In our experiment, ROUGE uses 2\-gram \(i\.e\. ROUGE\-2\) and BLEU uses 4\-gram\. Both BLEU and ROUGE\-2 scores are 100% when the generated hints are the same as the true hints and 0 when they are totally different\. The larger the value, the better the generated hints will be\.
### 4\.3Experimental Results for RQ1
The three retrieval baselines are given in Table[1](https://arxiv.org/html/2609.16060#S4.T1)for the first three rows\. It shows the performance results of all the experimented methods for different answer lengths respectively, where50,100,150,20050,100,150,200are the answer tokens we truncated\. The SimCSE outperforms the AnswerBot \(that is based on hand\-crafted features\) a lot which demonstrate that the capturing the semantics of question and passage through BERT is critical\. The BERT\+\+ outperforms the SimCSE, which demonstrates that the importance of sentences in paragraphs concerns a lot and thus it’s natural to apply some attention mechanism to better select proper sentences\. The HintMiner here directly select text\-spans with rather than coarse sentence\-level granularity, which significantly outperforms all the baselines\.
For the generation based baselines, as shown with middle three rows in Table[1](https://arxiv.org/html/2609.16060#S4.T1), HintMiner significantly outperforms the three baselines using the proposed MiningNet\. The MiningNet obtains a BLEU score of 36\.17% and a ROUGE\-2 score of 36\.29% when the maximum answer length is set to 100 words\. The HintMiner outperforms all the baselines given the length of the generated answers varies from 50 to 200 words, which shows the effectiveness of the proposed pre\-training objectives\. The public models such as GPT2, BART and UniLM is trained using common language modeling objectives, which is not a good solution for the professional situations in our research\. Through the unsupervised learning with the huge amount of programming posts, the MiningNet is able to select the needed text spans from the context to generate answers that are semantically similar to the true answers\.
We manually checked about 100 questions and hints generated from those methods\. The results showed that the sentences that are similar to the question are not necessary the sentences that can form the hints that need to useful for the question rather than just repeat the meaning of question again\. Also, the hints are not necessary to be one complete sentence because some contents in paragraphs are not necessary\. Therefore, selecting from sentence\-level is not rational \(i\.e\. the 3 retrieval baselines\)\. It is also sub\-optimal to only consider the common language semantics from Wikipedia or Bookcorpus to pre\-train a language model, which lack of knowledge for a specific domains and are not trained to distinguish useful contents as hints in our research problem\. In conclusion, these baseline methods cannot effectively generate proper hints given noisy relative posts, which leads to lower performance\.
### 4\.4Experimental Result for RQ2
Table[2](https://arxiv.org/html/2609.16060#S4.T2)shows the effectiveness of HintMiner when using different methods to find the relevant posts\. Using the Linked Posts manually marked by the Q&A community, HintMiner can achieve the best performance, but there are only around 19% of posts marked with links and newly posted questions lack these user marked links\. When using Google Custom Search, the search service by Stack Exchange and Elastic Search, both BLEU score and ROUGE\-2 score drop a little, but the performance is still acceptable\. This experiment also demonstrates the stability and scalability of HintMiner for dealing with different sources retrieved as context passage\.
### 4\.5Examples of the Generated Hints
Table[3](https://arxiv.org/html/2609.16060#S4.T3)shows some of the hints generated by HintMiner, where the1st1^\{st\}is a solved question \(i\.e\., questions with accepted answers\) and the rest are not solved yet\. We omit the detailed description of questions and provide the link to the corresponding Stack Overflow page\. The HintMiner can propoerly select text spans of accepted answers against noise paragraphs of sentences \(Section[2\.2](https://arxiv.org/html/2609.16060#S2.SS2)\)\. Although the results may contain grammatical errors they are generally readable and useful\. Currently, the generated hints do not contain code or mathematical expression \(i\.e\. represented with symbols such as\[NUM\]and\[CODE\]\)\. For example, the accepted answer of the2nd2^\{nd\}question is attached with a code snippet while our generated hints just shows the presence of code here \(\[CODE\]\)\.
To further evaluate the effectiveness of HintMiner, we also randomly sampled some without accepted answers\. We can see that HintMiner is able to generate meaningful and useful hints even without gold hints in the passage\. Taking the2nd2^\{nd\}hint as an example, the question is about ”usage of simulator” and the answer provides some tips for the question\. Those answers further confirm the usefulness of HintMiner\. These results are encouraging\.
## 5Related Work
In recent years, question answering \(Q&A\) has been receiving a lot of attention in natural language processing\.Generally, there are mainly three kinds Q&A systems, including IR based Q&A, KBase based Q&A, and MRC based Q&A\. Antonio et al\.\[[25](https://arxiv.org/html/2609.16060#bib.bib1)\]conducted a literature review in the 130 out of 1842 papers on Q&A systems, and found that 28\.57% of the surveyed papers are based on IR and 34\.9% on KBase\. IR\[[7](https://arxiv.org/html/2609.16060#bib.bib9)\]is widely studied in Q&A, and combining its with KBases based methods to fetch answers by searching knowledge bases\[[9](https://arxiv.org/html/2609.16060#bib.bib10),[39](https://arxiv.org/html/2609.16060#bib.bib11)\]are gaining momentum as some established knowledge bases like FreeBase and DBpedia are publicly available\. Recently, many MRC based Q&A methods\[[6](https://arxiv.org/html/2609.16060#bib.bib12),[31](https://arxiv.org/html/2609.16060#bib.bib16),[30](https://arxiv.org/html/2609.16060#bib.bib33)\]have been proposed\. For example, Miller et al\.\[[19](https://arxiv.org/html/2609.16060#bib.bib13)\]used MRC on wikipedia to find the text spans for questions\. Currently, these research mainly focus on general open domain Q&A and lack support for the utilization of Q&A forums resources\. To help better understand HintMiner, we introduce the MRC and Copy Mechanism here briefly\.
### 5\.1Copying Mechanism
Copying mechanism is inspired by pointer network\[[28](https://arxiv.org/html/2609.16060#bib.bib15)\]that is proposed for OOV problems\. It is widely used in many Seq2Seq models\[[13](https://arxiv.org/html/2609.16060#bib.bib39),[41](https://arxiv.org/html/2609.16060#bib.bib38)\]\. The copy mechanism selects a set of input tokens to the output\. The CopyNet based on the pointer\-network\[[28](https://arxiv.org/html/2609.16060#bib.bib15),[41](https://arxiv.org/html/2609.16060#bib.bib38)\]directly leverages the attention between the source encoder and the target decoder to generate an output probability distribution over the source tokens\.
### 5\.2Machine Reading Comprehension
Machine reading comprehension \(MRC\) comprehends a natural language question and then selects a text span \(usually not longer than 40 tokens\) from a given passage\[[23](https://arxiv.org/html/2609.16060#bib.bib14),[31](https://arxiv.org/html/2609.16060#bib.bib16),[30](https://arxiv.org/html/2609.16060#bib.bib33),[6](https://arxiv.org/html/2609.16060#bib.bib12)\]as the answer to the question\. For each question, the task is to select a text span to answer it by outputting a start index and an end index of the input sequence tokens of the passage\. For example, DrQA\[[6](https://arxiv.org/html/2609.16060#bib.bib12)\]select spans over millions of Wikipedia pages to answer general questions such as “who is the current president of USA?”\. Wang et al\. proposed a multi\-granularity attention fusion networks\[[31](https://arxiv.org/html/2609.16060#bib.bib16)\]to encode the question and the article via multi\-granularity attention to select proper text\-span\. Microsoft researchers also proposed R\-Net\[[30](https://arxiv.org/html/2609.16060#bib.bib33)\]to predict the answer text\-span, which leverages the pointer\-network for text\-span selection\.
## 6Conclusion
In this paper, we have proposed HintMiner, a machine comprehension and generation based approach to automatic mining hints for users’ questions\. Given a new question, HintMiner first selects relevant posts and filters away unimportant sentences in the retrieved posts from ES\. It then utilizes MiningNet to generate hints to the question from the paragraphs of the relevant posts\. MiningNet is an effective self\-supervised learning based model, which is able to distinguish proper contents from relevant posts to generate hints\. We conduct extensive experiments to evaluate the model effectiveness\. The evaluation results show that HintMiner outperforms several important Q&A methods and MiningNet is an effective hints generation neural network\. Our tool and experimental data are publicly availablehttps://github\.com/AnonymousAuthor2013/HintMiner\.
In the future, we plan to build a link prediction tool that can better find more relevant posts for a given question\. To further improve the capacity of HintMiner, we will also investigate models to comprehend code\(\[CODE\]\) and numerical expressions \(\[NUM\]\)in posts\. We will also explore more effective metric and perform user studies to evaluate the usefulness of our tool in practice\.
## References
- \[1\]\(2020\)Unilmv2: pseudo\-masked language models for unified language model pre\-training\.InInternational conference on machine learning,pp\. 642–652\.Cited by:[2nd item](https://arxiv.org/html/2609.16060#S4.I1.i2.p1.1)\.
- \[2\]D\. Britz, A\. Goldie, T\. Luong, and Q\. Le\(2017\)Massive Exploration of Neural Machine Translation Architectures\.ArXiv e\-prints\.External Links:1703\.03906Cited by:[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p1.1)\.
- \[3\]F\. Calefato, F\. Lanubile, and N\. Novielli\(2018\)How to ask for technical help? evidence\-based guidelines for writing questions on stack overflow\.Information and Software Technology94,pp\. 186–207\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p1.1)\.
- \[4\]J\. Carbonell and J\. Goldstein\(1998\)The use of mmr, diversity\-based reranking for reordering documents and producing summaries\.InProceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval,pp\. 335–336\.Cited by:[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[5\]C\. Chen, X\. Chen, J\. Sun, Z\. Xing, and G\. Li\(2018\)Data\-driven proactive policy assurance of post quality in community q&a sites\.Proceedings of the ACM on Human\-Computer Interaction2\(CSCW\),pp\. 33\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p1.1)\.
- \[6\]D\. Chen, A\. Fisch, J\. Weston, and A\. Bordes\(2017\)Reading wikipedia to answer open\-domain questions\.arXiv preprint arXiv:1704\.00051\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p2.1),[§1](https://arxiv.org/html/2609.16060#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.16060#S3.SS1.p2.1),[§5\.2](https://arxiv.org/html/2609.16060#S5.SS2.p1.1),[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[7\]W\. B\. Croft, D\. Metzler, and T\. Strohman\(2010\)Search engines: information retrieval in practice\.Vol\.283,Addison\-Wesley Reading\.Cited by:[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[8\]J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova\(2018\)Bert: pre\-training of deep bidirectional transformers for language understanding\.arXiv preprint arXiv:1810\.04805\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p3.1),[§3\.2\.1](https://arxiv.org/html/2609.16060#S3.SS2.SSS1.p1.1),[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p1.1),[2nd item](https://arxiv.org/html/2609.16060#S4.I1.i2.p1.1)\.
- \[9\]L\. Dong, F\. Wei, M\. Zhou, and K\. Xu\(2015\)Question answering over freebase with multi\-column convolutional neural networks\.InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),Vol\.1,pp\. 260–269\.Cited by:[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[10\]G\. Erkan and D\. R\. Radev\(2004\)Lexrank: graph\-based lexical centrality as salience in text summarization\.Journal of artificial intelligence research22,pp\. 457–479\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2.p2.1),[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[11\]T\. Gao, X\. Yao, and D\. Chen\(2021\)SimCSE: simple contrastive learning of sentence embeddings\.CoRRabs/2104\.08821\.External Links:[Link](https://arxiv.org/abs/2104.08821),2104\.08821Cited by:[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[12\]I\. Goodfellow, Y\. Bengio, A\. Courville, and Y\. Bengio\(2016\)Deep learning\.Vol\.1\.Cited by:[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p2.1)\.
- \[13\]J\. Gu, Z\. Lu, H\. Li, and V\. O\. Li\(2016\)Incorporating copying mechanism in sequence\-to\-sequence learning\.arXiv preprint arXiv:1603\.06393\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p3.1),[§3\.2\.1](https://arxiv.org/html/2609.16060#S3.SS2.SSS1.p2.1),[§5\.1](https://arxiv.org/html/2609.16060#S5.SS1.p1.1)\.
- \[14\]A\. Haghighi and L\. Vanderwende\(2009\)Exploring content models for multi\-document summarization\.InProceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics,pp\. 362–370\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2.p2.1)\.
- \[15\]D\. He, Y\. Xia, T\. Qin, L\. Wang, N\. Yu, T\. Liu, and W\. Ma\(2016\)Dual learning for machine translation\.InAdvances in Neural Information Processing Systems,pp\. 820–828\.Cited by:[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p2.1)\.
- \[16\]P\. Koehn and R\. Knowles\(2017\)Six challenges for neural machine translation\.InProceedings of the First Workshop on Neural Machine Translation, NMT@ACL 2017, Vancouver, Canada, August 4, 2017,pp\. 28–39\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2.p2.1),[§3\.2\.1](https://arxiv.org/html/2609.16060#S3.SS2.SSS1.p2.1)\.
- \[17\]C\. Lin\(2004\)Rouge: a package for automatic evaluation of summaries\.Text Summarization Branches Out\.Cited by:[§4\.2](https://arxiv.org/html/2609.16060#S4.SS2.p1.1)\.
- \[18\]R\. Mihalcea and P\. Tarau\(2004\)Textrank: bringing order into text\.InProceedings of the 2004 conference on empirical methods in natural language processing,Cited by:[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[19\]A\. Miller, A\. Fisch, J\. Dodge, A\. Karimi, A\. Bordes, and J\. Weston\(2016\)Key\-value memory networks for directly reading documents\.arXiv preprint arXiv:1606\.03126\.Cited by:[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[20\]K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu\(2002\)BLEU: a method for automatic evaluation of machine translation\.InProceedings of the 40th annual meeting on association for computational linguistics,pp\. 311–318\.Cited by:[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p2.1),[§4\.2](https://arxiv.org/html/2609.16060#S4.SS2.p1.1)\.
- \[21\]A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,et al\.\(2019\)Language models are unsupervised multitask learners\.OpenAI blog1\(8\),pp\. 9\.Cited by:[2nd item](https://arxiv.org/html/2609.16060#S4.I1.i2.p1.1)\.
- \[22\]C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu\(2020\)Exploring the limits of transfer learning with a unified text\-to\-text transformer\.The Journal of Machine Learning Research21\(1\),pp\. 5485–5551\.Cited by:[2nd item](https://arxiv.org/html/2609.16060#S4.I1.i2.p1.1)\.
- \[23\]P\. Rajpurkar, R\. Jia, and P\. Liang\(2018\)Know what you don’t know: unanswerable questions for squad\.arXiv preprint arXiv:1806\.03822\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p2.1),[§5\.2](https://arxiv.org/html/2609.16060#S5.SS2.p1.1)\.
- \[24\]M\. Ranzato, S\. Chopra, M\. Auli, and W\. Zaremba\(2015\)Sequence level training with recurrent neural networks\.arXiv preprint arXiv:1511\.06732\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.16060#S3.SS1.p2.1),[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p2.1)\.
- \[25\]M\. A\. C\. Soares and F\. S\. Parreiras\(2018\)A literature review on question answering techniques, paradigms and systems\.Journal of King Saud University\-Computer and Information Sciences\.Cited by:[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[26\]J\. Steinberger and K\. Jezek\(2004\)Using latent semantic analysis in text summarization and summary evaluation\.Proc\. ISIM4,pp\. 93–100\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2.p2.1)\.
- \[27\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems,pp\. 5998–6008\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p3.1),[§3\.2\.1](https://arxiv.org/html/2609.16060#S3.SS2.SSS1.p2.1)\.
- \[28\]O\. Vinyals, M\. Fortunato, and N\. Jaitly\(2015\)Pointer networks\.InAdvances in Neural Information Processing Systems,pp\. 2692–2700\.Cited by:[§5\.1](https://arxiv.org/html/2609.16060#S5.SS1.p1.1)\.
- \[29\]S\. Wang, T\. Chen, and A\. E\. Hassan\(2018\)Understanding the factors for fast answers in technical q&a websites\.Empirical Software Engineering23\(3\),pp\. 1552–1593\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.16060#S2.SS1.p1.1)\.
- \[30\]W\. Wang, N\. Yang, F\. Wei, B\. Chang, and M\. Zhou\(2017\)R\-net: machine reading comprehension with self\-matching networks\.Natural Lang\. Comput\. Group, Microsoft Res\. Asia, Beijing, China, Tech\. Rep5\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p2.1),[§5\.2](https://arxiv.org/html/2609.16060#S5.SS2.p1.1),[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[31\]W\. Wang, M\. Yan, and C\. Wu\(2018\)Multi\-granularity hierarchical attention fusion networks for reading comprehension and question answering\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vol\.1,pp\. 1705–1714\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p2.1),[§5\.2](https://arxiv.org/html/2609.16060#S5.SS2.p1.1),[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[32\]R\. S\. Wills\(2006\)Google’s pagerank\.The Mathematical Intelligencer28\(4\),pp\. 6–11\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2.p2.1),[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[33\]T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. L\. Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. M\. Rush\(2020\)Transformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Online,pp\. 38–45\.External Links:[Link](https://www.aclweb.org/anthology/2020.emnlp-demos.6)Cited by:[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p1.1),[2nd item](https://arxiv.org/html/2609.16060#S4.I1.i2.p2.1)\.
- \[34\]Y\. Wu, M\. Schuster, Z\. Chen, Q\. V\. Le, M\. Norouzi, W\. Macherey, M\. Krikun, Y\. Cao, Q\. Gao, K\. Macherey,et al\.\(2016\)Google’s neural machine translation system: bridging the gap between human and machine translation\.arXiv preprint arXiv:1609\.08144\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16060#S2.SS2.SSS2.p1.1)\.
- \[35\]B\. Xu, D\. Ye, Z\. Xing, X\. Xia, G\. Chen, and S\. Li\(2016\)Predicting semantically linkable knowledge in developer online forums via convolutional neural network\.In2016 31st IEEE/ACM International Conference on Automated Software Engineering \(ASE\),pp\. 51–62\.Cited by:[§2\.2](https://arxiv.org/html/2609.16060#S2.SS2.p1.1)\.
- \[36\]B\. Xu, Z\. Xing, X\. Xia, and D\. Lo\(2017\)AnswerBot: automated generation of answer summary to developersź technical questions\.InProceedings of the 32nd IEEE/ACM International Conference on Automated Software Engineering,pp\. 706–716\.Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p2.1),[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[37\]Y\. Yan, R\. Li, S\. Wang, F\. Zhang, W\. Wu, and W\. Xu\(2021\)ConSERT: A contrastive framework for self\-supervised sentence representation transfer\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, \(Volume 1: Long Papers\), Virtual Event, August 1\-6, 2021,C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),pp\. 5065–5075\.External Links:[Link](https://doi.org/10.18653/v1/2021.acl-long.393),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.393)Cited by:[1st item](https://arxiv.org/html/2609.16060#S4.I1.i1.p1.1)\.
- \[38\]D\. Ye, Z\. Xing, and N\. Kapre\(2017\)The structure and dynamics of knowledge network in domain\-specific q&a sites: a case study of stack overflow\.Empirical Software Engineering22\(1\),pp\. 375–406\.Cited by:[§2\.2](https://arxiv.org/html/2609.16060#S2.SS2.p1.1)\.
- \[39\]S\. W\. Yih, M\. Chang, X\. He, and J\. Gao\(2015\)Semantic parsing via staged query graph generation: question answering with knowledge base\.Cited by:[§5](https://arxiv.org/html/2609.16060#S5.p1.1)\.
- \[40\]X\. Yuan, T\. Wang, C\. Gulcehre, A\. Sordoni, P\. Bachman, S\. Subramanian, S\. Zhang, and A\. Trischler\(2017\)Machine comprehension by text\-to\-text neural question generation\.arXiv preprint arXiv:1705\.02012\.Cited by:[§3\.4](https://arxiv.org/html/2609.16060#S3.SS4.p2.1)\.
- \[41\]Q\. Zhou, N\. Yang, F\. Wei, and M\. Zhou\(2018\)Sequential copying networks\.InThirty\-Second AAAI Conference on Artificial Intelligence,Cited by:[§1](https://arxiv.org/html/2609.16060#S1.p3.1),[§3\.2\.1](https://arxiv.org/html/2609.16060#S3.SS2.SSS1.p2.1),[§5\.1](https://arxiv.org/html/2609.16060#S5.SS1.p1.1)\.Similar Articles
Self-Evolving Visual Questioner
This paper introduces a self-evolving framework for vision-language models to improve their question-generation capabilities without external supervision, enhancing both question quality and answerer performance.
J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers
J-Miner recovers executable decision knowledge from fine-tuned language-model classifiers by mining named concepts and learning decision rules, enabling inspection and transfer to lightweight models with high fidelity.
Diff Mining: Logit Differences Reveal Finetuning Objectives
The paper introduces Diff Mining, a framework for identifying finetuning objectives in language models by analyzing logit differences between finetuned and base models, enabling interpretable auditing of learned behaviors.
A Heuristic Perspective on Debiasing Language Models
This paper proposes HEIMAT, a heuristic-style automatic debiasing framework for language models that uses heuristic prompts to reveal biases and fine-tunes the model to reduce bias while preserving NLU performance.
Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
This paper presents HeuristicEdu, a pipeline to align Qwen2.5-7B as a Socratic tutor using supervised warm-up and GRPO with heuristic rewards, evaluated on a new dataset SocraticEdu, showing improved scaffolding effectiveness and reduced keyword leakage.