SCoP: 用于时序知识图谱问答的证据空间控制结构化约束解析
摘要
SCoP 引入了一个以约束为中心的框架,用于时序知识图谱问答,该框架将时序决策外部化以过滤证据,并在 MultiTQ 和 TimelineCronQ-R 等基准数据集上提高答案准确性。
arXiv:2609.22213v1 Announce Type: new
Abstract: Temporal Knowledge Graph Question Answering (TKGQA) requires answer inference from evidence that is both structurally valid and temporally admissible. Existing methods often leave anchor-event binding, temporal admissibility, and ordinal selection implicit in model reasoning, task-specific training, or similarity-driven retrieval, allowing locally relevant but invalid facts to enter the answer context. We formulate complex TKGQA as evidence-space control and propose SCoP (Structured Constraint Parsing), a constraint-centric framework that externalizes temporal decisions before answer inference. Instead of treating retrieved facts as admissible evidence by default, SCoP separates answer-seeking event patterns from temporal anchor events, conservatively grounds them to canonical TKG entities and relations, and translates temporal intent into executable constraints with optional ranking requirements. These constraints operate over normalized point and interval ranges, enabling deterministic filtering of structurally compatible candidates and producing a compact evidence space for generation. Experiments on MultiTQ and TimelineCronQ-R assess SCoP across timestamped point-fact and interval-oriented settings with richer temporal relations and ordering dependencies. Without task-specific parameter updates, SCoP achieves 0.825 Hits@1 on MultiTQ and 0.761 Hits@1 on TimelineCronQ-R, with gains on constraint-intensive question types. These results support explicit evidence-space control over unconstrained retrieval or implicit temporal reasoning.
查看缓存全文
缓存时间: 2026/09/22 09:08
# SCoP: Structured Constraint Parsing for Evidence-Space Control in Temporal Knowledge Graph Question Answering Source: [https://arxiv.org/html/2609.22213](https://arxiv.org/html/2609.22213) Conference:Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 07–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management \(CIKM ’26\), November 07–11, 2026, Rome, ItalyDOI:[10\.1145/3799682\.3840716](https://doi.org/10.1145/3799682.3840716)ISBN:979\-8\-4007\-2539\-5/2026/11CCS:Information systems Question answeringCCS:Computing methodologies Temporal reasoningCCS:Computing methodologies Knowledge representation and reasoningCCS:Information systems Information retrievalXiaokun GuoAffiliation:Institute of Information Engineering, Chinese Academy of Sciences,Beijing,ChinaAffiliation:School of Cyber Security,University of Chinese Academy of Sciences,Beijing,Chinaemail:[guoxiaokun@iie\.ac\.cn](mailto:[email protected])Zhen XuAffiliation:Institute of Information Engineering, Chinese Academy of Sciences,Beijing,ChinaAffiliation:School of Cyber Security,University of Chinese Academy of Sciences,Beijing,Chinaemail:[xuzhen@iie\.ac\.cn](mailto:[email protected]),Dongdong HuoNote:Corresponding author\.Affiliation:Institute of Information Engineering, Chinese Academy of Sciences,Beijing,ChinaAffiliation:School of Cyber Security,University of Chinese Academy of Sciences,Beijing,Chinaemail:[huodongdong@iie\.ac\.cn](mailto:[email protected]),Yanqiu ZhangAffiliation:Institute of Information Engineering, Chinese Academy of Sciences,Beijing,ChinaAffiliation:School of Cyber Security,University of Chinese Academy of Sciences,Beijing,Chinaemail:[zhangyanqiu@iie\.ac\.cn](mailto:[email protected]),Dongjin YuAffiliation:Institute of Information Engineering, Chinese Academy of Sciences,Beijing,ChinaAffiliation:School of Cyber Security,University of Chinese Academy of Sciences,Beijing,Chinaemail:[yudongjin@iie\.ac\.cn](mailto:[email protected])andYu WangAffiliation:Institute of Information Engineering, Chinese Academy of Sciences,Beijing,Chinaemail:[wangyu@iie\.ac\.cn](mailto:[email protected]) © cc ###### Abstract\. Temporal Knowledge Graph Question Answering \(TKGQA\) requires answer inference from evidence that is both structurally valid and temporally admissible\. Existing methods often leave anchor\-event binding, temporal admissibility, and ordinal selection implicit in model reasoning, task\-specific training, or similarity\-driven retrieval, allowing locally relevant but invalid facts to enter the answer context\. We formulate complex TKGQA as evidence\-space control and proposeSCoP\(StructuredConstraintParsing\), a constraint\-centric framework that externalizes temporal decisions before answer inference\. Instead of treating retrieved facts as admissible evidence by default, SCoP separates answer\-seeking event patterns from temporal anchor events, conservatively grounds them to canonical TKG entities and relations, and translates temporal intent into executable constraints with optional ranking requirements\. These constraints operate over normalized point and interval ranges, enabling deterministic filtering of structurally compatible candidates and producing a compact evidence space for generation\. Experiments on MultiTQ and TimelineCronQ\-R assess SCoP across timestamped point\-fact and interval\-oriented settings with richer temporal relations and ordering dependencies\. Without task\-specific parameter updates, SCoP achieves 0\.825 Hits@1 on MultiTQ and 0\.761 Hits@1 on TimelineCronQ\-R, with gains on constraint\-intensive question types\. These results support explicit evidence\-space control over unconstrained retrieval or implicit temporal reasoning\. ###### Keywords: Temporal Knowledge Graph Question Answering, Temporal Reasoning, Retrieval\-Augmented Generation, Constraint Parsing, Evidence Space Control, Knowledge Graphs ††cc\-license:byFigure 1\.Two failure modes in TKGQA caused by implicit temporal constraints: model\-internal decision errors and similarity\-induced noisy evidence\.Two\-panel illustration using a temporal question about negotiations with Ali Jafari\. The upper panel contrasts a constraint\-aware reasoning path that identifies an anchor event, filters candidate events by the before relation, and selects the latest valid event with an unconstrained reasoning path that focuses on local entity relevance and returns an invalid result\. The lower panel shows similarity\-based top\-k retrieval containing one valid candidate together with several candidates that occur after the anchor or fail the required ordering condition\.## 1\.Introduction Temporal Knowledge Graphs \(TKGs\) extend conventional knowledge graphs by attaching temporal annotations to facts, thereby enabling the representation of event order, temporal spans, and dynamic dependencies in evolving real\-world knowledge\([Trivedi et al\., 2017](https://arxiv.org/html/2609.22213#bib.bib3)\)\. Temporal Knowledge Graph Question Answering \(TKGQA\) aims to answer time\-sensitive questions based on such structured evidence\. Unlike conventional KGQA, TKGQA requires a system to jointly satisfy structural conditions over entities and relations, as well as temporal conditions over timestamps, time intervals, event ordering, and temporal granularity\. Real\-world temporal questions often involve multi\-hop dependencies, dynamically changing factual states, and diverse temporal expressions such as*before*,*during*, and*between*\([Jia et al\., 2021](https://arxiv.org/html/2609.22213#bib.bib4);[Zhang et al\., 2024](https://arxiv.org/html/2609.22213#bib.bib7);[Liu et al\., 2024](https://arxiv.org/html/2609.22213#bib.bib6)\)\. Therefore, a correct answer often depends on evidence that is admissible both structurally and temporally; even if a fact appears locally relevant, it may still lead to an incorrect conclusion if its structural or temporal conditions are not satisfied\. Although temporal annotations are already encoded in TKGs, selecting admissible evidence for complex temporal questions remains challenging\. Existing approaches may improve TKGQA through task\-specific fine\-tuning, model\-generated reasoning traces, or similarity\-driven retrieval, yet key temporal admissibility decisions are often still internalized in learned model behavior or approximated through retrieval relevance\. As shown in Fig\.[1](https://arxiv.org/html/2609.22213#S0.F1)\(a\), when anchor events, temporal conditions, temporal offsets, and ordering constraints remain implicit, the system may fail to bind them reliably to concrete event facts or execute the required temporal comparisons consistently\. As shown in Fig\.[1](https://arxiv.org/html/2609.22213#S0.F1)\(b\), similarity\-based retrieval may return facts that are locally related to the question in terms of entities, relations, or timestamps, but fail to satisfy its complete structural or temporal requirements\([Gade et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib5)\)\. Such unconstrained top\-kkcontexts can enlarge the evidence space and increase the risk that the answer model reasons over inadmissible evidence\([Fayyaz et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib24)\)\. We therefore study complex TKGQA from an evidence\-space control perspective governed by question\-induced constraints, rather than relying solely on model\-internal temporal reasoning\. In this work, evidence\-space control refers to constructing, before answer inference, a compact candidate set that is structurally compatible with the question and temporally admissible under its constraints\. To this end, we proposeSCoP\(StructuredConstraintParsing\), a constraint\-centric TKGQA framework that externalizes key temporal decisions into structured executable constraints before answer generation\.111Code, prompts, and datasets are available at:[https://github\.com/gxiaokun/scop](https://github.com/gxiaokun/scop)\.Unlike full logical\-form generation, SCoP does not specify a complete procedure for deriving the answer; instead, it structures the event and temporal conditions needed to determine which graph facts are admissible evidence\. Without task\-specific parameter updates, SCoP identifies answer\-seeking events and temporal anchor events, conservatively aligns them to canonical TKG elements, and applies executable temporal constraints to construct a compact, constraint\-compliant evidence space for final answer inference\. We evaluate SCoP in two progressively more demanding TKGQA settings\. We first conduct experiments onMultiTQ\([Chen et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib17)\), a standard multi\-granularity benchmark that primarily evaluates question answering over timestamped point facts\. We then evaluate SCoP onTimelineCronQ\-R, a verified evaluation variant derived from the CronQuestions\-KG subset generated byTimelineKGQA\([Sun et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib18)\)\. In this benchmark, facts are associated with temporal intervals, and questions involve richer temporal relations, temporal operations, and ordering dependencies\. Results on both datasets show that SCoP remains consistently effective as temporal reasoning moves from point\-based temporal facts to more complex interval\-based and relation\-intensive scenarios\. Overall, our main contributions are summarized as follows: - •Structured constraint parsing for evidence\-space control\.We instantiate evidence\-space control by externalizing question\-induced structural and temporal admissibility conditions before answer inference\. Rather than specifying a complete logical form for answer derivation, SCoP uses these structured conditions to determine which KG facts form the admissible evidence space for downstream answer inference\. - •Structured decomposition of temporal questions into graph\-compatible retrieval targets\.SCoP separates answer\-seeking event patterns from temporal anchor events and conservatively grounds the resulting structures to canonical TKG entities and relations, providing graph\-compatible retrieval targets without introducing unsupported facts\. This separation keeps event retrieval structurally grounded while leaving temporal intent to the subsequent constraint layer\. - •A structured, executable temporal constraint schema supporting deterministic filtering\.SCoP represents event\-referenced, explicit\-time, and interval conditions as executable constraints over normalized point\-or\-interval ranges, and compiles them into deterministic admissibility predicates for candidate evidence\. The schema supports conjunctive multi\-anchor constraints, temporal granularity, and day\-level offsets, while handling ordinal ranking separately from admissibility filtering\. ## 2\.Related Work Earlier work on temporal and complex knowledge\-base question answering explored semantic parsing and symbolic execution to explicitly represent temporal or compositional semantics through structured intermediate representations, including TEQUILA\([Jia et al\., 2018](https://arxiv.org/html/2609.22213#bib.bib8)\), SYGMA\([Neelam et al\., 2022](https://arxiv.org/html/2609.22213#bib.bib9)\), SF\-TQA\([Ding et al\., 2022](https://arxiv.org/html/2609.22213#bib.bib10)\), and Prog\-TQA\([Chen et al\., 2024b](https://arxiv.org/html/2609.22213#bib.bib2)\)\. These approaches instantiate structured reasoning in different forms: TEQUILA decomposes temporal questions into non\-temporal subquestions and temporal constraints; SF\-TQA uses a semantic framework of temporal constraints to guide query\-graph generation; Prog\-TQA generates symbolic program drafts that are aligned to the TKG and executed; and SYGMA produces a KB\-agnostic logical representation with high\-level reasoning constructs before KB\-specific question mapping and answering\. SCoP is related to this line of work in making temporal reasoning conditions explicit, but uses its structured representation specifically to control which TKG facts form the admissible evidence space before downstream answer inference, rather than directly using it for query construction or answer execution\. A major line of TKGQA research encodes questions, entities, relations, and temporal annotations into continuous vector spaces, and predicts answers through learned scoring or similarity functions\. EmbedKGQA\([Saxena et al\., 2020](https://arxiv.org/html/2609.22213#bib.bib19)\)uses KG embeddings for multi\-hop KGQA, while CronKGQA\([Saxena et al\., 2021](https://arxiv.org/html/2609.22213#bib.bib11)\)and TempoQR\([Mavromatis et al\., 2022](https://arxiv.org/html/2609.22213#bib.bib12)\)further adapt representation\-based reasoning to temporal QA over knowledge graphs by combining temporal KG embeddings with a question encoder\. MultiQA\([Chen et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib17)\)further studies multi\-granularity temporal QA by modeling temporal semantics at different granularities\. Graph\-enhanced variants, such as LGQA\([Liu et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib14)\)and TwiRGCN\([Sharma et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib15)\), also incorporate temporal structure or temporally weighted message passing\. These methods improve temporal representation learning and candidate ranking, especially in sparse or multi\-hop settings\. However, their evidence selection is still primarily governed by learned representations or similarity\-based scoring\. For complex temporal questions, semantically related facts may still violate anchor\-event conditions, interval boundaries, temporal granularity, offset constraints, or ordering requirements\. Recent work has also adapted LLMs, retrievers, or agent policies to temporal QA through task\-specific training\. TimeR4\([Qian et al\., 2024](https://arxiv.org/html/2609.22213#bib.bib21)\)integrates rewriting, temporal retrieval, reranking, and generation for RAG\-based TKGQA\. GenTKGQA\([Gao et al\., 2024](https://arxiv.org/html/2609.22213#bib.bib16)\)follows a two\-stage framework of subgraph retrieval and answer generation, using LLM\-guided temporal and structural cues together with graph signals\. PoK\([Qian et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib25)\)decomposes complex temporal questions into planned sub\-objectives and retrieves temporally aligned facts from a temporal knowledge store\. Beyond TKGQA, Search\-R1\([Jin et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib26)\)shows that reinforcement learning can train LLMs to interact with external search environments during step\-by\-step reasoning\. More recently, Temp\-R1\([Gong et al\., 2026](https://arxiv.org/html/2609.22213#bib.bib27)\)formulates complex TKGQA as an autonomous agent problem and improves temporal reasoning through an expanded action space and reverse curriculum reinforcement learning\. These methods demonstrate that task\-specific supervision, fine\-tuned retrievers, instruction tuning, or reinforcement learning can improve temporal reasoning and tool\-use behavior\. However, their effectiveness is often coupled with trained model parameters, supervised trajectories, or learned action policies, and temporal evidence selection is largely internalized into model behavior\. Another line of work uses LLMs as planners, decomposers, or reasoning controllers without necessarily updating model parameters\. ARI\([Chen et al\., 2024c](https://arxiv.org/html/2609.22213#bib.bib20)\)induces abstract reasoning procedures from historical temporal QA examples and separates knowledge\-agnostic reasoning guidance from knowledge\-based answering\. TempAgent\([Hu et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib28)\)adapts the ReAct paradigm to TKGQA and introduces temporal\-domain tools for interactive reasoning\. RTQA\([Gong et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib13)\)recursively decomposes complex temporal questions into sub\-problems, solves them bottom\-up, and aggregates multiple reasoning paths to mitigate error propagation\. MemoTime\([Tan et al\., 2026](https://arxiv.org/html/2609.22213#bib.bib29)\)augments temporal reasoning with memory of verified reasoning traces and operator\-aware retrieval strategies\. These methods improve over direct LLM answering by introducing planning, decomposition, tool invocation, or memory\-based reuse\. However, key temporal decisions often remain in natural\-language reasoning traces, including anchor\-event selection, temporal\-condition verification, and ordering\-based candidate selection\. When such decisions are not compiled into explicit predicates, they can be difficult to verify or execute consistently, and noisy retrieved evidence may still enter the final answering context\. Overall, SCoP differs from task\-trained temporal reasoners and prompt\-based LLM strategies by treating complex TKGQA as explicit evidence\-space control\. Rather than relying on learned policies or natural\-language reasoning traces to resolve temporal admissibility, it compiles question\-induced constraints into executable filtering conditions before answer inference\. ## 3\.Method Figure 2\.Overview of the SCoP framework\. SCoP consists of one offline retrieval\-store preparation stage, three query\-time evidence\-control stages, and a final answer inference stage: \(I\) Temporal Graph and Retrieval Store Preparation; \(II\) Event Triple Identification, which extracts anchor triples and an optional search triple; \(III\) Conservative Graph Alignment, which grounds extracted triples to canonical TKG elements; \(IV\) Constraint\-guided Evidence Construction, which retrieves structurally compatible facts and filters them using parsed temporal constraints and ranking requirements; and \(V\) Answer Inference over the constrained evidence space\.Pipeline diagram of SCoP with five stages\. An offline temporal knowledge graph and retrieval store provide entities, relations, and timestamps\. At query time, a sample question is decomposed into two anchor triples and one search triple, which are conservatively aligned to canonical graph entities and relations\. Temporal constraints and a ranking requirement are then parsed and applied to structurally compatible candidate triples\. The filtered and ranked evidence is finally supplied to the answer inference stage\.Given a temporal knowledge graphGGcomposed of temporal factsxx, SCoP performs temporal question answering by first constructing a controlled evidence space and then conducting answer inference over the constrained evidence context\. Specifically, SCoP consists of one offline retrieval\-store preparation stage and three query\-time evidence\-control stages: event triple identification, conservative graph alignment, and constraint\-guided evidence construction\. The query\-time stages first extract question\-relevant event structures, then ground them to canonical TKG elements, and finally execute question\-induced temporal constraints and optional ranking requirements over structurally compatible candidate facts\. The resulting constrained evidence is passed to the answer generation module, so that final inference is performed over a temporally controlled evidence context rather than an unconstrained retrieval set\. Figure[2](https://arxiv.org/html/2609.22213#S3.F2)presents an overview of the SCoP framework\. ### 3\.1\.Temporal Graph and Retrieval Store Preparation Before question\-specific reasoning, SCoP organizes the temporal knowledge graph into a graph\-backed retrieval store\. We represent it as a directed temporal graph G=\(V,E\),x=\(s,r,o,ts,te\)∈E,G=\(V,E\),\\qquad x=\(s,r,o,t\_\{s\},t\_\{e\}\)\\in E,whereVVis the entity set, and eachxxis a directed edge fromsstooo, labeled by its canonical relationrrand temporal scope\(ts,te\)\(t\_\{s\},t\_\{e\}\)\. To uniformly support point events and interval\-valued facts, SCoP preserves the temporal scope of each edge in a normalized point\-or\-interval form\. For a factxx, we denote this stored temporal scope byτx\\tau\_\{x\}, wherets=tet\_\{s\}=t\_\{e\}corresponds to a point event andts<tet\_\{s\}<t\_\{e\}corresponds to an interval fact\. This representation is later used to resolve anchor\-event times and execute temporal filtering\. To support schema grounding and candidate retrieval, SCoP builds dense indices over the canonical entity and relation vocabularies using an embedding encoderϕ\(⋅\)\\phi\(\\cdot\)and FAISS\([Johnson et al\., 2021](https://arxiv.org/html/2609.22213#bib.bib22)\)\. At query time, these indices and the graph\-backed retrieval store provide candidate entities, relations, and graph\-compatible fact patterns for subsequent alignment and constrained evidence construction\. Structural validity and temporal admissibility are determined only in later stages through conservative graph alignment and explicit constraint execution\. ### 3\.2\.Event Triple Identification At this stage, SCoP transforms the question into retrieval\-oriented structured event representations, and further maps it into an event pattern consisting of a set of known anchor events and at most one event to be retrieved: \(𝒜0\(q\),u0\(q\)\)=Parseevt\(q\),\(\\mathcal\{A\}\_\{0\}\(q\),u\_\{0\}\(q\)\)=\\mathrm\{Parse\}\_\{\\mathrm\{evt\}\}\(q\),where 𝒜0\(q\)=\[a0,…,am−1\],ai=\(si,ri,oi\),\\mathcal\{A\}\_\{0\}\(q\)=\[a\_\{0\},\\ldots,a\_\{m\-1\}\],\\qquad a\_\{i\}=\(s\_\{i\},r\_\{i\},o\_\{i\}\),denotes the extracted anchor triples, andu0\(q\)u\_\{0\}\(q\)denotes the extracted search triple\. Each anchor triple is a complete event description and never contains the unknown slot “?”, whereas the search triple, if present, is the only structure that may contain it\. For entity\-seeking questions, SCoP constructs u0\(q\)=\(s,r,o\),u\_\{0\}\(q\)=\(s,r,o\),with exactly one unknown slot in either the subject or object position, while keeping the relation slot explicit\. For questions whose target is not an unknown entity, such as pure timestamp lookup or closed\-target temporal\-computation queries, SCoP sets u0\(q\)=∅,u\_\{0\}\(q\)=\\varnothing,and retains the relevant complete events as anchor triples for downstream timestamp lookup, temporal comparison, or answer realization\. For example, given the question*“When was Bob Rock nominated for the JR\-Producer Award?”*, SCoP retains a0=\(Bob Rock,nominated for,JR\-Producer Award\)a\_\{0\}=\(\\text\{Bob Rock\},\\text\{nominated for\},\\text\{JR\-Producer Award\}\)as an anchor triple and setsu0\(q\)=∅u\_\{0\}\(q\)=\\varnothing\. A key design principle is that extracted triples encode only event content, not temporal intent\. Temporal expressions and operators, including*before*,*after*, explicit dates or intervals,*same day/month/year*, and ordinal modifiers such as*first*or*last*, are excluded from triple slots and handled later as executable temporal constraints or ranking requirements\. This separation keeps the extracted triples suitable for graph retrieval and prevents temporal conditions from being conflated with entities or relations\. Anchor triples are introduced only when a complete event is explicitly stated in the question or can be minimally recovered from a local parallel reference\. Pure temporal expressions, such as*before 2005*,*after July 2014*, or*in May 2009*, do not form anchors by themselves\. At this stage, relation phrases remain close to the question wording and are not required to match canonical TKG relation names; canonicalization is performed in the subsequent conservative graph alignment stage\. ### 3\.3\.Conservative Graph Alignment Triples extracted from natural language may use surface\-form entity names or relation phrases that do not exactly match the canonical schema of the temporal knowledge graph\. SCoP therefore performs conservative graph alignment before constraint execution, normalizing extracted event structures into graph\-compatible forms while preserving their original query semantics rather than rewriting the question or inferring missing facts\. Given\(𝒜0\(q\),u0\(q\)\)\(\\mathcal\{A\}\_\{0\}\(q\),u\_\{0\}\(q\)\), SCoP constructs a local candidate schema space 𝒱q=\(ℰq,ℛq,ℱq\),\\mathcal\{V\}\_\{q\}=\(\\mathcal\{E\}\_\{q\},\\mathcal\{R\}\_\{q\},\\mathcal\{F\}\_\{q\}\),whereℰq\\mathcal\{E\}\_\{q\},ℛq\\mathcal\{R\}\_\{q\}, andℱq\\mathcal\{F\}\_\{q\}denote retrieved candidate entities, relations, and fact triples, respectively\. For each extracted triple pattern, SCoP builds graph\-backed candidate facts via slot\-level retrieval\. Complete anchor triples retrieve candidates for all three slots; search triples retrieve candidates only for the observed slots and complete the unknown slot by enumerating compatible triples already present inGG, ensuring that all retained candidates correspond to valid graph facts\. For a candidate factf=\(s,r,o\)f=\(s,r,o\), its retrieval score is computed as Score\(f\)=sim\(s,s~\)\+sim\(r,r~\)\+sim\(o,o~\),\\mathrm\{Score\}\(f\)=\\mathrm\{sim\}\(s,\\tilde\{s\}\)\+\\mathrm\{sim\}\(r,\\tilde\{r\}\)\+\\mathrm\{sim\}\(o,\\tilde\{o\}\),where\(s~,r~,o~\)\(\\tilde\{s\},\\tilde\{r\},\\tilde\{o\}\)denotes the extracted surface\-form triple\. For the unknown slot, SCoP setssim\(z,z~\)=1\\mathrm\{sim\}\(z,\\tilde\{z\}\)=1whenz~=?\\tilde\{z\}=\\texttt\{?\}\. This constant contribution does not affect the relative ranking among candidates from the same search pattern\. Candidate facts are then ranked byScore\(f\)\\mathrm\{Score\}\(f\)and retained inℱq\\mathcal\{F\}\_\{q\}for downstream alignment\. The alignment module then outputs \(𝒜\(q\),u\(q\)\)=Align\(q,𝒜0\(q\),u0\(q\),𝒱q\),\(\\mathcal\{A\}\(q\),u\(q\)\)=\\mathrm\{Align\}\(q,\\mathcal\{A\}\_\{0\}\(q\),u\_\{0\}\(q\),\\mathcal\{V\}\_\{q\}\),where 𝒜\(q\)=\[a^0,…,a^m−1\]\\mathcal\{A\}\(q\)=\[\\hat\{a\}\_\{0\},\\ldots,\\hat\{a\}\_\{m\-1\}\]is the aligned anchor list, andu\(q\)u\(q\)is the aligned search triple or null\. SCoP adopts a conservative normalization strategy\. The alignment output preserves the number and order of anchor triples, represents every aligned triple as a three\-field structure\(s,r,o\)\(s,r,o\), and never fills or removes the unknown slot “?”\. By default, the original head–tail orientation and the position of “?” are preserved\. Only the search triple may undergo conservative orientation repair: conditioned on the original question and anchor\-supported candidate evidence, the alignment model may reverse its extracted direction when the reversed orientation is judged to better preserve the intended retrieval semantics\. The original question is used only for semantic disambiguation, such as resolving local coreference, role mentions, or relation\-action ambiguity, and is not used to introduce new entities, relations, or facts\. When no reliable normalization is available, the original surface form is retained\. Anchor and search triples are aligned differently\. Because anchor triples later serve as temporal references, SCoP prioritizes complete event\-level grounding for each ai=\(si,ri,oi\)\.a\_\{i\}=\(s\_\{i\},r\_\{i\},o\_\{i\}\)\.When a reliable candidate fact inℱq\\mathcal\{F\}\_\{q\}matches the complete event structure, SCoP adopts that fact\-level grounded triple\. Otherwise, it applies conservative slot\-level normalization overℰq\\mathcal\{E\}\_\{q\}andℛq\\mathcal\{R\}\_\{q\}: each entity or relation slot is canonicalized only when its mapping is sufficiently reliable, while uncertain slots retain their original surface forms\. This design avoids fabricating fully grounded events by combining partial matches from unrelated graph facts\. Because the search triple is an open event pattern rather than a closed fact, SCoP normalizes only its known entity slots and relation phrase: u0\(q\)=\(s,r,o\)⟹u\(q\)=\(s^,r^,o^\),u\_\{0\}\(q\)=\(s,r,o\)\\quad\\Longrightarrow\\quad u\(q\)=\(\\hat\{s\},\\hat\{r\},\\hat\{o\}\),while keeping the unknown slot open\. Across both anchor and search alignment, SCoP enforces entity\-type consistency and relation\-action consistency, preventing surface similarity from altering the intended event semantics\. In the running example of Figure[2](https://arxiv.org/html/2609.22213#S3.F2), conservative alignment normalizes*JR\-Producer Award*,*was a member of*, and*nominated*to the graph\-compatible forms*JRP Award*,*member of*, and*nominated for*, respectively, while preserving the open answer slot\. ### 3\.4\.Constraint\-guided Evidence Construction After event triple identification and conservative graph alignment, SCoP converts the question into graph\-executable structural retrieval conditions\. If a search triple is present, the aligned triple defines an open but structurally restricted candidate space; otherwise, aligned complete event triples serve as closed retrieval targets for timestamp lookup or downstream temporal processing\. For search\-triple questions, temporal conditions that cannot be encoded in retrieval triples themselves, such as explicit dates, event\-relative ordering, intervals, and offsets, are parsed into executable constraints that remove structurally matched but temporally inadmissible facts\. Rather than assigning a single temporal\-relation label to the whole question, SCoP parses an ordered constraint list together with an optional post\-filter ranking object: \(𝒦\(q\),ρ\(q\)\)=Parsetmp\(q,𝒜\(q\),u\(q\)\),\(\\mathcal\{K\}\(q\),\\rho\(q\)\)=\\mathrm\{Parse\}\_\{\\mathrm\{tmp\}\}\(q,\\mathcal\{A\}\(q\),u\(q\)\),where 𝒦\(q\)=\[κ0,…,κn\]\\mathcal\{K\}\(q\)=\[\\kappa\_\{0\},\\ldots,\\kappa\_\{n\}\]denotes temporal admissibility constraints, and ρ\(q\)∈\(\{asc,desc\}×ℕ\>0\)∪\{∅\}\\rho\(q\)\\in\\left\(\\\{\\text\{asc\},\\text\{desc\}\\\}\\times\\mathbb\{N\}\_\{\>0\}\\right\)\\cup\\\{\\varnothing\\\}denotes an optional ordinal ranking requirement over admissible candidates\. When a search triple is present,Parsetmp\\mathrm\{Parse\}\_\{\\mathrm\{tmp\}\}instantiates temporal admissibility constraints and, when needed, ordinal ranking requirements over the open answer\-bearing candidate space\. Whenu\(q\)=∅u\(q\)=\\varnothing, the same parsing stage resolves the question into an anchor\-only closed\-target retrieval mode: the aligned complete event triples directly specify the facts to be retrieved for timestamp lookup or downstream temporal answer realization\. Since this mode introduces no open answer\-bearing candidate space that requires additional temporal admissibility filtering, it yields𝒦\(q\)=\[\]\\mathcal\{K\}\(q\)=\[\]andρ\(q\)=∅\\rho\(q\)=\\varnothing\. Each temporal constraint has the form κi=\(αi,ωi,ηi,gi,δi\),\\kappa\_\{i\}=\(\\alpha\_\{i\},\\omega\_\{i\},\\eta\_\{i\},g\_\{i\},\\delta\_\{i\}\),whereαi\\alpha\_\{i\}specifies the reference type,ωi\\omega\_\{i\}the temporal operator,ηi\\eta\_\{i\}the reference content,gig\_\{i\}the comparison granularity, andδi\\delta\_\{i\}an optional day\-level offset\. Specifically, αi∈\{event,explicit\_time,explicit\_interval\},\\alpha\_\{i\}\\in\\\{\\text\{event\},\\text\{explicit\\\_time\},\\text\{explicit\\\_interval\}\\\},ωi∈\{before,after,equal,inside,overlap\}\.\\omega\_\{i\}\\in\\\{\\text\{before\},\\text\{after\},\\text\{equal\},\\text\{inside\},\\text\{overlap\}\\\}\.The reference payloadηi\\eta\_\{i\}depends onαi\\alpha\_\{i\}: ηi=\{j,αi=event,j∈\{0,…,\|𝒜\(q\)\|−1\},τ,αi=explicit\_time,\(τs,τe\),αi=explicit\_interval\.\\eta\_\{i\}=\\begin\{cases\}j,&\\alpha\_\{i\}=\\text\{event\},\\quad j\\in\\\{0,\\ldots,\|\\mathcal\{A\}\(q\)\|\-1\\\},\\\\ \\tau,&\\alpha\_\{i\}=\\text\{explicit\\\_time\},\\\\ \(\\tau\_\{s\},\\tau\_\{e\}\),&\\alpha\_\{i\}=\\text\{explicit\\\_interval\}\.\\end\{cases\}Here,jjindexes an aligned anchor event in𝒜\(q\)\\mathcal\{A\}\(q\),τ\\taudenotes an explicit temporal expression, and\(τs,τe\)\(\\tau\_\{s\},\\tau\_\{e\}\)denotes an explicit temporal interval\. Granularity is used only forequal, with gi∈\{day,month,year\},g\_\{i\}\\in\\\{\\text\{day\},\\text\{month\},\\text\{year\}\\\},and is null otherwise\. The offset fieldδi\\delta\_\{i\}is used for relative day\-offset expressions such as “NNdays before” or “NNdays after”; it stores a positive day count, with the temporal direction determined byωi\\omega\_\{i\}\. Ranking is separated from admissibility filtering:\(asc,k\)\(\\text\{asc\},k\)and\(desc,k\)\(\\text\{desc\},k\)indicate earlier\- and later\-oriented ordinal preferences, respectively, and retain a compact rank\-focused subset to accommodate ties, temporal granularity ambiguity, and multiple valid answers\. The operator inventory is designed as a compact set of executable primitives rather than an exhaustive taxonomy of temporal expressions\. More complex conditions are represented through their combination with reference types, granularity, offsets, and ranking requirements: before and after express directional comparisons; in and during correspond to interval inclusion; overlap captures temporal overlap; and between can be represented by an explicit interval or paired after–before constraints\. Ordinal modifiers such as first and last are handled separately through post\-filter ranking\. Table[1](https://arxiv.org/html/2609.22213#S3.T1)summarizes the structured temporal constraint and ranking schema used by SCoP\. Table 1\.Structured temporal constraint and ranking schema\.SCoP next defines the retrieval targets used to construct the initial candidate evidence space\. For entity\-seeking questions, the aligned search tripleu\(q\)u\(q\)is the target pattern; for time\-seeking questions,u\(q\)=∅u\(q\)=\\varnothing, and aligned complete event triples are used as closed target patterns for timestamp retrieval: 𝒰\(q\)=\{\{u\(q\)\},u\(q\)≠∅,𝒜\(q\),u\(q\)=∅∧\|𝒜\(q\)\|\>0\.\\mathcal\{U\}\(q\)=\\begin\{cases\}\\\{u\(q\)\\\},&u\(q\)\\neq\\varnothing,\\\\ \\mathcal\{A\}\(q\),&u\(q\)=\\varnothing\\land\|\\mathcal\{A\}\(q\)\|\>0\.\\end\{cases\}Given𝒰\(q\)\\mathcal\{U\}\(q\), SCoP retrieves structurally compatible candidate facts: C\(q\)=\{x∈E∣∃u∈𝒰\(q\),Γ\(x,u\)\},C\(q\)=\\\{x\\in E\\mid\\exists u\\in\\mathcal\{U\}\(q\),\\Gamma\(x,u\)\\\},where each temporal factx∈Ex\\in Eis represented asx=\(sx,rx,ox,τx\)x=\(s\_\{x\},r\_\{x\},o\_\{x\},\\tau\_\{x\}\)\. Foru=\(su,ru,ou\)u=\(s\_\{u\},r\_\{u\},o\_\{u\}\), structural compatibility is defined as Γ\(x,u\)=𝕀\[rx≡ru\]⋅𝕀\[su=?∨sx≡su\]⋅𝕀\[ou=?∨ox≡ou\],\\Gamma\(x,u\)=\\mathbb\{I\}\[r\_\{x\}\\equiv r\_\{u\}\]\\cdot\\mathbb\{I\}\[s\_\{u\}=\\texttt\{?\}\\vee s\_\{x\}\\equiv s\_\{u\}\]\\cdot\\mathbb\{I\}\[o\_\{u\}=\\texttt\{?\}\\vee o\_\{x\}\\equiv o\_\{u\}\],where≡\\equivdenotes schema\-level equivalence after alignment\. Each parsed constraintκi\\kappa\_\{i\}is then compiled into an executable admissibility predicate: Θκi\(x,G\)∈\{0,1\}\.\\Theta\_\{\\kappa\_\{i\}\}\(x;G\)\\in\\\{0,1\\\}\.During execution, each candidate or reference timestamp is normalized into a closed date range τ¯\(z\)=\[lz,uz\]\.\\bar\{\\tau\}\(z\)=\[l\_\{z\},u\_\{z\}\]\.Day\-level timestamps satisfylz=uzl\_\{z\}=u\_\{z\}, while month\- and year\-level expressions are expanded to their corresponding calendar spans\. For event\-referenced constraints, the indexed anchor triple may retrieve multiple graph facts; SCoP retains all resolvable anchor time ranges and regards a candidate as valid if it satisfies the constraint with respect to at least one such range\. If no resolvable anchor time is available, the corresponding constraint yields no admissible candidates\. Explicit time expressions and explicit intervals are normalized in the same form\. Given a candidate rangeX=\[lx,ux\]X=\[l\_\{x\},u\_\{x\}\]and a reference rangeA=\[la,ua\]A=\[l\_\{a\},u\_\{a\}\], SCoP evaluates before\(X,A\):ux<la,inside\(X,A\):lx≥la∧ux≤ua\\text\{before\}\(X,A\):\\ u\_\{x\}<l\_\{a\},\\quad\\text\{inside\}\(X,A\):\\ l\_\{x\}\\geq l\_\{a\}\\land u\_\{x\}\\leq u\_\{a\}after\(X,A\):lx\>ua,overlap\(X,A\):lx≤ua∧ux≥la\.\\text\{after\}\(X,A\):\\ l\_\{x\}\>u\_\{a\},\\quad\\text\{overlap\}\(X,A\):\\ l\_\{x\}\\leq u\_\{a\}\\land u\_\{x\}\\geq l\_\{a\}\.Forequal, both ranges are first coarsened to the required granularitygig\_\{i\}, and the predicate holds when the coarsened ranges overlap\. For offset constraints,afterderives target dates from \{la\+δi,ua\+δi\},\\\{l\_\{a\}\+\\delta\_\{i\},\\;u\_\{a\}\+\\delta\_\{i\}\\\},whilebeforederives target dates from \{la−δi,ua−δi\}\.\\\{l\_\{a\}\-\\delta\_\{i\},\\;u\_\{a\}\-\\delta\_\{i\}\\\}\.A candidate satisfies the offset constraint if its normalized range covers at least one derived target date\. Conjunctive execution of all parsed constraints yields the filtered evidence space: C~\(q\)=\{x∈C\(q\)∣⋀κi∈𝒦\(q\)Θκi\(x,G\)\}\.\\widetilde\{C\}\(q\)=\\\{x\\in C\(q\)\\mid\\bigwedge\_\{\\kappa\_\{i\}\\in\\mathcal\{K\}\(q\)\}\\Theta\_\{\\kappa\_\{i\}\}\(x;G\)\\\}\.This formulation naturally supports multi\-anchor questions, where different constraints may refer to different temporal references and must hold simultaneously\. For the running example, the temporal parser produces κ0=\(event,before,0,∅,∅\),κ1=\(event,overlap,1,∅,∅\),\\kappa\_\{0\}=\(\\text\{event\},\\text\{before\},0,\\varnothing,\\varnothing\),\\qquad\\kappa\_\{1\}=\(\\text\{event\},\\text\{overlap\},1,\\varnothing,\\varnothing\),together with the ranking requirementρ\(q\)=\(asc,1\)\\rho\(q\)=\(\\text\{asc\},1\)\. The two constraints jointly enforce the anchor\-relative temporal conditions, while the ranking object captures the ordinal modifier*first*\. After temporal filtering, SCoP applies the parsed ranking object: C∗\(q\)=\{RankSelect\(C~\(q\),ρ\(q\)\),ρ\(q\)≠∅,C~\(q\),ρ\(q\)=∅\.C^\{\*\}\(q\)=\\begin\{cases\}\\mathrm\{RankSelect\}\(\\widetilde\{C\}\(q\);\\rho\(q\)\),&\\rho\(q\)\\neq\\varnothing,\\\\ \\widetilde\{C\}\(q\),&\\rho\(q\)=\\varnothing\.\\end\{cases\}Here,RankSelect\\mathrm\{RankSelect\}sorts admissible candidates by the start date of their normalized temporal ranges and retains rank\-oriented subsets according toρ\(q\)\\rho\(q\)\. To handle ties and temporal ambiguity,*first*and*last*preserve the top three facts, while explicit ordinal constraints \(e\.g\., “the fifth”\) retain candidates up to the requested rank\. Finally, SCoP formats the retrieved temporal facts into a compact answer contextℋ\(q\)\\mathcal\{H\}\(q\), and invokes a=LLM\(q,ℋ\(q\)\)\.a=\\mathrm\{LLM\}\(q,\\mathcal\{H\}\(q\)\)\.For questions with a search triple,ℋ\(q\)\\mathcal\{H\}\(q\)contains the retrieved anchor evidence together with the filtered and ranked search facts\. For questions without a search triple, it contains the retrieved anchor facts used for downstream answer realization\. As a result, final inference is performed over a structurally controlled evidence context, with explicit temporal admissibility filtering applied whenever an open answer\-bearing candidate space is present\. ## 4\.Experiment #### Datasets\. We evaluate SCoP on two TKGQA datasets:MultiTQ\([Chen et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib17)\)andTimelineCronQ\-R\.MultiTQis a large\-scale multi\-granularity benchmark constructed from ICEWS05–15\([García\-Durán et al\., 2018](https://arxiv.org/html/2609.22213#bib.bib23)\), and we use it to evaluate SCoP under a standard temporal QA setting where facts are primarily represented by timestamps at different granularities\.TimelineCronQ\-Rreconstructs the CronQuestions\-KG subset generated byTimelineKGQA\([Sun et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib18)\), which is derived from CronQuestions\([Saxena et al\., 2021](https://arxiv.org/html/2609.22213#bib.bib11)\)\. Compared withMultiTQ, it is built on an interval\-oriented temporal KG, where facts are associated with temporal scopes defined by start and end times\. In addition,TimelineKGQAcharacterizes temporal question complexity along four dimensions: context complexity, answer focus, temporal relations, and required temporal capabilities\. In particular, its temporal relation space covers all 13 Allen interval relations, together with time\-range set operations, duration\-based operations, and temporal ranking\. Therefore,TimelineCronQ\-Rallows us to examine SCoP in a substantially richer interval\-oriented setting, involving diverse interval relations, temporal arithmetic, and ordering dependencies\. Dataset statistics are summarized in Table[2](https://arxiv.org/html/2609.22213#S4.T2)\. #### TimelineCronQ\-RReconstruction\. We found that the QA annotations ofTimelineKGQA\([Sun et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib18)\)contain duplicated or fragmented records, inconsistent answer representations, weakly grounded supporting events, and instances that cannot be reliably verified against the underlying temporal KG, which may compromise evaluation reliability\. We therefore conservatively reconstruct the QA annotation layer while keeping the temporal KG unchanged, using deterministic temporal execution, KG\-backed completion where unambiguously resolvable, and consistency filtering to repair answer sets, remove unverifiable or invalid cases, resolve duplication, and normalize answer and event representations\. Surface\-form rewriting is applied only when needed for naturalness and never changes answers, temporal relations, or supporting\-event semantics\. Full reconstruction rules, prompts, verification procedures, and stage\-wise statistics are provided in our repository\.[1](https://arxiv.org/html/2609.22213#footnote1) Table 2\.Statistics of KGs and QA datasets\. TCronQ\-R denotes TimelineCronQ\-R\. #### Baselines and Settings\. OnMultiTQ, we compare SCoP with three groups of baselines: TKG embedding\-based methods, including EmbedKGQA\([Saxena et al\., 2020](https://arxiv.org/html/2609.22213#bib.bib19)\), CronKGQA\([Saxena et al\., 2021](https://arxiv.org/html/2609.22213#bib.bib11)\), and MultiQA\([Chen et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib17)\); training\-based LLM methods, including Search\-R1\([Jin et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib26)\), TimeR4\([Qian et al\., 2024](https://arxiv.org/html/2609.22213#bib.bib21)\), PoK\([Qian et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib25)\), and Temp\-R1\([Gong et al\., 2026](https://arxiv.org/html/2609.22213#bib.bib27)\); and strategy\-based LLM methods, including ARI\([Chen et al\., 2024c](https://arxiv.org/html/2609.22213#bib.bib20)\), TempAgent\([Hu et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib28)\), MemoTime\([Tan et al\., 2026](https://arxiv.org/html/2609.22213#bib.bib29)\), and RTQA\([Gong et al\., 2025](https://arxiv.org/html/2609.22213#bib.bib13)\)\. Together, these baselines cover task\-specific training, retrieval\-augmented reasoning, agentic tool use, memory\-augmented reasoning, and recursive question decomposition\. OnTimelineCronQ\-R, where the complexity of its temporal relations makes straightforward adaptation infeasible for most existing methods, we employ a controlled evaluation setting for strategy\-based retrieval methods\. We compare SCoP with three groups of baselines: static single\-turn retrieval methods, including RAG\([Lewis et al\., 2020](https://arxiv.org/html/2609.22213#bib.bib30)\)and its filtering variant; retrieval methods enhanced by query transformation, including HyDE\([Gao et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib33)\)and Query2doc \(Q2D\)\([Wang et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib34)\); and dynamic interleaved retrieval methods, including ReAct\([Yao et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib31)\), IRCoT\([Trivedi et al\., 2023](https://arxiv.org/html/2609.22213#bib.bib32)\), and their filtering variants\. All methods use the same final evidence budget oftop\-k=20top\\text\{\-\}k=20\. Filtering variants first retrieve the top\-50 candidates and then apply model\-based filtering to retain 20 contexts\. Iterative retrieval methods use at most 5 reasoning steps, with filtering applied at each step\. SCoP is likewise restricted to a maximum retained evidence budget of 20\. For all dense retrieval components, we use BGE\-M3\([Chen et al\., 2024a](https://arxiv.org/html/2609.22213#bib.bib1)\)as the unified embedding encoder, including the FAISS\-based schema candidate retrieval in SCoP and the retrieval modules of the comparison methods onTimelineCronQ\-R\. For bothMultiTQandTimelineCronQ\-R, we use Qwen2\.5\-14B\-Instruct as the backbone answer model with temperature set to 0\. The SCoP results reported in Tables[3](https://arxiv.org/html/2609.22213#S4.T3)and[4](https://arxiv.org/html/2609.22213#S4.T4)are obtained on the full test set of each dataset\. OnMultiTQ, baseline results are collected from\([Gong et al\., 2026](https://arxiv.org/html/2609.22213#bib.bib27)\), whereas onTimelineCronQ\-R, all compared methods are evaluated under the same controlled setting\. ### 4\.1\.Main Results #### Performance comparison on MultiTQ Table[3](https://arxiv.org/html/2609.22213#S4.T3)compares SCoP with existing methods onMultiTQ\. SCoP achieves the best overall Hits@1 score of 0\.825, outperforming the strongest fine\-tuning\-based baseline Temp\-R1 by 0\.045 \(\+5\.8%\) and the strongest strategy\-based LLM baseline RTQA by 0\.060 \(\+7\.8%\)\. Its advantage is most pronounced on the*Multiple*question type, where SCoP reaches 0\.653, exceeding Temp\-R1 by 0\.103 \(\+18\.7%\) and RTQA by 0\.229 \(\+54\.0%\)\. By contrast, on*Single*questions, SCoP remains comparable to the strongest prior methods, indicating that its main gains arise in settings with richer temporal dependencies\. These results suggest that SCoP is particularly effective for constraint\-intensive questions requiring the joint satisfaction of multiple structural and temporal conditions\. SCoP also maintains strong performance across answer types, achieving the best score on*Entity*questions while remaining competitive on*Time*questions\. Notably, these gains are obtained with Qwen2\.5\-14B\-Instruct, whereas several previously reported strategy\-based baselines rely on larger backbones, such as GPT\-4\-Turbo and DeepSeek\-V3\. Table 3\.Performance comparison on the MultiTQ dataset\. Results are reported as Hits@1\. The best results are highlighted in bold, and the second\-best results are underlined\. #### Performance comparison on TimelineCronQ\-R Table[4](https://arxiv.org/html/2609.22213#S4.T4)summarizes the results onTimelineCronQ\-R\. SCoP achieves the best overall Hits@1 score of 0\.761, outperforming the strongest*dynamic interleaved retrieval*baseline, ReAct, by 0\.130 \(\+20\.6%\)\. The gains are concentrated on the more challenging*Medium*and*Complex*questions\. Specifically, SCoP reaches 0\.795 on*Medium*questions, exceeding the strongest baseline RAG\+Filter\{\}\_\{\\text\{\+Filter\}\}by 0\.211 \(\+36\.1%\), and achieves 0\.673 on*Complex*questions, surpassing ReAct\+Filter\{\}\_\{\\text\{\+Filter\}\}by 0\.173 \(\+34\.6%\)\. These results indicate that SCoP is particularly effective in settings requiring the joint enforcement of multiple temporal conditions\. On*Simple*questions, SCoP remains competitive but trails direct retrieval methods, suggesting that explicit evidence\-space control provides smaller marginal benefits when temporal constraint resolution is less demanding\. Table 4\.Performance comparison on TimelineCronQ\-R\. Results are reported as Hits@1\. Baseline methods are evaluated both in their standard form and with a post\-retrieval filter \(\+Filter\{\}\_\{\\text\{\+Filter\}\}\) to analyze the impact of noise reduction\. The best results are highlighted in bold, and the second\-best results are underlined\.Across both datasets, SCoP’s strongest gains consistently appear in settings with higher temporal\-constraint complexity:*Multiple*questions onMultiTQand*Medium*/*Complex*questions onTimelineCronQ\-R\. This pattern supports our central claim that complex TKGQA benefits from explicitly controlling structurally and temporally admissible evidence before answer inference\. ### 4\.2\.Influence of Backbone Models To examine the sensitivity of SCoP to backbone choice and facilitate a more direct comparison with representative strategy\-based methods under comparable backbone configurations, we evaluate SCoP with Qwen2\.5\-14B\-Instruct, Gemini\-3\-Flash, as well as GPT\-3\.5\-Turbo and DeepSeek\-V3, which are adopted by representative strategy\-based baselines\. We conduct experiments on 5,000 randomly sampled test instances fromMultiTQwhile preserving its original data distribution, and additionally evaluate on the full test set ofTimelineCronQ\-R\. Figure[3](https://arxiv.org/html/2609.22213#S4.F3)and Table[5](https://arxiv.org/html/2609.22213#S4.T5)show that all alternative backbones achieve higher overall Hits@1 than the Qwen2\.5\-14B\-Instruct setting, although the magnitude of improvement varies across question types\. OnMultiTQ, the gains are substantially larger for*Multiple*than for*Single*questions; onTimelineCronQ\-R, improvements are generally more pronounced on*Medium*and*Complex*questions than on*Simple*ones\. Table 5\.Detailed performance comparison \(Hits@1\) across distinct question complexity levels\.Table 6\.Candidate\-space control across question types on MultiTQ and TimelineCronQ\-R\.This pattern is consistent with the main results, suggesting that backbone capability becomes more influential as event parsing, grounding, and temporal constraint interpretation become more demanding\. At the same time, no single backbone dominates across all question categories, indicating complementary strengths across different temporal reasoning regimes\. These results also clarify the role of the backbone in SCoP: structured constraints externalize temporal admissibility for deterministic evidence filtering, while the quality of the parsed events, anchors, and constraints still depends on the language model\. Thus, SCoP is not tied to a specific backbone, while its absolute performance still depends on backbone capability\. Figure 3\.Effect of different backbone models on the overall Hits@1 performance of SCoP\.Horizontal bar chart comparing SCoP with four backbone models on MultiTQ and TimelineCronQ\-R\. All three alternative backbones improve over Qwen2\.5\-14B\-Instruct in overall Hits@1\. Gemini\-3\-Flash achieves the highest result on MultiTQ, while GPT\-3\.5\-Turbo and Gemini\-3\-Flash achieve the highest overall results on TimelineCronQ\-R\. ### 4\.3\.Evidence\-Space Control Analysis To examine whether SCoP effectively controls the candidate evidence space, we compare the initial candidate evidence before constraint execution with the final evidence passed to the answer model after constraint\-guided evidence construction\. The quantified evaluation results are summarized in Table[6](https://arxiv.org/html/2609.22213#S4.T6)\. Specifically,Initial Evid\.andFinal Evid\.report the average evidence counts per evaluable sample within each temporal reasoning type\. For questions with a search triple,Initial Evid\.denotes the number of answer\-bearing search\-fact candidates before temporal constraint execution\.Final Evid\.denotes the number of retained search facts after executable constraint filtering, optional ranking, and final context truncation\. For questions without a search triple,Initial Evid\.denotes the summed candidate count retrieved for the aligned complete anchor triples\.Final Evid\.denotes the summed number of anchor facts retained in the bounded closed\-target answer context\. We quantify the overall compression effect within each subset using theEvidence Compression Rate: ECR=1−∑iFinal Evid\.i∑iInitial Evid\.i,\\mathrm\{ECR\}=1\-\\frac\{\\sum\_\{i\}\\textit\{Final Evid\.\}\_\{i\}\}\{\\sum\_\{i\}\\textit\{Initial Evid\.\}\_\{i\}\},whereiiranges over all evaluable samples in each subset\. After excluding samples with incomplete retrieval statistics, the evaluation covers 99\.4% ofMultiTQand 98\.8% ofTimelineCronQ\-R\. The type\-wise results further show that compression varies with temporal selectivity\. InMultiTQ, equality\-, ranking\-, and before/after\-related questions are strongly compressed, whereas first\-last questions exhibit lower reduction because their initial candidate spaces are already smaller\. InTimelineCronQ\-R, quantitative\- and relative\-temporal questions show the strongest compression, while temporal\-operation questions retain more evidence for downstream answer realization\. These differences are consistent with the selectivity of the executable temporal constraints: tighter temporal conditions eliminate more structurally compatible candidates, whereas less selective conditions preserve a broader evidence context\. Table 7\.Gold\-event and gold\-answer retention in the final evidence space onTimelineCronQ\-R\.The compression analysis raises a further question: does the reduced evidence space still preserve the contextual support required for answer inference? BecauseMultiTQlacks gold event annotations for direct event\-level retention analysis, we conduct this evaluation onTimelineCronQ\-Racross different question difficulty levels\. We report retention at both the gold\-event and gold\-answer levels using two metrics:\(1\)\(1\)Any Retention Rate, the proportion of evaluable samples whose final evidence context contains at least one gold event or gold answer; and\(2\)\(2\)Full Retention Rate, the proportion preserving all gold events or gold answers\. As shown in Table[7](https://arxiv.org/html/2609.22213#S4.T7), SCoP retains a large majority of gold events and gold answers after compression, with overall Any/Full Retention rates of 85\.7%/81\.6% and 84\.9%/81\.3%, respectively\. Full Retention generally decreases with question complexity, with*Complex*questions showing the lowest retention for both gold events and gold answers\. The gap between Any and Full Retention indicates that SCoP may preserve partial support while losing complementary evidence required jointly for answer inference, particularly on*Complex*questions involving multiple temporally related facts\. Taken together, these results show that SCoP substantially contracts the evidence space while preserving most task\-relevant support, although complete evidence preservation remains more difficult for compositionally demanding questions\. ### 4\.4\.Ablation Studies Having shown that SCoP compresses the candidate evidence space while preserving most task\-relevant support, we further examine the contribution of its three core components on bothMultiTQandTimelineCronQ\-R\. All variants use Qwen2\.5\-14B\-Instruct as the unified backbone, and the results are summarized in Table[8](https://arxiv.org/html/2609.22213#S4.T8)\. We consider three variants\. For*w/o triple*, we remove structured event\-triple identification and replace the explicit answer\-anchor event separation with a retrieval\-based pseudo\-decomposition strategy, where the original question is directly used to retrieve pseudo anchor triples and construct a pseudo search triple\. For*w/o alignment*, we remove conservative graph alignment and directly use the raw extracted triples for downstream retrieval and constraint execution\. For*w/o constraint*, we retain event\-triple identification and graph alignment but disable temporal constraint parsing and execution, so that structurally compatible but temporally unfiltered candidate facts are passed to the answer model\. Together, these variants isolate whether SCoP’s gains arise from explicit event\-role decomposition, KG\-compatible grounding, and executable temporal admissibility control\. Table 8\.Ablation study of SCoP under the Overall Hits@1 metric\.Δ\\Deltadenotes the relative performance change compared with the full model\.As shown in Table[8](https://arxiv.org/html/2609.22213#S4.T8), all ablated variants underperform the full model on both datasets, showing that the three components contribute to different stages of evidence\-space control\. OnMultiTQ, removing constraint execution causes the largest drop, reducing Hits@1 from 0\.825 to 0\.449, while removing graph alignment leads to a comparable decline to 0\.463\. This indicates thatMultiTQrelies heavily on both temporal admissibility filtering and KG\-compatible schema grounding: without constraints, structurally relevant but temporally invalid facts enter the answer context; without alignment, extracted event structures cannot be reliably matched to canonical entities and relations\. OnTimelineCronQ\-R, removing event\-triple identification causes the largest degradation, from 0\.761 to 0\.420, suggesting that explicitly separating answer\-seeking events from temporal anchors is particularly important in interval\-oriented and relation\-intensive questions\. Disabling constraint execution also produces a substantial drop to 0\.541, confirming that executable temporal constraints remain necessary for filtering admissible evidence\. The smaller impact of removing alignment onTimelineCronQ\-Rmay reflect that its entity and relation mentions are often closer to KG\-compatible forms after extraction\. Overall, the ablation results support the coordinated design of SCoP: event\-triple identification defines the answer and anchor evidence targets, graph alignment grounds them to the TKG schema, and constraint execution controls which structurally retrieved facts are temporally admissible for answer inference\. ## 5\.Conclusion We present SCoP, a structured constraint parsing framework that controls the admissible evidence space for temporal knowledge graph question answering before answer inference\. By separating event identification, conservative graph alignment, and executable temporal constraint filtering, SCoP consistently improves performance onMultiTQandTimelineCronQ\-R, with its strongest gains on constraint\-intensive questions\. Further analysis shows that SCoP substantially reduces candidate evidence while preserving most task\-relevant support, supporting explicit evidence\-space control over unconstrained retrieval or implicit temporal reasoning\. These results validate the effectiveness of explicit evidence\-space control, while also revealing several limitations of the current framework\. SCoP relies on off\-the\-shelf language models for event identification, constraint parsing, and graph alignment, so errors in these stages may propagate to downstream evidence construction\. Its current constraint schema does not yet cover all ambiguous or compositional temporal phenomena, and complete evidence preservation remains more challenging for complex questions requiring multiple supporting facts to survive jointly\. Future work will therefore focus on extending the constraint schema, improving multi\-evidence preservation, and evaluating SCoP across broader temporal knowledge graphs and reasoning settings\. ## GenAI Usage Disclosure ChatGPT \(GPT\-5\.5\) was used for language polishing, translation, and assisting with refining and optimizing code written by the authors\. Additionally, Gemini\-3\-Flash was used during the reconstruction of theTimelineCronQ\-Rdataset for constrained surface\-form rewriting of benchmark questions\. Neither tool was involved in the design of the core algorithms or research methodology, which were developed by the authors\. All AI\-assisted code, data, and manuscript content were reviewed and verified by the authors, who take full responsibility for the final research outputs\. ## References - Chenet al\.\(2024a\)J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. LiuM3\-embedding: multi\-linguality, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 2318–2335\.External Links:[Link](https://aclanthology.org/2024.findings-acl.137/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.137)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p2.1)\. - Chenet al\.\(2024b\)Z\. Chen, Z\. Zhang, Z\. Li, F\. Wang, Y\. Zeng, X\. Jin, and Y\. XuSelf\-improvement programming for temporal knowledge graph question answering\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20\-25 May, 2024, Torino, Italy,pp\. 14579–14594\.External Links:[Link](https://aclanthology.org/2024.lrec-main.1270)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p1.1)\. - Chenet al\.\(2024c\)Z\. Chen, D\. Li, X\. Zhao, B\. Hu, and M\. ZhangTemporal knowledge question answering via abstract reasoning induction\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,pp\. 4872–4889\.External Links:[Link](https://doi.org/10.18653/v1/2024.acl-long.267),[Document](https://dx.doi.org/10.18653/V1/2024.ACL-LONG.267)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p4.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Chenet al\.\(2023\)Z\. Chen, J\. Liao, and X\. ZhaoMulti\-granularity temporal question answering over knowledge graphs\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2023, Toronto, Canada, July 9\-14, 2023,pp\. 11378–11392\.External Links:[Link](https://doi.org/10.18653/v1/2023.acl-long.637),[Document](https://dx.doi.org/10.18653/V1/2023.ACL-LONG.637)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p4.1),[§2](https://arxiv.org/html/2609.22213#S2.p2.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Dinget al\.\(2022\)W\. Ding, H\. Chen, H\. Li, and Y\. QuSemantic framework based query generation for temporal question answering over knowledge graphs\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7\-11, 2022,pp\. 1867–1877\.External Links:[Link](https://doi.org/10.18653/v1/2022.emnlp-main.122),[Document](https://dx.doi.org/10.18653/V1/2022.EMNLP-MAIN.122)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p1.1)\. - Fayyazet al\.\(2025\)M\. Fayyaz, A\. Modarressi, H\. Schütze, and N\. PengCollapse of dense retrievers: short, early, and literal biases outranking factual evidence\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,pp\. 9136–9152\.External Links:[Link](https://aclanthology.org/2025.acl-long.447/)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p2.1)\. - Gadeet al\.\(2025\)A\. Gade, J\. G\. Jetcheva, and H\. TrivediIt’s about time: incorporating temporality in retrieval augmented language models\.InIEEE Conference on Artificial Intelligence, CAI 2025, Santa Clara, CA, USA, May 5\-7, 2025,pp\. 75–82\.External Links:[Link](https://doi.org/10.1109/CAI64502.2025.00019),[Document](https://dx.doi.org/10.1109/CAI64502.2025.00019)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p2.1)\. - Gaoet al\.\(2023\)L\. Gao, X\. Ma, J\. Lin, and J\. CallanPrecise zero\-shot dense retrieval without relevance labels\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2023, Toronto, Canada, July 9\-14, 2023,pp\. 1762–1777\.External Links:[Link](https://doi.org/10.18653/v1/2023.acl-long.99),[Document](https://dx.doi.org/10.18653/V1/2023.ACL-LONG.99)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Gaoet al\.\(2024\)Y\. Gao, L\. Qiao, Z\. Kan, Z\. Wen, Y\. He, and D\. LiTwo\-stage generative question answering on temporal knowledge graph using large language models\.InFindings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11\-16, 2024,pp\. 6719–6734\.External Links:[Link](https://doi.org/10.18653/v1/2024.findings-acl.401),[Document](https://dx.doi.org/10.18653/V1/2024.FINDINGS-ACL.401)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p3.1)\. - García\-Duránet al\.\(2018\)A\. García\-Durán, S\. Dumančić, and M\. NiepertLearning sequence encoders for temporal knowledge graph completion\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 \- November 4, 2018,pp\. 4816–4821\.External Links:[Link](https://aclanthology.org/D18-1516/)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px1.p1.1)\. - Gonget al\.\(2025\)Z\. Gong, J\. Li, Z\. Liu, L\. Liang, H\. Chen, and W\. ZhangRTQA: recursive thinking for complex temporal knowledge graph question answering with large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 9853–9870\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.499/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.499)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p4.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Gonget al\.\(2026\)Z\. Gong, Z\. Liu, S\. Li, X\. Guo, Y\. Liu, X\. Deng, Z\. Liu, L\. Liang, H\. Chen, and W\. ZhangTemp\-r1: A unified autonomous agent for complex temporal KGQA via reverse curriculum reinforcement learning\.CoRRabs/2601\.18296\.External Links:[Link](https://doi.org/10.48550/arXiv.2601.18296),[Document](https://dx.doi.org/10.48550/ARXIV.2601.18296)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p3.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p2.1)\. - Huet al\.\(2025\)Q\. Hu, X\. Tu, C\. Guo, and S\. ZhangTime\-aware react agent for temporal knowledge graph question answering\.InFindings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 \- May 4, 2025,Findings of ACL,pp\. 6028–6039\.External Links:[Link](https://doi.org/10.18653/v1/2025.findings-naacl.334),[Document](https://dx.doi.org/10.18653/V1/2025.FINDINGS-NAACL.334)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p4.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Jiaet al\.\(2018\)Z\. Jia, A\. Abujabal, R\. S\. Roy, J\. Strötgen, and G\. WeikumTEQUILA: temporal question answering over knowledge bases\.InProceedings of the 27th ACM International Conference on Information and Knowledge Management, CIKM 2018, Torino, Italy, October 22\-26, 2018,pp\. 1807–1810\.External Links:[Link](https://doi.org/10.1145/3269206.3269247),[Document](https://dx.doi.org/10.1145/3269206.3269247)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p1.1)\. - Jiaet al\.\(2021\)Z\. Jia, S\. Pramanik, R\. S\. Roy, and G\. WeikumComplex temporal question answering on knowledge graphs\.InCIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 \- 5, 2021,pp\. 792–802\.External Links:[Link](https://doi.org/10.1145/3459637.3482416),[Document](https://dx.doi.org/10.1145/3459637.3482416)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p1.1)\. - Jinet al\.\(2025\)B\. Jin, H\. Zeng, Z\. Yue, D\. Wang, H\. Zamani, and J\. HanSearch\-r1: training llms to reason and leverage search engines with reinforcement learning\.CoRRabs/2503\.09516\.External Links:[Link](https://doi.org/10.48550/arXiv.2503.09516),[Document](https://dx.doi.org/10.48550/ARXIV.2503.09516)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p3.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Johnsonet al\.\(2021\)J\. Johnson, M\. Douze, and H\. JégouBillion\-scale similarity search with gpus\.IEEE Transactions on Big Data7\(3\),pp\. 535–547\.External Links:[Document](https://dx.doi.org/10.1109/TBDATA.2019.2921572)Cited by:[§3\.1](https://arxiv.org/html/2609.22213#S3.SS1.p2.1)\. - Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 9459–9474\.External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Liuet al\.\(2024\)J\. Liu, Z\. Liu, X\. Lyu, P\. Jin, and J\. XuTowards multi\-relational multi\-hop reasoning over dense temporal knowledge graphs\.InFindings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11\-16, 2024,pp\. 14367–14378\.External Links:[Link](https://doi.org/10.18653/v1/2024.findings-acl.853),[Document](https://dx.doi.org/10.18653/V1/2024.FINDINGS-ACL.853)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p1.1)\. - Liuet al\.\(2023\)Y\. Liu, D\. Liang, M\. Li, F\. Giunchiglia, X\. Li, S\. Wang, W\. Wu, L\. Huang, X\. Feng, and R\. GuanLocal and global: temporal question answering via information fusion\.InProceedings of the Thirty\-Second International Joint Conference on Artificial Intelligence, IJCAI 2023, 19th\-25th August 2023, Macao, SAR, China,pp\. 5141–5149\.External Links:[Link](https://doi.org/10.24963/ijcai.2023/571),[Document](https://dx.doi.org/10.24963/IJCAI.2023/571)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p2.1)\. - Mavromatiset al\.\(2022\)C\. Mavromatis, P\. L\. Subramanyam, V\. N\. Ioannidis, A\. Adeshina, P\. R\. Howard, T\. Grinberg, N\. Hakim, and G\. KarypisTempoQR: temporal question reasoning over knowledge graphs\.InThirty\-Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, Thirty\-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI 2022, The Twelfth Symposium on Educational Advances in Artificial Intelligence, EAAI 2022 Virtual Event, February 22 \- March 1, 2022,pp\. 5825–5833\.External Links:[Link](https://doi.org/10.1609/aaai.v36i5.20526),[Document](https://dx.doi.org/10.1609/AAAI.V36I5.20526)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p2.1)\. - Neelamet al\.\(2022\)S\. Neelam, U\. Sharma, H\. Karanam, S\. Ikbal, P\. Kapanipathi, I\. Abdelaziz, N\. Mihindukulasooriya, Y\. Lee, S\. Srivastava, C\. Pendus, S\. Dana, D\. Garg, A\. Fokoue, G\. P\. S\. Bhargav, D\. Khandelwal, S\. Ravishankar, S\. Gurajada, M\. Chang, R\. Uceda\-Sosa, S\. Roukos, A\. Gray, G\. Lima, R\. Riegel, F\. Luus, L\. V\. Subramaniam, Z\. Kozareva, and Y\. ZhangSYGMA: a system for generalizable and modular question answering over knowledge bases\.InFindings of the Association for Computational Linguistics: EMNLP 2022,pp\. 3866–3879\.External Links:[Link](https://aclanthology.org/2022.findings-emnlp.284/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-emnlp.284)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p1.1)\. - Qianet al\.\(2025\)X\. Qian, Y\. Zhang, Y\. Zhao, B\. Zhou, X\. Sui, and X\. YuanPlan of knowledge: retrieval\-augmented large language models for temporal knowledge graph question answering\.CoRRabs/2511\.04072\.External Links:[Link](https://doi.org/10.48550/arXiv.2511.04072),[Document](https://dx.doi.org/10.48550/ARXIV.2511.04072)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p3.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Qianet al\.\(2024\)X\. Qian, Y\. Zhang, Y\. Zhao, B\. Zhou, X\. Sui, L\. Zhang, and K\. SongTimeR4: time\-aware retrieval\-augmented large language models for temporal knowledge graph question answering\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12\-16, 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),pp\. 6942–6952\.External Links:[Link](https://doi.org/10.18653/v1/2024.emnlp-main.394),[Document](https://dx.doi.org/10.18653/V1/2024.EMNLP-MAIN.394)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p3.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Saxenaet al\.\(2021\)A\. Saxena, S\. Chakrabarti, and P\. P\. TalukdarQuestion answering over temporal knowledge graphs\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, \(Volume 1: Long Papers\), Virtual Event, August 1\-6, 2021,pp\. 6663–6676\.External Links:[Link](https://doi.org/10.18653/v1/2021.acl-long.520),[Document](https://dx.doi.org/10.18653/V1/2021.ACL-LONG.520)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p2.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Saxenaet al\.\(2020\)A\. Saxena, A\. Tripathi, and P\. P\. TalukdarImproving multi\-hop question answering over knowledge graphs using knowledge base embeddings\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5\-10, 2020,pp\. 4498–4507\.External Links:[Link](https://doi.org/10.18653/v1/2020.acl-main.412),[Document](https://dx.doi.org/10.18653/V1/2020.ACL-MAIN.412)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p2.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Sharmaet al\.\(2023\)A\. Sharma, A\. Saxena, C\. Gupta, S\. M\. Kazemi, P\. P\. Talukdar, and S\. ChakrabartiTwiRGCN: temporally weighted graph convolution for question answering over temporal knowledge graphs\.InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2\-6, 2023,pp\. 2049–2060\.External Links:[Link](https://doi.org/10.18653/v1/2023.eacl-main.150),[Document](https://dx.doi.org/10.18653/V1/2023.EACL-MAIN.150)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p2.1)\. - Sunet al\.\(2025\)Q\. Sun, S\. Li, D\. Huynh, M\. Reynolds, and W\. LiuTimelineKGQA: A comprehensive question\-answer pair generator for temporal knowledge graphs\.InCompanion Proceedings of the ACM on Web Conference 2025, WWW 2025, Sydney, NSW, Australia, 28 April 2025 \- 2 May 2025,pp\. 797–800\.External Links:[Link](https://doi.org/10.1145/3701716.3715308),[Document](https://dx.doi.org/10.1145/3701716.3715308)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p4.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px2.p1.1)\. - Tanet al\.\(2026\)X\. Tan, X\. Wang, Q\. Liu, X\. Xu, X\. Yuan, L\. Zhu, and W\. ZhangMemoTime: memory\-augmented temporal knowledge graph enhanced large language model reasoning\.InProceedings of the ACM Web Conference 2026,pp\. 4220–4231\.External Links:[Link](https://doi.org/10.1145/3774904.3792581),[Document](https://dx.doi.org/10.1145/3774904.3792581)Cited by:[§2](https://arxiv.org/html/2609.22213#S2.p4.1),[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Trivediet al\.\(2023\)H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. SabharwalInterleaving retrieval with chain\-of\-thought reasoning for knowledge\-intensive multi\-step questions\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2023, Toronto, Canada, July 9\-14, 2023,pp\. 10014–10037\.External Links:[Link](https://doi.org/10.18653/v1/2023.acl-long.557),[Document](https://dx.doi.org/10.18653/V1/2023.ACL-LONG.557)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Trivediet al\.\(2017\)R\. Trivedi, H\. Dai, Y\. Wang, and L\. SongKnow\-evolve: deep temporal reasoning for dynamic knowledge graphs\.InProceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6\-11 August 2017,Proceedings of Machine Learning Research, Vol\.70,pp\. 3462–3471\.External Links:[Link](http://proceedings.mlr.press/v70/trivedi17a.html)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p1.1)\. - Wanget al\.\(2023\)L\. Wang, N\. Yang, and F\. WeiQuery2doc: query expansion with large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6\-10, 2023,pp\. 9414–9423\.External Links:[Link](https://doi.org/10.18653/v1/2023.emnlp-main.585),[Document](https://dx.doi.org/10.18653/V1/2023.EMNLP-MAIN.585)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1\-5, 2023,External Links:[Link](https://openreview.net/forum?id=WE\_vluYUL-X)Cited by:[§4](https://arxiv.org/html/2609.22213#S4.SS0.SSS0.Px3.p1.1)\. - Zhanget al\.\(2024\)T\. Zhang, J\. Wang, Z\. Li, J\. Qu, A\. Liu, Z\. Chen, and H\. ZhiMusTQ: A temporal knowledge graph question answering dataset for multi\-step temporal reasoning\.InFindings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11\-16, 2024,pp\. 11688–11699\.External Links:[Link](https://doi.org/10.18653/v1/2024.findings-acl.696),[Document](https://dx.doi.org/10.18653/V1/2024.FINDINGS-ACL.696)Cited by:[§1](https://arxiv.org/html/2609.22213#S1.p1.1)\.
相似文章
SABET-QA:时序知识图谱问答
SABET-QA 提出了一个用于时序知识图谱问答的迭代框架,通过双向实体-时序评分和上下文化增强了多跳推理,并在 CronQuestions 和 TimeQuestions 等基准测试中显示出相对于基线的一致改进。
TRACE: 基于时间证据图的对话数据状态感知查询处理
本文提出 TRACE,一种查询处理框架,将对话数据建模为时间证据图,以支持对不断演进的用户状态进行状态感知推理,从而提高长对话问答中的时间推理和多跳推理能力。
S3Mem:面向长周期交互式问答的结构化时空场景事件记忆
S3Mem 提出了一种用于长周期交互式问答的结构化时空场景事件记忆框架,采用锚点敏感检索和令牌预算感知的证据接口,在多个环境中优于标准 RAG。
多跳知识图谱问答的本体引导证据路径推理
提出 OPI,一种面向多跳知识图谱问答的本体引导框架,利用以关系为中心的本体图进行双向检索和迭代精炼,在多个基准上取得了最先进的结果。
TCAR-Gen:面向知识基础生成的时间图检索与证据融合
TCAR-Gen 提出了一种结合查询条件图神经网络、时间证据融合和树链推理的框架,用于知识基础生成中的时间图检索。在 Victorian Crime Diaries 基准测试中,它在多种查询类型上实现了改进的召回率。