From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search
Summary
The paper proposes a task-adapted retrieval approach for mapping free-form location phrases to geographic entities in people search, using a prompt-asymmetric bi-encoder that improves relevance, especially for non-canonical queries, as demonstrated in production and benchmark evaluations.
View Cached Full Text
Cached at: 09/01/26, 12:44 PM
# From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search Source: [https://arxiv.org/html/2608.28965](https://arxiv.org/html/2608.28965) CCS:Information systems Retrieval models and rankingCCS:Information systems Users and interactive retrievalCCS:Computing methodologies Learning latent representations,Chujie ZhengNote:Work done while at LinkedIn\.Affiliation:LinkedIn,Sunnyvale,CA,USA,Jiahao XuAffiliation:LinkedIn,Sunnyvale,CA,USA,Chetan BholeAffiliation:LinkedIn,Sunnyvale,CA,USAemail:[yanbli@linkedin\.com](mailto:[email protected]),Lingyu ZhangAffiliation:LinkedIn,Sunnyvale,CA,USA,Puneet Singh AhluwaliaAffiliation:LinkedIn,Sunnyvale,CA,USA,Kevin NguyenAffiliation:LinkedIn,Sunnyvale,CA,USA,Raghavan MuthuregunathanAffiliation:LinkedIn,Sunnyvale,CA,USAemail:[clzhang@linkedin\.com](mailto:[email protected]),Santhosh SachindranAffiliation:LinkedIn,Sunnyvale,CA,USA,Sachin AhujaAffiliation:LinkedIn,Sunnyvale,CA,USAandFedor BorisyukNote:Corresponding author\.Affiliation:LinkedIn,Sunnyvale,CA,USAemail:[ssachindran@linkedin\.com](mailto:[email protected]) ###### Abstract\. People search must map free\-form location phrases to geographic entities used as structured retrieval filters\. Lexical standardizers handle canonical names well but are brittle to aliases, misspellings, metropolitan expressions, and same\-name ambiguity\. We formulate this task as graded, set\-valued entity retrieval over a fixed ontology\. We identify three coupled design requirements: distinguishing identity\-preserving variation from knowledge\-dependent aliases, controlling false negatives among valid same\-name entities, and separating stable transformations from mutable entity knowledge\. We realize them in a prompt\-asymmetric bi\-encoder with calibrated alias support, bounded ambiguity\-aware negatives, and editable entity documents that support localized updates without retraining\. Across a fixed production\-derived development benchmark and a public GeoNames transfer task, task adaptation improves substantially over frozen encoders and standard token baselines\. Controlled development ablations show that specialized supervision contributes beyond standard task fine\-tuning and encoder scaling\. On GeoNames, the adapted model improves known\-target Recall@1 throughout zero\-to\-moderate character overlap, while characternn\-grams retain a small aggregate Target Recall@5 advantage\. In a blinded human comparison on a stratified production challenge set, our model raises relevant P@1 from 28\.0% to 46\.0% \(p=0\.012p=0\.012\)\. Fixed\-query endpoint estimates improve on non\-canonical queries and remain close to control on frequent queries; a randomized live experiment detects no engagement regression\. These results support task\-adapted geographic entity retrieval as a practical replacement for the incumbent taxonomy\-based standardizer, with the largest relevance gains on non\-canonical queries\. ###### Keywords: location grounding, geographic entity retrieval, dense retrieval, toponym resolution, editable entity representations, people search ## 1\.Introduction People search queries combine unstructured intent, such as skills, titles, names, and companies, with structured constraints\. A query such as “engineers in the greater Boston area” requires the location phrase to be mapped to the canonical entities consumed by candidate generation\. This mapping is commonly implemented as a*geographic standardizer*: given a location span extracted by query understanding, it returns one or more identifiers from a curated geographic ontology\. Production standardizers have traditionally combined gazetteers, taxonomy aliases, typeahead indices, and manually specified string rules\. These systems are interpretable and accurate on high\-frequency, well\-formed queries, but coverage declines on less conventional expressions\. Abbreviations \(“la”, “WPB”\), informal names \(“pink city”\), metropolitan phrasing \(“philly metro”\), misspellings, and variations in order or punctuation may not have an exact entry in the alias table\. Adding rules improves individual cases, but requires continuing maintenance and can introduce collisions among places that share a surface form\. Dense retrieval maps varied surface forms and canonical entities into a shared space, but this setting raises three coupled design questions\. Which query transformations preserve entity identity, and which aliases require external validation? How can training expose confusable same\-name entities without penalizing valid alternatives as false negatives? Which transformations belong in shared model parameters, and which mutable facts should remain in entity documents? These questions concern positive support, negative sampling, and memory placement rather than encoder scale alone\. We present a geographic standardizer deployed in the people search stack of a large professional social network\. The system fine\-tunes a 0\.6B\-parameter instruction embedding model as a shared\-weight bi\-encoder\. Query and entity inputs use different prompts, while entity representations are precomputed\. For the experiments in this paper, we evaluate a compact representation for fixed\-inventory retrieval\. Our formulation jointly specifies three objects: the support of valid query forms, the distribution of confusable negatives, and the boundary between stable transformations in model parameters and mutable knowledge in entity documents\. We formalize these objects in a single fixed\-ontology retrieval objective\. The design does not require encoder\-specific architectural changes and also applies to fixed\-inventory retrieval tasks with non\-canonical surface forms, ambiguous names, and evolving entity metadata\. Contributions\. 1. \(1\)We formulate short\-query geographic grounding as graded, set\-valued retrieval over a fixed, multi\-granularity ontology\. This formulation captures the downstream semantics in which ordered Top\-kkentities become a disjunctive candidate\-generation constraint\. 2. \(2\)We instantiate a unified supervision design with three matched components: identity\-preserving views expand invariant support; a generate–verify–calibrate procedure adds knowledge\-dependent aliases without overwhelming canonical forms; and dominance\-gated, capped same\-name negatives increase confusability while controlling false negatives\. 3. \(3\)We separate stable cross\-entity transformations in model parameters from mutable facts in entity documents, and compare entity\-only updates with lexical insertion\. The comparison exposes complementary update paths: entity documents support localized neural updates without retraining, while lexical insertion remains stronger for known deterministic mappings\. 4. \(4\)We evaluate the complete design against standard surface\-form fine\-tuning and frozen encoders up to 8B, test public transfer against lexical baselines on GeoNames, and compare directly with the production system in a blinded five\-annotator study\. Online evaluation separates fixed\-query relevance from live engagement guardrails\. Under a fixed\-seed controlled development study, the selected task\-adapted system raises nDCG@10 from 0\.8099 under standard surface\-form fine\-tuning to 0\.9201; larger frozen encoders do not close the gap\. Public GeoNames transfer provides independent, judge\-free target metrics: the adapted model exceeds exact, prefix, and BM25 retrieval, improves Target Recall@1 overall and throughout zero\-to\-moderate character overlap, and trails characternn\-grams slightly on aggregate Target Recall@5 because of high\-overlap aliases\. Entity\-document edits improve held\-out\-alias retrieval with unrelated controls unchanged, but incur a canonical\-retention cost; direct lexical insertion remains stronger for curated mappings\. On a stratified production challenge set, the model raises human\-judged relevant P@1 from 28\.0% to 46\.0%\. Fixed\-query endpoint point estimates rise for non\-canonical queries, while the live experiment detects no engagement regression\. ## 2\.Related Work People search and structured query understanding\.Professional people search combines free\-text intent with structured constraints and must retrieve and rank profiles at scale\([Geyik et al\., 2018](https://arxiv.org/html/2608.28965#bib.bib19)\)\. Recent systems move beyond lexical matching: Gupta et al\. simplify queries and member documents, fine\-tune an embedding model, and use Matryoshka representations for semantic profile retrieval\([Gupta et al\., 2025](https://arxiv.org/html/2608.28965#bib.bib20)\); Borisyuk et al\. combine query understanding, embedding retrieval, an LLM relevance judge, and a distilled reranker\([Borisyuk et al\., 2026](https://arxiv.org/html/2608.28965#bib.bib21)\)\. These systems retrieve profiles from the user’s overall intent\. We study a complementary query\-understanding stage: after a location span is extracted, it must be grounded to a fixed geographic ontology whose ordered Top\-kkentity IDs become a disjunctive retrieval constraint\. The output is therefore an entity set rather than a member ranking, and several geographic granularities can be simultaneously relevant\. Toponym resolution and geographic representation\.Toponym resolution conventionally detects place mentions in documents and links them to gazetteer entries, using context to resolve names such as “Melbourne”\([Gritta et al\., 2018](https://arxiv.org/html/2608.28965#bib.bib6);[Weissenbacher et al\., 2019](https://arxiv.org/html/2608.28965#bib.bib5)\)\. GeoNorm improves candidate generation and transformer reranking with ontology and population signals\([Zhang and Bethard, 2023](https://arxiv.org/html/2608.28965#bib.bib7)\); GeoPLACE predicts geographic attributes before constraining deterministic ontology lookup\([Zhang et al\., 2024](https://arxiv.org/html/2608.28965#bib.bib22)\)\. Complementary work learns spatial representations from coordinates, map relations, or nearby entities\([Li et al\., 2022](https://arxiv.org/html/2608.28965#bib.bib23);[Li et al\., 2023](https://arxiv.org/html/2608.28965#bib.bib8)\)\. Recent retrieval formulations contrastively encode point\-of\-interest mentions and gazetteer entries\([Nakatani et al\., 2025](https://arxiv.org/html/2608.28965#bib.bib24)\), while Masis and O’Connor study noisy, multilingual, user\-provided location strings\([Masis and O’Connor, 2024](https://arxiv.org/html/2608.28965#bib.bib9)\)\. Our inputs often contain little context because they consist only of an extracted search span\. The correct output can include several granularities or same\-name locations\. We therefore evaluate ordered sets of ontology entities, not only a single coordinate or gazetteer entry\. Entity normalization over alias\-rich ontologies\.Dense entity retrieval provides the closest non\-geographic precedent\. Dual encoders retrieve knowledge\-base entities without an alias\-table candidate generator\([Gillick et al\., 2019](https://arxiv.org/html/2608.28965#bib.bib25)\), and BLINK scales this formulation to million\-entity linking\([Wu et al\., 2020](https://arxiv.org/html/2608.28965#bib.bib4)\)\. Biomedical normalization makes the alias problem explicit: BioSyn learns from incomplete synonym sets through synonym marginalization\([Sung et al\., 2020](https://arxiv.org/html/2608.28965#bib.bib26)\), whereas SapBERT aligns aliases belonging to the same ontology concept\([Liu et al\., 2021](https://arxiv.org/html/2608.28965#bib.bib27)\)\. Recent work also systematically studies label verbalization and negative sampling in dual\-encoder disambiguation\([Rücker and Akbik, 2025](https://arxiv.org/html/2608.28965#bib.bib28)\)\. Our setting combines three properties: queries are short and frequently misspelled, knowledge\-dependent aliases may be absent from the ontology, and one surface can have several valid outputs with different relevance grades\. This motivates separate supervision for identity\-preserving variation, validated alias acquisition, and bounded same\-name disambiguation instead of treating every unlabeled entity as negative\. Synthetic supervision and negative construction\.LLM\-generated queries can turn demonstrations into task\-specific retriever training data; Promptagator additionally filters generated pairs using round\-trip consistency\([Dai et al\., 2023](https://arxiv.org/html/2608.28965#bib.bib29)\)\. Dense retrieval also benefits from mined negatives\([Xiong et al\., 2021](https://arxiv.org/html/2608.28965#bib.bib10)\), but unlabeled positives make unfiltered hard\-negative mining unreliable\. RocketQA addresses this issue with cross\-encoder\-denoised negatives\([Qu et al\., 2021](https://arxiv.org/html/2608.28965#bib.bib30)\)\. Our generate–verify–calibrate procedure instead targets entity\-specific geographic aliases: generation expands knowledge support, verification checks the referent, and calibrated sampling prevents synthetic forms from displacing canonical queries\. Likewise, our negatives use ontology identity, hierarchy, and a bounded prominence relation rather than equating retrieval rank with irrelevance\. Compact retrieval and editable entity memory\.Dual encoders make candidate representations precomputable\([Huang et al\., 2013](https://arxiv.org/html/2608.28965#bib.bib1);[Karpukhin et al\., 2020](https://arxiv.org/html/2608.28965#bib.bib2);[Reimers and Gurevych, 2019](https://arxiv.org/html/2608.28965#bib.bib3)\), and Matryoshka training makes vector prefixes useful at multiple dimensions\([Kusupati et al\., 2022](https://arxiv.org/html/2608.28965#bib.bib11)\)\. Description\-based zero\-shot entity linking shows that independently encoded content can support entities unseen during task supervision\([Logeswaran et al\., 2019](https://arxiv.org/html/2608.28965#bib.bib31)\); retrieval\-augmented models more broadly separate parametric behavior from external memory\([Lewis et al\., 2020](https://arxiv.org/html/2608.28965#bib.bib32)\)\. DynamicER studies emerging mentions and evolving entities through continual adaptation\([Kim et al\., 2024](https://arxiv.org/html/2608.28965#bib.bib33)\)\. We investigate a narrower update operation for a fixed ontology: add a held\-out alias to the affected entity document, re\-embed only that entity, and leave model parameters fixed\. This makes the knowledge boundary operational rather than merely architectural\. LLM relevance assessment\.LLM judges reduce assessment cost but require explicit validity evidence\([Faggioli et al\., 2023](https://arxiv.org/html/2608.28965#bib.bib13)\)\. Calibrated assessors can reproduce searcher preferences, yet remain sensitive to prompt paraphrases\([Thomas et al\., 2024](https://arxiv.org/html/2608.28965#bib.bib14)\); UMBRELA reports strong correlations between system rankings induced by LLM and human judgments\([Upadhyay et al\., 2024](https://arxiv.org/html/2608.28965#bib.bib15)\)\. Such run\-level agreement does not imply interchangeable query–result labels\. Across eight assessors, Fröbe et al\. find stronger LLM–LLM than LLM–human agreement and potential preference for LLM\-based rankers\([Fröbe et al\., 2025](https://arxiv.org/html/2608.28965#bib.bib16)\)\. LARA and LLM\-Rubric use human calibration or explicit multidimensional rubrics to reduce this gap\([Takehi et al\., 2025](https://arxiv.org/html/2608.28965#bib.bib17);[Hashemi et al\., 2024](https://arxiv.org/html/2608.28965#bib.bib18)\)\. SAGE follows this domain\-calibrated direction for People Search\([Le et al\., 2026](https://arxiv.org/html/2608.28965#bib.bib12)\), but does not validate our prompted geographic\-entity judge\. We therefore freeze the judge’s model, prompt, rubric, and input fields across systems and directly measure its agreement with a blinded five\-annotator study\. Taken together, prior work addresses contextual toponym linking, alias\-rich entity normalization, synthetic retriever supervision, editable entity memory, and LLM assessment largely as separate problems\. Our contribution is not a new base encoder or judge\. It is the fixed\-ontology retrieval formulation that couples graded, set\-valued relevance with validated alias support, bounded same\-name contrast, and editable entity documents, then tests that formulation through entity\-only alias updates, public transfer, blinded production comparison, and online serving evidence\. This combination distinguishes the work from both end\-to\-end profile retrieval and single\-referent gazetteer linking\. ## 3\.Problem Formulation Letℰ\\mathcal\{E\}be a curated set of geographic entities\. Each entityeehas a canonical name, alternate names, type \(e\.g\., city, metropolitan area, administrative division, or country\), containment hierarchy, and an aggregate member\-count signal\. Given a location spanqqextracted from a people search query, the standardizer returns an ordered listRk\(q\)=\(e1,…,ek\)∈ℰkR\_\{k\}\(q\)=\(e\_\{1\},\\ldots,e\_\{k\}\)\\in\\mathcal\{E\}^\{k\}of distinct entities\. Candidate generation then uses the corresponding entity IDs as structured filters\. The embedding model receives only the extracted spanqq, not the full people search query\. A location span can admit multiple relevant entities\. “SF”, for example, may refer to both a city and a surrounding market area\. Conversely, an unqualified name such as “Alexandria” requires the dominant interpretation to precede less probable same\-name entities\. We therefore model relevance as graded and evaluate both ranking quality and downstream result coverage\. In the reported experiments, we retrievek=5k=5entities\. Their identifiers are passed to candidate generation as a disjunction \(an OR filter\); the embedding scores determine their retrieval order\. For binary Top\-kkevaluation, a query is counted as correct when at least one relevant entity appears inRk\(q\)R\_\{k\}\(q\)\. Graded ranking metrics additionally reward placing the most relevant interpretation earlier while allowing multiple entities to be valid\. Formally, lety\(q,e\)∈\{0,1,2,3,4\}y\(q,e\)\\in\\\{0,1,2,3,4\\\}be the relevance grade defined by Table[1](https://arxiv.org/html/2608.28965#S6.T1)\. For a returned listRk\(q\)R\_\{k\}\(q\), we compute \(1\)DCG@k\(q\)=∑i=1k2y\(q,ei\)−1log2\(i\+1\),\\operatorname\{DCG@\}k\(q\)=\\sum\_\{i=1\}^\{k\}\\frac\{2^\{y\(q,e\_\{i\}\)\}\-1\}\{\\log\_\{2\}\(i\+1\)\},For each reported comparison, letGqG\_\{q\}be the deduplicated union of the Top\-10 entities returned by all compared systems for queryqq\. Every pair in this shared pool is judged once\.IDCG@10\(q\)\\operatorname\{IDCG@10\}\(q\)is the DCG of the ten highest grades inGqG\_\{q\}, while each system’s DCG uses its own returned ordering\. Thus all systems share the same denominator\. nDCG@10 is defined as zero when the ideal gain is zero\. For thresholdτ∈\{3,4\}\\tau\\in\\\{3,4\\\}, Success@5 is the fraction of queries satisfyingmaxe∈R5\(q\)y\(q,e\)≥τ\\max\_\{e\\in R\_\{5\}\(q\)\}y\(q,e\)\\geq\\tau\. This query\-level success metric asks whether the OR filter contains at least one usable interpretation, whereas nDCG distinguishes a dominant referent from a valid but unlikely same\-name entity and rewards placing it earlier\. Reporting both avoids collapsing a multi\-entity operating point into single\-label accuracy\. We use a shared encoderhθh\_\{\\theta\}with prompt\-dependent inputs: \(2\)𝐮q=hθ\(pq,q\),𝐯e=hθ\(pe,d\(e\)\),\\mathbf\{u\}\_\{q\}=h\_\{\\theta\}\(p\_\{q\},q\),\\qquad\\mathbf\{v\}\_\{e\}=h\_\{\\theta\}\(p\_\{e\},d\(e\)\),wherepqp\_\{q\}is a retrieval instruction,pep\_\{e\}is an identity prompt, andd\(e\)d\(e\)is an entity document\. Vectors areℓ2\\ell\_\{2\}normalized and scored bys\(q,e\)=𝐮q⊤𝐯es\(q,e\)=\\mathbf\{u\}\_\{q\}^\{\\top\}\\mathbf\{v\}\_\{e\}\. Retrieval returns thekkhighest\-scoring entities, optionally after a type or product filter\. For the experiments reported here, we evaluate a compactD=64D=64representation withk=5k=5over a large fixed ontology\. Learning requires specifying both the positive support and the exclusion boundary\. Letp\+\(q∣e;Se,Ae,𝒯\)p\_\{\+\}\(q\\mid e;S\_\{e\},A\_\{e\},\\mathcal\{T\}\)be the distribution of query forms for entityee, induced by canonical formsSeS\_\{e\}, validated aliasesAeA\_\{e\}, and identity\-preserving transformations𝒯\\mathcal\{T\}\. Letp−\(𝒩∣q,e\)p\_\{\-\}\(\\mathcal\{N\}\\mid q,e\)generate a negative set from three sources: hierarchy\-overlap entities, dominance\-gated same\-name entities, and in\-batch entities\. The training problem is \(3\)minθ𝔼e∼pℰ,q∼p\+,𝒩∼p−\[ℓctr\(q,e,𝒩,θ,Ae\)\]\.\\min\_\{\\theta\}\\;\\mathbb\{E\}\_\{e\\sim p\_\{\\mathcal\{E\}\},\\,q\\sim p\_\{\+\},\\,\\mathcal\{N\}\\sim p\_\{\-\}\}\\left\[\\ell\_\{\\mathrm\{ctr\}\}\(q,e,\\mathcal\{N\};\\theta,A\_\{e\}\)\\right\]\.The method in Section[5](https://arxiv.org/html/2608.28965#S5)specifies these two distributions and the entity documentd\(e,Ae\)d\(e,A\_\{e\}\)\. Surface invariance and alias verification shapep\+p\_\{\+\}; hierarchy overlap and bounded dominance shapep−p\_\{\-\}; the mutable fieldAeA\_\{e\}determines which facts can be changed without updatingθ\\theta\. This separates the proposed learning formulation from the underlying contrastive bi\-encoder architecture\. Two constraints shape the design\. First, the lexical incumbent has its highest accuracy on the high\-volume head, whereas the proposed method targets a smaller tail; evaluation must report these regimes separately\. Second, entity documents change slowly and can be encoded offline\. Retrieval therefore requires one query encoding followed by exhaustive scoring over compact entity vectors, rather than online encoding of both sides or cross\-encoder inference\. ## 4\.Retrieval Architecture Figure[1](https://arxiv.org/html/2608.28965#S4.F1)separates the weekly entity pipeline from the per\-query path\. Offline, we assemble one text document per geographic entity, encode it, truncate and normalize the resulting vector, and store the vectors in a compact exact\-search index\. A compact, low\-precision representation permits exact scoring over the fixed entity ontology\. Online, upstream query understanding supplies only the extracted location span; the full natural\-language people search query is not passed to the embedding encoder\. The serving path encodes the short string into the same compact serving space, scans the filtered corpus exactly, and returns the Top\-kkcanonical entity IDs, not free text\. Atk=5k=5, all returned IDs are applied as OR filters by downstream candidate retrieval; downstream people ranking otherwise remains unchanged\. offline refreshonline retrievalCanonical geo corpus\+ validated aliasesShared encoderhθh\_\{\\theta\}identity prompt, offlineNormalized compactentity vectorsExact\-search indexprecomputed vectorsExtracted location spanqqShared encoderhθh\_\{\\theta\}query promptNormalized compactquery vectorFiltered exact Top\-kkGeographic entity IDsFigure 1\.Prompt\-asymmetric, shared\-weight bi\-encoder\. Entity representations are precomputed; each query requires one encoding and an exhaustive filtered search\.A shared bi\-encoder precomputes geographic entity vectors offline and encodes each location query online before exact filtered retrieval\.Two properties enable exhaustive retrieval\. First, the fixed entity side is precomputed during the weekly refresh, leaving only the query representation to be computed per request\. Second, the training loss directly optimizes the compact serving subspace\. This representation supports exact scoring without approximation\-induced recall loss or an index\-specific tuning variable\. Aliases are included in the entity document as well as in query\-side training examples\. Section[5\.5](https://arxiv.org/html/2608.28965#S5.SS5)defines the resulting entity\-only update, and Section[6\.3](https://arxiv.org/html/2608.28965#S6.SS3)evaluates it: a targeted alias addition requires re\-embedding only the affected entity\. Compatibility with the production contract\.The embedding path replaces only geographic standardization\. It consumes the same extracted span and returns the same geographic entity identifier type as the incumbent taxonomy\-based system, so candidate generation and people ranking do not require model\-specific features\. Retrieval depth is an explicit treatment: Top\-1 passes one entity filter, while Top\-5 passes five retrieved IDs as a disjunction\. This stable boundary supports three comparisons without changing the rest of the stack: entity\-level offline evaluation, paired calls to the production endpoint, and a randomized live experiment\. It also limits the scope of an alias edit, because refreshing an entity vector changes geographic matching but not the downstream ranking model\. ## 5\.Learning Robust and Editable Geo Representations ### 5\.1\.Design principles A generic bi\-encoder objective collapses three distinct error channels in geographic grounding\.*Surface variation*changes form without changing place identity;*knowledge\-dependent aliases*such as “pink city” cannot be derived from string transformations; and*ambiguity*arises when one string legitimately denotes several entities\. Treating them as undifferentiated augmentation either misses knowledge\-bearing expressions or creates false negatives for ambiguous names\. We assign each channel to a different part of Equation[3](https://arxiv.org/html/2608.28965#S3.E3)\. Identity\-preserving views and validated aliases define positive support underp\+p\_\{\+\}; calibrated sampling allocates that mass\. Hierarchy\-overlap and bounded same\-name construction shapep−p\_\{\-\}\. A mutable entity\-side alias set provides a separate update path: re\-embedding an edited entity changes entity\-specific knowledge without updating the encoder\. This allocation assigns each failure mode to a corresponding supervision or memory mechanism\. It does not require encoder\-specific architectural changes and avoids treating the three channels as interchangeable augmentation\. For entityee, letHeH\_\{e\}denote its canonical name, type, and containment hierarchy; letAeA\_\{e\}be a validated, mutable alias set\. We render an entity document \(4\)d\(e,Ae\)=render\(He,b\(e\),Ae\),d\(e,A\_\{e\}\)=\\operatorname\{render\}\(H\_\{e\},b\(e\),A\_\{e\}\),whereb\(e\)=Bucket\(me\)b\(e\)=\\operatorname\{Bucket\}\(m\_\{e\}\)andmem\_\{e\}is an aggregate prominence signal foree\. The resulting text contains the canonical name; city, administrative divisions, and country when available; entity type; a token representing the popularity bucketb\(e\)b\(e\); and anAlso Known Asfield\. For example: > Name: Mountain View, Administrative Division 1: California, Country: United States, Popularity bucket: b, Type: City, Also Known As: mv ca\. The bucket is a coarse platform\-specific prominence proxy rather than a census population estimate\. We use natural language for inspectability and compatibility with the encoder’s pretraining format; a later JSON training variant did not improve development\-set retrieval metrics\. The final entity document excludes geohashes: the query contains no corresponding spatial token, and removing the document\-side fields did not materially change development\-set point estimates\. ### 5\.2\.Identity\-preserving query views For each entity document, we construct a setSeS\_\{e\}of query forms conditioned on entity type\. A populated place may yield its name alone, name plus country, or name plus first\-level administrative division; thus “San Jose”, “San Jose US”, and “San Jose California” are distinct anchors for the same entity\. We then sample transformationst∈𝒯t\\in\\mathcal\{T\}known to preserve the entity identity and constructq=t\(s\)q=t\(s\)fors∈Ses\\in S\_\{e\}\. The transformation set includes country and US\-state abbreviation, reversed component order, comma insertion, and lowercasing\. Unlike unrestricted textual augmentation, these operations have an explicit invariance contract: they alter formatting or hierarchy expression without changing denotation\. This is important for geography, where a small character edit can produce another valid place name\. Common misspellings and orthographic variants are therefore admitted only through the alias validation path below, rather than through unconstrained character corruption\. ### 5\.3\.Generate–verify–calibrate alias learning Invariant views cannot derive knowledge\-bearing names such as “NYC”, “the Bay Area”, or “pink city”\. An instruction\-tuned model first proposes abbreviations, nicknames, alternative spellings, and colloquial forms fromHeH\_\{e\}\. An independent verifier evaluates each proposal against the canonical entity and its hierarchy\. We retain \(5\)Ae=\{a:a∈G\(He\),Accept\(a,e,He\)=1\},A\_\{e\}=\\\{a:a\\in G\(H\_\{e\}\),\\;\\operatorname\{Accept\}\(a,e,H\_\{e\}\)=1\\\},and remove case\-insensitive copies of the canonical name\. This generate–verify separation expands candidate coverage while preventing unverified generations from entering the supervision set\. Uniform alias weighting improved conditional ranking quality but reduced Top\-5 recall: verbose and lower\-confidence forms shifted probability mass away from minimal names\. We therefore use calibrated non\-uniform weights that prioritize minimal stylized names, then validated aliases, then other stylized forms\. A smooth prominence\-dependent sampler additionally increases the sampling rate for prominent entities\. The resulting positive distribution can be written as \(6\)p\(q∣e\)∝w\(s\)p\(t\),q=t\(s\),s∈Se∪Ae,t∈𝒯\.p\(q\\mid e\)\\propto w\(s\)\\,p\(t\),\\quad q=t\(s\),\\quad s\\in S\_\{e\}\\cup A\_\{e\},\\;t\\in\\mathcal\{T\}\.Alias generation therefore expands the support of the positive distribution, while sampling calibration bounds its effect on the learned geometry\. ### 5\.4\.Ambiguity\-aware negative construction Uniformly sampled negatives are dominated by lexically and geographically unrelated entities, yielding low\-confusability comparisons with limited signal for geographic disambiguation\. We instead construct negatives along two axes\. First,*hierarchy\-overlap negatives*share a subset of the query’s name and hierarchy components\. For “San Jose US”, San Jose, Costa Rica preserves the city name but conflicts with the country constraint, whereas an unrelated country shares neither component\. We stratify candidates by component overlap and sample round\-robin across strata so that high\-frequency patterns do not dominate the constructed\-negative distribution\. Second, same\-name entities require false\-negative control\. A less prominent entity is not automatically wrong; “Alexandria” may validly refer to several places\. We use same\-name entitye−e^\{\-\}as a negative fore\+e^\{\+\}only when \(7\)me\+\+1≥γ\(me−\+1\),m\_\{e^\{\+\}\}\+1\\geq\\gamma\(m\_\{e^\{\-\}\}\+1\),whereγ\>1\\gamma\>1is a fixed dominance threshold; we also cap negative exposure relative to positive frequency\. The location\-popularity token supplies an explicit tie\-breaker among same\-name geographic entities\. Together, the dominance threshold and cap bound the contrastive penalty assigned to valid secondary interpretations\. This design addresses a failure mode of the earlier objective, which improved the dominant entity’s rank by suppressing legitimate secondary entities\. Each positive receives a bounded set of constructed negatives; other batch positives provide global in\-batch negatives\. ### 5\.5\.Two\-timescale alias memory Aliases have two distinct roles\. Query\-side aliases supervise a semantic mapping that transfers across entities\. Entity\-side aliases in Equation[4](https://arxiv.org/html/2608.28965#S5.E4)explicitly condition the representation on entity\-specific or newly discovered knowledge\. After training with theAlso Known Asfield, an updateΔAe\\Delta A\_\{e\}changes only \(8\)𝐯e′=norm\(Pdhθ\(pe,d\(e,Ae∪ΔAe\)\)\),θ′=θ\.\\mathbf\{v\}^\{\\prime\}\_\{e\}=\\operatorname\{norm\}\\\!\\left\(P\_\{d\}h\_\{\\theta\}\(p\_\{e\},d\(e,A\_\{e\}\\cup\\Delta A\_\{e\}\)\)\\right\),\\qquad\\theta^\{\\prime\}=\\theta\.The system re\-embeds the affected entity and replaces its index vector; it does not regenerate training data or update model parameters\. This factorization yields two update timescales: model parameters encode cross\-entity regularities, while entity documents store mutable, entity\-specific facts\. Section[6\.3](https://arxiv.org/html/2608.28965#S6.SS3)evaluates the resulting update along three axes: new\-alias acquisition, canonical\-form retention, and locality on unaffected queries\. As a concrete intervention, addingSLCto the Salt Lake City document makes the previously missed query retrieve the intended entity after re\-embedding\. We restrict the document field to validated aliases; including lower\-confidence taxonomy abbreviations reduced retrieval quality in an earlier version\. ### 5\.6\.Training\-instance assembly The preceding components define a single training\-instance generator\. We first sample an entity using the smoothed prominence\-aware distribution, then select a minimal name, validated alias, or other stylized form using the calibrated ordering above\. An identity\-preserving transformation produces the query view\. The positive is the document of the same ontology entity, including its validated entity\-side aliases\. We then attach a bounded set of constructed negatives by cycling through hierarchy\-overlap strata and dominance\-qualified same\-name candidates; positive documents belonging to other batch items supply global in\-batch negatives\. Consequently, lexical proximity alone never makes an entity negative: ontology identity determines the positive, hierarchy creates targeted contrast, and Equation[7](https://arxiv.org/html/2608.28965#S5.E7)gates ambiguous same\-name contrast\. This assembly also separates two uses of synthetic text\. Generated aliases can enter query\-side supervision only after referent verification, where they teach a mapping intended to transfer across entities\. The entity\-sideAlso Known Asfield stores the validated subset as editable content\. Sampling calibration controls how often aliases affect parameter updates, whereas editing the document changes only one entity vector\. The distinction is important operationally: increasing alias sampling mass can change the global embedding geometry, while an entity edit is local to the refreshed vector\. It also motivates reporting retrieval quality and editability as related but different properties of the system\. ### 5\.7\.Objective and compact\-space training We fine\-tune an open\-source 0\.6B\-parameter embedding model as a shared\-weight, prompt\-asymmetric bi\-encoder\. For positive pair\(qi,ei\+\)\(q\_\{i\},e\_\{i\}^\{\+\}\)and the union𝒩i\\mathcal\{N\}\_\{i\}of constructed and in\-batch negatives, the cached multiple\-negatives objective is \(9\)ℒi=−logexp\(s\(qi,ei\+\)/τ\)exp\(s\(qi,ei\+\)/τ\)\+∑e∈𝒩iexp\(s\(qi,e\)/τ\)\.\\mathcal\{L\}\_\{i\}=\-\\log\\frac\{\\exp\(s\(q\_\{i\},e\_\{i\}^\{\+\}\)/\\tau\)\}\{\\exp\(s\(q\_\{i\},e\_\{i\}^\{\+\}\)/\\tau\)\+\\sum\_\{e\\in\\mathcal\{N\}\_\{i\}\}\\exp\(s\(q\_\{i\},e\)/\\tau\)\}\.For all reported experiments, we use full fine\-tuning for two epochs \(batch size 256, learning rate10−410^\{\-4\}, maximum sequence length 512, and bfloat16 precision\)\. Frozen 0\.6B, 4B, and 8B baselines come from one public embedding family; adaptation starts from its 0\.6B checkpoint, with architecture and tokenizer unchanged\. A Matryoshka objective\([Kusupati et al\., 2022](https://arxiv.org/html/2608.28965#bib.bib11)\)applies the retrieval loss to the leading 64 coordinatesPdP\_\{d\}in Equation[8](https://arxiv.org/html/2608.28965#S5.E8), rather than evaluating a compact post\-hoc projection, thereby aligning training and evaluation\. Query and entity templates are fixed\. These are evaluated study settings, not deployed\-system specifications; the exact checkpoint identity is omitted under organizational disclosure constraints\. Controlled comparisons fix initialization, training budget, the 64\-dimensional objective, templates, and evaluation sets, varying only the stated supervision or document component\. Section[6\.2](https://arxiv.org/html/2608.28965#S6.SS2)varies frozen encoder size; the GeoNames transfer holds the selected 0\.6B checkpoint fixed\. We do not apply unrestricted character\-level typo corruption\. Alternative spellings enter supervision only through the same validation path used for other aliases, preventing arbitrary edits from changing the intended referent\. ## 6\.Experiments We evaluate four questions: whether the specialized design contributes beyond standard fine\-tuning and encoder scale; how entity\-only updates trade alias acquisition against canonical retention and compare with lexical insertion; whether the selected model transfers to public GeoNames; and what relevance and engagement evidence is provided by blinded production comparisons, fixed\-query endpoint evaluation, and a live A/B test\. ### 6\.1\.Evaluation protocol Leakage control and split integrity\.The benchmark holds out surface forms, not entities, from a fixed inventory, and evaluation queries are frozen before training\. After Unicode normalization, case folding, and punctuation and whitespace normalization, we remove every evaluation\-form match from query supervision, generated aliases, and entity documents\. We define near duplicates by Jaccard similarity≥0\.8\\geq 0\.8over boundary\-padded 3–5\-character\-gram unions and keep each resulting group within one split\. Filtering is repeated after alias generation\. Consequently, no normalized evaluation form is used as a training query or document alias, and no near\-duplicate group crosses splits\. Fixed prompted geo judge\.For entity retrieval, a prompted LLM judge receives only the extracted location text and the candidate’s name, type, country, city, and administrative hierarchy; geohash and retrieval scores are excluded\. It scores each pair\(q,e\)\(q,e\)from 0 to 4 using the rubric in Table[1](https://arxiv.org/html/2608.28965#S6.T1)\. The rubric separates the dominant referent from technically valid but unlikely same\-name places, and penalizes a nearby place that does not denote the query\. We aggregate the grades with nDCG@10 and thresholded Success@5 and precision at the reported cutoffs; unless a stricter threshold is shown explicitly, grades 3–4 count as relevant\. The same metric implementation and ideal\-gain convention in Section[3](https://arxiv.org/html/2608.28965#S3)are used for every compared variant\. Specifically, for each query we pool and deduplicate the Top\-10 entities from all systems in the reported comparison, judge each pooled query–entity pair, and derive a shared IDCG@10 from that pool\. The judge model, prompt, decoding configuration, rubric, and input fields are fixed across all compared retrieval variants\. We do not assume its validity from prior work; the blinded five\-annotator study in Section[6\.5](https://arxiv.org/html/2608.28965#S6.SS5)directly measures its agreement with human consensus on our geo\-specific outputs\. Table 1\.Geo\-relevance rubric, illustrated for the query “la”\. ### 6\.2\.Task adaptation versus encoder scale We first test whether pretrained model scale can replace task adaptation\. All variants are evaluated on the same large, fixed production\-derived development set against the same entity inventory\. The prompted LLM judge, its 0–4 rubric, and the retrieval pipeline are fixed across variants\. The frozen checkpoints come from a single open\-source embedding family at 0\.6B, 4B, and 8B parameters\. We apply the same prompts and retrieval path to each checkpoint without task\-specific parameter updates\. Table[2](https://arxiv.org/html/2608.28965#S6.T2)compares them with the selected task\-adapted 0\.6B model\. Success@5 \(≥t\\geq t\) is the fraction of queries for which at least one Top\-5 entity receives judge gradettor higher; nDCG@10 retains the graded relevance signal\. Table 2\.Frozen open\-source models versus task adaptation\. Success@5 counts queries with at least one Top\-5 result at or above the indicated grade threshold; bold marks the best result\.Task adaptation raises nDCG@10 by 0\.4620 \(100\.9%\) over the same\-size frozen encoder and by 0\.2077 \(29\.2%\) over the frozen 8B encoder\. The Success@5 gaps are similarly large\. Thus, increasing pretrained model scale alone does not recover the task\-specific supervision\. Because this benchmark is used for model selection, these are controlled development\-set point estimates rather than held\-out test results; GeoNames and the production evaluations provide external evidence\. Controlled ablation study\.All unlisted factors are held fixed: initialization, random seed, optimization schedule, training budget, entity inventory, and evaluation set\. The first row is standard surface\-form fine\-tuning with the shared encoder and objective; each subsequent comparison changes only the named intervention\. Table 3\.Controlled ablations of supervision and entity representation\. In the first block, eachΔ\\Deltais relative to the preceding row\. In the second, JSON and validated\-only compare with the NL reference; no\-geohash compares with validated\-only\. Bold denotes the offline best; the dagger marks the deployed variant\.ConfigurationΔ\\DeltanDCG@10Cumulative supervision and ambiguity handlingStandard surface\-form fine\-tuning—0\.8099\+\+Initial ambiguity\-aware sampling\+0\.0511\+0\.05110\.8610\+\+Refined ambiguity\-aware sampling\+0\.0087\+0\.00870\.8697\+\+Validated query\-side aliases\+0\.0127\+0\.01270\.8824\+\+Calibrated source sampling\+0\.0067\+0\.00670\.8891\+\+Popularity\-bounded same\-name negatives\+0\.0197\+0\.01970\.9088Controlled entity\-representation changesNL documents with editable aliases \(reference\)—0\.9133JSON serialization only−0\.0092\-0\.00920\.9041Validated\-only entity aliases\+0\.0118\\mathbf\{\+0\.0118\}0\.9251Validated aliases, no geohashes†\\dagger−0\.0050\-0\.00500\.9201The cumulative supervision sequence raises nDCG@10 from 0\.8099 to 0\.9088\. Initial ambiguity\-aware sampling gives the largest single increase \(\+0\.0511\+0\.0511\); validated query aliases add 0\.0127, and popularity\-bounded same\-name negatives add 0\.0197, consistent with their intended roles\. Relative to natural\-language documents, JSON serialization reduces nDCG@10 by 0\.0092 and validated\-only entity aliases improve it by 0\.0118\. Removing geohashes from that offline best costs only 0\.0050 \(0\.54%\) and caused no meaningful endpoint change while simplifying document updates, so we deployed the no\-geohash variant\. The first row is the direct control for ordinary task fine\-tuning: the specialized supervision and selected representation raise nDCG@10 by a further 0\.1102, from 0\.8099 to 0\.9201, under the fixed conditions above\. These deltas are conditional component effects, not training\-seed variance estimates\. ### 6\.3\.Editable alias memory We evaluate tens of thousands of held\-out alias–entity mappings across tens of thousands of entities\. Their query forms are absent from task training and the corresponding pre\-update entity documents; 3\.4% of unique surfaces map to multiple geographic entity IDs\. Acquisition uses every selected mapping, retention uses canonical queries available for the updated entities, and locality uses disjoint alias and canonical controls whose targets are not updated\. For the entity\-document update, we add aliases only to the target entity’sAlso Known Asfield and recompute that vector with the encoder and all other vectors frozen\. The lexical update inserts the same mappings into an exact alias table\. Each mechanism is paired with its own pre\-update state\. We report target Recall@1 and Recall@5 for acquisition, canonical target Recall@5 change for retention, and control nDCG@10 change for locality\. Confidence intervals use 10,000 entity\-level paired bootstrap replicates\. Table 4\.Incremental alias acquisition\. Each mechanism is compared with its own pre\-update state;Δ\\Deltacells report entity\-level paired\-bootstrap 95% confidence intervals\. Unshown disjoint controls are unchanged at four decimals\.The entity\-document update raises target Recall@1 from 0\.1819 to 0\.6435 and target Recall@5 from 0\.2869 to 0\.7664\. Controls remain flat, but canonical Recall@5 decreases by 0\.0211 \(95% CI \[−0\.0231\-0\.0231,−0\.0193\-0\.0193\]\), making the update local rather than retention\-neutral\. Exact lexical insertion reaches 0\.9369 Recall@1 and 0\.9949 Recall@5; its canonical Recall@5 decreases by only 0\.0018 \(95% CI \[−0\.0023\-0\.0023,−0\.0014\-0\.0014\]\) and controls remain unchanged\. The mechanisms are therefore complementary: lexical tables suit validated deterministic mappings, whereas entity documents offer a localized neural update without encoder retraining, at a larger retention cost\. ### 6\.4\.Transfer to a public ontology We test transfer on a frozen GeoNames snapshot released under CC BY 4\.0\([GeoNames, 2026](https://arxiv.org/html/2608.28965#bib.bib34)\)\. The candidate inventory retains country, administrative\-area, and populated\-place entries \(feature classesAandP\)\. Without production fields, each document contains canonical and ASCII names, feature code, country and administrative hierarchy, and a population bucket; coordinates remain excluded\. Evaluation\.A 3,002,271\-query canonical\-name slice serves only as a harness sanity check\. Alias evidence uses two fixed sets: known\-target metrics use 860,582 source\-linked pairs, whereas pooled graded metrics use 716,322 held\-out queries\. Every method shares the queries and candidate inventory within each metric group\. Query forms are excluded from the corresponding training pairs and entity documents\. In the source\-linked set, 29\.5% of pairs have a surface associated with multiple entities and are scored against their own source entity\. Lexical systems share Unicode\-aware query/document normalization: Latin diacritics are removed after decomposition, text is lowercased, and runs of non\-alphanumeric, non\-mark characters collapse to spaces\. Character retrieval unions boundary\-padded whole\-string 3–5\-grams, scores them with sublinear TF–IDF cosine, and drops grams occurring in more than 1% of documents; BM25 usesk1=1\.2k\_\{1\}=1\.2andb=0\.75b=0\.75\. All systems use the same public name fields and inventory, produce Top\-10 before labels, and exclude each held\-out form from its source entity document\. The proprietary production standardizer cannot be ported to GeoNames without replacing its inventory and taxonomy\-dependent logic\. We instead compare exact and prefix matching, BM25, characternn\-grams, a frozen open\-source 0\.6B encoder, and our checkpoint, which is selected on the production\-derived development set and not tuned on GeoNames\. Target Recall@1, Target Recall@5, and MRR@10 follow directly from the source entity identifier and require no judge; MRR@10 is zero when the target is absent from the Top\-10\. Pooled judged nDCG@10 and P@1 allow another entity sharing the surface to receive graded relevance\. Table 5\.GeoNames alias transfer\. Known\-target metrics use each alias’s source entity; pooled metrics allow graded relevance to other valid entities\. Bold marks the best result\.Table 6\.Character\-overlap analysis on GeoNames\. Bold marks the higher Target Recall@1 per slice; ours leads for allJ≤0\.50J\\leq 0\.50bands \(79\.1% of pairs\) and overall, while characternn\-gram leads forJ\>0\.50J\>0\.50\. Deltas are ours minus characternn\-gram; 95% CIs use 10,000 query\-clustered bootstrap replicates\.For known\-target metrics, the aggregate task\-adapted\-minus\-character\-nn\-gram differences are\+0\.0101\+0\.0101at Recall@1 \(95% CI \[\+0\.0094\+0\.0094,\+0\.0108\+0\.0108\]\),−0\.0055\-0\.0055at Recall@5 \(\[−0\.0064\-0\.0064,−0\.0046\-0\.0046\]\), and\+0\.0032\+0\.0032at MRR@10 \(\[\+0\.0025\+0\.0025,\+0\.0039\+0\.0039\]\)\. The intervals use 10,000 query\-clustered paired bootstrap replicates and none spans zero\. Fixed overlap bands localize this trade\-off\. Adaptation significantly improves Target Recall@1 in allJ≤0\.50J\\leq 0\.50bands \(79\.1% of pairs\), with the gain rising from 0\.0163 atJ=0J=0to 0\.0453 for moderate overlap; characternn\-grams win forJ\>0\.50J\>0\.50\. At Target Recall@5, gains in the zero\- and low\-overlap bands are outweighed by the high\-overlap loss, producing the aggregate−0\.0055\-0\.0055\. Both systems improve in absolute Recall@1 as overlap rises, so the result indicates lower, not zero, dependence on surface overlap\. GeoNames is a transfer benchmark, not a production\-selection proxy: the overlap analysis diagnoses lexical regimes, while deployment selection rests on the production\-derived, human, endpoint, and live evidence below\. On pooled judged point estimates, adaptation improves over the frozen encoder by 65\.1% in nDCG@10 and 166\.4% in P@1, exceeds exact, prefix, and BM25 on both measures, and narrowly leads characternn\-grams by 0\.0003 nDCG@10 and 0\.0062 P@1\. On the canonical sanity slice, exact matching and the adapted model reach 0\.9699 and 0\.9885 P@1\. ### 6\.5\.Blinded human evaluation and judge validation The 50\-query challenge set comprises 30 high\-frequency production location queries, 10 queries selected by product managers from quality tickets filed before the present evaluation, and 10 queries sampled at random from historical search impressions\. We intentionally restrict the set to queries for which the systems’ Top\-1 entities differ\. It targets actionable cases rather than estimating a traffic\-wide effect\. Five internal domain experts from engineering and product management, all familiar with people search and geographic standardization, independently graded both outputs for every query on the 0–4 scale in Table[1](https://arxiv.org/html/2608.28965#S6.T1)\. The 100 query–document pairs were presented in random order without system identifiers, yielding 500 ratings\. For each output, grade@1 is the mean of its five ratings\. We call an output relevant when at least three annotators assign grade 3 or 4\. Mean\-grade differences use the 50 complete production–embedding pairs; the confidence interval is obtained by paired query\-level bootstrap and thepp\-value by a pairedtt\-test\. Relevant P@1 uses an exact McNemar test\. Win/tie/loss compares the two mean grades within each query and uses an exact sign test after removing ties\. For judge validation, the human consensus for itemiiis the median of its five ordinal ratings,hiconsensus=median\(hi1,…,hi5\)h\_\{i\}^\{\\mathrm\{consensus\}\}=\\operatorname\{median\}\(h\_\{i1\},\\ldots,h\_\{i5\}\)\. Spearman’sρ\\rhomeasures rank association between the judge grade and this consensus; weighted Cohen’sκ\\kappameasures agreement while accounting for the distance between ordinal grades\. Table 7\.Blinded Top\-1 comparison on the stratified production challenge set\. W/T/L denotes embedding wins, ties, and losses\.Paired inference\.Grade@1: 95% CI \[\+0\.136\+0\.136,\+0\.456\+0\.456\],p=0\.001p=0\.001; relevant P@1: \[\+6\.0\+6\.0,\+30\.0\+30\.0\] pp,p=0\.012p=0\.012; W/T/L: exact sign testp=0\.011p=0\.011\. The embedding system raises average grade@1 by 0\.296 \(p=0\.001p=0\.001\) and relevant P@1 by 18\.0 percentage points, from 28\.0% to 46\.0%\. Ten queries are relevant only for embedding and one only for production \(p=0\.012p=0\.012, exact McNemar\); graded outcomes give 26/14/10 wins/ties/losses \(p=0\.011p=0\.011\)\. Thus both graded and thresholded human measures favor embedding on this stratified Top\-1 challenge set\. #### People\-result human audit\. Three experts independently scored blinded Top\-1 outputs for 50 paired endpoint queries \(100 items; 300 ratings\)\. Median grade defines consensus, grades 3–4 define relevance, and wins compare paired consensus grades\. Table 8\.Blinded endpoint audit and judge validation against median human consensus\.\(a\) Endpoint Top\-1 comparison \(b\) Agreement by evaluation distribution Panel \(a\): excluding 24 ties, the two\-sided exact sign test givesp=0\.845p=0\.845\. Exp\. denotes experts; QWK and Spearman’sρ\\rhocompare each judge with median human consensus\. Treatment and control have mean consensus grades of 3\.04 and 3\.00 and relevant P@1 of 74% and 72%\. Treatment wins 14 queries, control wins 12, and 24 tie \(p=0\.845p=0\.845\); this supports similar sampled Top\-1 relevance, not equivalence\. For production geo, expert agreement isα=0\.712\\alpha=0\.712and judge alignment is QWK=0\.761=0\.761,ρ=0\.781\\rho=0\.781; for people results, the corresponding values are 0\.790, 0\.680, and 0\.730\. Exact and within\-one\-level judge agreement on people results are 64% and 93%; binary agreement at grade 3 is 92% \(κ=0\.83\\kappa=0\.83\)\. ### 6\.6\.Online relevance and engagement evaluation Paired production\-endpoint calls use traffic\-derived frequent queries and non\-canonical\-location queries with randomly sampled member identifiers\. Control uses the taxonomy standardizer; treatment enables Top\-1 or Top\-5 geo\-embedding filters with the remaining stack fixed\. The human\-validated people\-result judge in Table[8](https://arxiv.org/html/2608.28965#S6.T8)grades the returned Top\-10 people results\. Geo Top\-kkdenotes filter depth, not people\-result rank\. These fixed\-query calls measure serving relevance rather than live\-user effects\. Table 9\.Fixed\-query endpoint relevance\. Values are mean judged grade@10; Geo Top\-kkis geographic retrieval depth\. Relative changes are computed from unrounded means\. Descriptive point estimates only; no confidence intervals are available\.Frequent\-query estimates stay within 0\.4% of control; non\-canonical estimates increase by 2\.1% at Top\-1 and 5\.0% at Top\-5\. Without query\-level intervals, we make no significance or non\-inferiority claim\. Separately, a one\-week, 50/50 member\-randomized A/B test detects no regression in top\-level engagement and search\-health guardrails, including query volume and search success rate; exact values are withheld\. Limitations\.GeoNames differs from production, and its aliases may occur in base\-model pretraining\. Its pooled judged results are point estimates, while paired intervals cover only known\-target comparison with characternn\-grams\. The update benchmark tests local edits rather than globally novel knowledge\. Finite human audits do not establish judge validity across all regions, languages, query types, or ranks; character overlap diagnoses lexical evidence, not semantic reasoning\. Fixed\-query endpoint results lack randomization and intervals, while the live test establishes only an engagement guardrail\. ## 7\.Discussion Ordinary task fine\-tuning reaches 0\.8099 nDCG@10 versus 0\.9201 for the selected task\-adapted system under fixed conditions; Table[3](https://arxiv.org/html/2608.28965#S6.T3)thus reports conditional effects under one seed\. On GeoNames, adaptation improves Target Recall@1 forJ≤0\.50J\\leq 0\.50\(79\.1% of pairs\), whereas characternn\-grams dominate at high overlap, making aggregate Recall@5 a regime\-composition result\. This public baseline does not change the production role: the learned standardizer replaces the incumbent taxonomy\-based path\. Other evaluations also answer distinct questions\. Lexical insertion is strongest for deterministic mappings; entity documents enable neural updates but reduce canonical retention\. The challenge set measures Top\-1 disagreements, endpoint gains are descriptive, and the live experiment supplies only an engagement guardrail\. These estimands should not be conflated\. ## 8\.Conclusion We formulate geographic standardization as graded, set\-valued retrieval\. Fixed\-seed development ablations show gains beyond ordinary task fine\-tuning and encoder scale, tracing them to ambiguity\-aware sampling, verified aliases, and mutable entity documents\. On GeoNames, adaptation leads Target Recall@1 and MRR@10 overall and Recall@1 forJ≤0\.50J\\leq 0\.50; characternn\-gram leads on high\-overlap forms and aggregate Target Recall@5\. The gain is therefore regime\-dependent\. Human relevant P@1 rises from 28\.0% to 46\.0%; endpoint estimates improve on non\-canonical queries with no live engagement regression\. Together, these results support replacing taxonomy\-based standardization with task\-adapted retrieval\. ## 9\.Ethical and Responsible AI Considerations The geographic model is non\-personalized: it resolves location phrases against a curated ontology and does not encode, profile, or rank people, or use individual sensitive attributes as model features\. The member\-count signal is aggregated at the location level and used only to order geographic candidates for same\-name disambiguation; it is never used to rank people\. Reported production results are aggregate changes rather than individual query or member records\. The audit of the stratified production challenge set used five internal domain experts from engineering and product management; the people\-result audit used three internal expert raters\. All raters were familiar with people search and used the predefined ordinal relevance rubric\. Outputs were randomized and system\-blinded, and every expert independently rated every item in the corresponding audit\. The expert audits and member\-randomized experiment followed the organization’s applicable experimentation, privacy, and human\-participant review requirements\. Safeguards are applied at the geographic\-entity level\. Popularity is bucketed, and dominance gating with a per\-entity cap constrains its use in same\-name negative construction\. Aliases are validated before entering training or entity documents\. Public\-catalog transfer, the audit of the stratified production challenge set, and the people\-result audit are reported separately\. The challenge set’s purposive sampling and limited scope are disclosed rather than used as evidence of production\-wide, regional, or linguistic quality\. Our evaluation draws from a global location inventory and is not restricted to a particular language or region\. Aggregate metrics nevertheless do not establish uniform performance across region–language strata\. Future evaluation should report region\-by\-language slices, false\-broadening rates, and performance for low\-population entities, with human review of a stratified sample, before making parity claims\. ## References - Borisyuket al\.\(2026\)F\. Borisyuk, S\. Vasudevan, M\. Wu, G\. Li, B\. Le,et al\.Semantic search at LinkedIn\.arXiv preprint arXiv:2602\.07309\.External Links:[Link](https://arxiv.org/abs/2602.07309)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p1.1)\. - Daiet al\.\(2023\)Z\. Dai, V\. Y\. Zhao, J\. Ma, Y\. Luan, J\. Ni, J\. Lu, A\. Bakalov, K\. Guu, K\. B\. Hall, and M\. ChangPromptagator: few\-shot dense retrieval from 8 examples\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p4.1)\. - Faggioliet al\.\(2023\)G\. Faggioli, L\. Dietz, C\. L\. A\. Clarke, G\. Demartini, M\. Hagen, C\. Hauff, N\. Kando, E\. Kanoulas, M\. Potthast, B\. Stein, and H\. WachsmuthPerspectives on large language models for relevance judgment\.InProceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval,pp\. 39–50\.External Links:[Document](https://dx.doi.org/10.1145/3578337.3605136)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - Fröbeet al\.\(2025\)M\. Fröbe, A\. Parry, F\. Schlatt, S\. MacAvaney, B\. Stein, M\. Potthast, and M\. HagenLarge language model relevance assessors agree with one another more than with human assessors\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,External Links:[Document](https://dx.doi.org/10.1145/3726302.3730218)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - GeoNames \(2026\)GeoNamesGeoNames gazetteer data export\.Note:[https://download\.geonames\.org/export/dump/](https://download.geonames.org/export/dump/)CC BY 4\.0; accessed July 30, 2026Cited by:[§6\.4](https://arxiv.org/html/2608.28965#S6.SS4.p1.1)\. - Geyiket al\.\(2018\)S\. C\. Geyik, Q\. Guo, B\. Hu, C\. Ozcaglar, K\. Thakkar, X\. Wu, and K\. KenthapadiTalent search and recommendation systems at LinkedIn: practical challenges and lessons learned\.InProceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1353–1354\.External Links:[Document](https://dx.doi.org/10.1145/3209978.3210205)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p1.1)\. - Gillicket al\.\(2019\)D\. Gillick, S\. Kulkarni, L\. Lansing, A\. Presta, J\. Baldridge, E\. Ie, and D\. Garcia\-OlanoLearning dense representations for entity retrieval\.InProceedings of the 23rd Conference on Computational Natural Language Learning,pp\. 528–537\.External Links:[Document](https://dx.doi.org/10.18653/v1/K19-1049)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p3.1)\. - Grittaet al\.\(2018\)M\. Gritta, M\. T\. Pilehvar, and N\. CollierWhich Melbourne? augmenting geocoding with maps\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics,Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Guptaet al\.\(2025\)R\. Gupta, C\. Zheng, and H\. LiRetrieval for semantic people search\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 4229–4233\.External Links:[Document](https://dx.doi.org/10.1145/3726302.3731938)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p1.1)\. - Hashemiet al\.\(2024\)H\. Hashemi, J\. Eisner, C\. Rosset, B\. Van Durme, and C\. KedzieLLM\-rubric: a multidimensional, calibrated approach to automated evaluation of natural language texts\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 13806–13834\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.745)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - Huanget al\.\(2013\)P\. Huang, X\. He, J\. Gao, L\. Deng, A\. Acero, and L\. HeckLearning deep structured semantic models for web search using clickthrough data\.InProceedings of the 22nd ACM International Conference on Information and Knowledge Management,pp\. 2333–2338\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1)\. - Karpukhinet al\.\(2020\)V\. Karpukhin, B\. Oğuz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. YihDense passage retrieval for open\-domain question answering\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing,pp\. 6769–6781\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1)\. - Kimet al\.\(2024\)J\. Kim, D\. Ko, and G\. KimDynamicER: resolving emerging mentions to dynamic entities for RAG\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 13752–13770\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.762)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1)\. - Kusupatiet al\.\(2022\)A\. Kusupati, G\. Bhatt, A\. Rege, M\. Wallingford, A\. Sinha, V\. Ramanujan, W\. Howard\-Snyder, K\. Chen, S\. Kakade, P\. Jain, and A\. FarhadiMatryoshka representation learning\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1),[§5\.7](https://arxiv.org/html/2608.28965#S5.SS7.p1.2)\. - Leet al\.\(2026\)B\. H\. Le, X\. Lu, N\. Stern, W\. Liu, I\. Lapchuk, X\. Li, B\. Zheng, K\. Rosenberg, J\. Huang, Z\. Zhang, A\. Cabangbang, S\. M\. Wagle, J\. Shen, R\. Muthuregunathan, A\. Gupta, M\. Teoh, A\. J\. N\. Kirk, T\. Kwan, J\. Wu, and W\. ZhangSAGE: scalable AI governance & evaluation\.InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2,External Links:[Document](https://dx.doi.org/10.1145/3770855.3818476)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems,Vol\.33\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1)\. - Liet al\.\(2022\)Z\. Li, J\. Kim, Y\. Chiang, and M\. ChenSpaBERT: a pretrained language model from geographic data for geo\-entity representation\.InFindings of the Association for Computational Linguistics: EMNLP 2022,pp\. 2757–2769\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.findings-emnlp.200)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Liet al\.\(2023\)Z\. Li, W\. Zhou, Y\. Chiang, and M\. ChenGeoLM: empowering language models for geospatially grounded language understanding\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 5227–5240\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Liuet al\.\(2021\)F\. Liu, E\. Shareghi, Z\. Meng, M\. Basaldella, and N\. CollierSelf\-alignment pretraining for biomedical entity representations\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 4228–4238\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.334)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p3.1)\. - Logeswaranet al\.\(2019\)L\. Logeswaran, M\. Chang, K\. Lee, K\. Toutanova, J\. Devlin, and H\. LeeZero\-shot entity linking by reading entity descriptions\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 3449–3460\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1335)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1)\. - Masis and O’Connor \(2024\)T\. Masis and B\. O’ConnorWhere on earth do users say they are?: geo\-entity linking for noisy multilingual user input\.InProceedings of the Sixth Workshop on Natural Language Processing and Computational Social Science,pp\. 86–98\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.nlpcss-1.7)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Nakataniet al\.\(2025\)H\. Nakatani, H\. Teranishi, S\. Higashiyama, Y\. Sawada, H\. Ouchi, and T\. WatanabeA text embedding model with contrastive example mining for point\-of\-interest geocoding\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 7279–7291\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Quet al\.\(2021\)Y\. Qu, Y\. Ding, J\. Liu, K\. Liu, R\. Ren, W\. X\. Zhao, D\. Dong, H\. Wu, and H\. WangRocketQA: an optimized training approach to dense passage retrieval for open\-domain question answering\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 5835–5847\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.466)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p4.1)\. - Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence\-BERT: sentence embeddings using siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing,pp\. 3982–3992\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p5.1)\. - Rücker and Akbik \(2025\)S\. Rücker and A\. AkbikEvaluating design decisions for dual encoder\-based entity disambiguation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics,pp\. 15685–15701\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.764)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p3.1)\. - Sunget al\.\(2020\)M\. Sung, H\. Jeon, J\. Lee, and J\. KangBiomedical entity representations with synonym marginalization\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 3641–3650\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.335)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p3.1)\. - Takehiet al\.\(2025\)R\. Takehi, E\. M\. Voorhees, T\. Sakai, and I\. SoboroffLLM\-assisted relevance assessments: when should we ask LLMs for help?\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 95–105\.External Links:[Document](https://dx.doi.org/10.1145/3726302.3729916)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - Thomaset al\.\(2024\)P\. Thomas, S\. Spielman, N\. Craswell, and B\. MitraLarge language models can accurately predict searcher preferences\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1930–1940\.External Links:[Document](https://dx.doi.org/10.1145/3626772.3657707)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - Upadhyayet al\.\(2024\)S\. Upadhyay, R\. Pradeep, N\. Thakur, N\. Craswell, and J\. LinUMBRELA: UMbrela is the \(open\-source reproduction of the\) Bing relevance assessor\.arXiv preprint arXiv:2406\.06519\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p6.1)\. - Weissenbacheret al\.\(2019\)D\. Weissenbacher, A\. Magge, K\. O’Connor, M\. Scotch, and G\. Gonzalez\-HernandezSemEval\-2019 task 12: toponym resolution in scientific papers\.InProceedings of the 13th International Workshop on Semantic Evaluation,pp\. 907–916\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Wuet al\.\(2020\)L\. Wu, F\. Petroni, M\. Josifoski, S\. Riedel, and L\. ZettlemoyerScalable zero\-shot entity linking with dense entity retrieval\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing,pp\. 6397–6407\.Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p3.1)\. - Xionget al\.\(2021\)L\. Xiong, C\. Xiong, Y\. Li, K\. Tang, J\. Liu, P\. Bennett, J\. Ahmed, and A\. OverwijkApproximate nearest neighbor negative contrastive learning for dense text retrieval\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p4.1)\. - Zhang and Bethard \(2023\)Z\. Zhang and S\. BethardImproving toponym resolution with better candidate generation, transformer\-based reranking, and two\-stage resolution\.InProceedings of the 12th Joint Conference on Lexical and Computational Semantics,pp\. 48–60\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.starsem-1.6)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\. - Zhanget al\.\(2024\)Z\. Zhang, E\. Laparra, and S\. BethardImproving toponym resolution by predicting attributes to constrain geographical ontology entries\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Short Papers\),pp\. 35–44\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-short.3)Cited by:[§2](https://arxiv.org/html/2608.28965#S2.p2.1)\.
Similar Articles
Follow the Entities: A Corpus Map for Agentic Search
CorpusMap is an entity-based navigation layer for agentic search that organizes large document collections around recurring entities to improve evidence discovery and answer quality while reducing token usage.
Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent
The paper introduces PersonaLink, a training-free method that distills user history into a bounded persona, matching retrieval on classification tasks but not on regression, highlighting a task-type asymmetry.
Location-Aware Language Models via Secondary Embeddings
The paper proposes a lightweight, model-agnostic method to enhance language models with geo-spatial awareness by augmenting embeddings with location data, improving spatial alignment while maintaining standard NLP performance.
EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development
The paper introduces EAR, an entity-aware partitioning approach for retrieval-augmented generation in multiple-choice question answering, which extracts anchors from questions and corpora to retrieve compact windows instead of fixed-size chunks.
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
This paper introduces Pi-Serini, a BM25-based agentic search system that demonstrates lexical retrieval can suffice for deep search when agents refine queries, achieving high accuracy and reducing costs compared to default settings.