generative-retrieval

Tag

Cards List
#generative-retrieval

@omarsar0: Great RAG paper from IBM. There are some really good ideas on how to solve common RAG issues. It's well known that retr…

X AI KOLs Following · 2026-09-06 Cached

IBM researchers introduce STAIR, a generative retriever that uses table of contents to preserve document structure, achieving 82.6% Recall@1 and reducing hallucination in RAG systems.

0 favorites 0 likes
#generative-retrieval

Off-Policy Evaluation for Semantic ID Recommenders: Does the Model's Own Code Hierarchy Help?

arXiv cs.LG · 2026-09-01 Cached

This paper explores using a model's semantic ID hierarchy for off-policy evaluation in generative recommenders, showing that coarsening to code-prefix clusters improves estimation accuracy under production logging constraints.

0 favorites 0 likes
#generative-retrieval

PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization

arXiv cs.LG · 2026-08-26 Cached

PRQ-KMeans is a post-hoc tokenization method that improves residual quantization for semantic identifiers by removing global-mean components, refining centroids, and using projection residuals, achieving significant performance gains in industrial search and public recommendation benchmarks.

0 favorites 0 likes
#generative-retrieval

UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval

arXiv cs.AI · 2026-08-05 Cached

Kuaishou researchers propose UniGD, a unified generative-discriminative framework for industrial retrieval that integrates retrieval and relevance scoring into a single model, with techniques like CAGE and CAM to improve effectiveness and reduce latency. Online A/B tests show a 5.78% ad revenue increase and 33.1% inference latency reduction.

0 favorites 0 likes
#generative-retrieval

FlashTrie: A GPU-Accelerated Constrained Beam Search for Generative Retrieval

arXiv cs.LG · 2026-07-14 Cached

FlashTrie presents a GPU-accelerated constrained beam search for generative retrieval, using a succinct trie layout and cooperative CUDA kernels to reduce decoding latency and enable real-time serving at scale, achieving up to 24× speedup and a 0.71% revenue lift in a commercial search engine.

0 favorites 0 likes
#generative-retrieval

ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging

Hugging Face Daily Papers · 2026-05-12 Cached

ORBIT proposes a method to mitigate catastrophic forgetting in large language models fine-tuned for generative retrieval by tracking parameter distances and using weight averaging, outperforming common continual learning baselines.

0 favorites 0 likes
← Back to home

Submit Feedback