llm

Tag

Cards List
#llm

AI has no intent and no motivation

Hacker News Top · 53m ago Cached

The article argues that AI, particularly large language models, lack intent and motivation, which undermines concerns about AI posing threats or going rogue.

0 favorites 0 likes
#llm

Feedback on a Privacy Middleware for Sending Sensitive Data to LLMs

Reddit r/AI_Agents · 4h ago

A concept for a privacy middleware that anonymizes sensitive data before sending to LLMs to protect privacy, with local placeholder mapping, seeking feedback on feasibility.

0 favorites 0 likes
#llm

Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

arXiv cs.LG · 6h ago Cached

This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.

0 favorites 0 likes
#llm

SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection

arXiv cs.LG · 6h ago Cached

SR-Fraud is an outcome-supervised reflective LLM agent framework for non-stationary payment fraud detection, improving detection metrics over traditional methods on a production benchmark.

0 favorites 0 likes
#llm

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

arXiv cs.LG · 6h ago Cached

The paper proposes LLMAE, a method to repurpose pre-trained decoder-only LLMs as continuous text autoencoders using a latent bottleneck, achieving high-fidelity reconstruction and enabling downstream tasks like image captioning.

0 favorites 0 likes
#llm

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

arXiv cs.LG · 6h ago Cached

COPE is a novel optimization framework for continual personalization of large language models under sparse user feedback, using learnable user embeddings and self-evaluation calibration. Experiments show it outperforms training-free and training-based baselines and remains robust in various settings.

0 favorites 0 likes
#llm

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

arXiv cs.AI · 6h ago Cached

This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.

0 favorites 0 likes
#llm

EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine

arXiv cs.CL · 6h ago Cached

EviStreams is an open-source, no-code web platform that uses human-in-the-loop AI for data extraction in systematic reviews, demonstrating that extraction quality is more dependent on field specification than model choice.

0 favorites 0 likes
#llm

MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression

arXiv cs.CL · 6h ago Cached

This paper introduces MORSE, a compression-aware method for evidence-preserving context ordering in large language models that improves evidence retention and downstream QA performance.

0 favorites 0 likes
#llm

Planned Test-Time Scaling with Coordinated Reasoning Paths

arXiv cs.CL · 6h ago Cached

This paper introduces Planned Test-Time Scaling (PTTS), a method that coordinates reasoning branches to enhance performance on challenging tasks, achieving significant gains over repeated sampling in mathematical reasoning benchmarks.

0 favorites 0 likes
#llm

Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs

arXiv cs.CL · 6h ago Cached

The paper investigates fine-tuning strategies for customer support LLMs, comparing multi-task training, sequential updates, and model merging across multiple model families. It concludes that multi-task full fine-tuning is the most robust default, while specialist models degrade off-task and require reliable routing.

0 favorites 0 likes
#llm

Provably Complete Generalized Planning with LLMs

arXiv cs.AI · 6h ago Cached

This paper introduces an approach for automatically generating generalized plans in Lean with formal proofs of completeness using LLMs, evaluated on benchmark domains showing significant advancements in automatic plan verification.

0 favorites 0 likes
#llm

LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning

arXiv cs.CL · 6h ago Cached

LEGO is a dual-module framework that synergizes Expert GraphRAG and Expert Chain-of-Thought to enhance complex legal reasoning in large language models, achieving superior performance on benchmarks like LawExamQA_Civil.

0 favorites 0 likes
#llm

Mercury 2.5 LLM hits 770 tokens per second

Hacker News Top · 11h ago Cached

The Mercury 2.5 LLM achieves a speed of 770 tokens per second, as evaluated by Artificial Analysis through various intelligence benchmarks and capability indexes.

0 favorites 0 likes
#llm

@paulg: A CS undergrad asked me where he could have most effect in the AI age. I said probably at either extreme: either close …

X AI KOLs Timeline · yesterday

Paul Graham advises a CS undergrad that in the AI age, one can have the most impact by either developing LLMs or using AI to directly serve customer needs.

0 favorites 0 likes
#llm

@akshay_pachaar: Redis built a cache that cuts LLM costs by 70%! Production LLM apps often receive different versions of the same questi…

X AI KOLs Timeline · yesterday Cached

Redis LangCache is a semantic caching tool that reduces LLM costs by up to 70% by storing and reusing similar question-response pairs, making AI applications faster and more cost-effective.

0 favorites 0 likes
#llm

Some people seems to not get this: LLMs do not need to be the ASI themselves

Reddit r/singularity · yesterday

The article argues that LLMs do not need to become ASI themselves; they can research and develop new AI models that might not be LLMs, leading to ASI through narrow super-intelligence.

0 favorites 0 likes
#llm

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

arXiv cs.LG · yesterday Cached

This paper provides a mechanistic comparison of knowledge-conflict circuits in LLMs under instruction tuning, finding that tuning gates rather than rewires these circuits across multiple model families, with implications for interpretability.

0 favorites 0 likes
#llm

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

arXiv cs.LG · yesterday Cached

The study reveals that chat templates control whether language models adopt a disclaimer voice (e.g., 'I'm just an AI') or an experiential voice (e.g., 'I feel'), and identifies an activation direction that can steer this behavior, impacting AI safety research.

0 favorites 0 likes
#llm

Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees

arXiv cs.AI · yesterday Cached

This research presents a Structured Knowledge Tree architecture to control narrative reliability and epistemic pacing in LLM-driven detective games, reducing hallucinations by 64.78% and preventing premature information disclosure.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback