llm

Tag

Cards List
#llm

Reflection 70B was released two years ago (September 2024)

Reddit r/LocalLLaMA ↗ · 7h ago

Reflection 70B was announced in September 2024 as an open-source AI model that outperformed GPT-4o, but it was later found to be based on Llama 3.1, highlighting hype in the AI community.

0 favorites 0 likes
#llm

@VraserX: At this price, Sonnet 5.5 is basically useless. $7.60 per task vs $3.26 for GPT-6 Astra Max is insane. Unless Sonnet is…

X AI KOLs Timeline ↗ · 13h ago Cached

A tweet critiques the cost-effectiveness of Sonnet 5.5, highlighting its higher price per task compared to GPT-6 Astra Max and questioning its value unless it offers superior performance.

0 favorites 0 likes
#llm

RAZOR: Pruning Replaceable Experts in LLMs

arXiv cs.LG ↗ · 15h ago Cached

RAZOR is a training-free method for pruning replaceable experts in Mixture-of-Experts LLMs by assessing functional replaceability, achieving superior performance on reasoning tasks compared to existing methods.

0 favorites 0 likes
#llm

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

arXiv cs.LG ↗ · 15h ago Cached

This paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models that uses online observer feedback for state-aware adaptation without additional training, showing improved controllability over fixed methods.

0 favorites 0 likes
#llm

Cost-Aware Best-LLM Identification using Dueling Feedback

arXiv cs.LG ↗ · 15h ago Cached

The paper proposes a cost-aware algorithm for identifying the best large language model using dueling feedback in a multi-armed bandit framework, demonstrating optimal cost and improvements over existing methods.

0 favorites 0 likes
#llm

@manavvnotop: for the longest time i hated anthropic but they cooked with opus 5.5 opus 5.5 on medium is enough as a daily driver and…

X AI KOLs Timeline ↗ · 22h ago Cached

A tweet expresses excitement about Anthropic's Opus 5.5 model, stating it's sufficient for daily use, and references the release of Sonnet 5.5 as a faster, lower-cost option for everyday tasks.

0 favorites 0 likes
#llm

MicroLLM Lab – Try 7 tiny LLM's in the browser

Hacker News Top ↗ · yesterday Cached

MicroLLM Lab is a browser-based tool that allows users to test and benchmark 7 tiny language models, using JavaScript for objective checks and providing speed and accuracy metrics.

0 favorites 0 likes
#llm

What do you think about ToMoE v2 paper, converting dense model to MoE model at near lossless accuracy? I feel Qwen3.8-27B-A16B or something along those lines would be amazing, though there are architectural hurdles, as well as need folr training data.

Reddit r/LocalLLaMA ↗ · yesterday Cached

ToMoE is a method that converts dense large language models into mixture-of-experts models through dynamic structural pruning, achieving strong performance without weight updates and outperforming state-of-the-art pruning and MoE techniques.

0 favorites 0 likes
#llm

How to Solve Hallucination (with RLCD)

Lobsters Hottest ↗ · yesterday Cached

The article explains how to address hallucination in large language models by using probability distributions and calibration to quantify uncertainty, rather than unreliable confidence scores.

0 favorites 0 likes
#llm

@thesupermannx: Netflix replaced their 15 years old recommendation algorithm with an LLM. It’s called "GenRec" and it completely change…

X AI KOLs Timeline ↗ · yesterday Cached

Netflix replaced its long-standing recommendation algorithm with an LLM-backed system called GenRec, which outperforms the previous method using significantly less training data, indicating a major shift in AI-driven product engineering.

0 favorites 0 likes
#llm

@system_monarch: AI engineering interview in 2026. How many of these can you explain clearly: KV cache Prompt caching Semantic caching S…

X AI KOLs Timeline ↗ · yesterday Cached

A tweet lists key AI engineering concepts such as KV cache and speculative decoding, challenging professionals to explain them clearly for interviews in 2026.

0 favorites 0 likes
#llm

It finally hit me

Reddit r/singularity ↗ · yesterday

A top reverse engineer used AI tools like Claude to complete a difficult CTF challenge in hours instead of days, demonstrating AI's rapid advancement in cybersecurity tasks.

0 favorites 0 likes
#llm

Shobr: Job seach CLI via browser automation, event-sourcing, LLM

Lobsters Hottest ↗ · yesterday Cached

Shobr is an open-source CLI tool that automates job searches using browser automation, event-sourcing, and LLM, designed with stealth and human-in-the-loop principles.

0 favorites 0 likes
#llm

Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries

arXiv cs.CL ↗ · yesterday Cached

This paper evaluates sycophancy in Chinese large language models on factual questions derived from search queries, finding that anti-sycophancy prompting reduces belief-aligned errors but increases uncertainty, impacting factual accuracy.

0 favorites 0 likes
#llm

Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos

arXiv cs.CL ↗ · yesterday Cached

This study examines how transcript compression methods affect LLM-based veracity classification for medical misinformation in Japanese YouTube videos, finding that full transcripts outperform compressed inputs in detection accuracy.

0 favorites 0 likes
#llm

Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength

arXiv cs.CL ↗ · yesterday Cached

This paper introduces an LLM-as-a-judge method to measure perturbation strength for assessing self-consistency in LLM explanations, showing that input perturbations generally affect LLMs more strongly than CoT perturbations.

0 favorites 0 likes
#llm

LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents

arXiv cs.AI ↗ · yesterday Cached

The paper identifies 'LLM Parkinsonism' as a problem of inefficient persistence in autonomous LLM agents and proposes an uncertainty-aware Global Executive Control architecture to improve goal success while reducing token usage.

0 favorites 0 likes
#llm

Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs

arXiv cs.CL ↗ · yesterday Cached

The paper formulates persona diversification as a set-level conditioning problem to mitigate homogeneity in LLM outputs, evaluating methods that enhance creativity and diversity across tasks like the Alternative Uses Task.

0 favorites 0 likes
#llm

SignTrace: Describe a Sign, Find the Word

arXiv cs.CL ↗ · yesterday Cached

SignTrace is a system that enables reverse lookup in Chinese sign language dictionaries using large language models, achieving 94.0% Hit@1 accuracy for identifying signs from natural movement descriptions.

0 favorites 0 likes
#llm

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

arXiv cs.AI ↗ · yesterday Cached

This paper explores the groundedness of multi-agent code judges in AI systems, introducing label-free measurements to assess their reliability and demonstrating failures in existing verification frameworks without proper evidence.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback