large-language-models

Tag

Cards List
#large-language-models

Natural Language Processing Psychometrics

arXiv cs.CL · 5h ago Cached

This paper introduces NLP Psychometrics, a framework that treats psychological prediction from text as a psychometric problem. Using LLM personas, emotional profiles, and syntactic-semantic networks with random forest regressors, it explains up to 76% of variance in mental health scores and shows promise and limits of synthetic data for psychometric prediction.

0 favorites 0 likes
#large-language-models

Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models

arXiv cs.CL · 5h ago Cached

This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.

0 favorites 0 likes
#large-language-models

FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding

arXiv cs.CL · 5h ago Cached

FutureBridge introduces a token reranker for collaborative decoding that ranks LLM-SLM candidates based on how well the SLM can continue reasoning from them, improving the Qwen3-1.7B SLM's math accuracy by 35.1% over greedy decoding.

0 favorites 0 likes
#large-language-models

Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation

arXiv cs.LG · 5h ago Cached

A research paper introducing a multi-stage forward-query method to cryptanalytically extract isolated bias-free GLU feed-forward block weights, demonstrating sub-percent recovery accuracy on Qwen, Llama, and Gemma components while noting end-to-end model API attacks remain unsolved.

0 favorites 0 likes
#large-language-models

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

arXiv cs.AI · 5h ago Cached

ReQuant introduces a backpropagation-free, fixed-grid discrete refinement stage for post-training quantization (PTQ) that iteratively improves initial quantized models while preserving the quantized format, showing consistent gains across various LLMs and bit-widths.

0 favorites 0 likes
#large-language-models

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents

arXiv cs.AI · 5h ago Cached

Gated-BEPO is a new credit assignment method for LLM agents that derives step-level credit from empirical rollout graphs using Bellman fixed-point estimation and adaptively fuses it with episode-level credit via a confidence gate. Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements over existing critic-free methods.

0 favorites 0 likes
#large-language-models

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

arXiv cs.AI · 5h ago Cached

Introduces AgentPatch, a training-free coarse-to-fine repair framework for merging agentic multimodal large language models, addressing asymmetric capability preservation and behavior-critical forgetting.

0 favorites 0 likes
#large-language-models

Data poisoning and RAG manipulation

Reddit r/ArtificialInteligence · 12h ago

A discussion of how data poisoning and RAG manipulation pose a silent, dangerous threat to AI systems, arguing that security must extend beyond input filtering to memory, data pipelines, and multi-agent logic.

0 favorites 0 likes
#large-language-models

@antpalkin: A finance professor manages $200M with AI agents, and he told everyone why: "Large language models are at the level of …

X AI KOLs Timeline · 2d ago Cached

Promotional post about Horizon, an AI-driven trading strategy platform that lets users backtest and deploy strategies in plain English, citing a professor managing $200M with AI agents and returning 56% last year.

0 favorites 0 likes
#large-language-models

Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies

arXiv cs.CL · 3d ago Cached

This paper surveys clinical communication processing using LLM-generated synthetic data and presents 13 case studies across EMS reports, nurse handoffs, and more, showing that synthetic data can bootstrap clinical NLP systems.

0 favorites 0 likes
#large-language-models

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

arXiv cs.LG · 3d ago Cached

Introduces GROM, a gradient-free one-shot machine unlearning method that computes a closed-form additive weight update via ridge-regularized least squares, achieving state-of-the-art forgetting-utility trade-offs on benchmarks like TOFU and WMDP, and resisting quantization-based recovery attacks.

0 favorites 0 likes
#large-language-models

Human-Like Anaphor Resolution in Large Language Models

arXiv cs.CL · 3d ago Cached

This paper investigates whether five open-weight LLMs exhibit human-like sensitivity to psycholinguistic factors in anaphor resolution, using surprisal and comprehension accuracy as behavioral measures. Results show selective cognitive alignment, with some models matching human discourse sensitivity but not semantic interference effects.

0 favorites 0 likes
#large-language-models

Example-Guided Prompting for Document-Level Text Simplification

arXiv cs.CL · 3d ago Cached

This paper investigates using retrieved document-simplification examples to guide LLM prompting for document-level text simplification, showing improvements over prompt-only generation on the OneStopEnglish corpus.

0 favorites 0 likes
#large-language-models

Safe Evolution with Circuit Anchors

arXiv cs.CL · 3d ago Cached

This paper proposes Circuit-Anchored Evolution (CAE), a method that uses mechanistic interpretability to identify and anchor a tiny safety circuit in LLMs during self-evolution, preventing models from misevolving into capable but dangerous systems while preserving capability.

0 favorites 0 likes
#large-language-models

Disentangling 3D Modeling from Spatial Reasoning

arXiv cs.LG · 3d ago Cached

This paper proposes DiSR, a framework that separates 3D perception from reasoning by using off-the-shelf perception models to reconstruct explicit 3D evidence and fine-tuning an LLM with LoRA for spatial reasoning, achieving competitive performance with improved interpretability and efficiency.

0 favorites 0 likes
#large-language-models

Counterfactual Analysis via Large Language Models

arXiv cs.AI · 3d ago Cached

This paper investigates using GPT-3.5 for counterfactual analysis in online lending, showing that prompt engineering improves prediction accuracy and enables coherent counterfactual ROI generation under alternative interest rates.

0 favorites 0 likes
#large-language-models

OpenAI's latest math breakthroughs commit research misconduct, experts say

Reddit r/ArtificialInteligence · 3d ago Cached

OpenAI released AI-generated math breakthroughs that experts are calling research misconduct due to lack of academic rigor.

0 favorites 0 likes
#large-language-models

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

arXiv cs.CL · 4d ago Cached

Presents RESPClinBench, a real-world scenario benchmark for respiratory clinical decision-making, evaluating seven LLMs on COPD and pulmonary nodule cases. Finds task-specific limitations including imaging hallucination and medication-safety risks.

0 favorites 0 likes
#large-language-models

Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation

arXiv cs.CL · 4d ago Cached

A PhD proposal outlining a unified end-to-end framework for multilingual metaphor processing, integrating metaphor detection, translation evaluation, and joint modeling using linguistic theory and large language models.

0 favorites 0 likes
#large-language-models

TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering

arXiv cs.AI · 4d ago Cached

The paper presents TourSynbio-Search, an LLM-driven agent framework for unified protein engineering search across literature and biological databases, powered by the TourSynbio-7B multimodal model with dual PaperSearch and ProteinSearch components.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback