large-language-model

Tag

Cards List
#large-language-model

By opus 5.5.

Reddit r/singularity · 3h ago

The title references 'opus 5.5', likely referring to Anthropic's AI model, suggesting a version update or release.

0 favorites 0 likes
#large-language-model

Hi3D launches an AI agent for 3D creation

Reddit r/artificial · yesterday

Hi3D launched Hi3D Agent, an AI agent that combines a large language model with 3D tools to streamline workflows from model generation to final production in game, animation, and design.

0 favorites 0 likes
#large-language-model

yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash

Reddit r/LocalLLaMA · yesterday

Yandex has released AliceAI Foundation 80B A3B Base, a custom AI model competing with Qwen 35B and DeepSeek V4 Flash, featuring its own architecture without post-training.

0 favorites 0 likes
#large-language-model

Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

arXiv cs.LG · 2d ago Cached

This paper presents Jarvis, an open-source voice assistant framework for autonomous racecars that uses a fine-tuned Mistral 7B model for offline intent recognition, achieving 97.63% accuracy with low latency and outperforming larger online models.

0 favorites 0 likes
#large-language-model

DAPO: An Open-Source RL System from ByteDance Seed and Tsinghua Air

Hacker News Top · 2d ago Cached

DAPO is an open-source reinforcement learning system for large language models, developed by ByteDance Seed and Tsinghua AIR, which achieves state-of-the-art performance on the AIME 2024 benchmark with the Qwen2.5-32B model.

0 favorites 0 likes
#large-language-model

@liulicheng10: Proud of what we’ve achieved!

X AI KOLs Following · 3d ago Cached

StepFun introduces Step 5 Preview, a new flagship AI model for agentic work with frontier-level performance in software engineering and finance, featuring a 600B parameter scale.

0 favorites 0 likes
#large-language-model

@RohanNayak2: Introducing Sherpa: the most advanced fiction writing AI We accelerated from $250M in ARR to $500M because Sherpa helpe…

X AI KOLs Timeline · 4d ago Cached

Sherpa is an advanced AI fiction writing tool by Pocket FM that uses specialized narrative models to enhance serialized content creation, boosting production and enabling significant earnings for creators.

0 favorites 0 likes
#large-language-model

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

arXiv cs.AI · 5d ago Cached

This paper empirically investigates the susceptibility of LLM-based GUI agents to digital nudges, finding that reasoning configuration redirects rather than reduces nudge effects, positioning interface design as a governance concern for autonomous AI.

0 favorites 0 likes
#large-language-model

IFM/K2-Horizon-7B-Uno · Hugging Face - 5200tps with no quality loss

Reddit r/LocalLLaMA · 5d ago Cached

The article presents K2-Horizon-7B-Uno, a diffusion-augmented LLM that combines autoregressive and diffusion pathways to achieve 5200 tokens per second throughput without quality loss, with benchmarks showing competitive performance across various tasks.

0 favorites 0 likes
#large-language-model

@hooshaaii: Can LLMs achieve Recursive Self-Improvement (RSI)? The highly-cited STaR paper proved models can bootstrap their own re…

X AI KOLs Timeline · 5d ago Cached

The STaR paper introduces a technique for large language models to bootstrap their own reasoning by iteratively generating rationales, filtering for correctness, and fine-tuning on successful examples.

0 favorites 0 likes
#large-language-model

XingChen-AGI/Xing4.0-29B-A4B MoE

Reddit r/LocalLLaMA · 6d ago

Xing4.0-29B-A4B is a next-generation MoE large language model developed by China Telecom, featuring 29B total parameters with 4B active per token, native support for 256K context length, and optimization for Ascend NPU with agent-oriented architecture for complex engineering tasks.

0 favorites 0 likes
#large-language-model

1.3B token used recently, qwen3.8 27b

Reddit r/LocalLLaMA · 6d ago

A Reddit post highlights the strong performance of the Qwen 3.8 27B model, with mention of 1.3B tokens used recently.

0 favorites 0 likes
#large-language-model

@Kay2289123: That ghost story about AI development slowing down—read it, have a laugh, and move on; don't fool even yourself. Grok4.…

X AI KOLs Timeline · 2026-09-16 Cached

The article refutes the idea that AI development is slowing down, highlighting that Grok4.8, a 2.5T parameter model, has finished training and that major tech companies are still actively training AI models.

0 favorites 0 likes
#large-language-model

XingChen-AGI/Xing4.0-29B-A4B

Hugging Face Models Trending · 2026-09-16 Cached

Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.

0 favorites 0 likes
#large-language-model

LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems

arXiv cs.AI · 2026-09-15 Cached

This paper proposes an LLM-enhanced multi-agent reinforcement learning framework to simultaneously optimize electric vehicle charging scheduling, station profitability, and grid stability, using LLMs for feature selection and adaptive weighting, outperforming state-of-the-art methods with reduced training time.

0 favorites 0 likes
#large-language-model

A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies

arXiv cs.AI · 2026-09-12 Cached

This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms that uses large language models to enable natural language device control, focusing on architecture design and key technologies like KV Cache optimization and progressive prompt disclosure to reduce latency in sterile environments.

0 favorites 0 likes
#large-language-model

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

arXiv cs.AI · 2026-09-11 Cached

OntologyAligner presents a three-stage framework combining ontology-aligned retrieval and LLM reranking for biomedical ontology normalization, achieving state-of-the-art performance on HPO tasks and introducing the PhenoNormBench benchmark.

0 favorites 0 likes
#large-language-model

A Dynamic Fusion Large Language Model for Traffic Flow Prediction

arXiv cs.LG · 2026-09-11 Cached

The paper proposes a Dynamic Fusion Large Language Model (DF-LLM) for traffic flow prediction, integrating spatiotemporal embedding, fusion modules, and an LLM backbone with adaptation strategies to enhance performance over traditional methods in intelligent transportation systems.

0 favorites 0 likes
#large-language-model

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

arXiv cs.LG · 2026-09-11 Cached

REVA is a framework that mines historical attention traces to create reusable evidence views for context-efficient RAG serving, improving generation quality while reducing compression overhead and latency.

0 favorites 0 likes
#large-language-model

Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

Hacker News Top · 2026-09-10 Cached

Cognition's SWE-2 model, post-trained from Kimi K3, achieves a score of 92.8 on Terminal-Bench 2.1, offering competitive performance with lower cost compared to frontier models like Fable 5.1 and GPT-6 Astra.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback