transformer

Tag

Cards List
#transformer

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

arXiv cs.LG ↗ · 7h ago Cached

This paper proposes PolicyAttention, a method using causal softmax attention to implement policy mirror descent for closed-loop control in reinforcement learning, achieving lower losses than existing adaptations in experiments.

0 favorites 0 likes
#transformer

T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation

arXiv cs.AI ↗ · yesterday Cached

The paper introduces T-RoPE, a time-aware modification to Rotary Position Embedding for sequential recommendation systems, demonstrating significant performance improvements on benchmarks and real-world impact in e-commerce.

0 favorites 0 likes
#transformer

Naive-N0.5-Flash - 309B-A15.5B

Reddit r/LocalLLaMA ↗ · yesterday Cached

Naive-N0.5-Flash is an open-weight 309B Mixture-of-Experts AI model with 15.5B active parameters, optimized for coding and AI research, featuring a native 1M-token context window and high inference speeds.

0 favorites 0 likes
#transformer

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Hugging Face Daily Papers ↗ · 2d ago Cached

YuE2 is a unified AI model for symbolic and audio music generation that uses symbolic planning to achieve high-quality results, outperforming previous baselines and competing with proprietary systems like Suno.

0 favorites 0 likes
#transformer

@KirkDBorne: AI and ML Papers Explained (HUGE list): https://github.com/dair-ai/ML-Papers-Explained… Compiled by @omarsar0 @dair_ai …

X AI KOLs Timeline ↗ · 3d ago Cached

A tweet sharing a GitHub repository that compiles explanations for key machine learning papers, such as Transformer, BERT, and GPT, for educational purposes.

0 favorites 0 likes
#transformer

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

arXiv cs.LG ↗ · 2026-09-22 Cached

The paper proposes CS-MoE, a novel Transformer architecture that shares neural experts across layers to improve parameter utilization, achieving lower perplexity with only 55% of parameters activated.

0 favorites 0 likes
#transformer

DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction

arXiv cs.LG ↗ · 2026-09-22 Cached

The paper presents DPTM-DT, a dual-pretrained Transformer framework for multitask drug-target prediction that combines molecular graph embeddings, protein language-model embeddings, and physicochemical descriptors, achieving state-of-the-art performance on benchmark datasets.

0 favorites 0 likes
#transformer

A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification

arXiv cs.CL ↗ · 2026-09-22 Cached

This paper introduces a multi-agent system called CB-MAS for document sensitivity classification, which addresses the limitation of fixed input length in transformer models like BERT by using iterative consultation and channel boosting.

0 favorites 0 likes
#transformer

@_michael_yu_: I think Stanford's CS336 is one of the best materials to learn LLMs from scratch. I just finished working through it (t…

X AI KOLs Timeline ↗ · 2026-09-22 Cached

Michael Yu shares his detailed implementation and learnings from Stanford's CS336 course on building LLMs from scratch, covering tokenizers, transformers, training systems, and alignment techniques.

0 favorites 0 likes
#transformer

Transformers Explained Visually

Hacker News Top ↗ · 2026-09-21 Cached

An interactive tool that visually explains the architecture of Transformer models, using GPT-2 to illustrate key components like embedding, attention mechanisms, and output predictions.

0 favorites 0 likes
#transformer

16GB (and in many cases 12GB) is the max vram most people will ever reasonably have

Reddit r/LocalLLaMA ↗ · 2026-09-21

The article discusses how 16GB of VRAM is the realistic high-end limit for most users due to financial constraints, but recent AI model improvements like Qwen 27B quants are enabling more capabilities on such hardware, with hopes for future architectural innovations.

0 favorites 0 likes
#transformer

A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters

arXiv cs.LG ↗ · 2026-09-21 Cached

This paper introduces a lightweight pre-encoder gate mechanism for Transformer-based time-series forecasting models to regulate covariate admission, with experiments showing competitive performance against baselines on multiple datasets.

0 favorites 0 likes
#transformer

MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

arXiv cs.LG ↗ · 2026-09-21 Cached

MOSAIC-SR introduces a Transformer-guided approach for symbolic regression, using initial sketches to improve equation recovery and predictive accuracy on scientific datasets.

0 favorites 0 likes
#transformer

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

arXiv cs.LG ↗ · 2026-09-21 Cached

A comparative study of temporal deep learning architectures for physiological emotion recognition using multimodal wearable datasets, evaluating LSTM, TCN, and Transformer models under different sensing configurations.

0 favorites 0 likes
#transformer

Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions

arXiv cs.CL ↗ · 2026-09-21 Cached

Reviser is a novel decoder-only Transformer model that enables revision-capable text generation via autoregressive cursor actions, achieving competitive performance with lower inference compute compared to baselines.

0 favorites 0 likes
#transformer

World Models From Scratch 2: Model Training and Dreaming [P]

Reddit r/MachineLearning ↗ · 2026-09-19 Cached

This article details building a world model for Super Mario Land using a transformer to predict next game frames from tokenized inputs, and demonstrates autoregressive 'dreaming' to generate future frames.

0 favorites 0 likes
#transformer

@antiAIvo: When learning Transformer The hardest part to handle is multi-dimensional matrix operations The human brain can only mo…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

The article introduces a visualization tool for the attention mechanism in Transformers, allowing users to customize input matrices and see step-by-step computations of attention scores. It is an initial version with plans for future updates based on user feedback.

0 favorites 0 likes
#transformer

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

Hugging Face Daily Papers ↗ · 2026-09-19 Cached

The paper introduces Neural Spectral Capacity (NSC), a training-free metric based on the singular-value spectrum to evaluate and optimize neural network architectures, with a dynamic programming method for optimal design under constraints.

0 favorites 0 likes
#transformer

@Sumanth_077: Train your own LLM from scratch! A step-by-step repo that walks you through building and training a transformer model f…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

A step-by-step repository guiding users to build and train a transformer model from scratch using PyTorch, with comprehensive coverage of data processing, training, and post-training techniques like SFT and RLHF.

0 favorites 0 likes
#transformer

YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers

arXiv cs.CL ↗ · 2026-09-18 Cached

This paper presents the YNU-HPCC team's approach for SemEval-2025 Task 11 on text-based emotion recognition, using a RoBERTa model with enhanced output headers and achieving a ranking score of 0.44 through English-translated datasets.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback