Tag
This paper proposes PolicyAttention, a method using causal softmax attention to implement policy mirror descent for closed-loop control in reinforcement learning, achieving lower losses than existing adaptations in experiments.
The paper introduces T-RoPE, a time-aware modification to Rotary Position Embedding for sequential recommendation systems, demonstrating significant performance improvements on benchmarks and real-world impact in e-commerce.
Naive-N0.5-Flash is an open-weight 309B Mixture-of-Experts AI model with 15.5B active parameters, optimized for coding and AI research, featuring a native 1M-token context window and high inference speeds.
YuE2 is a unified AI model for symbolic and audio music generation that uses symbolic planning to achieve high-quality results, outperforming previous baselines and competing with proprietary systems like Suno.
A tweet sharing a GitHub repository that compiles explanations for key machine learning papers, such as Transformer, BERT, and GPT, for educational purposes.
The paper proposes CS-MoE, a novel Transformer architecture that shares neural experts across layers to improve parameter utilization, achieving lower perplexity with only 55% of parameters activated.
The paper presents DPTM-DT, a dual-pretrained Transformer framework for multitask drug-target prediction that combines molecular graph embeddings, protein language-model embeddings, and physicochemical descriptors, achieving state-of-the-art performance on benchmark datasets.
This paper introduces a multi-agent system called CB-MAS for document sensitivity classification, which addresses the limitation of fixed input length in transformer models like BERT by using iterative consultation and channel boosting.
Michael Yu shares his detailed implementation and learnings from Stanford's CS336 course on building LLMs from scratch, covering tokenizers, transformers, training systems, and alignment techniques.
An interactive tool that visually explains the architecture of Transformer models, using GPT-2 to illustrate key components like embedding, attention mechanisms, and output predictions.
The article discusses how 16GB of VRAM is the realistic high-end limit for most users due to financial constraints, but recent AI model improvements like Qwen 27B quants are enabling more capabilities on such hardware, with hopes for future architectural innovations.
This paper introduces a lightweight pre-encoder gate mechanism for Transformer-based time-series forecasting models to regulate covariate admission, with experiments showing competitive performance against baselines on multiple datasets.
MOSAIC-SR introduces a Transformer-guided approach for symbolic regression, using initial sketches to improve equation recovery and predictive accuracy on scientific datasets.
A comparative study of temporal deep learning architectures for physiological emotion recognition using multimodal wearable datasets, evaluating LSTM, TCN, and Transformer models under different sensing configurations.
Reviser is a novel decoder-only Transformer model that enables revision-capable text generation via autoregressive cursor actions, achieving competitive performance with lower inference compute compared to baselines.
This article details building a world model for Super Mario Land using a transformer to predict next game frames from tokenized inputs, and demonstrates autoregressive 'dreaming' to generate future frames.
The article introduces a visualization tool for the attention mechanism in Transformers, allowing users to customize input matrices and see step-by-step computations of attention scores. It is an initial version with plans for future updates based on user feedback.
The paper introduces Neural Spectral Capacity (NSC), a training-free metric based on the singular-value spectrum to evaluate and optimize neural network architectures, with a dynamic programming method for optimal design under constraints.
A step-by-step repository guiding users to build and train a transformer model from scratch using PyTorch, with comprehensive coverage of data processing, training, and post-training techniques like SFT and RLHF.
This paper presents the YNU-HPCC team's approach for SemEval-2025 Task 11 on text-based emotion recognition, using a RoBERTa model with enhanced output headers and achieving a ranking score of 0.44 through English-translated datasets.