triton-kernels

Tag

Cards List
#triton-kernels

Block Sparse Attention with Log-Linear Complexity

Hugging Face Daily Papers ↗ · 5d ago Cached

The paper proposes PISA, a block sparse attention mechanism using a pyramid Top-K selection strategy to achieve O(N log N) complexity, enhancing efficiency for long-context language models with comparable performance on benchmarks and better results on retrieval tasks.

0 favorites 0 likes
#triton-kernels

@PyTorch: AMD has been upstreaming optimizations for improved FP8 training support in PyTorch/TorchTitan and PyTorch/TorchAO, mak…

X AI KOLs Timeline ↗ · 2026-08-13 Cached

AMD upstreamed optimizations to PyTorch/TorchTitan and TorchAO for FP8 training on AMD Instinct GPUs, achieving up to 13.4% throughput gains on Llama3-8B and recovering 89% of FP8 quantization overhead on DeepSeek-V3 via fused Triton kernels.

0 favorites 0 likes
#triton-kernels

I released a softmax-free attention model at GPT-2 Medium scale (~354M params, 11.5B tokens): structural sparsity + tile-skipping kernels for long-context VRAM savings. Open weights + custom Triton kernels [R]

Reddit r/MachineLearning ↗ · 2026-06-21 Cached

Released RRT-355M, a softmax-free attention model at GPT-2 Medium scale with 354M parameters trained from scratch on 11.5B tokens, using structural sparsity and tile-skipping kernels for long-context efficiency, achieving comparable performance to GPT-2 Medium on a 22-task benchmark.

0 favorites 0 likes
#triton-kernels

@raphaelsrty: Computing max similarity (scoring step of colbert, colpali) on gpus can be optimized and this is what @tonywu_71 did. I…

X AI KOLs Following ↗ · 2026-06-10 Cached

Tony Wu released late-interaction-kernels (LIK): fused Triton kernels for MaxSim, the scoring step behind ColBERT and ColPali, integrated into PyLate and colpali-engine, offering memory efficiency and performance gains.

0 favorites 0 likes
#triton-kernels

Wall Attention (GitHub Repo)

TLDR AI ↗ · 2026-06-03 Cached

Wall Attention is a new attention variant with per-channel, per-timestep multiplicative decay, providing content-dependent forgetting rates and efficient training/decode kernels implemented in Triton.

0 favorites 0 likes
#triton-kernels

@akshay_pachaar: PyTorch Autograd vs. Unsloth Triton Kernels. The core engineering behind UnslothAI has always been impressive! Instead …

X AI KOLs Following ↗ · 2026-04-20 Cached

Technical explanation comparing PyTorch's default autograd with UnslothAI's custom backpropagation kernels written in OpenAI's Triton language for faster LLM fine-tuning.

0 favorites 0 likes
#triton-kernels

An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU

Papers with Code Trending ↗ · 2026-03-17 Cached

SlideFormer introduces a heterogeneous co-design for full-parameter LLM fine-tuning on a single GPU, leveraging GPU/CPU/RAM/NVMe with a layer-sliding engine and optimized Triton kernels, enabling fine-tuning of 123B+ models on a single RTX 4090 with significant throughput improvements.

0 favorites 0 likes
← Back to home

Submit Feedback