Tag
SharpMoE is a post-training framework that improves routing in diffusion mixture-of-experts models by using clean latent features to identify salient tokens and a trajectory routing loss to allocate compute precisely, achieving state-of-the-art visual generation.
AttnGen is an attention-guided training framework that embeds interpretability into the optimization of deep neural networks for genomic sequence classification, achieving improved accuracy and encouraging models to focus on informative nucleotide positions.