mxfp4-quantization

Tag

Cards List
#mxfp4-quantization

ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression

arXiv cs.CL · 2d ago Cached

This paper describes six submissions to the WMT26 Model Compression Shared Task, using routing-informed expert pruning and MXFP4 quantization to compress GPT-OSS-20B into smaller translation models with parameters ranging from 4.186B to 7.770B.

0 favorites 0 likes
← Back to home

Submit Feedback