Tag
This paper describes six submissions to the WMT26 Model Compression Shared Task, using routing-informed expert pruning and MXFP4 quantization to compress GPT-OSS-20B into smaller translation models with parameters ranging from 4.186B to 7.770B.