Tag
GAS introduces a generation-guided training framework that improves visual understanding in multimodal models by using auxiliary generation tasks with no inference overhead, via a decoupled mixture-of-transformers architecture.