Tag
MOSS-VL is an open vision-language model family designed for real-time interaction, using gated cross-attention to process vision during generation and achieving top performance in streaming benchmarks among open-source models.