Tag
This paper presents ZooClaw-FashionSigLIP2, a fashion-specialized vision-language model that achieves superior retrieval performance through full fine-tuning with knowledge distillation and weight interpolation, outperforming larger backbones and LoRA. It also introduces a new high-quality benchmark, ZooClaw-Fashion, and analyzes structural biases in existing datasets.