Tag
LensVLM-9B is a 9B-parameter Vision Language Model from Apple that scans compressed images of text and selectively expands relevant pages using learned tools, with paper, code, and usage instructions provided.