Tag
A new vision-encoder-decoder model, fine-tuned from TrOCR, makes OCR accessible for the Saraiki language.
A new vision-encoder-decoder model is introduced that can process both images and text to generate human-like responses, suitable for tasks like image captioning and visual question answering.