Google Gemini can now transcribe Sign Language via video
Summary
Google Gemini now supports sign language transcription from video, a notable accessibility advancement for AI models.
Similar Articles
Gemini
Google's Gemini AI model represents a significant advancement in multimodal AI capabilities.
Google announces Gemini 3.5 Live Translate for instant voice-to-voice translation
Google announces Gemini 3.5 Live Translate, a speech-to-speech model that provides instant voice translation in over 70 languages, rolling out across Google ecosystem.
@_philschmid: Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs! 5 modalities in a …
Google releases Gemini Embedding 2 for general availability, offering a single model that embeds text, images, video, audio, and PDFs into one unified space across 100+ languages without needing audio transcription.
Improved Gemini audio models for powerful voice experiences
Google has updated Gemini 2.5 Flash Native Audio to improve live voice agent capabilities, including sharper function calling, better instruction following, and smoother conversation context retrieval. The update also introduces live speech translation in the Google Translate app beta, preserving intonation across 70+ languages.
Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start
Google announces Gemini Omni, a family of multimodal models that can generate video from images, audio, and text, reasoning across inputs to produce consistent, high-quality outputs. The first model, Gemini Omni Flash, rolls out at Google I/O to the Gemini app, YouTube Shorts, and Flow.