Tag
Stowaway is a real-time web experience that lets you pick any aircraft or satellite overhead and view a simulated window-seat perspective from it, rendered over live terrain with WebGL.
VibeVoice 1.5B runs locally on an iPhone with ~2.2 GB memory and up to 1.28× real-time speed; the author plans to release the xcframework and code for audio.cpp.
This arXiv paper presents a unified LLMOps architecture for real-time, enterprise-ready LLM deployments, integrating data ingestion, continual learning, RAG, and feedback loops. It introduces components like AIPO, STAR+FAR, and SAGE to address knowledge staleness, hallucination, and latency-cost trade-offs in regulated sectors.
NVIDIA open-sourced PersonaPlex 7B, a real-time conversational model that listens and speaks simultaneously, handling natural interruptions and overlaps unlike most voice models.
JoyAI-Video-Edit is a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing, achieving 720p editing at ~30 FPS on a single NVIDIA B200 GPU.
ssh.place is a collaborative drawing canvas that anyone can draw on over SSH, with no account or install required, featuring a shared 200x60 grid and a 15-second cooldown per placement.
A 2010 guide to getting started with Google Wave, covering how to create a Wave, add participants, reply, embed content, use extensions, and replay history.
Pathway's open-source llm-app is a framework for building enterprise-grade RAG systems. It supports real-time data sync, built-in vector retrieval, and comes with ready-made cloud templates. It has earned over 59,000 stars on GitHub.
Discusses 200 milliseconds as a critical threshold for response time, highlighting its role in perceived performance and user experience in tech systems.
Smallest.ai raises $13M in Series A funding to develop small, specialized voice models enabling real-time, human-like conversation for AI agents, aiming to make voice interactions indistinguishable from human speech.
The Gemini Live API enables low-latency, real-time voice and vision interactions with Gemini, supporting continuous streams of audio, images, and text for building natural conversational agents.
Arize Phoenix announces customizable visualizations for agent traces, enabling real-time tracking of cache hits, online eval degradations, and tool call errors in production.
OpenAI and Microsoft released GPT-transcribe and GPT-live-transcribe in Microsoft Foundry, offering high-accuracy asynchronous transcription and low-latency streaming transcription for recorded and live audio, with features like background noise handling, accent robustness, and alphanumeric perception.
Fish Audio has made its S2.1 Pro voice cloning service free for a month, featuring 10-15 second voice cloning, ~90ms response time, support for 83 languages, word-level control, and open-weight models at 1/6th the cost of ElevenLabs.
Vivix.AI introduces Vivix-A1, a real-time interactive multimodal model that enables AI characters to move, manipulate objects, and respond with low latency, expanding beyond static talking faces.
A tutorial on building a real-time object detection and tracking pipeline for robotics using ROS 2 and YOLOv11, covering threaded inference, ByteTrack integration, confidence validation, and ONNX export for edge deployment.
Uvilox AI introduces a system that translates American Sign Language into real-time emergency alerts, improving accessibility for deaf individuals in emergencies.
Omnigent introduces a queue and steer feature for agent runs, allowing users to line up follow-up messages while the agent works, and edit, reorder, delete, or steer queued messages without stopping the agent.
LeapTalk proposes a novel framework that overcomes the latency-quality trade-off in talking head generation via single-step bridge distillation, enabling stable real-time generation at up to 200 FPS with reduced identity drift.
TurboVLA introduces a new Vision-Language-Action paradigm that directly maps vision and language to action, achieving 97.7% success on LIBERO with only 0.2B parameters and real-time inference at 32 Hz on consumer GPUs, significantly reducing computational cost.