Tag
This paper presents SkinAgent AI, a multimodal agentic framework for non-diagnostic skincare support that combines visual analysis with safe, auditable orchestration using large language models.
GLM-5.3 outperforms Space Bunny Alpha in a 5-scene physics test conducted by AI/ML API, highlighting strengths and weaknesses in AI physics simulation.
The paper introduces a two-model architecture called Summarize-Judge-Refine (SJR) for multimodal content moderation, which decouples content understanding and policy learning via natural language summaries, enabling significant performance gains and few-shot policy adaptation.
This paper investigates offline multimodal large language models as decision support tools for air operations, detailing a modular retrieval-augmented architecture and presenting a pilot study with the Brazilian Air Force that demonstrates reduced cognitive workload and improved efficiency in doctrinal assessment tasks.
This paper proposes ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection that achieves state-of-the-art performance by decoupling explanation and detection tasks.
TrioRAG is a graph-free multimodal RAG framework that uses multi-signal late fusion for efficient cross-document question answering, matching or outperforming graph-based systems. It also introduces AutoQA, an automotive benchmark with web-sourced images.
The paper proposes NarraLite, an efficient multimodal generative recommendation framework that uses latent narrative reasoning to improve episodic content prediction with better accuracy and efficiency.
SHIFT-M3 is a lightweight pre-fusion screen that measures alignment consistency between LLM-generated and clinical report summaries of ECG records, achieving high accuracy in detecting data integrity issues.
Sierra launches multimodal AI agents that dynamically switch between voice, text, and visuals to enhance customer experience.
Elon Musk tweets that Grok 5 may be better than any existing AI model, comparing Grok versions to other models like Opus 5.0 and highlighting improvements in upcoming versions.
This paper proposes an exploration-guided prompt scaffolding framework that dynamically adjusts training prompts for multimodal reinforcement learning, achieving up to 9.7% relative improvement in performance on benchmarks.
This paper introduces a synthetic benchmark to evaluate information-theoretic metrics for assessing the informativeness of text annotations in multimodal time series forecasting, demonstrating their utility for annotation auditing without model training.
Abhinav Anand, a 20-year-old independent builder from Bihar, released Arcle V1, an open-weight 5.84B-parameter unified omni AI model that outperforms Apple's AFM 3B model on multiple benchmarks including MATH-500.
SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, and expert optimization.
A researcher shares preliminary results demonstrating a method that reduces image-processing token usage by approximately 95% compared to GPT-4o while maintaining similar accuracy, and seeks feedback on its significance.
The paper introduces ConflictGUI, a benchmark for conflict-aware termination in GUI agents, and proposes ConflictGuard, an inference-time framework to reduce over-compliance and improve performance on conflicting instructions.
Nunchux AI has launched Modelverse, a multimodal generative AI inference service providing fast, affordable access to over 30 image, video, and avatar models through a single API.
The PyTorch Conference North America will be held in San Jose, featuring sessions on agentic search, hardware-guided workflows, and AI performance optimization with speakers from major tech companies.
PhoenixNest-Video introduces an evidence-grounded multimodal agent framework for automated video interview assessment, achieving 91.50% grade-level accuracy on a benchmark by using structured video graphs and reinforcement learning.
A case study demonstrates an AI classroom using fal's H3 Max to generate real-time video explanations based on user questions, merging interactive tutoring with multimodal content creation.