I asked DeepSeek-V4-Flash to work with Muse-Glimmer for Vision ability in PI agent and it produced this

Reddit r/LocalLLaMA News

Summary

A user demonstrates an AI-agent workflow where DeepSeek-V4-Flash teams with the Muse-Glimmer vision model to iteratively build a realistic car-driving HTML canvas animation, using screenshots for visual feedback. In a follow-up run, DeepSeek ditches the vision model and instead uses PIL to inspect images, producing a stunning 'goldenhour' scene in under 10 minutes.

Same old prompt, just appended a TIP in the end: "Write a single HTML file with a full-page canvas and no libraries. Simulate a realistic side-view of a moving car as the main subject. Keep the car visible in the foreground while the background landscape scrolls continuously to create the feeling that the car is driving forward. Use layered scenery for depth: nearby ground, roadside elements, trees, poles, and distant hills or mountains should move at different speeds for a natural parallax effect. Animate the wheels spinning realistically and add subtle body motion so the car feels connected to the road. Let the environment pass smoothly behind it, with repeating but varied scenery that makes the movement feel believable. Use cinematic lighting and a cohesive sky, such as sunset, dusk, or daylight, to enhance atmosphere. The overall motion should feel calm, immersive, and realistic, with a seamless looping animation. TIPS: You don't have vision abilities so don't try it yourself. If you feel in need of vision ability, you can access http://xxx:8080/v1, model id: Muse-Glimmer for help, it will see the picture, and describe it for you." Then the PI agent started spinning, round and round, every round deepseek wrote or modify something, then called google-chrome for a screenshot of the page, then sent it to the muse-glimmer model to check, then modify according to the reply from muse. It took way longer than deepseek alone. After approximately 30~60 minutes(I left for an hour), finally muse felt satisfied, deepseek then stopped and spat out this. *** what impressed me is the original deepseek alone version, it has a "intro scene", that's a fade-in effect: starting from full darkness and gets bright smoothly. this is the deepseek alone version: deepseek_alone.gif and this is the deepseek+muse vision version: deepseek_plus_muse.gif EDIT: It's Non Reproducible So I followed your suggestions and tried one more time. This time, very soon, deepseek said Muse's response is "not reliable" and decided not to please Muse anymore. After ditching the Muse, deepseek suddenly decided to use PIL library to inspect the scene. This time it was really quick, deepseek committed its work in less than 10 minutes. Here's what it gave me: goldenhour.gif I'm really happy that deepseek has discovered new skills for himself (to use PIL to inspect the image), I was about to ask explicitly (inspired by the commenter). And check the animation, it's just astonishing! The light ray from behind the mountain, the golden river, even the filename is "goldenhour.html"!
Original Article

Similar Articles

DeepSeek-V4-Flash-Vision-Exp

Reddit r/LocalLLaMA

DeepSeek-V4-Flash-Vision-Exp is an experimental or updated AI model from DeepSeek focusing on vision capabilities.

DeepSeek-v4-flash-vision-exp

Hacker News Top

The article provides documentation for DeepSeek's vision model 'deepseek-v4-flash-vision-exp', explaining how to use the API to process images with text prompts via methods like base64 encoding, URLs, or file references.

DeepSeek V4 Flash Vision is now live !

Reddit r/ArtificialInteligence

DeepSeek has released vision capabilities for its V4 Flash AI model, providing a cheaper inference option through DeepInfra compared to the official API.

DeepSeek V4 Flash 0731

Hacker News Top

DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.