@ArtificialAnlys: Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audi…

X AI KOLs Timeline Models

Summary

Wan 3.0 is Alibaba's new all-in-one video generation and editing model that debuts at #1 on the Artificial Analysis Video Editing Leaderboard, featuring native audio and multimodal inputs, available in public preview via Alibaba Cloud.

Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audio Wan 3.0 is Alibaba's new all-in-one video generation and editing model, positioned as a single system for turning multimodal creative direction into video. It generates up to 30 seconds at 1080p with native audio and accepts text, images, video, audio, documents, and web pages as creative references. The same model supports Text to Video, Image to Video, reference-based generation, and instruction-led editing, including changes to visuals, plot, dialogue, and sound. In the Artificial Analysis Video Arena, Wan 3.0 ranks #1 in Video Editing with Audio, #2 in Text to Video with Audio, and #5 in Image to Video with Audio. Wan 3.0 marks a large generational improvement: against the most recent Wan 2.7 version on each leaderboard, it rises from #5 to #1 in Video Editing with Audio, #6 to #2 in Text to Video with Audio, and #12 to #5 in Image to Video with Audio. Wan 3.0 is available now in public preview through Alibaba Cloud Model Studio. Pricing starts at $0.05 per second for 480p, increasing to $0.10 for 720p and $0.20 for 1080p. Congratulations to @Alibaba_Wan and @alibaba_cloud on the release! See below for comparisons between Wan 3.0 and other leading models in the Artificial Analysis Video Arena 🧵
Original Article
View Cached Full Text

Cached at: 09/03/26, 04:08 AM

Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audio

Wan 3.0 is Alibaba’s new all-in-one video generation and editing model, positioned as a single system for turning multimodal creative direction into video. It generates up to 30 seconds at 1080p with native audio and accepts text, images, video, audio, documents, and web pages as creative references. The same model supports Text to Video, Image to Video, reference-based generation, and instruction-led editing, including changes to visuals, plot, dialogue, and sound.

In the Artificial Analysis Video Arena, Wan 3.0 ranks #1 in Video Editing with Audio, #2 in Text to Video with Audio, and #5 in Image to Video with Audio.

Wan 3.0 marks a large generational improvement: against the most recent Wan 2.7 version on each leaderboard, it rises from #5 to #1 in Video Editing with Audio, #6 to #2 in Text to Video with Audio, and #12 to #5 in Image to Video with Audio.

Wan 3.0 is available now in public preview through Alibaba Cloud Model Studio. Pricing starts at $0.05 per second for 480p, increasing to $0.10 for 720p and $0.20 for 1080p.

Congratulations to @Alibaba_Wan and @alibaba_cloud on the release!

See below for comparisons between Wan 3.0 and other leading models in the Artificial Analysis Video Arena 🧵

Text to Video (With Audio) Prompt: A cat staring at its own reflection in a toaster, paw tapping the chrome surface. The distorted cat reflection taps back. Audio: Paw taps, confused meow.

Image to Video (With Audio) Prompt: Dragon spreads its wings, lifts off, and flies across the sky with powerful wingbeats. Tail trails behind as it soars into the distance. Heavy wings flapping, deep roar, and rushing wind.

Video Editing (With Audio) Prompt: Restyle the cat’s fisheye POV into a hand-drawn cartoon look while keeping the first-person motion.

Check out Wan 3.0 on the Artificial Analysis Video Leaderboards:

Text to Video: https://artificialanalysis.ai/video/leaderboard/text-to-video…

Image to Video: https://artificialanalysis.ai/video/leaderboard/image-to-video…

Video Editing: https://artificialanalysis.ai/video/leaderboard/video-editing…

Or vote in the Video Arena: https://artificialanalysis.ai/video/arena

Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities

Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62)

Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score

Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release!

Key Takeaways:

➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh)

➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta’s unchanged 1.25/4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and 0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 (0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available

➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh)

➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh)

Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: 1.25/4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta’s first-party API and Muse Code

Similar Articles