Alibaba's Wan 3.0 Enters Public Beta, Pushing Boundaries in AI Video Generation

Alibaba has launched public testing for Wan 3.0, its latest AI model for video generation. This move signals a shift from producing short clips to enabling more substantial and practical creative applications.

Key Advancement: 30-Second Coherence and Intent Understanding

The most noticeable upgrade is in output length. Wan 3.0 can now generate a coherent 30-second video in a single pass, allowing for the depiction of brief narratives or detailed processes. The model focuses on comprehending and fully expressing the user's creative intent, ensuring thematic and logical consistency throughout the clip.

A New Creative Gateway: From Documents to Video

Beyond standard inputs like text, images, audio, and video, Wan 3.0 introduces a groundbreaking feature: direct document support. Users can now feed the model with files such as Word documents, Excel spreadsheets, PowerPoint presentations, PDFs, and Markdown notes. The AI parses the key information and structure from these documents to construct a video narrative. This significantly lowers the technical barrier for video production, opening new avenues for education, business reporting, and knowledge sharing.

Enhanced Realism: Evolution in Characters and Scenes

In terms of visual quality, Wan 3.0 prioritizes authenticity and credibility. The model generates diverse, non-repetitive human characters, addressing a common flaw in earlier AI video systems. For scenes, it strives to accurately replicate real-world physical details and lighting, aiming for every frame to appear plausible and professionally crafted.

The public beta provides developers and creators hands-on access to test these capabilities. The model's real-world performance and its potential to unlock new creative workflows are now open for broader exploration and validation.