Text or image input
Start from a written scene or animate one or two key frames.
Wan 2.7 on Cinel makes 2–15 s videos from text, first and end frames, or image, audio and video references. 720p or 1080p with voice matching, from 160 credits.
See how a reference frame becomes a moving shot.
Start from a written scene or animate one or two key frames.
Use up to three image references when the scene needs more visual guidance.
Choose 720p or 1080p and set any duration from 2 to 15 seconds.
An engineering and creative breakdown of model performance, architecture, and practical production workflows.
Wan 2.7 is the most controllable model in Alibaba's Tongyi Wanxiang lineup on Cinel. It can start from a prompt, animate between two keyframes you choose, or take guidance from a mix of reference images, a voice clip and a short video. Clips run 2 to 15 seconds at 720p or 1080p. If Wan 2.5 is the everyday clip and Wan 2.6 the storyteller, Wan 2.7 is the director's tool — for when you need the shot to start, look and sound a particular way. On Cinel it starts from 160 credits.
Three capabilities set it apart from its siblings. First, keyframe control: supply a first frame, a last frame or both, and the model works out the motion path between them while keeping the subject stable. Wan 2.6 only accepts a first frame.
Second, multimodal references. Beyond images, Wan 2.7 can take an audio clip and a reference video. On Cinel you can add up to 3 reference images to keep characters and visual details consistent, and the model can match a character's speech to the voice in a 1–10 second sample — useful for recurring hosts or spokescharacters.
Third, finer prompt control: negative prompts tell the model what to avoid (extra fingers, text overlays, a cluttered background), and optional prompt enhancement expands short ideas into fuller briefs. At model level Wan 2.7 also supports instruction-based editing of an existing 2–10 second clip, such as swapping the background for a rainy street without regenerating everything.
| Item | Wan 2.7 |
|---|---|
| Developer | Alibaba — Tongyi Wanxiang team |
| Modes | Text-to-video, image-to-video with first/last frame, reference video; model-level video edit |
| Duration | 2–15 s (Cinel default 5 s) |
| Resolution | 720p / 1080p at 30 fps (1080p ≈ 1.67× per second) |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4 |
| References on Cinel | Up to 3 images · 15 s audio · 15 s video (audio/video member-only) |
| Voice sample | 1–10 s for voice matching |
| Prompt tools | Prompt up to 5,000 characters; negative prompt up to 500; optional prompt enhancement |
| Cinel cost | From 160 credits |
Compared with Wan 2.6 (from 145 credits), Wan 2.7 adds last-frame control, audio and video references, voice matching and negative prompts; choose Wan 2.6 when automatic multi-shot storytelling matters more. Compared with Kling O1 (from 230 credits, 5/10 s only), Wan 2.7 offers first-and-last-frame animation at a lower starting price with any duration from 2 to 15 seconds, plus the option to drive the clip with your own audio track.
Wan 2.7 adds last-frame control, audio and video references, voice matching from a short sample, negative prompts and model-level video editing. Wan 2.6 supports a first frame only.
Upload one image for the opening frame, one for the closing frame, or both. The model infers the movement between them and keeps the subject consistent.
In reference mode the model can match a character's speech to the vocal characteristics of a 1–10 second audio sample. Only use voices you have permission to use.
A negative prompt lists things you do not want, such as subtitles, watermarks or extra people. It helps keep marketing footage clean.
Up to 3 reference images, plus one audio or video reference of up to 15 seconds for members.
It starts from 160 credits on Cinel. Duration, resolution and references change the estimate shown before you generate.