Wan 2.7 AI Video Generator
Wan 2.7 generates 2 to 15 second 30 fps MP4 clips at 720p or 1080p with a native soundtrack and multi-shot narrative in one pass. Alibaba's Wan2.7-Video family also locks first and last frames and up to five mixed image and video references.
What is Wan 2.7
Wan 2.7 is Alibaba's Wan2.7-Video generation family that outputs 2 to 15 second 30 fps MP4 clips at 720p or 1080p with a native soundtrack. On Model Studio it covers text-to-video (wan2.7-t2v), first-and-last-frame image-to-video (wan2.7-i2v), and multi-entity reference-to-video (wan2.7-r2v).
2 to 15 seconds
Issuer T2V and I2V duration is an integer from 2 to 15 seconds. CamArt Standard picker is 5 to 10 seconds. Pro I2V reaches 15.
Native soundtrack
Without an uploaded track the model writes matching music or effects. I2V can take driving audio for lip-sync.
Multi-shot narrative
One prompt can schedule shots, including timestamped Shot 1 blocks, instead of a single locked camera.
Omni refs, five mixed
Reference-to-video takes up to five images and videos combined, addressed as Image 1 and Video 1.
First and last frame
Start-End locks the opening and closing stills on image-to-video.
720p or 1080p
Resolution tiers control pixel count. T2V ratios include 16:9 and 9:16.
Examples of Wan 2.7 from X
Community clips on X made with Wan 2.7. Copy a prompt and try it in the editor.
-
Pets & Animals
Spell Portrait With Cat
PROMPTHyper-realistic cinematic, photorealistic 8K, dramatic fantasy style. Close-up portrait of a beautiful young woman with long dark braided hair, flawless skin, dramatic black winged eyeliner, glossy nude-pink lips, wearing a red leather zip-up jacket. She stands in a dark mystical teal-blue environment with soft volumetric lighting and floating dust particles. A realistic gray tabby cat with bright golden-yellow eyes sits calmly in the lower right foreground, staring at her. The woman slowly raises her left hand, revealing a faint glowing cyan-blue magical orb forming from thin air. The light softly illuminates her face and reflects in the catβs eyes. Subtle camera push-in, shallow depth of field, soft rim lighting, cinematic teal-orange color grading. The cat blinks slowly, unimpressed. Smooth 24fps motion, atmospheric, suspenseful tone. π¬ Shot 2 (15β30s) β Action / Magical Interaction Prompt: Hyper-realistic cinematic 15-second vertical . The cyan-blue magical orb in the womanβs hand intensifies, emitting glowing particles and soft energy trails. She smirks playfully and makes a gentle flicking motion toward the gray tabby cat. In slow motion, magical energy wraps around the cat and lifts it gracefully into the air. The cat follows an elegant floating arc across the frame with glowing particle trails. Ultra-detailed fur physics, soft slow-motion movement, cinematic lighting, dynamic camera tracking the motion smoothly. The woman watches with amusement, slightly tilting her head. Fantasy realism, high detail textures, smooth cinematic movement. π¬ Shot 3 (30β45s) β Humor / Resolution Prompt: Hyper-realistic cinematic 15-second conclusion. The gray tabby cat lands softly and gracefully on the ground, perfectly balanced. The magical glow fades. The cat pauses, looks up at the woman with a completely unimpressed, slightly judgmental expression. After a brief beat, the cat casually turns and walks away toward the right edge of the frame with slow confident steps, tail slightly raised. The womanβs expression shifts from playful confidence to mild confusion and amusement. Subtle camera pan following the cat as it exits. Soft particles fade in the background, cinematic teal-orange grading, shallow depth of field, smooth 24fps motion, humorous fantasy ending, viral meme energy.
-
Cinematic
Twin-Moon Glass Canyon
PROMPTA cinematic 15-second sci-fi fantasy sequence, photorealistic hyper-detailed CGI. [0-4s] Wide establishing shot of a fractured glass canyon under a twin-moon sky, aurora ribbons casting cool indigo and silver light over bioluminescent frost. [4-9s] Three warriors in iridescent scale-mail and ash-dyed cloaks stand before a humming geode altar, flanked by two massive obsidian-furred wolves with silver-tipped tails and eyes glowing like captured starlight. Close-up on the lead warrior, breath pluming in the cold, mouth moving in deliberate speech as he says: "The sky fractures. We answer with fire." [9-15s] He strikes the altar with a reinforced gauntlet; a visible shockwave of golden light ripples outward, warping the air. The wolves snap to attention, fur bristling with static. Energy arcs from the altar into their collars, igniting geometric runes across the warriors' armor. Fluid motion, shallow depth of field, cinematic color grading with deep indigos, molten golds, and stark silver highlights, volumetric aurora lighting, 4K ultra-realistic, dramatic lighting, no text, no subtitles. A cinematic 15-second sci-fi fantasy sequence, photorealistic hyper-detailed CGI. [0-4s] Sudden cut to a ruined crystalline observatory under a blood-crimson sky, floating debris suspended in pockets of warped gravity, ash snow drifting like embers. [4-9s] The warriors and wolves stand back-to-back amid the wreckage, armor flickering with residual kinetic energy. The second warrior turns, eyes locked forward, lips parting as she speaks: "Hold the line. Let the old song wake them." [9-15s] She slams her spear into the fractured ground; a resonant pulse detonates upward. The wolves howl in unison, jaws parting to reveal crystalline fangs, as a blinding silver-blue aura erupts around the pack. Shockwaves lift debris into slow motion, armor plates realign with mechanical precision, faces contorted in fierce resolve. Epic low-angle tracking shot, high contrast, shallow focus on the howling lead wolf, cinematic color grading with crimson shadows, cool metallic tones, and electric blue highlights, professional VFX particle systems, movie trailer pacing, 4K ultra-realistic, dramatic lighting, no text, no subtitles. A cinematic 15-second sci-fi fantasy sequence, photorealistic hyper-detailed CGI. [0-4s] The silver-blue aura collapses inward, pulling ash and floating debris into a slow-spinning vortex above the warriors. The two obsidian wolves leap upward, their bodies unraveling into prismatic light threads that spiral into the storm. [4-9s] The lead warrior lowers his gauntlet, chest rising, lips parting as he speaks calmly: "The sky remembers." The vortex splits open, revealing a massive arch of crystallized starlight bridging the twin moons, casting soft indigo-gold rays across the fractured terrain. [9-15s] Wide pull-back as the warriors stand shoulder-to-shoulder, armor settling into a steady golden hum. The lead warrior turns his head slightly, eyes catching the new light, and finishes: "We walk the dawn." Final frame: the healed glass canyon blooms with bioluminescent flora, the wolves now resting as translucent spectral guardians beside them, light washing over everything in a slow, serene fade. Fluid realistic motion, cinematic color grading with deep violets, radiant golds, and cool silver highlights, volumetric god rays, high contrast, shallow depth of field on final faces, professional VFX light convergence, movie trailer pacing, 4K ultra-realistic, dramatic lighting, no text, no subtitles.
-
Cute & Wholesome
Aura Surge Warrior
PROMPTA young anime warrior standing still as energy begins to surge around him, glowing aura expanding, hair and clothes lifting in the wind, cracks forming on the ground, camera movement: slow circular dolly around the character, slight shake during power surge, lighting: bright glowing aura, dramatic contrast, sparks and particles, style: anime, cinematic, ultra detailed, dynamic lighting, 4K, dramatic scene Made in @openart_ai OpenArt CPP
-
Camera Magic
Neon Motorbike Chase
PROMPTA lone rider on a futuristic motorbike speeding through a neon-lit cyberpunk city at night, rain falling and reflecting on wet streets, flying drones chasing overhead, camera movement: fast tracking shot from behind, slight handheld shake during turns, lighting: neon lights, reflections, high contrast, glowing rain particles, style: cinematic, ultra realistic, motion blur, 4K, anamorphic lens, dynamic action Made in @openart_ai OpenArt CPP
-
FanCam & Sports
FA Cup Broadcast Beat
PROMPTA hyper-realistic 15-second 4K live sports broadcast clip of a dramatic FA Cup match between Manchester City and Chelsea under bright stadium floodlights, packed roaring stadium at night, authentic ESPN+ football coverage aesthetic. Scoreboard overlay reads: βCHE 0 - 0 MNC | 71:00β. Opening shot: high-angle cinematic wide shot of the lush vibrant green pitch as Manchester City players in sky-blue kits surge forward on a dangerous counterattack. Chelsea defenders in dark-blue kits scramble to recover position. Dynamic tracking camera follows the attack with realistic handheld sports-broadcast movement, subtle motion blur, natural stadium lighting, intense crowd atmosphere. Middle sequence: quick fast-paced passing near the penalty area, crowd volume building, player shouts and boot impacts audible. The ball is whipped into the box, chaotic cluster of players fighting for possession. At exactly 7 seconds, the ball smashes into the back of the net. Stadium erupts instantly. The commentator, ORHAN , shouts with raw emotional excitement: βAt last the breakthrough! Manchester City have finally found the net!β Hard cut immediately to a tight emotional close-up of the goal scorer: a white female footballer with curly blonde hair bouncing naturally as she celebrates wildly. She wears a bright sky-blue Manchester City Puma jersey, sweat glistening realistically on her face under stadium lights. She explodes into a huge joyful smile with teeth showing, laughing breathlessly, turning her head side to side in disbelief and pure ecstasy. Eyes shining with emotion. Background: blurred colorful crowd roaring and waving scarves with shallow depth of field, Chelsea defenders disappointed in the distance, goal net still shaking subtly. FA Cup branding and ESPN+ graphics visible throughout. Authentic cinematic sports realism, dramatic zoom-ins, realistic skin texture, vibrant colors, documentary broadcast feel, high contrast lighting, immersive crowd audio, 24fps, premium football television production quality.
-
ASMR & Satisfying
Gothic Mansion Interior
PROMPTA surreal gothic luxury mansion interior, golden dust floating in sunbeams, wide static shot of an opulent room with antique dΓ©cor. A lifelike doll-woman stands motionless on a pedestal, porcelain skin, glassy blue eyes, elegant gown. Extreme close-up of antique clock ticking. Her eye flickers subtly, wrist joint twitches with a faint mechanical click. A man appears as dark silhouette in doorway, slow footsteps across Persian rug. He approaches and gently touches her cold cheek. He carefully tilts her head and adjusts her posture like an art object. Sudden cut: she lies on the rug, dress spread like petals, staring upward. Low heartbeat bass rises. He lifts her back onto pedestal, fixing chin, shoulders, fabric folds. Slow dolly out as he retreats into shadow. Final extreme close-up: her eye moves slightly. Hard cut to black. Ultra cinematic, psychological tension, soft golden light, rich textures, shallow depth of field, eerie realism, masterpiece sound design.
Key Features of Wan 2.7 AI Video Generator
What Alibaba documents for Wan2.7-Video, written as controls you can actually prompt: native soundtrack, multi-shot, first and last frames, five mixed refs, and the 720p to 1080p ladder.
-
Native soundtrack in one pass
Without an uploaded track, Wan 2.7 writes matching music or effects into the same 30 fps MP4. CamArt lists Sound for this model. Review the take before you publish.
-
Multi-shot narrative in the prompt
Describe Shot 1 and later beats in one Wan 2.7 prompt. CamArt does not ship a 9-grid control. Use timestamped shots or Omni refs instead.
-
First and last frame
Upload a start still and an optional last still on CamArt Start-End. Official docs do not mix first/last frame with Omni refs in one request.
-
Five mixed identity refs
Alibaba reference-to-video allows up to five images and videos combined. Mention them as Image 1 and Video 1. CamArt Omni uses the same combined cap, with no separate audio slot.
-
720p or 1080p, honest length
Issuer T2V and I2V run 2 to 15 seconds at 720p or 1080p. On CamArt, WAN 2.7 Standard duration is 5 to 10 seconds. WAN 2.7 Pro image-to-video is 5 to 15 seconds.
-
Motion that holds under 1080p
Alibaba positions Wan 2.7 as a motion upgrade over Wan 2.6. Ranking pages talk about cleaner 1080p that holds under motion. We do not quote a success percent.
Name Image 1 and Video 1, or lock first and last frames.
Wan 2.7 follows camera, action, and reference mentions when you keep Standard, Omni, or Start-End consistent with the assets you attached.
Wan 2.7 Model Comparison
While Wan 2.7 is a 5 to 10 second native-audio clip with mixed refs capped at 5, Kling 3.0, Seedance 2.0, and MiniMax H3 all stretch to 15 seconds.
| Fact | Wan 2.7 | Kling 3.0 | Seedance 2.0 | MiniMax H3 |
|---|---|---|---|---|
| Issuer | Alibaba Wan. Native audio on a shorter Standard card. | Kuaishou Kling 3.0 Standard, Pro, and 4K SKUs. | ByteDance Seed. Unified multimodal audio-video. | MiniMax Hailuo family. H3 is the native-stereo card. |
| Best for | Shorter native-audio clips with mixed stills and clips. | 3 to 15 second cinematic clips with a 4K SKU. | Mixed-ref multimodal clips up to 15 seconds. | Native stereo clips drafted at 768p or finished at 2K. |
| Strength | Multi-shot soundtrack with mixed refs capped at 5. | Optional Sound and Multi-Shot from a written shot list. | Joint picture and sound from text, image, audio, and video. | 32 kHz stereo with picture. No Omni on CamArt. |
| Duration on CamArt | Standard 5 to 10 seconds. Issuer docs also list 2 to 15 seconds. | 3 to 15 seconds per Generate job. | 4 to 15 seconds per Generate job. | 4 to 15 seconds per Generate job. |
| Resolution on CamArt | 720p / 1080p on Standard. Pro I2V can pick 1080p to 4K. | 720p to 4K in the picker, including a dedicated 4K SKU. | 480p to 1080p on Fast and Standard. | 768p or 2K. A 1080p / 4K picker value maps to 2K. |
| Native audio | Yes. Native soundtrack on connected SKUs. | Optional Sound. Start and end frames sit on the same 3.0 card. | Yes. Picture and sound on connected Generate SKUs. | Yes. Native 32 kHz stereo on H3. |
| Modes and refs | T2V, I2V, mixed refs combined up to 5. | T2V, I2V, Start-End. No Omni on CamArt 3.0 Generate. | T2V, I2V, Omni 9 images / 3 videos / 3 audio, Start-End. | T2V and I2V. No Omni on CamArt H3. |
| CamArt credits | Standard is 12 credits per second (60 at 5 seconds). | Pro is 11.2 credits per second (56 at 5 seconds). | Fast is 10 credits per second (50 at 5 seconds). Standard is 12 per second. | 10 credits per second at 768p (50 at 5 seconds). 14 per second at 2K. |
Sources: Alibaba Wan 2.7; Kling VIDEO 3.0 guide; ByteDance Seedance 2.0; MiniMax H3.
Why Choose Wan 2.7 AI Video Generator in CamArt
Choose WAN 2.7 when the job is a 5 to 10 second Standard clip with native sound and up to five mixed refs. Credits show before Generate.
Ship picture and soundtrack together
Wan 2.7 writes matching music or effects into the same 30 fps MP4 when you do not pass an audio file. You do not have to add a track later unless you want to.
- Leave Sound on unless the brief is silent.
- Preview the clip with audio in Assets before you download.
Keep a face or product on-model
Alibaba reference-to-video takes up to five mixed images and videos. CamArt Omni Reference uses that combined cap. Name Image 1 and Video 1 in the prompt.
- Attach face, wardrobe, or pack shots before you generate.
- Stay in Omni or Start-End. Do not mix both in one official request.
Set a Standard length you can actually buy
Issuer T2V and I2V run 2 to 15 seconds. On CamArt, WAN 2.7 Standard duration is 5 to 10 seconds at 720p or 1080p. WAN 2.7 Pro image-to-video reaches 15 seconds.
- Pick 5s or 10s on Standard before you Generate.
- Switch to Pro image-to-video only when you need 15 seconds or 4K.
See credits on Generate before you spend
CamArt bills 1 credit as $0.01 COGS. WAN 2.7 Standard is 12 credits per second. A 5 second job is 60 credits. The button shows the charge before you start.
- Draft at 720p, finish at 1080p when the take holds.
- Starter credits cover low-cost models, not every 1080p job.
How to Use Wan 2.7 AI Video Generator in CamArt
Three steps from a 5 to 10 second brief to a saved clip.
-
STEP 01
Write the beat
Keep WAN 2.7 selected. Write the action, camera, and sound, or upload a start frame and a last frame for Start-End.
-
STEP 02
Attach Omni refs and set length
Add images and clips up to five combined. Set duration (5 to 10 seconds on Standard) and 720p or 1080p. Name Image 1 and Video 1 in the prompt.
-
STEP 03
Generate, preview, and save
Hit Generate. Credits on the button deduct on start. Preview picture and sound, then save from Assets. Switch to Pro image-to-video when you need 15 seconds or 4K.
Who Should Use Wan 2.7 on CamArt
WAN 2.7 fits briefs that need a 5 to 10 second Standard clip, native sound, and mixed refs. Use Wan 3.0 when you need a 30-second pass.
-
Brand and performance marketers
A 5 to 10 second Standard clip with native sound can carry a product beat without a second audio pass.
-
Filmmakers and previs teams
Lock first and last frames on Start-End, or schedule Shot 1 timestamps in the prompt for a short narrative test.
-
UGC and social editors
9:16 on CamArt T2V plus a native soundtrack fits talking-head and lifestyle briefs at 720p or 1080p.
-
Product explainers
Omni refs keep pack shots and faces on-model across up to five mixed stills and clips.
-
Agencies pitching campaigns
Credits show on Generate, so a 5 second 720p draft and a 10 second 1080p finish are priced before you spend.
-
Education teams
Prompt plus Omni refs can walk a short lesson. CamArt does not expose issuer video-edit or 9-grid controls.
Explore More AI Video Models on CamArt
FAQs
How long can Wan 2.7 videos be?
Alibaba documents 2 to 15 seconds for Wan 2.7 text-to-video and image-to-video. Reference-to-video with a video file is 2 to 10 seconds. On CamArt, WAN 2.7 Standard duration is 5 to 10 seconds. WAN 2.7 Pro image-to-video is 5 to 15 seconds.
Does Wan 2.7 generate audio with the picture?
Yes. On Model Studio, Wan 2.7 writes matching music or effects when you do not pass an audio file. Image-to-video can take driving audio for lip-sync and timing. CamArt lists Sound for this model. Review the take before you publish.
How is Wan 2.7 different from Wan 2.6?
Wan 2.7 adds documented first and last frame control, multi-entity reference-to-video, native soundtrack as the default, and instruction-based video edit on the issuer. Wan 2.6 is the prior synced-audio generation. This page runs WAN 2.7.
How is Wan 2.7 different from Wan 3.0?
Wan 3.0 is the 30-second all-in-one model with up to 10 images, 5 videos, and 5 audios. Wan 2.7 tops out at 15 seconds on T2V/I2V and five mixed image and video refs, with no CamArt Omni audio slot. Use the Wan 3.0 page when you need the longer pass.
Can Wan 2.7 lock first and last frames?
Yes. Upload a start still and an optional last still on CamArt Start-End. Official docs do not mix first/last frame with reference-to-video in one request. The CamArt picker follows that split.
How many reference images and videos does Wan 2.7 take?
Alibaba reference-to-video allows up to five images and videos combined. Mention them as Image 1 and Video 1. CamArt Omni Reference uses the same combined cap. Per-entity voice files are an issuer field, not a CamArt Omni audio slot.
What is 9-grid image-to-video on Wan 2.7?
Ranking pages describe a 3x3 still layout. Model Studio documents a multi-panel storyboard image on reference-to-video. CamArt does not ship a 9-grid control. Use Omni refs or a prompt-scheduled multi-shot instead.
Is Wan 2.7 free or open source on CamArt?
CamArt is not an unlimited free Wan 2.7 host. Credits on Generate are the charge (12 credits per second on Standard). Starter credits cover low-cost models, not every 1080p job. Do not treat wrapper open source claims as Model Studio 2.7 video facts.
Does CamArt run Wan 2.7 Pro or 4K video?
WAN 2.7 Pro is CamArt's image-to-video SKU at 5 to 15 seconds and 1080p to 4K. Default on this page is WAN 2.7 Standard text-to-video at 720p or 1080p. Issuer T2V is not a 4K model.
How do I generate a Wan 2.7 clip on CamArt?
Keep WAN 2.7 selected, write the prompt, add Omni or Start-End refs if you need them, set duration and resolution, then Generate. Sign in with Google or an email magic link if asked. Credits shown on the button deduct on start. Commercial use follows CamArt terms and third-party model terms.
Generate a Wan 2.7 clip with native sound
Describe the beat or upload a starting frame. Note camera, sound, and Image 1 / Video 1, then hit Generate.