Kling O1 AI Video Generator

Animate one or two keyframes with Kling O1 on Cinel. Set the opening and closing image, choose 5 or 10 seconds, and get a controlled, consistent shot.

Image to videoFirst / last frame5s / 10s
5s · Default

See Kling O1 in action

See how a reference frame becomes a moving shot.

01

Controlled frames

Guide the opening and closing image of the shot.

02

Natural motion

Turn a still image into a focused moving shot.

03

Simple setup

Choose a five- or ten-second image-to-video output.

Deep Dive & Practical Guide

What Sets Kling O1 Apart

An engineering and creative breakdown of model performance, architecture, and practical production workflows.

Kling O1 is the first "Omni" model from Kuaishou's Kling team, announced as a single engine that understands text, images, video and subjects together. On Cinel we use it for the thing it does with particular care: keyframe animation. Upload a starting image — and, if you like, an ending image — and Kling O1 builds the motion that travels between them, keeping the character, product or scene consistent from the first frame to the last. Choose a 5-second or 10-second clip; pricing on Cinel starts from 230 credits.

Key Section

What makes Kling O1 different

Most image-to-video models only take a first frame and improvise what happens next. Kling O1 lets you pin both ends of the shot. That turns animation from a gamble into a plan: a closed flower in frame one and the open bloom in frame two; a product in its box, then on the table; a character facing away, then turning to camera. The model fills the in-between with natural, focused motion.

Consistency was the stated goal of the O1 launch, and it shows in how subjects hold their identity through the move — faces stay the same face, logos stay the same logo. The settings are intentionally minimal: two durations and the frames you supply. That simplicity is a feature for storyboard artists, e-commerce teams and editors who need a predictable transition rather than a surprise. Kling O3, the later Omni model, adds text-to-video and higher resolutions; O1 remains the specialist for frame-to-frame control.

Key Section

Key features

  • First-and-last-frame control — define where the shot begins and where it lands.
  • Single-image animation — or supply just one frame and let the motion unfold.
  • Identity consistency — characters, products and scenes stay recognisable.
  • Two fixed durations — 5 seconds or 10 seconds.
  • Three aspect ratios — 16:9, 9:16 and 1:1.
  • Omni foundation — built on Kling's unified multimodal architecture.
Key Section

Specs

Item Kling O1
Developer Kuaishou (Kling AI)
Model family Kling Omni (first generation)
Mode on Cinel Image-to-video with first / last frame
Duration 5 s or 10 s (10 s ≈ 2× the 5 s price)
Aspect ratios 16:9, 9:16, 1:1
References on Cinel Up to 2 images (start and end frame)
Cinel cost From 230 credits
Key Section

How to use Kling O1 on Cinel

  1. Open Kling O1 in the video generator (a membership or credit pack is required for video).
  2. Upload your opening frame. Add a closing frame if you want the shot to land on a specific image.
  3. Make the two frames compatible — same character, similar lighting and lens — so the transition can be smooth.
  4. Describe the motion between them: "the camera slowly circles as she turns toward us and smiles".
  5. Pick 5 s for a quick beat or 10 s for a gentler transformation, confirm the credit estimate and generate.
Key Section

Prompt tips for Kling O1

  • Make the start and end frames share the same lighting, lens feel and background. The closer they match, the smoother and more believable the motion between them.
  • Describe the journey, not the destinations: "the camera arcs left as the lid lifts and the product rises" tells O1 how to travel from frame A to frame B.
  • Keep the subject in a similar position and scale in both frames unless you want a deliberate zoom; large jumps can force awkward in-between motion.
  • Choose 5 seconds for a quick reveal and 10 seconds for slower, more elegant transitions — the 10-second option costs about twice as much, so test at 5 first.
  • Generate the two keyframes with an image model on Cinel first (for example Nano Banana 2 or Seedream 4.5), then bring them into Kling O1 for animation.
Key Section

Use cases

  • Before-and-after transformations — room makeovers, outfit changes, product unboxing.
  • Storyboard animatics — move between two approved storyboard panels.
  • E-commerce product motion — from packaging to in-use, keeping the item identical.
  • Character turns and reveals — controlled head turns or expressions for trailers.
  • Logo and brand transitions — land precisely on a final branded frame.
  • Art and illustration loops — animate between two painted states.
Key Section

Kling O1 vs Kling O3 and Kling 3.0

Kling O3 (from 93 credits) is the newer Omni model with text-to-video, 3–15 second durations and 720p/1080p/4K output — choose it when you need flexible length or no starting image. Kling 3.0 suits fast, energetic text- or image-driven clips. Choose Kling O1 when the exact opening and closing frames matter more than anything else.

FAQ

Before you create

What does "first and last frame" mean in Kling O1?

You supply the image the video starts on and, optionally, the image it ends on. Kling O1 generates the motion connecting the two.

Can I use Kling O1 with only one image?

Yes. A single starting frame works too. Adding a final frame simply gives you more control over where the shot finishes.

How long are Kling O1 videos?

Either 5 or 10 seconds. A 10-second clip costs about twice as much as a 5-second clip.

Does Kling O1 do text-to-video on Cinel?

On Cinel, Kling O1 is set up for image-to-video. For text-only prompts, use Kling O3, Kling 3.0 or Kling 3.0 Turbo.

How do I get a smooth transition between frames?

Keep the subject, lighting and camera angle consistent between the two images and describe one clear motion. Big jumps in setting make the in-between harder.

How many credits does Kling O1 cost?

It starts from 230 credits on Cinel. The 10-second option and other settings change the estimate shown before generation.