Wan NSFW
Glossary

AI Video Glossary

Plain-language definitions of every technical term you'll encounter when using Wan AI video models. Click any related term or model link to go deeper.

T2V (Text-to-Video)
#

Generate a video from a text prompt alone.

Text-to-Video (T2V) takes a written prompt and synthesizes a video clip from scratch — no input image required. You describe a scene, pose, lighting, and camera move, and the model renders motion directly from that description. All Wan video models support T2V. Quality improves significantly with detailed, specific prompts.
I2V (Image-to-Video)
#

Animate a still image into a video clip.

Image-to-Video (I2V) takes a source image as the visual starting point and generates motion from it. In Wan video models, the image usually works as the first frame, and newer workflows can also use first-and-last frame control to guide how the clip begins and ends. Wan 2.2, Wan 2.5, Wan 2.6, Wan 2.7, and Wan 3.0 all support I2V workflows in different forms.
R2V (Reference-to-Video)
#

Generate a video with help from one or more reference images.

Reference-to-Video (R2V) is a generation mode where a Wan video model uses reference images as visual guidance. The reference helps define what should stay recognizable, such as the subject, style, or composition, while the text prompt tells the model what motion, scene, lighting, and camera direction to generate. R2V is different from T2V because it does not rely on text alone, and different from basic I2V because it is usually used when consistency across outputs matters.
Diffusion Model
#

The core AI architecture behind Wan video generation.

A diffusion model generates content by learning to reverse a noise process: it starts from random noise and progressively refines it into a coherent image or video, guided by a text or image condition. Wan AI models are built on diffusion architectures. The quality and coherence of the output depends heavily on how many denoising steps are run and the guidance scale used.
CFG Scale (Classifier-Free Guidance)
#

How closely the model follows your prompt vs. adds its own variation.

CFG (Classifier-Free Guidance) is a number that controls how strictly the model follows your prompt. A low CFG (e.g. 3–5) gives the model more creative freedom and tends to produce smoother motion but may drift from the prompt. A high CFG (e.g. 10–15) forces the model to stick closer to your words but can produce over-saturated or artifact-heavy frames. Most generators on this site use a sensible default — adjust only if you see obvious prompt drift or color blowout.
Denoise Steps (Sampling Steps)
#

Number of refinement passes the model makes before producing the final frame.

Each denoising step refines the output from noise toward a coherent result. More steps (e.g. 50) generally produce sharper, more detailed frames at the cost of longer generation time. Fewer steps (e.g. 20) are faster but may leave the output slightly blurry or inconsistent. For NSFW video generation, 20–30 steps is typically a good balance. The generator on this site manages step count automatically.
Prompt
#

The text description that drives what the model generates.

A prompt is the text you write to tell the model what to generate. For video generation, a strong prompt typically includes: subject description (appearance, pose), action or motion, setting and lighting, camera angle or movement, and style cues. Wan models respond well to explicit, concrete descriptions. Vague prompts produce generic output; specific prompts produce controlled results. See the Prompt Guide for model-specific tips.
Negative Prompt
#

Words that tell the model what to avoid generating.

A negative prompt is a second text input where you list things you do not want in the output — e.g. "blurry, watermark, extra limbs, distorted face". The model pushes generation away from these concepts. Effective negative prompts for NSFW video typically include common artifact descriptions. Not all interfaces expose the negative prompt field; the Wan generator on this site includes it under advanced settings.
Subject Consistency
#

How stable a subject stays across frames, shots, and clips.

Subject consistency describes whether the same subject remains recognizable across frames, shots, or generated clips. Weak consistency can cause identity drift, where appearance changes as the video moves. Wan 2.2 was an early improvement for stable subject handling, and Wan 2.5 and 2.6 improved it further. I2V helps by starting from a source image, while R2V and newer Wan 2.7 / Wan 3.0 reference-based workflows can use reference media to guide consistency across more complex outputs.
FPS (Frames Per Second)
#

How many video frames are rendered each second.

FPS measures how many frames appear in one second of video. Higher FPS usually makes motion look smoother, but it also means more frames are generated for the same clip length. Older Wan video specs may describe output around ~24 fps, such as Wan 2.6 generating up to 15 seconds, or about ~360 frames per generation. Newer Wan 2.7 and Wan 3.0 video models are documented with 30 fps MP4 output; a 30-second Wan 3.0 clip at 30 fps contains about 900 frames.
Resolution
#

The pixel size and clarity tier of generated video output.

Resolution is the width × height of each video frame, expressed in pixels or as output tiers such as 720p and 1080p. Higher resolution captures finer detail but can require more compute, memory, and generation time. Wan 2.2, Wan 2.5, and Wan 2.6 are commonly presented as 1080p native video models. Wan 2.7 video supports 720p and 1080p output, while Wan 3.0 supports 480p, 720p, and 1080p, with 1080p as the highest output tier. Actual perceived sharpness also depends on the model's detail capacity, prompt quality, source media, and motion complexity.
Prompt Adherence
#

How closely the output follows your written prompt.

Prompt adherence describes how well a Wan model follows the scene, subject, motion, camera, lighting, and style you write in the prompt. Strong prompt adherence means the output stays close to your requested direction instead of adding unrelated details or ignoring key instructions. Wan 2.2 improved prompt control in the earlier video lineup, while newer Wan 2.7 and Wan 3.0 workflows can also use prompt extension and reference media to guide the result more clearly.
Motion Smoothness
#

How natural and stable movement looks across frames.

Motion smoothness describes whether movement feels continuous from one frame to the next. Weak motion can look jittery, stiff, or inconsistent, especially when the prompt asks for too many actions at once. FPS affects smoothness, but it is not the only factor: prompt clarity, subject consistency, scene complexity, and model version also matter. Wan 2.5 and Wan 2.6 improved motion stability over earlier models, while Wan 2.7 and Wan 3.0 add stronger video workflows for longer and more controlled motion.
Aspect Ratio
#

The shape of the generated frame.

Aspect ratio describes the width-to-height shape of an image or video frame, such as 16:9 for landscape, 9:16 for vertical video, or 1:1 for square output. It affects framing, composition, and where the subject fits inside the shot. Wan 2.6 specs commonly use 9:16, 16:9, and 1:1, while Wan 2.7 and Wan 3.0 also support wider ratio options such as 4:3 and 3:4. In first-frame or reference-based workflows, the output may follow the shape of the input media instead of using a manually selected ratio.
Reference Media
#

Input media used to guide what the model generates.

Reference media means any input material used to guide a Wan generation beyond text alone. Depending on the model and workflow, this can include images, videos, audio, files, or web-page links. Reference media helps the model understand identity, scene direction, visual style, motion, or sound more clearly than a prompt by itself. I2V usually starts from a source image, while R2V and Wan 3.0 reference-based workflows can use broader media inputs for stronger control across more complex outputs.