AI 视频模型

如何提示 Gemini Omni:带示例的实用指南

学习如何为 Gemini Omni 编写 AI 视频创作和编辑提示词,包含提示公式、可复用模板、参考输入技巧、迭代提示、常见错误和示例。

Gemini Omni
发布于 2026年8月31日

如何提示 Gemini Omni:带示例的实用指南

展示如何用参考素材、分镜面板和视频预览提示 Gemini Omni 的编辑工作区

最好的 Gemini Omni 提示词读起来像一份压缩版制作 brief:目标、参考、主体、运动、镜头和约束。

如果你想学习如何提示 Gemini Omni,首先要放下一个想法:不存在一句万能咒语。

Gemini Omni 的价值在于它围绕多模态、视频优先创作而设计。你的提示词不只是描述一个画面,还可以定义片段目的、说明如何保留图片或视频参考中的内容、控制镜头运动、塑造情绪,并用自然语言继续指导后续编辑。

弱提示通常像这样:

Make a cinematic video of my product.

更强的写法是一份紧凑创意 brief:

Create a 6-second horizontal product reveal video for the uploaded image.
Use the product shape, color, and front-facing angle from the reference.
Start with the product in soft shadow, then use a slow push-in as warm studio light reveals the surface texture.
Style: premium editorial product film, clean background, realistic reflections.
Keep the logo position unchanged. Avoid warped text, extra labels, sci-fi effects, and camera shake.

第二个提示给了 Gemini Omni 最需要的六件事:结果、输入说明、主体、运动、镜头和约束。

这篇指南会展示怎样写出足够具体但不臃肿的 Gemini Omni 提示词,包含提示词公式、可复用模板、编辑提示、参考输入技巧,以及产品视频、社交短片、分镜、UI 演示和视频编辑示例。

快速答案:如何提示 Gemini Omni

要更好地提示 Gemini Omni,请把提示词写成视频制作 brief:

  1. 说明结果:片段用于什么、要完成什么。
  2. 命名输入:解释 Gemini Omni 应如何使用每张图片、每段视频、音频或文字参考。
  3. 锁定主体:定义必须保持一致的内容。
  4. 描述场景:环境、光线、情绪和视觉风格。
  5. 指导运动:从第一帧到最后一帧发生什么变化。
  6. 指导镜头:角度、镜头感、移动方式、景别和节奏。
  7. 添加约束:时长、画幅、需要保留的细节和需要避免的问题。
  8. 用范围很窄的后续提示迭代,而不是每次重写全部内容。

好的 Gemini Omni 提示词不是为了长而长,而是清楚说明屏幕上应该发生什么。

Gemini Omni 提示有什么不同?

提示 Gemini Omni 不同于提示普通聊天机器人,因为输出不是段落,而是基于时间的媒体。

这改变了提示词的任务。文本任务可以要求引言、改写或摘要;视频任务则必须描述时间中的变化:第一帧出现什么,片段中发生什么,镜头如何运动,哪些内容必须保持一致。

Google 将 Gemini Omni 定位为从不同输入类型创建和编辑内容,并且从视频开始。实际使用中,这让三个提示习惯尤其重要:

  • 你需要给参考输入分配职责。
  • 你需要描述运动,而不只是描述外观。
  • 你需要通过有针对性的编辑来修订结果。

最大的升级不是把提示词写得更长,而是换一个心智模型:你不是让模型“做点酷的东西”,你是在导演一段小型制作。

Gemini Omni 提示词公式

展示 Gemini Omni 提示词六个组成部分的图示:目标、输入、主体、运动、镜头和约束

强提示会把创意 brief 拆成可控制的部分。

需要稳定第一版生成时,用这个公式:

Create a [duration] [aspect ratio] video for [purpose].
Use [reference input] to preserve [specific details].
The subject is [who or what].
The scene is [setting, lighting, mood, style].
During the clip, [action or transformation].
Camera: [shot type, angle, movement, pacing].
Keep [details that must remain consistent].
Avoid [specific failure modes].

完整示例:

Create an 8-second vertical video for a TikTok product teaser.
Use the uploaded product image to preserve the bottle shape, label placement, cap color, and matte texture.
The subject is a minimalist skincare bottle on a wet stone surface.
The scene is a clean bathroom counter with soft morning light and subtle steam in the background.
During the clip, water droplets form on the bottle while the camera slowly pushes in from a medium shot to a close-up.
Camera: stable macro-product style, shallow depth of field, no handheld shake.
Keep the label layout unchanged.
Avoid misspelled text, extra logos, distorted packaging, floating objects, and neon sci-fi effects.

这个提示有效,是因为每句话都有制作功能,没有浪费在不会改变输出的形容词上。

先写任务,而不是先写风格

很多人用风格词开头:

Cinematic, high quality, ultra realistic, viral, beautiful...

这些词不是完全没用,但它们不是 brief。它们没有告诉 Gemini Omni 视频要做什么。

请先写任务:

Create a 6-second video ad showing how a messy desk becomes a calm productivity setup.

这句话给了模型目的、主体和变化。之后再补风格:

Style: bright editorial lifestyle video, realistic lighting, clean modern workspace, calm morning mood.

差别很关键。“电影感”是情绪,“凌乱桌面变成平静高效的工作区”才是事件。Gemini Omni 需要先理解事件。

告诉 Gemini Omni 如何使用每个参考

参考输入只有在被解释时才强大。如果上传了产品图、logo、分镜和情绪板,不要假设 Gemini Omni 知道哪个最重要。你要明确说明:

Use Image 1 for product identity: preserve the shape, color, logo position, and front-facing angle.
Use Image 2 for visual style only: match the lighting and background mood, but do not copy the objects.
Use Video 1 for motion reference: copy the slow camera push-in and pacing, not the subject.
Use the text brief for claims and audience context, but do not render the text inside the video.

这样可以减少常见失败:模型把参考素材混在一起,看起来合理,却丢掉你真正需要保留的东西。

每个参考都回答一个问题:Gemini Omni 应该从这个文件里保留什么?

好的保留目标包括:

  • 角色身份
  • 产品形状
  • Logo 位置
  • 色彩调色板
  • 镜头运动
  • 光线风格
  • 场景布局
  • 时间节奏
  • 情绪
  • 分镜顺序

模糊指令不好:

Use these images as inspiration.

更好:

Use Image 1 for the exact product silhouette and label position. Use Image 2 only for the warm studio lighting and beige background.

像时间线一样描述运动

静态图提示描述构图,Gemini Omni 提示应该描述时间。使用第一帧、中段和最后一帧:

First frame: the product is centered on a clean desk beside a closed laptop.
Middle: the laptop opens and the screen softly lights the product edge.
Final frame: the camera lands on a close-up of the product with the desk blurred in the background.

这比“做一个有笔记本电脑和好看光线的酷产品视频”清楚得多。你不需要完整剧本,只需要足够的顺序让 Gemini Omni 理解推进。

简单片段可以用三拍结构:

Beat 1: establish the subject.
Beat 2: show the change.
Beat 3: hold the final result.

复杂片段可以用镜头语言:

Shot 1: wide establishing shot of the studio desk.
Shot 2: medium shot as the product rotates slightly.
Shot 3: close-up of the texture and label.
Shot 4: final hero frame with clean negative space on the right.

Gemini Omni 越理解时间顺序,就越少需要自行发明。

使用真正有意义的镜头指令

镜头指令最有用的地方,是控制观众注意力。

有效指令包括:slow push-in、locked tripod shot、top-down flat lay、macro close-up、low-angle reveal、smooth tracking shot、over-the-shoulder view、wide establishing shot、shallow depth of field、handheld documentary feel。

弱指令包括:make it cinematic、great camera、professional angle、viral movement。

“电影感”可以设定口味,但不能说明镜头如何移动。把它放在具体镜头语言之后,而不是替代具体指令:

Camera: start with a locked medium shot, then use a slow push-in over 5 seconds until the product fills the center third of the frame. Cinematic but restrained, no fast cuts.

这给了模型节奏、构图和约束。

添加负面约束,但不要和模型对抗

负面提示很有用,但应当具体。过长的避免列表会让方向混乱。围绕最可能毁掉片段的问题写约束:

Avoid misspelled text, changing logo placement, extra fingers, warped product geometry, flickering labels, fast camera shake, and random sci-fi effects.

产品视频:保留形状、颜色、标签布局和 logo 位置;避免额外标签、虚假声明、包装变形和不可读文字。

角色视频:保留脸型、发型、服装颜色和身体比例;避免身份漂移、多余肢体、塑料皮肤和突然换装。

UI 演示:保持界面布局干净可读;避免伪造小字、破碎图标、变形浏览器框和随机按钮。

电影场景:避免过度炫光、抖动、极端慢动作、霓虹赛博朋克光线和不真实物理。

目标不是列出所有 AI 可能犯的错,而是阻止会毁掉这个特定片段的失败。

如何不从头开始地迭代

Gemini Omni 提示工作流:初稿、有针对性的编辑、一致性检查和最终导出

好的 Gemini Omni 提示是迭代式的:生成、检查、做一个有针对性的修改,再检查。

第一版输出只是方向,不是最终资产。错误做法是每次都重新生成,那会丢掉已经有效的部分。更好的流程是保留有用内容,一次只改一个问题。

Keep the same subject and background. Make the camera movement 50% slower and remove the handheld shake.
Preserve the product shape and label position. Change only the lighting from dramatic side light to soft morning light.
Keep the first three seconds. Replace the final shot with a cleaner close-up where the object is centered and fully in frame.
Use the same scene and action, but make the motion more physically realistic. Reduce floating objects and keep the product touching the table.
Preserve the composition. Remove the fake text from the background and keep the interface panels visually clean.

好的修订提示包含三部分:保留什么、改变什么、避免什么。视频里一个元素变化可能牵动很多东西,所以这个结构尤其重要。

可复用提示模板

把下面模板当起点,替换方括号里的内容。

产品揭示提示

Create a [duration] [aspect ratio] product reveal video for [platform or use case].
Use the uploaded product image to preserve [shape, color, label position, material, angle].
First frame: [initial setup].
Middle: [reveal action or camera movement].
Final frame: [hero shot].
Style: [premium studio, editorial lifestyle, clean SaaS, documentary, etc.].
Camera: [locked shot, slow push-in, macro close-up, tracking shot].
Keep [brand/product details] consistent.
Avoid [distorted packaging, misspelled text, extra labels, unwanted style].

示例:

Create a 7-second horizontal product reveal video for a landing page hero.
Use the uploaded device render to preserve the exact screen shape, silver frame, and keyboard layout.
First frame: the laptop is closed on a clean white desk.
Middle: the laptop opens and the screen glows softly.
Final frame: a close-up of the screen with clean negative space around it.
Style: premium SaaS editorial, bright neutral lighting, realistic reflections.
Camera: stable slow push-in, no handheld shake.
Keep the device proportions and keyboard layout consistent.
Avoid fake readable UI text, extra logos, warped edges, and sci-fi glow.

社交视频提示

Create a [duration] vertical video for [TikTok, Reels, Shorts, X, LinkedIn].
The hook is: [visual hook in the first second].
The subject is [person, product, object, scene].
Show [specific action or transformation].
Style: [creator-style, documentary, studio, educational, editorial].
Camera: [phone-style handheld, locked tripod, overhead, close-up].
Pacing: [fast, calm, single take, three quick cuts].
Avoid [failure modes].

示例:

Create an 8-second vertical video for a short-form productivity post.
The hook is a chaotic desk instantly becoming organized in the first second.
The subject is a notebook, laptop, coffee cup, and task cards on a real desk.
Show the clutter moving into a clean weekly planning layout.
Style: bright creator-style desk video, realistic morning light, satisfying but not magical.
Camera: top-down locked shot with subtle natural movement.
Pacing: three clean beats, no fast glitch transitions.
Avoid floating objects, warped text, fake brand logos, and unrealistic hands.

分镜提示

Create a [duration] video concept based on this shot list:
Shot 1: [wide scene].
Shot 2: [subject action].
Shot 3: [detail or reaction].
Shot 4: [final hero frame].
Style: [visual direction].
Camera: [movement and shot types].
Keep [continuity details].
Avoid [continuity failures].

示例:

Create a 12-second cinematic concept video based on this shot list:
Shot 1: wide shot of a small design studio at night, rain on the window.
Shot 2: over-the-shoulder shot of a designer dropping a rough sketch into an AI image tool.
Shot 3: close-up as the rough sketch becomes a polished campaign image.
Shot 4: final hero frame of the designer reviewing three clean variations.
Style: grounded editorial tech, warm desk light, realistic workspace, no sci-fi holograms.
Camera: slow controlled movement, shallow depth of field, no fast cuts.
Keep the same desk, laptop, and designer across all shots.
Avoid fake readable UI text, extra screens, random neon effects, and identity drift.

现有视频编辑提示

Edit the uploaded video.
Preserve [subject, timing, composition, identity, product details].
Change only [background, lighting, camera pacing, style, object, mood].
Keep the original [motion, face, product shape, scene layout] consistent.
Avoid [specific artifacts].

示例:

Edit the uploaded 6-second product clip.
Preserve the bottle shape, label placement, rotation timing, and camera angle.
Change only the background from a dark studio to a bright marble bathroom counter with soft morning light.
Keep the bottle physically grounded on the surface.
Avoid changing the logo, adding extra text, warping the cap, or introducing floating reflections.

UI 演示提示

Create a [duration] video demo for [app or feature].
Use the uploaded screenshot as the interface reference.
Show [specific interaction or transformation].
Camera: [screen recording style, slight push-in, over-the-shoulder, product mockup].
Keep the layout simple and readable.
Avoid fake detailed text, broken UI, extra buttons, distorted browser elements, and unreadable labels.

示例:

Create a 6-second horizontal video demo for an AI image editor.
Use the uploaded screenshot as the interface layout reference.
Show a rough product photo turning into three clean campaign variations inside the workspace.
Camera: screen-recording style with a very slight push-in, stable and readable.
Keep the panels, toolbar, and preview area simple.
Avoid fake tiny text, random extra buttons, distorted browser chrome, and sci-fi effects.

前后对比提示示例

示例 1:旅行片段

弱提示:

Make a beautiful travel video of the coast.

更好:

Create an 8-second horizontal travel video of a quiet coastal road at sunrise.
First frame: wide aerial view of the road along the cliffs.
Middle: the camera glides forward as a single car follows the curve of the road.
Final frame: the coastline opens into a warm sunrise over the ocean.
Style: natural documentary travel film, warm light, realistic waves, no exaggerated saturation.
Camera: smooth drone-style movement, slow and stable.
Avoid crowded traffic, fake landmarks, shaky motion, and overdone lens flare.

为什么有效:它给了 Gemini Omni 主体、地点、时间线、镜头路径、情绪和避免列表。

示例 2:时尚角色片段

弱提示:

Make this person walk in a cool city.

更好:

Create a 6-second vertical fashion video using the uploaded portrait as the identity reference.
Preserve the person's face shape, hairstyle, jacket color, and overall proportions.
Scene: quiet city street after rain, soft reflections on the pavement, early evening light.
Action: the person walks slowly toward camera for three steps, then turns slightly to the side.
Camera: chest-height tracking shot, stable, shallow depth of field.
Style: editorial streetwear film, realistic skin texture, natural motion.
Avoid identity drift, extra fingers, changing clothing color, plastic skin, and distorted background signs.

为什么有效:它把身份保留和新场景清楚分开。

示例 3:应用发布片段

弱提示:

Make a video of my app looking futuristic.

更好:

Create a 7-second horizontal launch video for an AI research app.
Use the uploaded screenshot for layout only: preserve the sidebar, search panel, and result cards.
Show a user query turning into three organized research summaries.
Style: clean professional SaaS demo, bright neutral background, restrained blue and green accents.
Camera: screen-recording style with a subtle push-in, no dramatic 3D rotation.
Keep the interface readable as a layout, but do not generate detailed fake text.
Avoid cyberpunk colors, holograms, extra buttons, warped browser chrome, and unreadable tiny labels.

为什么有效:它避免了常见的“未来感 UI 糊成一团”,并给模型明确的界面边界。

常见 Gemini Omni 提示错误

错误 1:要求氛围,而不是事件

“做得电影感”不够,要说发生什么。

Show the product emerging from shadow as the camera slowly pushes in, ending on a stable close-up.

错误 2:上传参考却不说明用途

Gemini Omni 可能理解文件,但仍需要优先级。参考可以代表身份、风格、光线、构图、运动或情绪。告诉它是哪一种。

错误 3:一个提示塞太多任务

不要让一个短片同时做产品广告、角色动画、UI 演示、logo 动效和电影故事。选择主要任务。

错误 4:相信生成文本

AI 视频系统很难稳定生成精确字体、字幕、logo 和 UI 标签。严肃资产请生成后在确定性编辑器里加文字:

Leave clean negative space for a headline. Do not render the headline inside the video.

错误 5:后续提示改变一切

如果第一版有一个好镜头,就保留它。后续提示应当窄:

Keep the composition and product position. Change only the camera speed and reduce the background clutter.

错误 6:忘记最终检查

短 AI 视频正常速度下可能很好看,但逐帧有问题。检查身份漂移、手部变形、产品不一致、logo 闪烁、UI 破碎、物理错误、不想要的文字、意外品牌标记和过度平滑的人工运动。

商业工作中,提示词只是流程一部分,人工审核仍然必须存在。

实用 Gemini Omni 提示工作流

  1. 写一句创意目标。
  2. 收集参考,并标注每个参考要保留什么。
  3. 用上面的公式起草提示。
  4. 生成一个版本。
  5. 看两遍:一遍看整体感觉,一遍看细节。
  6. 写有针对性的修订提示。
  7. 重复直到主要问题解决。
  8. 检查瑕疵、声明和品牌细节后再导出。
检查看什么
主体一致性人、产品或界面是否保持可识别?
运动动作是否按预期顺序发生?
镜头移动是否稳定并且有用?
参考Gemini Omni 是否保留了正确细节?
文本标签、logo 和 UI 元素是否可安全使用?
瑕疵是否有变形、闪烁或不可能的物理?
用例片段是否真的解决了原始任务?

如果最后一个答案是否定的,不要继续打磨,重写 brief。

提升 Gemini Omni 提示的高级技巧

使用“preserve”和“change only”

Preserve the subject, composition, and lighting. Change only the background from a kitchen to a clean studio desk.

要求更少剪辑

Use one continuous shot. No fast cuts. The camera should move slowly enough that the product remains readable.

用留白代替生成文字

Leave clean negative space on the right side for a headline. Do not render text inside the video.

把风格和内容分开

Use Image 2 for lighting and color mood only. Do not copy its objects, layout, or people.

直接控制真实感

Keep the product resting on the table throughout the clip. Reflections should follow the surface naturally. No floating, bending, or melting.

把时长当成约束

短片需要简单动作。5 到 8 秒用一个核心运动;10 到 15 秒用小镜头清单;更长内容用分镜,不要写一个巨型提示。

不适合只靠 Gemini Omni 提示完成的事

Gemini Omni 适合概念探索、创意迭代和视频生成,但提示词不能替代生产控制。遇到精确法律文案、医疗/金融/安全声明、像素级产品标签、完美 logo、精确 UI 文本、受监管广告、最终排版和合同级品牌合规时要谨慎。

这些任务应使用 Gemini Omni 做视觉生成或探索,再用编辑、设计或合规工具完成最终资产。

可复制的最终提示

Create a [duration] [aspect ratio] video for [specific use case].
Use the uploaded references as follows:
- Image 1: preserve [identity/product/layout].
- Image 2: use only for [style/lighting/mood].
- Video 1: use only for [motion/pacing/camera reference].

The subject is [main subject].
The scene is [setting, lighting, mood].
First frame: [what we see at the start].
Middle: [what changes].
Final frame: [where the clip should land].

Camera: [shot type, movement, angle, pacing].
Style: [visual direction].
Keep [must-preserve details] consistent.
Avoid [specific artifacts, unwanted styles, text problems, and brand issues].

修订版:

Keep [what worked] exactly the same.
Change only [one specific issue].
Make the result more [desired direction].
Avoid [the failure seen in the last version].

最终结论

提示 Gemini Omni 的最好方式,是停止像提示词收藏者一样思考,开始像导演一样思考。

定义任务,解释参考,描述第一帧、运动和最后一帧,让镜头承担角色,添加符合真实风险的约束,然后用聚焦编辑迭代,而不是每次从头开始。

这不会让每次生成都完美,但会让失败更容易诊断,也让成功结果更容易复现。

FAQ

最好的 Gemini Omni 提示结构是什么?

结构是:结果、参考输入说明、主体、场景、运动、镜头、风格、保留规则和避免列表。它同时给 Gemini Omni 创意方向和生产约束。

Gemini Omni 提示应该多长?

足够定义任务即可,不要长到需要模型满足互不相关的想法。多数短视频 120 到 250 个英文词等量内容就够;分镜可以更长。

应该使用负面提示吗?

应该,但要具体。不要列所有可能瑕疵,只写会毁掉这个片段的问题,如标签变化、身份漂移、UI 变形、镜头抖动或不想要的科幻效果。

Gemini Omni 可以使用图片或视频参考吗?

Google 将 Gemini Omni 定位为多模态创建和编辑模型,所以基于参考的提示是核心流程。具体输入和控制会因产品界面、账号、地区和发布阶段而不同。生产前请检查当前 Gemini、Flow 或 AI Studio 界面。

如何提示 Gemini Omni 保留角色?

上传最强身份参考,并明确保留脸型、发型、服装、身体比例和表情。后续编辑时,在要求场景或动作变化前重复保留指令。

如何提示产品视频?

从参考图保留产品形状、材质、标签位置、颜色和镜头角度。使用简单运动,如慢推近或转台展示。避免额外 logo、包装变形、虚假声明和不可读文字。

Gemini Omni 适合精确文本或 logo 吗?

要谨慎。专业工作更安全的做法是生成干净留白,再在确定性编辑器中加入精确文字、字幕、法律文案和 logo。

最大的提示错误是什么?

最大的错误是要求风格而不是屏幕事件。“电影感产品视频”很弱;“瓶子从阴影中出现,镜头推近并落在稳定特写上”强得多。

已检查来源

本文基于 2026 年 8 月 31 日检查的公开产品信息撰写: