ساخت ویدیو واقعی با هوش مصنوعی؛ آموزش کامل از عکس تا ویدیو حرفه‌ای

آخرین به‌روزرسانی: 30 شهریور 1405, 12:28 ب.ظ
هوشا 30 شهریور 1405 آموزش ۳۷ دقیقه زمان مطالعه 0 دیدگاه ( ۰ امتیاز )

ساخت ویدیو واقعی با هوش مصنوعی به شما امکان می‌دهد تنها با یک پرامپت و تصاویر مرجع، صحنه‌هایی طبیعی و باورپذیر را بدون نیاز به فیلم‌برداری یا تجهیزات حرفه‌ای بسازید. در این مقاله، نحوه ساخت ویدیو با هوش مصنوعی را آموزش می‌دهیم و نمونه ویدیوهای واقع‌گرا را همراه با پرامپت و نتیجه نهایی بررسی می‌کنیم.

در این آموزش با موارد زیر همراه ما باشید:

  • آموزش مرحله‌به‌مرحله ساخت ویدیوی واقعی در هوشا
  • معرفی مدل‌های مناسب برای ساخت ویدیوی طبیعی 
  • ۴ نمونه استایل آماده برای ساخت ویدیوی واقعی به همراه نتیجه

ساخت ویدیو واقعی با هوش مصنوعی؛ تبدیل عکس به ویدیو

تصویر شاخص مقاله ساخت ویدیوی واقعی با هوش مصنوعی در هوشا

ویدیوی واقعی زیر با استفاده از استایل آماده در سرویس ساخت ویدیوی هوشا ساخته شده است؛ در ادامه، مرحله‌به‌مرحله توضیح می‌دهیم که چطور می‌توانید دقیقاً همین ویدیو را بسازید. 

آموزش ساخت ویدیو واقعی در پلتفرم هوشا در ۵ گام

در این بخش با استفاده از یک استایل آماده و عکس مرجع، فرایند ساخت ویدیو واقعی با هوش مصنوعی را در پلتفرم ایرانی هوشا و با مدل Grok Imagine 1.5 از ابتدا تا انتها انجام می‌دهیم. در ادامه، مراحل را به‌صورت تصویری دنبال کنید.

۱. ورود به سرویس ساخت ویدیوی هوشا

وارد حساب کاربری خود در هوشا (Hoosha) شوید و در منوی سمت راست، از میان سرویس‌های اصلی، سرویس ساخت ویدیو را انتخاب کنید. 

اسکرین‌شات از صفحه اصلی سرویس ساخت ویدیو در هوشا، نمایش نحوه ورود به سرویس
برای ورود به سرویس ساخت ویدیو، از منوی سمت راست در هوشا، گزینه ویدیو را انتخاب کنید.

۲. انتخاب استایل آماده

در قسمت زیرِ کادر پرامپت، از میان استایل‌های آماده ساخت ویدیو، استایل دلخواه خود را انتخاب کنید. با کلیک روی گزینه «امتحان کن»، پرامپت، تصاویر مرجع و برخی تنظیمات مربوط به همان استایل مستقیماً در صفحه ساخت ویدیو اعمال می‌شوند و می‌توانید ویرایش پرامپت را در همین بخش انجام دهید. 

نکته: از قسمت «بر اساس مدل» می‌توانید مدلِ هوش مصنوعی مورد نظر را انتخاب کنید تا فقط استایل‌های آماده سازگار با همان مدل نمایش داده شوند.

اسکرین‌شات از انتخاب استایل آماده در سرویس ساخت ویدیوی هوشا، نمایش انواع مدل‌ها
با مرتب‌سازی بر اساس مدل‌ها، می‌توانید استایل‌های آماده برای هر مدل را جداگانه بررسی کنید.

هوش مصنوعی ساخت ویدیو ایرانی (با دیالوگ + امکانات حرفه‌ای)

۳. ذخیره پرامپت ساخت ویدیو

با انتخاب هر یک از استایل‌های آماده، صفحه زیر نمایش داده می‌شود. می‌توانید متن پرامپت را از گزینه «کپی پرامپت» کپی کرده و در یک فایل ورد برای ویرایش و شخصی‌سازی، در صورت نیاز، ذخیره کنید.

اسکرین‌شات از محل گزینه کپی پرامپت در سرویس ساخت ویدیوی هوشا
برای ذخیره پرامپت و ویرایش آن، متن را از گزینه کپی پرامپت، کپی کنید.

متن پرامپت ویدیوی تبلیغاتی گردنبند

برای ساخت ویدیوی تبلیغاتی گردنبند، پرامپت زیر را کپی کرده و عیناً در کادر ساخت ویدیو وارد کنید. این پرامپت حفظ مشخصات گردنبند در عکس مرجع و همچنین ساخت موسیقی پس‌زمینه را پوشش می‌دهد و نیازی به آپلود موسیقی جداگانه نیست. 

  • مدل: Grok Imagine 1.5
  • ابعاد: ۹:۱۶
Create a 15-second ultra-premium vertical 9:16 luxury jewelry commercial from the provided image. Use the uploaded image as the exact first-frame and style foundation. The hero subject is the gold necklace and pendant worn by the model. The necklace and pendant must remain the absolute visual focus in every shot. Preserve the jewelry precisely: same pendant shape, same chain design, same gold tone, same diamond placements, same proportions, same placement on the neck, and the same elegant luxury identity.


Style: high-end fine jewelry campaign, cinematic beauty ad, soft gold and champagne palette, refined editorial lighting, creamy highlights, luminous skin, shallow depth of field, polished reflections, minimal upscale environment, expensive and timeless mood. The result should feel like a multi-million-dollar global luxury jewelry advertisement.


Timeline:
0.0–3.0s: Open with an intimate macro shot of the pendant resting on the model’s collarbone. Slow cinematic push-in. Soft warm light glides across the gold pendant and chain. Natural breathing only. Keep the pendant razor-sharp and dominant.
3.0–6.0s: Camera gently glides upward and slightly around the neckline and jawline while keeping the necklace centered. Add subtle focus pull from chain details to pendant. Hair remains controlled and elegant with minimal movement.
6.0–9.0s: Move into a refined beauty shot from chest to lips and lower face. The model turns very slightly with graceful confidence. The pendant catches elegant highlights and remains the hero. No distracting pose or exaggerated expression.
9.0–12.0s: Extreme premium jewelry moment: closer macro emphasis on the pendant and upper chain. Show crisp gold edges, diamond sparkle, clean reflections, and soft skin texture. Add a smooth light sweep across the pendant to make it feel iconic and luxurious.
12.0–15.0s: Final hero packshot. Pull back slightly to a balanced chest-to-face composition, with the pendant perfectly centered and fully visible. End on a poised still-like luxury frame where the jewelry feels unforgettable, radiant, and premium.


Character and styling: elegant adult woman, refined editorial beauty, smooth luminous skin, tasteful natural makeup, sleek pulled-back hair, graceful posture, calm luxurious presence. Wardrobe should remain minimal, silky, and understated so the necklace stays dominant.


Environment: soft upscale editorial setting with warm natural light and subtle architectural shadows, clean premium background, no clutter, no distraction, no extra props.


Camera language: slow macro push-in, subtle orbit, graceful glide, precise focus pulls, stable movement, no shaky camera, no fast cuts, no random zooms.


Lighting: soft premium beauty lighting, warm champagne key light, delicate rim light, crisp specular highlights on the gold pendant and diamonds, flattering skin glow, elegant contrast.


Audio: sensual luxury soundtrack, soft refined music, subtle shimmer textures, polished whoosh transitions, delicate jewelry sparkle accents, no vocals, no dialogue, no harsh beats.


Motion rules: all movement must be slow, graceful, realistic, and controlled. The model moves minimally. The necklace remains stable, wearable, realistic, and hero-focused.


Avoid: changing the necklace design, pendant distortion, wrong chain shape, extra jewelry, cluttered styling, plastic skin, warped anatomy, distracting fashion poses, low-resolution textures, oversaturation, text overlays, captions, watermark, or background distractions.

۴. آپلود عکس مرجع و اعمال تنظیمات ساخت ویدیو

پس از اضافه کردن پرامپت در کادر مربوط، برای آپلود تصویر مرجع و اعمال تنظیمات ویدیو، مراحل زیر را طبق شماره‌گذاری تصویر انجام دهید:

  1. آپلود تصویر: تصویر مرجع واضح و با کیفیت بالا را برای ساخت ویدیو بارگذاری کنید.
  2. انتخاب مدل: مدل هوش مصنوعی موردنظر را انتخاب کنید.
  3. انتخاب کیفیت: کیفیت ویدیو را تعیین کنید. بدیهی است که کیفیت بالاتر به اعتبار بیشتری نیاز دارد.
  4. انتخاب مدت‌زمان: از گزینه «زمان» در بخش تنظیمات می‌توانید مدت ویدیو را تعیین کنید. از آنجا که استایل‌های آماده معمولاً زمان ویدیو و طول هر صحنه را مشخص می‌کنند، این گزینه باید با زمان‌بندی ذکرشده در پرامپت مطابقت داشته باشد.
  5. انتخاب ابعاد: بهتر است ابعاد ویدیو با تصویر مرجع هماهنگ باشد. برای مثال، اگر ویدیوی عمودی با نسبت ۹:۱۶ می‌خواهید، بهتر است تصویر مرجع نیز همین نسبت را داشته باشد. 
  6.  تولید ویدیو: با انتخاب گزینه «تولید کن»، ویدیوی شما تولید می‌شود.
اسکرین‌شات از مرحله اعمال تنظیمات ساخت ویدیو در هوشا با شماره‌گذاری هر مرحله
بهتر است عکس مرجع برای ساخت ویدیو، واضح، با کیفیت و در سایز هماهنگ با ویدیوی درخواستی باشد.

۵. دریافت ویدیو

در پایان فرایند ساخت ویدیو با هوش مصنوعی، می‌توانید نتیجه را بررسی و در صورت نیاز، مطابق شماره‌گذاری تصویر زیر، اقدامات لازم را انجام دهید:

  1. دانلود: ویدیو را با فرمت MP4 دانلود کنید.
  2. تغییرات: متن پرامپت را دوباره در صفحه ساخت ویدیو وارد کرده و تغییرات موردنظر را اعمال کنید. این گزینه برای ویرایش‌های جزئی مناسب است.
  3. بازآفرینی: ویدیو را با همان تنظیمات قبلی دوباره تولید کنید. از آنجا که نتیجه هر بار یکسان نیست، می‌توانید چند نسخه بسازید و نمونه دلخواه را انتخاب کنید.
  4. رفرنس: ویدیو را به‌عنوان رفرنس برای ساخت ویدیوی جدید اضافه کنید تا هوش مصنوعی با استفاده از پرامپت و رفرنس، نمونه‌ای مشابه تولید کند.
  5. اشتراک‌گذاری: لینک ویدیو را کپی کرده و با دیگران به اشتراک بگذارید.

آموزش ۰ تا ۱۰۰ ساخت ویدیو با گوگل فلو

اسکرین‌شات از مرحله پایانی ساخت ویدیو در هوشا و نمایش گزینه‌ها با شماره‌گذاری
در سرویس ساخت ویدیوی هوشا، امکان دانلود خروجی تولید شده در فرمت MP4 وجود دارد.

پلتفرم فارسی هوشا (Hoosha)؛ هوش مصنوعی همه کاره

هوشا یک پلتفرم هوش مصنوعی ایرانی با بیش از ۲ میلیون کاربر است که مجموعه‌ای از مدل‌ها و ابزارهای هوش مصنوعی را در یک حساب کاربری ارائه می‌دهد. کاربران ایرانی می‌توانند بدون VPN و با امکان پرداخت ریالی و همچنین پرداخت اقساطی از درگاه اسنپ‌پی، به مدل‌ها و سرویس‌های مختلف دسترسی داشته باشند. 

سرویس‌های هوش مصنوعی هوشا

از مهم‌ترین خدمات هوشا می‌توان موارد زیر را نام برد:  

  • چت و تولید متن: دسترسی به مدل‌هایی مانند ChatGPT، Claude و Gemini 
  • تولید موسیقی: ساخت موسیقی و آهنگ با مدل‌هایی مانند  Suno
  • تولید و ویرایش تصویر: استفاده از مدل‌هایی مانند ChatGPT Image 2.5، Nano Banana 2 و  Flux Kontext
  • تولید ویدیو: دسترسی به مدل‌هایی مانند Veo، Kling، Seedance و  Runway
  • تبدیل متن به صدا و دوبله: تولید صدا به زبان‌های مختلف و هماهنگ‌سازی آن با تصویر (لیپ‌سینک)
  • ساخت وب‌سایت: ایجاد وب‌سایت با کمک هوش مصنوعی
  • ابزارها و دستیارهای تخصصی: برای کارهایی مانند ترجمه، آموزش، ایده‌پردازی و برنامه‌نویسی
  • پرامپت و قالب آماده: استفاده از نمونه‌های آماده برای تولید متن، تصویر و ویدیو

بهترین مدل‌های هوش مصنوعی تولید ویدیو طبیعی در هوشا

در هوشا امکان ساخت ویدیو با استفاده از انواع مدل‌های پیشرفته فراهم شده است. در ادامه چند مدل ساخت ویدیو واقعی را معرفی می‌کنیم: 

  1. Google Omni (Gemini Omni Flash): مدل چندوجهی Google است که متن، تصویر، صدا و ویدیو را دریافت می‌کند و در هوشا امکان آپلود ویدیو، تغییر اجسام در آن و تولید ویدیو با پرامپت و دیالوگ فارسی را فراهم می‌کند.
  2. Seedance 2.5: نسل جدید از ByteDance و نسخه ارتقایافته Seedance 2.0 است که قابلیت‌های تولید و ویرایش ویدیو با استفاده از رفرنس‌های تصویری، ویدیویی و صوتی و حفظ چهره را بهبود داده است. 
  3. Veo 3.1: مدل تولید ویدیوی Google DeepMind است که بر واقع‌گرایی و دنبال کردن دقیق پرامپت تمرکز دارد و برای تولید ویدیوهای سینمایی با صدای همگام و دیالوگ طبیعی استفاده می‌شود.
  4. Grok Imagine 1.5: مدل تولید ویدیو از کمپانی xAI است که هم با تصویر ورودی و هم بدون آن، ویدیوهای واقع‌گرایانه با صدا تولید می‌کند و در هوشا از پرامپت و دیالوگ فارسی نیز پشتیبانی می‌کند.
  5. Kling Motion: ابزاری مناسب برای انتقال حرکت، ژست و حالت چهره از ویدیوی مرجع به تصویر یک کاراکتر؛ مدت خروجی با ویدیوی مرجع برابر است و بین ۳ تا ۳۰ ثانیه خواهد بود.

راهنمای کامل تولید ویدیو با هوش مصنوعی

۳ روش دسترسی به هوشا

برای استفاده از خدمات هوشا می‌توانید از ۳ مسیر زیر اقدام کنید: 

  1. وب‌سایت: دسترسی به تمام ابزارهای تولید ویدیو، تصویر و دیگر سرویس‌ها. 
  2. ربات تلگرام: تولید عکس و ویدیو در مدل‌های مختلف هوش مصنوعی.
  3. وب‌اپ: نسخه سازگار با مرورگر گوشی‌های آیفون و اندروید با قابلیت دسترسی به تمامی ابزارها.

پرامپت ساخت ویدیو واقعی

لازم به ذکر است برای ساخت ویدیو در هوشا فقط به استایل‌های آماده محدود نیستید؛ بلکه می‌توانید ویدیوی موردنظر خود را از صفر و با یک پرامپت شخصی بسازید. 

در ادامه نیز ۴ نمونه استایل آماده قرار گرفته است که می‌توانید برای ایده گرفتن از آن‌ها استفاده کنید یا با ویرایش پرامپت‌های آماده، ویدیوی جدیدی بسازید.

۱. استایل آماده آنباکسینگ محصول

با پرامپت آماده و تصویر استوری‌بورد آنباکسینگ محصول به‌عنوان رفرنس، ویدیوی زیر را در هوشا ساخته‌ایم. این استایل برای ساخت استوری‌های تبلیغاتی مناسب است و برای طراحی استوری‌بورد نیز می‌توانید از نانو بنانا در سرویس ساخت عکس هوشا استفاده کنید.

  • مدل: Google Omni
  • ابعاد: ۹:۱۶
First-person POV ASMR unboxing video, 10 seconds, cinematic 8K quality. A dark moody studio tabletop with soft dramatic key lighting. Male hands with a dark sweater enter frame holding a black gift box covered in a colorful mosaic pattern of red, blue, green, yellow and orange squares, with a golden World Cup trophy logo visible through a window cutout.
0–2s: Hands present the box to camera, tilting it slightly so light glints off the glossy print.
2–4s: A box cutter slices the seal with a crisp cutting sound; the lid flips open revealing kraft cardboard interior.
4–6s: Hands slide out a black ceramic mug wrapped in crinkling clear plastic — loud satisfying ASMR crinkle — then peel the wrap away.
6–9s: Slow 360-degree rotation of the mug close to camera: glossy black ceramic with a colorful mosaic tile pattern and a gold trophy emblem, smooth curved handle, light reflections sweeping across the glaze.
9–10s: Final hero shot — mug placed on the dark table beside the open box, a thumb-up enters frame.
Style: shallow depth of field, macro detail on ceramic gloss and plastic texture, warm rim lighting against black background, no music, only ASMR audio (cardboard flaps, blade cut, plastic crinkle, ceramic tap). No text overlays, no faces, hands only.

۲. استایل آماده ساخت محتوای داستانی

برای شخصی‌سازی این پرامپت، مشخصات سوژه را در بخش @CHARACTER و در صورت تمایل به حمل وسیله‌ای مانند کیف، مشخصات آن را در بخش [CARRIED ITEM] وارد کنید. اگر سوژه وسیله‌ای در دست ندارد، این بخش را از متن حذف کنید.

  • مدل: Seedance 2.5
  • ابعاد: ۹:۱۶
SCENE CONTEXT @CHARACTER walks a Dalmatian puppy down a sunlit city pedestrian street, an open silver flip phone in the right hand, a takeaway coffee cup in the left, the leash looped around the left wrist. A passing pedestrian’s elbow clips their arm. At the moment of contact time slows to a near-freeze and the camera flies through the suspended accident to inspect three objects, then time snaps back and everything hits the ground.
 
ACTIVE REFERENCES
 
@CHARACTER: [age + build + hair + headwear if any + upper garment + lower garment + footwear + visible jewellery/watch]. 100% matches the reference. Ignore the reference’s backdrop, pose and framing entirely — only identity, wardrobe and proportions carry over.
 
Filled example: mid-20s woman, slim, long wavy blonde hair worn loose past the shoulders, pale-blue floral chiffon headscarf tied under the chin, cropped pale-blue tweed jacket over a white tee, light-wash cropped straight jeans with raw hems, tan leather belt, ivory pointed-toe pumps, silver watch on left wrist, fine silver necklace.
 
[CARRIED ITEM]: [bag / backpack / case, and which shoulder]. Stays on the body for the entire take. (Delete this line if the character carries nothing.)
 
@PUPPY: Dalmatian puppy, young — soft short white coat with irregular black spots across the back, flanks and legs, large solid black patches over both floppy ears, black nose, dark round eyes, slim puppy legs, long thin tail. Wears a thin pink woven collar with a small silver tag. 100% matches the reference. Walks on a fine leather leash the whole time; it is never carried, never picked up, never released.
 
@SUNGLASSES: angular cat-eye sunglasses. Glossy bright-white acetate frame with thick chunky rims and a faceted geometric outer corner, dark charcoal-grey lenses, wide flat white temple arms each carrying one large polished gold interlocking-loop metal ornament mounted on the outer face near the hinge, small gold pin rivets, white temple tips. 100% matches the reference. This is the HERO OBJECT.
 
@FLIPPHONE: brushed-silver clamshell flip phone, open. Upper half holds a bright rectangular colour display; lower half holds a raised silver numeric keypad with a round four-way navigation control, green call key and red end key. Thick cylindrical metal hinge across the middle. 100% matches the reference. The display shows an incoming call: the word BOSS in white on a dark screen, a small “Incoming call” line beneath it, and a slide-to-answer control at the bottom.
 
@COFFEECUP: takeaway coffee cup, kraft-brown ripple-textured corrugated outer wall, plain white inner lip, white base ring, white domed plastic sip lid with a small tan tab. Filled roughly three-quarters full with black coffee — the cup carries a real, heavy, visible volume of liquid and is never an empty prop. 100% matches the reference.
 
LOCATION A clean upscale European city pedestrian street on a bright late-morning. Pale stone paving, low kerb, cream and grey stone facades with shop glass on both sides, a row of small street trees on the left, parked cars far back on the right beyond the kerb. Wide open walking surface at least 4 m across. Pedestrians and the puppy stay on the paving, vehicles stay beyond the kerb, architecture keeps realistic human scale. Only two people are on the paving: @CHARACTER and the passing pedestrian.
 
FORMAT MODE One single continuous take. One motion-control camera, no internal cuts, no montage, no dissolves. The take ends on a hard cut.
 
FIRST FRAME AND SPATIAL BLOCKING First frame is already occupied: an extreme low camera at ankle height, 0.3 m above the paving, looking up the street at @CHARACTER’s footwear mid-stride, with @PUPPY’s spotted legs trotting alongside in the same low frame. The street stretches deep behind them. No empty establishing frame.
 
@CHARACTER walks toward camera down the centre of the paving. @PUPPY trots slightly ahead and to their screen-left at the end of a taut-then-slack leash, staying on the ground at all times. Camera stays ahead of them and retreats as they advance, always on the same side of the walking line. The passing pedestrian enters from the far background walking toward them on their screen-right side and passes on that side, alone and without a dog. The camera never crosses to the opposite side of the walking line at any point in the take.
 
Before contact: right hand holds @FLIPPHONE open at chest height, screen turned up toward the face. Left hand holds @COFFEECUP by the ripple wall, with the leash loop hooked around that same left wrist. @SUNGLASSES are worn on the face. @CHARACTER is not on a call and does not speak — they are looking down at the ringing screen.
 
OPTICS The camera physically travels through the scene, so framing changes by distance, not by lens tricks. FOV is assigned per beat and holds inside each beat:
 
Reveal and walk: 84° diagonal field of view, classic wide rectilinear character. Camera physically close, environment readable to the frame edges, straight architectural lines stay straight, no fisheye curve.
Frozen tableau: 84°, camera 3–4 m back.
Travel legs between objects: 84° wide rectilinear, cinematic FPV drone character — the frame is physically small and agile, passing close to surfaces, with mild edge perspective stretch and deep readable space beyond.
Three object inspections: 29° short-telephoto macro character, camera physically 0.2–0.4 m from the object, razor-thin plane of focus on the object surface, background dissolved into soft bokeh while the street, @CHARACTER and @PUPPY remain faintly recognizable behind.
Return and ground impact: back to 84°.
 
No lens drift inside a beat. Every change of subject size comes from the camera moving.
 
FPV DRONE CHARACTER (frozen section only, 5.0–12.8s). Between the objects the camera behaves like a small cinematic FPV drone threading a physical obstacle course. It never cuts and never stops moving: it flies a continuous curved path through the suspended tableau, banking into each turn with a real 10–20° roll and levelling out on arrival, gaining and losing height as it passes over and under objects, with faint organic air-drift and micro-corrections rather than perfect rail-smooth motion. It passes between and around the suspended props — the phone on one side of frame, the lid on the other, the falling body behind — so that objects it is not inspecting sweep through the near foreground as large soft out-of-focus shapes and exit frame. That foreground sweep is what proves the objects are real solids sitting in three-dimensional space. On arrival at each hero object the drone decelerates hard into a near-hover, rolls level, and holds steady for the macro read; it accelerates away again to reach the next one. Fast confident travel, calm stable holds. No cuts, no whip pans, no chaotic spin, no bobbing during the holds.
 
CAMERA AND ACTION TIMING
 
0.0–2.5s— REVEAL. Camera rises in one unbroken crane from ankle height to chest height while retreating at walking pace, revealing the body from the ground up in this order: footwear and the puppy trotting beside it, lower garment, waistline, upper garment, shoulders and any carried strap, neck and jewellery, head and any headwear, face. Land on a near-full-body frontal frame with @CHARACTER centred, @PUPPY in the lower screen-left of frame, sunglasses on, phone screen glowing in the right hand.
 
2.5–4.2s— WALK. Camera holds the frontal retreating track. @CHARACTER walks with real weight: heel contact, hip shift, toe push-off, body mass settling on each step. @PUPPY trots with loose puppy gait, ears bouncing, head turning once toward a shop window, the leash alternately going slack and taut against the left wrist. Any carried bag swings on its strap against the body. Any loose fabric — coat tails, scarf ends, hem, open jacket — trails a few centimetres behind the motion. The passing pedestrian closes distance on the screen-right side. Exactly one puppy, one leash, one passing pedestrian.
 
4.2–5.0s — CONTACT AND RELEASE. As they pass, the pedestrian’s elbow clips @CHARACTER’s upper right arm. @CHARACTER recoils and their weight goes back over their heels. The leash snaps taut and @PUPPY plants its front paws and braces against the pull, staying on the ground. Four release events, all caused by that one impact:
 
@SUNGLASSES lift off the face and travel a short distance forward and up, tumbling slowly. They do not fly at the camera.
@COFFEECUP leaves the left hand with a short outward and downward rotational impulse and begins to tilt.
The white lid pops loose from the cup rim seam under the impact and separates as its own object.
@FLIPPHONE leaves the right hand with a short outward and upward impulse, still open.
 
The instant each object separates, that hand is completely empty and stays empty for the rest of the take. Nothing appears in either hand again. The leash stays looped around the left wrist and @PUPPY stays attached to it for the entire take — the leash is not released and the puppy never leaves the ground. Any carried bag stays on the body, swinging on its strap.
 
The coffee only escapes after the lid is gone and the cup has tilted far enough for the liquid surface to cross the open rim. First a small restrained lip of dark coffee over one side of the rim, then a thin irregular sheet with a few uneven droplets separating from it.
 
The cup is never empty at any point. Most of the coffee is still inside it — a heavy dark mass pooled against the lower inner wall, sloshed up one side by the rotation, its surface tilted opposite to the cup’s spin. The escaping coffee stays physically connected to that mass as a continuous dark column leaving the rim. The amount outside the cup only ever equals the amount that has already crossed the rim, and it grows slowly. The open rim always shows dark liquid filling it, never a dry white interior. As the cup rotates further the coffee inside stays weighted toward the low side and never sits flat with the cup.
 
5.0–6.0s— FROZEN TABLEAU. Time drops to roughly 2% of normal speed. The camera pulls back to a clear medium-wide of the whole accident before any close-up. @CHARACTER is mid-fall backward, both hands empty, sunglasses gone from the face. @PUPPY is caught mid-brace with paws planted, ears lifted, one loose fold of leash suspended in the air. The four objects hang at four separate points, each near where it left the body. These positions are now fixed anchors in the world.
 
6.0–8.0s — SUNGLASSES. The camera drops out of the wide tableau and flies in on @SUNGLASSES, banking once around @CHARACTER’s raised forearm and passing under the suspended phone so the phone sweeps across the top of frame in soft focus. Decelerate hard into a macro hover. Small 25° orbit around the glasses. Read the thick glossy white acetate and its faceted outer corner, the dark lens surface with a curved street reflection sliding across it, and the polished gold ornament on the temple arm catching one hard specular highlight that travels along the metal as the camera orbits. Real material edges, tiny gold pin rivets, a hairline of dust. Hold.
 
8.0–10.0s — COFFEE. The camera leaves the glasses hanging where they are and flies diagonally across the frozen space to @COFFEECUP, banking around the suspended lid on the way in, the lid sweeping through the near foreground out of focus. Decelerate into a macro hover. Read the kraft ripple texture, the white inner lip now stained dark and wet, and — through the tilted open rim — the coffee still inside, a heavy black mass shining with one hard specular highlight, pressed up the low inner wall. The escaping coffee reads as one continuous dark column leaving the rim and breaking into an irregular sheet with a bright edge and several droplets of different sizes, all still tethered back to the liquid in the cup. Opaque, dark, warm-sheened, never transparent, never a symmetrical splash. Small lateral drift around the cup so the depth between rim, column and droplets separates. Hold.
 
10.0–12.0s— PHONE. The camera peels away from the cup, climbs, banks across the front of the falling body and flies down to @FLIPPHONE from above, the suspended coffee droplets streaking past the lower frame edge in soft focus. Level out into a hover. First establish the whole open clamshell — brushed silver, thick hinge, raised keypad — then push closer until the display fills frame and BOSS / Incoming call reads clearly. Show brushed metal grain, precise button edges, faint fingerprints, a reflection curving across the screen glass. Hold on the screen.
 
12.0–12.8s— RETURN. The drone reverses hard, climbing and pulling back out through the same suspended objects it flew in past, and settles level into the full wide tableau. All four objects are exactly where they were left, in positions consistent with where they were released. @PUPPY is still mid-brace at the end of the leash. Nothing has followed the camera.
 
12.8–15.0s— TIME RETURNS. Speed snaps back to normal. Gravity and existing momentum resume from each object’s own position. @CHARACTER completes the backward fall and lands hard on hip and forearm. @SUNGLASSES skid across stone, gold ornament flashing once. @COFFEECUP hits, crumples on one side and throws the rest of the coffee across the paving in a real irregular splash. The lid lands separately and spins flat. @FLIPPHONE hits, bounces once, and stays open and undamaged, screen still lit. The four objects land at four separate places, not clustered beside the body. @PUPPY stumbles two steps from the slack leash, then turns and trots back toward the fallen @CHARACTER. Camera drops fast to paving level and lands on a low-angle final frame connecting the fallen @CHARACTER, the puppy at their side, the white sunglasses, the spilled coffee and the lit phone screen. Hard cut.
 
PHYSICS Real mass, gravity, friction and follow-through on everything. Light acetate glasses tumble slowly, the small gold ornament giving them slightly uneven rotation. The paper cup deforms on impact; it does not shatter. Coffee is opaque, dark and non-carbonated — it moves as sheets and fat droplets with surface tension, never as a decorative symmetrical splash.
 
Liquid volume is conserved. The coffee is a single body of liquid that starts inside the cup and leaves it gradually: what is outside the cup at any moment must have physically crossed the rim, and what has not crossed the rim is still visible inside. The cup carries real liquid weight — it rotates slower and hangs heavier than an empty paper cup would, and the mass inside shifts against the low wall as it turns. An inverted or steeply tilted cup with a dry white interior and nothing pouring out is wrong.
 
@PUPPY has real light puppy mass: loose skin and ears lag behind head turns, paws slip slightly on smooth stone under the leash pull, the tail counterbalances. The leash is a real flexible strap with weight — it goes slack in curves and snaps straight under tension, and it stays attached at both ends from first frame to hard cut.
 
Clothing carries real cloth delay through the recoil and the fall: structured garments hold their shape and crease at the joints, light and loose fabrics lag behind the body and settle late. Anything tied, buttoned, belted or fastened stays fastened — knots hold, straps stay on the shoulder, nothing comes undone from the elbow contact. Any headwear stays on the head. Hair moves with inertia delay and keeps its original length, volume and silhouette; it does not reshape, restyle or change colour at any point.
 
Objects become larger in frame only because the camera closes the distance. Nothing grows, teleports, or drifts toward the lens.
 
IDENTITY The face is anatomically complete underneath the sunglasses before they separate. When the glasses leave, the same face is already there — the same eyes, lids and lashes as the reference, same brows, same nose, same cheekbones, same skin, same facial proportions. The face does not regenerate, sharpen or change at the moment of separation. Identity, wardrobe and hair stay locked from the first frame to the hard cut. @PUPPY keeps the same spot pattern, the same ear patches and the same pink collar in every shot and at every camera distance.
 
LIGHTING Bright late-morning sun from high camera-right and slightly behind @CHARACTER, throwing a warm rim along the head contour, the shoulder line and any strap, and picking out the puppy’s white coat against the paving. Pale stone paving bounces soft fill up into the face so the eyes stay readable. Shop glass on both sides gives clean specular highlights that travel across the white acetate, the gold temple ornament, the brushed silver phone and the phone screen during the macro beats. Expose so the white sunglasses and white lid keep visible surface shading and edge detail rather than blowing out to flat white; the gold reads as warm polished metal, not a bright blob. No flat frontal key, no beauty fill, no studio look.
 
AUDIO Street ambience, footsteps on stone, light claw taps and a jingling collar tag, a muffled phone ringtone before contact. At the moment of impact all sound drops to a low suspended hush. Sound returns hard on the ground impacts, with one short puppy whine at the end. No dialogue, no music, no narration, no subtitles.
 
STYLE Ultra-photorealistic live action, high-budget international fashion campaign. Real actors, a real dog, real practical props, real high-speed cinema photography with a motion-control rig. Restrained editorial colour grade, natural skin tones, subtle 35mm grain, real optical depth of field. Not CGI product animation, not game-engine render.
 
QUALITY 4K, natural 180-degree shutter cadence — real directional motion blur on the fast walk, the impact and the final ground drop, resolving to sharp readable holds during the macro inspections. Maximum skin, hair, fur, fabric and material detail. Stable identity and stable object geometry across every distance change. No ghosting, no duplicated limbs, no frame stutter.
 
POSITIVE CONSTRAINTS Exactly one @CHARACTER, one passing pedestrian, one @PUPPY, one leash, one @SUNGLASSES, one @COFFEECUP, one lid, one volume of coffee, one @FLIPPHONE. Wide, medium and macro views show the same physical objects at different camera distances — same design, same colours, same materials, same wear. Both hands stay empty from the moment of release to the end of the take, while the leash stays on the left wrist throughout.

۳. استایل آماده ساخت ویدیو واقع‌گرا با چهره افراد مشهور

این پرامپت برای ساخت ویدیوهای تبلیغاتی محصولات کوچک هنری مناسب است و امکان شخصی‌سازی کاراکتر، لباس و جزئیات محصول را فراهم می‌کند. برای هماهنگ‌سازی پرامپت با محصول خود، می‌توانید قبل از تولید از چت جی پی تی فارسی کمک بگیرید.

  • مدل: Veo 3.1
  • ابعاد: ۱۶:۹
یک عکس سینمایی هایپررئال از یک بازیگر مرد (مهران مدیری) که پشت میز چوبی کارگاه هنری خود نشسته و در فضای آرام و نور ملایم استودیو مشغول کار است. بازیگر با دقت در حال مجسمه‌سازی و رنگ‌آمیزی یک فیگور کوچک از یکی از نقش‌های معروف خودش است — مجسمه‌ای زنده و واقع‌گرایانه از شخصیت سینمایی او که روی پایه‌ای مشکی با پلاک طلاییِ نام آن نقش قرار گرفته است.


بازیگر با چهره‌ای متمرکز و کمی نوستالژیک دیده می‌شود؛ پیراهن سفید تمیزی پوشیده و آستین‌ها را تا آرنج بالا زده است. در یک دستش قلم‌موی ظریفی دارد و با دست دیگر، پایه‌ی فیگور را نگه داشته و با دقت جزئیات چهره و لباس آن را رنگ می‌زند.


فضای کارگاه پر از نور گرم بعدازظهر است که از پنجره‌ای بلند با پرتوهای غبارآلود به داخل می‌تابد. محیط آرام و هنری است — در پس‌زمینه قفسه‌هایی پر از مجسمه‌های قدیمی، نیم‌تنه‌های گچی، قلم‌موها و شیشه‌های رنگ دیده می‌شود. میز چوبی پر از بطری‌های رنگ باز شده در رنگ‌های مختلف، قلم‌موهای قرارگرفته در ظرف‌های سفالی و پالت رنگ با ترکیب رنگ‌های مخلوط است.


دوربین روی چهره‌ی بازیگر و فیگور متمرکز است، تا ارتباط احساسی میان خالق و اثرش را نشان دهد — نگاهی که هنرمند به آفریده‌ی خودش دارد. دود یا مه ملایمی در فضا پخش شده تا عمق و حس سینمایی صحنه را افزایش دهد.


نورپردازی:
نور نرم و پخش‌شده‌ی عصر طلایی، پرتوهای حجمی نور، سایه‌های لطیف، عمق میدان کم، و تون‌مپینگ سینمایی.


زاویه و ترکیب‌بندی دوربین:
نمای نزدیک متوسط (Medium Close-up)، لنز ۳۵ میلی‌متری، پس‌زمینه‌ی کمی تار (بوکه)، فوکوس روی دست‌ها و فیگور.
نسبت تصویر: افقی ۱۶:۹
نوع نما: فریم ثابت سینمایی


لحن تصویری:
رنگ‌های گرم، واقع‌گرایی احساسی، بافت‌های دقیق و جزئی، حس دست‌ساز و ترکیب‌بندی نقاشی‌گونه.


برچسب‌های سبک:
رئالیسم سینمایی، روایت احساسی، هنر دست‌ساز، آتلیه ایرانی، رابطه‌ی خالق و اثر، نور عصر طلایی، نورپردازی پرتره، حال‌و‌هوای استودیو، داستان‌گویی دراماتیک، جزئیات دقیق با کیفیت 4K PBR

۴. استایل آماده انتقال ژست

علاوه بر تولید ویدیو از متن و تصویر، برخی مدل‌ها امکان انتقال حرکت و ژست از یک ویدیوی مرجع به تصویر را نیز فراهم می‌کنند.

این استایل نیز با استفاده از یک عکس سوژه و یک ویدیوی مرجع برای انتقال ژست، بدون نیاز به وارد کردن متن پرامپت ساخته شده است. برای تولید ویدیو با تصویر خودتان، وارد لینک ویدیو شوید، گزینه «امتحان کن» را انتخاب کنید و سپس تصویر دلخواهتان را جایگزین تصویر کاراکتر کنید.

  • مدل: Kling Motion
  • ابعاد: ۱:۱

جمع‌بندی

ساخت ویدیو واقعی با هوش مصنوعی، راهی برای تولید محتوای خلاقانه و حرفه‌ای بدون نیاز به فیلم‌برداری است و برای کسب‌وکارهای کوچک و ویدیوهای تبلیغاتی کاربرد دارد. مدل‌های مختلف قابلیت‌های متفاوتی دارند و انتخاب آن‌ها به نوع ویدیوی موردنظر بستگی دارد. برای مثال، Grok Imagine 1.5 از پرامپت و دیالوگ فارسی پشتیبانی می‌کند و امکان تبدیل عکس به ویدیو را دارد، در حالی که Kling Motion برای انتقال ژست مناسب است.

در این محتوا، مراحل تولید ویدیوی واقعی در هوشا را از انتخاب استایل مناسب تا استفاده از تصویر مرجع، تنظیمات و ویرایش خروجی، قدم‌به‌قدم بررسی کردیم. حال اگر قصد دارید ساخت ویدیو با هوش مصنوعی را امتحان کنید، می‌توانید از یکی از استایل‌های آماده شروع کرده و مراحل را طبق این آموزش در هوشا پیش ببرید.

ورود به سرویس ساخت ویدیوی هوشا و بررسی امکانات مختلف آن از لینک زیر:

نویسنده: سپیده بزازی

ساخت ویدیوی واقعی با هوش مصنوعی چه تفاوتی با ویدیوهای انیمیشنی دارد؟

ویدیوی واقعی با شبیه‌سازی نور، بافت، حرکت و چهره‌ها، ظاهری شبیه فیلم‌برداری واقعی ایجاد می‌کند؛ در حالی که ویدیوهای انیمیشنی از سبک‌های تصویری غیرواقعی استفاده می‌کنند.

آیا برای ساخت ویدیوی واقعی با هوش مصنوعی به دوربین یا بازیگر نیاز است؟

خیر. با یک پرامپت متنی و در صورت نیاز، تصویر یا ویدیوی مرجع می‌توانید بدون تجهیزات فیلم‌برداری ویدیو بسازید.

کدام مدل هوش مصنوعی برای ساخت ویدیوی واقعی مناسب است؟

مدل‌هایی مانند Grok Imagine 1.5 و Veo 3.1 برای تولید ویدیوهای واقع‌گرایانه گزینه‌های مناسبی هستند. انتخاب مدل به قابلیت‌های موردنیاز پروژه، مانند صدا، دیالوگ یا استفاده از تصویر مرجع بستگی دارد.

آیا برای ساخت ویدیوی واقعی باید پرامپت را به انگلیسی بنویسیم؟

خیر. برخی مدل‌ها مانند Grok Imagine 1.5 و Google Omni از پرامپت و دیالوگ فارسی پشتیبانی می‌کنند؛ با این حال، استفاده از اصطلاحات انگلیسی تخصصی می‌تواند برای توصیف دوربین یا نورپردازی مفید باشد.

آیا می‌توان از عکس شخصی برای ساخت ویدیوی واقعی استفاده کرد؟

بله. در مدل‌هایی که از تبدیل تصویر به ویدیو پشتیبانی می‌کنند، مانند Grok Imagine 1.5، می‌توانید عکس شخصی را به‌عنوان تصویر مرجع وارد کرده و آن را به ویدیو تبدیل کنید.

آیا می‌توان از عکس محصول برای ساخت ویدیوی تبلیغاتی استفاده کرد؟

بله. می‌توانید عکس محصول را به‌عنوان رفرنس وارد کنید و نحوه حرکت دوربین، محیط و نمایش محصول را در پرامپت مشخص کنید.

آیا می‌توان به ویدیوی ساخته‌شده با هوش مصنوعی صدا و دیالوگ اضافه کرد؟

بله. مدل‌هایی مانند Veo 3.1 و Grok Imagine 1.5 امکان تولید ویدیوی همراه با صدا را دارند و برخی مدل‌ها از دیالوگ نیز پشتیبانی می‌کنند.

چگونه می‌توان ویدیوی تولید شده با هوش مصنوعی را طبیعی‌تر کرد؟

توصیف دقیق نور، حرکت، زاویه دوربین و حالت چهره در پرامپت و استفاده از تصویر یا ویدیوی مرجع می‌تواند به طبیعی‌تر شدن خروجی کمک کند.

آیا می‌توان ویدیوی تولید شده با هوش مصنوعی را ویرایش کرد؟

بله. برخی مدل‌ها مانند Google Omni و Seedance 2.5 امکان ویرایش ویدیو یا استفاده از رفرنس برای ایجاد نسخه جدید را فراهم می‌کنند.

سوالات متداول این بخش
نظرات کاربران

دیدگاهتان را بنویسید

نشانی ایمیل شما منتشر نخواهد شد. بخش‌های موردنیاز علامت‌گذاری شده‌اند *

مقالات مشابه
ساخت پوستر با هوش مصنوعی؛ معرفی بهترین ابزارها + آموزش عملی 
ساخت پوستر با هوش مصنوعی یعنی به‌جای باز کردن فتوشاپ یا سپردن کار به یک طرا…
هوشا ( ۰ امتیاز )
تحول معماری با هوش مصنوعی؛ بررسی ابزارها و خلاقیت معماران با AI
معماری به‌عنوان هنری که همواره در جستجوی نوآوری و بهبود بوده، اکنون با ظهور…
تیم تحریریه ( ۰ امتیاز )
آیا هوش مصنوعی می‌تواند نتایج فوتبال را دقیق پیش‌بینی کند؟
پیش‌بینی فوتبال با هوش مصنوعی از الگوریتم‌های پیشرفته و تحلیل داده‌های گستر…
تیم تحریریه ( ۳ امتیاز )