ساخت ویدیو واقعی با هوش مصنوعی به شما امکان میدهد تنها با یک پرامپت و تصاویر مرجع، صحنههایی طبیعی و باورپذیر را بدون نیاز به فیلمبرداری یا تجهیزات حرفهای بسازید. در این مقاله، نحوه ساخت ویدیو با هوش مصنوعی را آموزش میدهیم و نمونه ویدیوهای واقعگرا را همراه با پرامپت و نتیجه نهایی بررسی میکنیم.
در این آموزش با موارد زیر همراه ما باشید:
- آموزش مرحلهبهمرحله ساخت ویدیوی واقعی در هوشا
- معرفی مدلهای مناسب برای ساخت ویدیوی طبیعی
- ۴ نمونه استایل آماده برای ساخت ویدیوی واقعی به همراه نتیجه
ساخت ویدیو واقعی با هوش مصنوعی؛ تبدیل عکس به ویدیو
ویدیوی واقعی زیر با استفاده از استایل آماده در سرویس ساخت ویدیوی هوشا ساخته شده است؛ در ادامه، مرحلهبهمرحله توضیح میدهیم که چطور میتوانید دقیقاً همین ویدیو را بسازید.
آموزش ساخت ویدیو واقعی در پلتفرم هوشا در ۵ گام
در این بخش با استفاده از یک استایل آماده و عکس مرجع، فرایند ساخت ویدیو واقعی با هوش مصنوعی را در پلتفرم ایرانی هوشا و با مدل Grok Imagine 1.5 از ابتدا تا انتها انجام میدهیم. در ادامه، مراحل را بهصورت تصویری دنبال کنید.
۱. ورود به سرویس ساخت ویدیوی هوشا
وارد حساب کاربری خود در هوشا (Hoosha) شوید و در منوی سمت راست، از میان سرویسهای اصلی، سرویس ساخت ویدیو را انتخاب کنید.

۲. انتخاب استایل آماده
در قسمت زیرِ کادر پرامپت، از میان استایلهای آماده ساخت ویدیو، استایل دلخواه خود را انتخاب کنید. با کلیک روی گزینه «امتحان کن»، پرامپت، تصاویر مرجع و برخی تنظیمات مربوط به همان استایل مستقیماً در صفحه ساخت ویدیو اعمال میشوند و میتوانید ویرایش پرامپت را در همین بخش انجام دهید.
نکته: از قسمت «بر اساس مدل» میتوانید مدلِ هوش مصنوعی مورد نظر را انتخاب کنید تا فقط استایلهای آماده سازگار با همان مدل نمایش داده شوند.

هوش مصنوعی ساخت ویدیو ایرانی (با دیالوگ + امکانات حرفهای)
۳. ذخیره پرامپت ساخت ویدیو
با انتخاب هر یک از استایلهای آماده، صفحه زیر نمایش داده میشود. میتوانید متن پرامپت را از گزینه «کپی پرامپت» کپی کرده و در یک فایل ورد برای ویرایش و شخصیسازی، در صورت نیاز، ذخیره کنید.

متن پرامپت ویدیوی تبلیغاتی گردنبند
برای ساخت ویدیوی تبلیغاتی گردنبند، پرامپت زیر را کپی کرده و عیناً در کادر ساخت ویدیو وارد کنید. این پرامپت حفظ مشخصات گردنبند در عکس مرجع و همچنین ساخت موسیقی پسزمینه را پوشش میدهد و نیازی به آپلود موسیقی جداگانه نیست.
- مدل: Grok Imagine 1.5
- ابعاد: ۹:۱۶
| Create a 15-second ultra-premium vertical 9:16 luxury jewelry commercial from the provided image. Use the uploaded image as the exact first-frame and style foundation. The hero subject is the gold necklace and pendant worn by the model. The necklace and pendant must remain the absolute visual focus in every shot. Preserve the jewelry precisely: same pendant shape, same chain design, same gold tone, same diamond placements, same proportions, same placement on the neck, and the same elegant luxury identity. Style: high-end fine jewelry campaign, cinematic beauty ad, soft gold and champagne palette, refined editorial lighting, creamy highlights, luminous skin, shallow depth of field, polished reflections, minimal upscale environment, expensive and timeless mood. The result should feel like a multi-million-dollar global luxury jewelry advertisement. Timeline: 0.0–3.0s: Open with an intimate macro shot of the pendant resting on the model’s collarbone. Slow cinematic push-in. Soft warm light glides across the gold pendant and chain. Natural breathing only. Keep the pendant razor-sharp and dominant. 3.0–6.0s: Camera gently glides upward and slightly around the neckline and jawline while keeping the necklace centered. Add subtle focus pull from chain details to pendant. Hair remains controlled and elegant with minimal movement. 6.0–9.0s: Move into a refined beauty shot from chest to lips and lower face. The model turns very slightly with graceful confidence. The pendant catches elegant highlights and remains the hero. No distracting pose or exaggerated expression. 9.0–12.0s: Extreme premium jewelry moment: closer macro emphasis on the pendant and upper chain. Show crisp gold edges, diamond sparkle, clean reflections, and soft skin texture. Add a smooth light sweep across the pendant to make it feel iconic and luxurious. 12.0–15.0s: Final hero packshot. Pull back slightly to a balanced chest-to-face composition, with the pendant perfectly centered and fully visible. End on a poised still-like luxury frame where the jewelry feels unforgettable, radiant, and premium. Character and styling: elegant adult woman, refined editorial beauty, smooth luminous skin, tasteful natural makeup, sleek pulled-back hair, graceful posture, calm luxurious presence. Wardrobe should remain minimal, silky, and understated so the necklace stays dominant. Environment: soft upscale editorial setting with warm natural light and subtle architectural shadows, clean premium background, no clutter, no distraction, no extra props. Camera language: slow macro push-in, subtle orbit, graceful glide, precise focus pulls, stable movement, no shaky camera, no fast cuts, no random zooms. Lighting: soft premium beauty lighting, warm champagne key light, delicate rim light, crisp specular highlights on the gold pendant and diamonds, flattering skin glow, elegant contrast. Audio: sensual luxury soundtrack, soft refined music, subtle shimmer textures, polished whoosh transitions, delicate jewelry sparkle accents, no vocals, no dialogue, no harsh beats. Motion rules: all movement must be slow, graceful, realistic, and controlled. The model moves minimally. The necklace remains stable, wearable, realistic, and hero-focused. Avoid: changing the necklace design, pendant distortion, wrong chain shape, extra jewelry, cluttered styling, plastic skin, warped anatomy, distracting fashion poses, low-resolution textures, oversaturation, text overlays, captions, watermark, or background distractions. |
۴. آپلود عکس مرجع و اعمال تنظیمات ساخت ویدیو
پس از اضافه کردن پرامپت در کادر مربوط، برای آپلود تصویر مرجع و اعمال تنظیمات ویدیو، مراحل زیر را طبق شمارهگذاری تصویر انجام دهید:
- آپلود تصویر: تصویر مرجع واضح و با کیفیت بالا را برای ساخت ویدیو بارگذاری کنید.
- انتخاب مدل: مدل هوش مصنوعی موردنظر را انتخاب کنید.
- انتخاب کیفیت: کیفیت ویدیو را تعیین کنید. بدیهی است که کیفیت بالاتر به اعتبار بیشتری نیاز دارد.
- انتخاب مدتزمان: از گزینه «زمان» در بخش تنظیمات میتوانید مدت ویدیو را تعیین کنید. از آنجا که استایلهای آماده معمولاً زمان ویدیو و طول هر صحنه را مشخص میکنند، این گزینه باید با زمانبندی ذکرشده در پرامپت مطابقت داشته باشد.
- انتخاب ابعاد: بهتر است ابعاد ویدیو با تصویر مرجع هماهنگ باشد. برای مثال، اگر ویدیوی عمودی با نسبت ۹:۱۶ میخواهید، بهتر است تصویر مرجع نیز همین نسبت را داشته باشد.
- تولید ویدیو: با انتخاب گزینه «تولید کن»، ویدیوی شما تولید میشود.

۵. دریافت ویدیو
در پایان فرایند ساخت ویدیو با هوش مصنوعی، میتوانید نتیجه را بررسی و در صورت نیاز، مطابق شمارهگذاری تصویر زیر، اقدامات لازم را انجام دهید:
- دانلود: ویدیو را با فرمت MP4 دانلود کنید.
- تغییرات: متن پرامپت را دوباره در صفحه ساخت ویدیو وارد کرده و تغییرات موردنظر را اعمال کنید. این گزینه برای ویرایشهای جزئی مناسب است.
- بازآفرینی: ویدیو را با همان تنظیمات قبلی دوباره تولید کنید. از آنجا که نتیجه هر بار یکسان نیست، میتوانید چند نسخه بسازید و نمونه دلخواه را انتخاب کنید.
- رفرنس: ویدیو را بهعنوان رفرنس برای ساخت ویدیوی جدید اضافه کنید تا هوش مصنوعی با استفاده از پرامپت و رفرنس، نمونهای مشابه تولید کند.
- اشتراکگذاری: لینک ویدیو را کپی کرده و با دیگران به اشتراک بگذارید.
آموزش ۰ تا ۱۰۰ ساخت ویدیو با گوگل فلو

پلتفرم فارسی هوشا (Hoosha)؛ هوش مصنوعی همه کاره
هوشا یک پلتفرم هوش مصنوعی ایرانی با بیش از ۲ میلیون کاربر است که مجموعهای از مدلها و ابزارهای هوش مصنوعی را در یک حساب کاربری ارائه میدهد. کاربران ایرانی میتوانند بدون VPN و با امکان پرداخت ریالی و همچنین پرداخت اقساطی از درگاه اسنپپی، به مدلها و سرویسهای مختلف دسترسی داشته باشند.
سرویسهای هوش مصنوعی هوشا
از مهمترین خدمات هوشا میتوان موارد زیر را نام برد:
- چت و تولید متن: دسترسی به مدلهایی مانند ChatGPT، Claude و Gemini
- تولید موسیقی: ساخت موسیقی و آهنگ با مدلهایی مانند Suno
- تولید و ویرایش تصویر: استفاده از مدلهایی مانند ChatGPT Image 2.5، Nano Banana 2 و Flux Kontext
- تولید ویدیو: دسترسی به مدلهایی مانند Veo، Kling، Seedance و Runway
- تبدیل متن به صدا و دوبله: تولید صدا به زبانهای مختلف و هماهنگسازی آن با تصویر (لیپسینک)
- ساخت وبسایت: ایجاد وبسایت با کمک هوش مصنوعی
- ابزارها و دستیارهای تخصصی: برای کارهایی مانند ترجمه، آموزش، ایدهپردازی و برنامهنویسی
- پرامپت و قالب آماده: استفاده از نمونههای آماده برای تولید متن، تصویر و ویدیو
بهترین مدلهای هوش مصنوعی تولید ویدیو طبیعی در هوشا
در هوشا امکان ساخت ویدیو با استفاده از انواع مدلهای پیشرفته فراهم شده است. در ادامه چند مدل ساخت ویدیو واقعی را معرفی میکنیم:
- Google Omni (Gemini Omni Flash): مدل چندوجهی Google است که متن، تصویر، صدا و ویدیو را دریافت میکند و در هوشا امکان آپلود ویدیو، تغییر اجسام در آن و تولید ویدیو با پرامپت و دیالوگ فارسی را فراهم میکند.
- Seedance 2.5: نسل جدید از ByteDance و نسخه ارتقایافته Seedance 2.0 است که قابلیتهای تولید و ویرایش ویدیو با استفاده از رفرنسهای تصویری، ویدیویی و صوتی و حفظ چهره را بهبود داده است.
- Veo 3.1: مدل تولید ویدیوی Google DeepMind است که بر واقعگرایی و دنبال کردن دقیق پرامپت تمرکز دارد و برای تولید ویدیوهای سینمایی با صدای همگام و دیالوگ طبیعی استفاده میشود.
- Grok Imagine 1.5: مدل تولید ویدیو از کمپانی xAI است که هم با تصویر ورودی و هم بدون آن، ویدیوهای واقعگرایانه با صدا تولید میکند و در هوشا از پرامپت و دیالوگ فارسی نیز پشتیبانی میکند.
- Kling Motion: ابزاری مناسب برای انتقال حرکت، ژست و حالت چهره از ویدیوی مرجع به تصویر یک کاراکتر؛ مدت خروجی با ویدیوی مرجع برابر است و بین ۳ تا ۳۰ ثانیه خواهد بود.
راهنمای کامل تولید ویدیو با هوش مصنوعی
۳ روش دسترسی به هوشا
برای استفاده از خدمات هوشا میتوانید از ۳ مسیر زیر اقدام کنید:
- وبسایت: دسترسی به تمام ابزارهای تولید ویدیو، تصویر و دیگر سرویسها.
- ربات تلگرام: تولید عکس و ویدیو در مدلهای مختلف هوش مصنوعی.
- وباپ: نسخه سازگار با مرورگر گوشیهای آیفون و اندروید با قابلیت دسترسی به تمامی ابزارها.
پرامپت ساخت ویدیو واقعی
لازم به ذکر است برای ساخت ویدیو در هوشا فقط به استایلهای آماده محدود نیستید؛ بلکه میتوانید ویدیوی موردنظر خود را از صفر و با یک پرامپت شخصی بسازید.
در ادامه نیز ۴ نمونه استایل آماده قرار گرفته است که میتوانید برای ایده گرفتن از آنها استفاده کنید یا با ویرایش پرامپتهای آماده، ویدیوی جدیدی بسازید.
۱. استایل آماده آنباکسینگ محصول
با پرامپت آماده و تصویر استوریبورد آنباکسینگ محصول بهعنوان رفرنس، ویدیوی زیر را در هوشا ساختهایم. این استایل برای ساخت استوریهای تبلیغاتی مناسب است و برای طراحی استوریبورد نیز میتوانید از نانو بنانا در سرویس ساخت عکس هوشا استفاده کنید.
- مدل: Google Omni
- ابعاد: ۹:۱۶
| First-person POV ASMR unboxing video, 10 seconds, cinematic 8K quality. A dark moody studio tabletop with soft dramatic key lighting. Male hands with a dark sweater enter frame holding a black gift box covered in a colorful mosaic pattern of red, blue, green, yellow and orange squares, with a golden World Cup trophy logo visible through a window cutout. 0–2s: Hands present the box to camera, tilting it slightly so light glints off the glossy print. 2–4s: A box cutter slices the seal with a crisp cutting sound; the lid flips open revealing kraft cardboard interior. 4–6s: Hands slide out a black ceramic mug wrapped in crinkling clear plastic — loud satisfying ASMR crinkle — then peel the wrap away. 6–9s: Slow 360-degree rotation of the mug close to camera: glossy black ceramic with a colorful mosaic tile pattern and a gold trophy emblem, smooth curved handle, light reflections sweeping across the glaze. 9–10s: Final hero shot — mug placed on the dark table beside the open box, a thumb-up enters frame. Style: shallow depth of field, macro detail on ceramic gloss and plastic texture, warm rim lighting against black background, no music, only ASMR audio (cardboard flaps, blade cut, plastic crinkle, ceramic tap). No text overlays, no faces, hands only. |
۲. استایل آماده ساخت محتوای داستانی
برای شخصیسازی این پرامپت، مشخصات سوژه را در بخش @CHARACTER و در صورت تمایل به حمل وسیلهای مانند کیف، مشخصات آن را در بخش [CARRIED ITEM] وارد کنید. اگر سوژه وسیلهای در دست ندارد، این بخش را از متن حذف کنید.
- مدل: Seedance 2.5
- ابعاد: ۹:۱۶
| SCENE CONTEXT @CHARACTER walks a Dalmatian puppy down a sunlit city pedestrian street, an open silver flip phone in the right hand, a takeaway coffee cup in the left, the leash looped around the left wrist. A passing pedestrian’s elbow clips their arm. At the moment of contact time slows to a near-freeze and the camera flies through the suspended accident to inspect three objects, then time snaps back and everything hits the ground. ACTIVE REFERENCES @CHARACTER: [age + build + hair + headwear if any + upper garment + lower garment + footwear + visible jewellery/watch]. 100% matches the reference. Ignore the reference’s backdrop, pose and framing entirely — only identity, wardrobe and proportions carry over. Filled example: mid-20s woman, slim, long wavy blonde hair worn loose past the shoulders, pale-blue floral chiffon headscarf tied under the chin, cropped pale-blue tweed jacket over a white tee, light-wash cropped straight jeans with raw hems, tan leather belt, ivory pointed-toe pumps, silver watch on left wrist, fine silver necklace. [CARRIED ITEM]: [bag / backpack / case, and which shoulder]. Stays on the body for the entire take. (Delete this line if the character carries nothing.) @PUPPY: Dalmatian puppy, young — soft short white coat with irregular black spots across the back, flanks and legs, large solid black patches over both floppy ears, black nose, dark round eyes, slim puppy legs, long thin tail. Wears a thin pink woven collar with a small silver tag. 100% matches the reference. Walks on a fine leather leash the whole time; it is never carried, never picked up, never released. @SUNGLASSES: angular cat-eye sunglasses. Glossy bright-white acetate frame with thick chunky rims and a faceted geometric outer corner, dark charcoal-grey lenses, wide flat white temple arms each carrying one large polished gold interlocking-loop metal ornament mounted on the outer face near the hinge, small gold pin rivets, white temple tips. 100% matches the reference. This is the HERO OBJECT. @FLIPPHONE: brushed-silver clamshell flip phone, open. Upper half holds a bright rectangular colour display; lower half holds a raised silver numeric keypad with a round four-way navigation control, green call key and red end key. Thick cylindrical metal hinge across the middle. 100% matches the reference. The display shows an incoming call: the word BOSS in white on a dark screen, a small “Incoming call” line beneath it, and a slide-to-answer control at the bottom. @COFFEECUP: takeaway coffee cup, kraft-brown ripple-textured corrugated outer wall, plain white inner lip, white base ring, white domed plastic sip lid with a small tan tab. Filled roughly three-quarters full with black coffee — the cup carries a real, heavy, visible volume of liquid and is never an empty prop. 100% matches the reference. LOCATION A clean upscale European city pedestrian street on a bright late-morning. Pale stone paving, low kerb, cream and grey stone facades with shop glass on both sides, a row of small street trees on the left, parked cars far back on the right beyond the kerb. Wide open walking surface at least 4 m across. Pedestrians and the puppy stay on the paving, vehicles stay beyond the kerb, architecture keeps realistic human scale. Only two people are on the paving: @CHARACTER and the passing pedestrian. FORMAT MODE One single continuous take. One motion-control camera, no internal cuts, no montage, no dissolves. The take ends on a hard cut. FIRST FRAME AND SPATIAL BLOCKING First frame is already occupied: an extreme low camera at ankle height, 0.3 m above the paving, looking up the street at @CHARACTER’s footwear mid-stride, with @PUPPY’s spotted legs trotting alongside in the same low frame. The street stretches deep behind them. No empty establishing frame. @CHARACTER walks toward camera down the centre of the paving. @PUPPY trots slightly ahead and to their screen-left at the end of a taut-then-slack leash, staying on the ground at all times. Camera stays ahead of them and retreats as they advance, always on the same side of the walking line. The passing pedestrian enters from the far background walking toward them on their screen-right side and passes on that side, alone and without a dog. The camera never crosses to the opposite side of the walking line at any point in the take. Before contact: right hand holds @FLIPPHONE open at chest height, screen turned up toward the face. Left hand holds @COFFEECUP by the ripple wall, with the leash loop hooked around that same left wrist. @SUNGLASSES are worn on the face. @CHARACTER is not on a call and does not speak — they are looking down at the ringing screen. OPTICS The camera physically travels through the scene, so framing changes by distance, not by lens tricks. FOV is assigned per beat and holds inside each beat: Reveal and walk: 84° diagonal field of view, classic wide rectilinear character. Camera physically close, environment readable to the frame edges, straight architectural lines stay straight, no fisheye curve. Frozen tableau: 84°, camera 3–4 m back. Travel legs between objects: 84° wide rectilinear, cinematic FPV drone character — the frame is physically small and agile, passing close to surfaces, with mild edge perspective stretch and deep readable space beyond. Three object inspections: 29° short-telephoto macro character, camera physically 0.2–0.4 m from the object, razor-thin plane of focus on the object surface, background dissolved into soft bokeh while the street, @CHARACTER and @PUPPY remain faintly recognizable behind. Return and ground impact: back to 84°. No lens drift inside a beat. Every change of subject size comes from the camera moving. FPV DRONE CHARACTER (frozen section only, 5.0–12.8s). Between the objects the camera behaves like a small cinematic FPV drone threading a physical obstacle course. It never cuts and never stops moving: it flies a continuous curved path through the suspended tableau, banking into each turn with a real 10–20° roll and levelling out on arrival, gaining and losing height as it passes over and under objects, with faint organic air-drift and micro-corrections rather than perfect rail-smooth motion. It passes between and around the suspended props — the phone on one side of frame, the lid on the other, the falling body behind — so that objects it is not inspecting sweep through the near foreground as large soft out-of-focus shapes and exit frame. That foreground sweep is what proves the objects are real solids sitting in three-dimensional space. On arrival at each hero object the drone decelerates hard into a near-hover, rolls level, and holds steady for the macro read; it accelerates away again to reach the next one. Fast confident travel, calm stable holds. No cuts, no whip pans, no chaotic spin, no bobbing during the holds. CAMERA AND ACTION TIMING 0.0–2.5s— REVEAL. Camera rises in one unbroken crane from ankle height to chest height while retreating at walking pace, revealing the body from the ground up in this order: footwear and the puppy trotting beside it, lower garment, waistline, upper garment, shoulders and any carried strap, neck and jewellery, head and any headwear, face. Land on a near-full-body frontal frame with @CHARACTER centred, @PUPPY in the lower screen-left of frame, sunglasses on, phone screen glowing in the right hand. 2.5–4.2s— WALK. Camera holds the frontal retreating track. @CHARACTER walks with real weight: heel contact, hip shift, toe push-off, body mass settling on each step. @PUPPY trots with loose puppy gait, ears bouncing, head turning once toward a shop window, the leash alternately going slack and taut against the left wrist. Any carried bag swings on its strap against the body. Any loose fabric — coat tails, scarf ends, hem, open jacket — trails a few centimetres behind the motion. The passing pedestrian closes distance on the screen-right side. Exactly one puppy, one leash, one passing pedestrian. 4.2–5.0s — CONTACT AND RELEASE. As they pass, the pedestrian’s elbow clips @CHARACTER’s upper right arm. @CHARACTER recoils and their weight goes back over their heels. The leash snaps taut and @PUPPY plants its front paws and braces against the pull, staying on the ground. Four release events, all caused by that one impact: @SUNGLASSES lift off the face and travel a short distance forward and up, tumbling slowly. They do not fly at the camera. @COFFEECUP leaves the left hand with a short outward and downward rotational impulse and begins to tilt. The white lid pops loose from the cup rim seam under the impact and separates as its own object. @FLIPPHONE leaves the right hand with a short outward and upward impulse, still open. The instant each object separates, that hand is completely empty and stays empty for the rest of the take. Nothing appears in either hand again. The leash stays looped around the left wrist and @PUPPY stays attached to it for the entire take — the leash is not released and the puppy never leaves the ground. Any carried bag stays on the body, swinging on its strap. The coffee only escapes after the lid is gone and the cup has tilted far enough for the liquid surface to cross the open rim. First a small restrained lip of dark coffee over one side of the rim, then a thin irregular sheet with a few uneven droplets separating from it. The cup is never empty at any point. Most of the coffee is still inside it — a heavy dark mass pooled against the lower inner wall, sloshed up one side by the rotation, its surface tilted opposite to the cup’s spin. The escaping coffee stays physically connected to that mass as a continuous dark column leaving the rim. The amount outside the cup only ever equals the amount that has already crossed the rim, and it grows slowly. The open rim always shows dark liquid filling it, never a dry white interior. As the cup rotates further the coffee inside stays weighted toward the low side and never sits flat with the cup. 5.0–6.0s— FROZEN TABLEAU. Time drops to roughly 2% of normal speed. The camera pulls back to a clear medium-wide of the whole accident before any close-up. @CHARACTER is mid-fall backward, both hands empty, sunglasses gone from the face. @PUPPY is caught mid-brace with paws planted, ears lifted, one loose fold of leash suspended in the air. The four objects hang at four separate points, each near where it left the body. These positions are now fixed anchors in the world. 6.0–8.0s — SUNGLASSES. The camera drops out of the wide tableau and flies in on @SUNGLASSES, banking once around @CHARACTER’s raised forearm and passing under the suspended phone so the phone sweeps across the top of frame in soft focus. Decelerate hard into a macro hover. Small 25° orbit around the glasses. Read the thick glossy white acetate and its faceted outer corner, the dark lens surface with a curved street reflection sliding across it, and the polished gold ornament on the temple arm catching one hard specular highlight that travels along the metal as the camera orbits. Real material edges, tiny gold pin rivets, a hairline of dust. Hold. 8.0–10.0s — COFFEE. The camera leaves the glasses hanging where they are and flies diagonally across the frozen space to @COFFEECUP, banking around the suspended lid on the way in, the lid sweeping through the near foreground out of focus. Decelerate into a macro hover. Read the kraft ripple texture, the white inner lip now stained dark and wet, and — through the tilted open rim — the coffee still inside, a heavy black mass shining with one hard specular highlight, pressed up the low inner wall. The escaping coffee reads as one continuous dark column leaving the rim and breaking into an irregular sheet with a bright edge and several droplets of different sizes, all still tethered back to the liquid in the cup. Opaque, dark, warm-sheened, never transparent, never a symmetrical splash. Small lateral drift around the cup so the depth between rim, column and droplets separates. Hold. 10.0–12.0s— PHONE. The camera peels away from the cup, climbs, banks across the front of the falling body and flies down to @FLIPPHONE from above, the suspended coffee droplets streaking past the lower frame edge in soft focus. Level out into a hover. First establish the whole open clamshell — brushed silver, thick hinge, raised keypad — then push closer until the display fills frame and BOSS / Incoming call reads clearly. Show brushed metal grain, precise button edges, faint fingerprints, a reflection curving across the screen glass. Hold on the screen. 12.0–12.8s— RETURN. The drone reverses hard, climbing and pulling back out through the same suspended objects it flew in past, and settles level into the full wide tableau. All four objects are exactly where they were left, in positions consistent with where they were released. @PUPPY is still mid-brace at the end of the leash. Nothing has followed the camera. 12.8–15.0s— TIME RETURNS. Speed snaps back to normal. Gravity and existing momentum resume from each object’s own position. @CHARACTER completes the backward fall and lands hard on hip and forearm. @SUNGLASSES skid across stone, gold ornament flashing once. @COFFEECUP hits, crumples on one side and throws the rest of the coffee across the paving in a real irregular splash. The lid lands separately and spins flat. @FLIPPHONE hits, bounces once, and stays open and undamaged, screen still lit. The four objects land at four separate places, not clustered beside the body. @PUPPY stumbles two steps from the slack leash, then turns and trots back toward the fallen @CHARACTER. Camera drops fast to paving level and lands on a low-angle final frame connecting the fallen @CHARACTER, the puppy at their side, the white sunglasses, the spilled coffee and the lit phone screen. Hard cut. PHYSICS Real mass, gravity, friction and follow-through on everything. Light acetate glasses tumble slowly, the small gold ornament giving them slightly uneven rotation. The paper cup deforms on impact; it does not shatter. Coffee is opaque, dark and non-carbonated — it moves as sheets and fat droplets with surface tension, never as a decorative symmetrical splash. Liquid volume is conserved. The coffee is a single body of liquid that starts inside the cup and leaves it gradually: what is outside the cup at any moment must have physically crossed the rim, and what has not crossed the rim is still visible inside. The cup carries real liquid weight — it rotates slower and hangs heavier than an empty paper cup would, and the mass inside shifts against the low wall as it turns. An inverted or steeply tilted cup with a dry white interior and nothing pouring out is wrong. @PUPPY has real light puppy mass: loose skin and ears lag behind head turns, paws slip slightly on smooth stone under the leash pull, the tail counterbalances. The leash is a real flexible strap with weight — it goes slack in curves and snaps straight under tension, and it stays attached at both ends from first frame to hard cut. Clothing carries real cloth delay through the recoil and the fall: structured garments hold their shape and crease at the joints, light and loose fabrics lag behind the body and settle late. Anything tied, buttoned, belted or fastened stays fastened — knots hold, straps stay on the shoulder, nothing comes undone from the elbow contact. Any headwear stays on the head. Hair moves with inertia delay and keeps its original length, volume and silhouette; it does not reshape, restyle or change colour at any point. Objects become larger in frame only because the camera closes the distance. Nothing grows, teleports, or drifts toward the lens. IDENTITY The face is anatomically complete underneath the sunglasses before they separate. When the glasses leave, the same face is already there — the same eyes, lids and lashes as the reference, same brows, same nose, same cheekbones, same skin, same facial proportions. The face does not regenerate, sharpen or change at the moment of separation. Identity, wardrobe and hair stay locked from the first frame to the hard cut. @PUPPY keeps the same spot pattern, the same ear patches and the same pink collar in every shot and at every camera distance. LIGHTING Bright late-morning sun from high camera-right and slightly behind @CHARACTER, throwing a warm rim along the head contour, the shoulder line and any strap, and picking out the puppy’s white coat against the paving. Pale stone paving bounces soft fill up into the face so the eyes stay readable. Shop glass on both sides gives clean specular highlights that travel across the white acetate, the gold temple ornament, the brushed silver phone and the phone screen during the macro beats. Expose so the white sunglasses and white lid keep visible surface shading and edge detail rather than blowing out to flat white; the gold reads as warm polished metal, not a bright blob. No flat frontal key, no beauty fill, no studio look. AUDIO Street ambience, footsteps on stone, light claw taps and a jingling collar tag, a muffled phone ringtone before contact. At the moment of impact all sound drops to a low suspended hush. Sound returns hard on the ground impacts, with one short puppy whine at the end. No dialogue, no music, no narration, no subtitles. STYLE Ultra-photorealistic live action, high-budget international fashion campaign. Real actors, a real dog, real practical props, real high-speed cinema photography with a motion-control rig. Restrained editorial colour grade, natural skin tones, subtle 35mm grain, real optical depth of field. Not CGI product animation, not game-engine render. QUALITY 4K, natural 180-degree shutter cadence — real directional motion blur on the fast walk, the impact and the final ground drop, resolving to sharp readable holds during the macro inspections. Maximum skin, hair, fur, fabric and material detail. Stable identity and stable object geometry across every distance change. No ghosting, no duplicated limbs, no frame stutter. POSITIVE CONSTRAINTS Exactly one @CHARACTER, one passing pedestrian, one @PUPPY, one leash, one @SUNGLASSES, one @COFFEECUP, one lid, one volume of coffee, one @FLIPPHONE. Wide, medium and macro views show the same physical objects at different camera distances — same design, same colours, same materials, same wear. Both hands stay empty from the moment of release to the end of the take, while the leash stays on the left wrist throughout. |
۳. استایل آماده ساخت ویدیو واقعگرا با چهره افراد مشهور
این پرامپت برای ساخت ویدیوهای تبلیغاتی محصولات کوچک هنری مناسب است و امکان شخصیسازی کاراکتر، لباس و جزئیات محصول را فراهم میکند. برای هماهنگسازی پرامپت با محصول خود، میتوانید قبل از تولید از چت جی پی تی فارسی کمک بگیرید.
- مدل: Veo 3.1
- ابعاد: ۱۶:۹
| یک عکس سینمایی هایپررئال از یک بازیگر مرد (مهران مدیری) که پشت میز چوبی کارگاه هنری خود نشسته و در فضای آرام و نور ملایم استودیو مشغول کار است. بازیگر با دقت در حال مجسمهسازی و رنگآمیزی یک فیگور کوچک از یکی از نقشهای معروف خودش است — مجسمهای زنده و واقعگرایانه از شخصیت سینمایی او که روی پایهای مشکی با پلاک طلاییِ نام آن نقش قرار گرفته است. بازیگر با چهرهای متمرکز و کمی نوستالژیک دیده میشود؛ پیراهن سفید تمیزی پوشیده و آستینها را تا آرنج بالا زده است. در یک دستش قلمموی ظریفی دارد و با دست دیگر، پایهی فیگور را نگه داشته و با دقت جزئیات چهره و لباس آن را رنگ میزند. فضای کارگاه پر از نور گرم بعدازظهر است که از پنجرهای بلند با پرتوهای غبارآلود به داخل میتابد. محیط آرام و هنری است — در پسزمینه قفسههایی پر از مجسمههای قدیمی، نیمتنههای گچی، قلمموها و شیشههای رنگ دیده میشود. میز چوبی پر از بطریهای رنگ باز شده در رنگهای مختلف، قلمموهای قرارگرفته در ظرفهای سفالی و پالت رنگ با ترکیب رنگهای مخلوط است. دوربین روی چهرهی بازیگر و فیگور متمرکز است، تا ارتباط احساسی میان خالق و اثرش را نشان دهد — نگاهی که هنرمند به آفریدهی خودش دارد. دود یا مه ملایمی در فضا پخش شده تا عمق و حس سینمایی صحنه را افزایش دهد. نورپردازی: نور نرم و پخششدهی عصر طلایی، پرتوهای حجمی نور، سایههای لطیف، عمق میدان کم، و تونمپینگ سینمایی. زاویه و ترکیببندی دوربین: نمای نزدیک متوسط (Medium Close-up)، لنز ۳۵ میلیمتری، پسزمینهی کمی تار (بوکه)، فوکوس روی دستها و فیگور. نسبت تصویر: افقی ۱۶:۹ نوع نما: فریم ثابت سینمایی لحن تصویری: رنگهای گرم، واقعگرایی احساسی، بافتهای دقیق و جزئی، حس دستساز و ترکیببندی نقاشیگونه. برچسبهای سبک: رئالیسم سینمایی، روایت احساسی، هنر دستساز، آتلیه ایرانی، رابطهی خالق و اثر، نور عصر طلایی، نورپردازی پرتره، حالوهوای استودیو، داستانگویی دراماتیک، جزئیات دقیق با کیفیت 4K PBR |
۴. استایل آماده انتقال ژست
علاوه بر تولید ویدیو از متن و تصویر، برخی مدلها امکان انتقال حرکت و ژست از یک ویدیوی مرجع به تصویر را نیز فراهم میکنند.
این استایل نیز با استفاده از یک عکس سوژه و یک ویدیوی مرجع برای انتقال ژست، بدون نیاز به وارد کردن متن پرامپت ساخته شده است. برای تولید ویدیو با تصویر خودتان، وارد لینک ویدیو شوید، گزینه «امتحان کن» را انتخاب کنید و سپس تصویر دلخواهتان را جایگزین تصویر کاراکتر کنید.
- مدل: Kling Motion
- ابعاد: ۱:۱
جمعبندی
ساخت ویدیو واقعی با هوش مصنوعی، راهی برای تولید محتوای خلاقانه و حرفهای بدون نیاز به فیلمبرداری است و برای کسبوکارهای کوچک و ویدیوهای تبلیغاتی کاربرد دارد. مدلهای مختلف قابلیتهای متفاوتی دارند و انتخاب آنها به نوع ویدیوی موردنظر بستگی دارد. برای مثال، Grok Imagine 1.5 از پرامپت و دیالوگ فارسی پشتیبانی میکند و امکان تبدیل عکس به ویدیو را دارد، در حالی که Kling Motion برای انتقال ژست مناسب است.
در این محتوا، مراحل تولید ویدیوی واقعی در هوشا را از انتخاب استایل مناسب تا استفاده از تصویر مرجع، تنظیمات و ویرایش خروجی، قدمبهقدم بررسی کردیم. حال اگر قصد دارید ساخت ویدیو با هوش مصنوعی را امتحان کنید، میتوانید از یکی از استایلهای آماده شروع کرده و مراحل را طبق این آموزش در هوشا پیش ببرید.
ورود به سرویس ساخت ویدیوی هوشا و بررسی امکانات مختلف آن از لینک زیر:
نویسنده: سپیده بزازی
ساخت ویدیوی واقعی با هوش مصنوعی چه تفاوتی با ویدیوهای انیمیشنی دارد؟
ویدیوی واقعی با شبیهسازی نور، بافت، حرکت و چهرهها، ظاهری شبیه فیلمبرداری واقعی ایجاد میکند؛ در حالی که ویدیوهای انیمیشنی از سبکهای تصویری غیرواقعی استفاده میکنند.
آیا برای ساخت ویدیوی واقعی با هوش مصنوعی به دوربین یا بازیگر نیاز است؟
خیر. با یک پرامپت متنی و در صورت نیاز، تصویر یا ویدیوی مرجع میتوانید بدون تجهیزات فیلمبرداری ویدیو بسازید.
کدام مدل هوش مصنوعی برای ساخت ویدیوی واقعی مناسب است؟
مدلهایی مانند Grok Imagine 1.5 و Veo 3.1 برای تولید ویدیوهای واقعگرایانه گزینههای مناسبی هستند. انتخاب مدل به قابلیتهای موردنیاز پروژه، مانند صدا، دیالوگ یا استفاده از تصویر مرجع بستگی دارد.
آیا برای ساخت ویدیوی واقعی باید پرامپت را به انگلیسی بنویسیم؟
خیر. برخی مدلها مانند Grok Imagine 1.5 و Google Omni از پرامپت و دیالوگ فارسی پشتیبانی میکنند؛ با این حال، استفاده از اصطلاحات انگلیسی تخصصی میتواند برای توصیف دوربین یا نورپردازی مفید باشد.
آیا میتوان از عکس شخصی برای ساخت ویدیوی واقعی استفاده کرد؟
بله. در مدلهایی که از تبدیل تصویر به ویدیو پشتیبانی میکنند، مانند Grok Imagine 1.5، میتوانید عکس شخصی را بهعنوان تصویر مرجع وارد کرده و آن را به ویدیو تبدیل کنید.
آیا میتوان از عکس محصول برای ساخت ویدیوی تبلیغاتی استفاده کرد؟
بله. میتوانید عکس محصول را بهعنوان رفرنس وارد کنید و نحوه حرکت دوربین، محیط و نمایش محصول را در پرامپت مشخص کنید.
آیا میتوان به ویدیوی ساختهشده با هوش مصنوعی صدا و دیالوگ اضافه کرد؟
بله. مدلهایی مانند Veo 3.1 و Grok Imagine 1.5 امکان تولید ویدیوی همراه با صدا را دارند و برخی مدلها از دیالوگ نیز پشتیبانی میکنند.
چگونه میتوان ویدیوی تولید شده با هوش مصنوعی را طبیعیتر کرد؟
توصیف دقیق نور، حرکت، زاویه دوربین و حالت چهره در پرامپت و استفاده از تصویر یا ویدیوی مرجع میتواند به طبیعیتر شدن خروجی کمک کند.
آیا میتوان ویدیوی تولید شده با هوش مصنوعی را ویرایش کرد؟
بله. برخی مدلها مانند Google Omni و Seedance 2.5 امکان ویرایش ویدیو یا استفاده از رفرنس برای ایجاد نسخه جدید را فراهم میکنند.
آیا سرعت پیشرفت هوش مصنوعی باید کاهش یابد؟ آنتروپیک، متا و Open AI چه میگویند؟
تعبیر خواب با هوش مصنوعی رایگان؛ آموزش با ۳ مثال واقعی
ساخت ویدیو واقعی با هوش مصنوعی؛ آموزش کامل از عکس تا ویدیو حرفهای
Grok Voice Transcribe 2.0؛ قابلیتها و کاربردهای مدل صوتی جدید xAI گراک
ساخت استوری با هوش مصنوعی؛ معرفی ابزارها + آموزش عملی با مثال