Rap Duo Prompts

You saw the viral orange booth trend and want to create your own clip without guessing the words. Copy 26 creator AI rap video prompts to turn two photos into a 12-second rap video in the style of the prompt you pick, or try our rap duo generator to start with no prompt to edit.

26 prompts · 23 creators · public posts on X

Turn Two Photos into Music Videos with Rapduo's Rap Duo Prompts

0 / 4,000

Photo 1 is what a prompt calls @Image1, image_1 or “the uploaded reference image”. Photo 2 is @Image2. Your clip is @Video1.

Photo 1 stands on the left and raps first. Both are optional.

Video shape
12s · 480p

Adults (18+) only. Upload only photos of people who agreed to be in the video.

Inputs
Made with

Hotel Lobby swaps2

Swap two faces into a performance clip. These prompts expect the source clip and a photo for each side.

Made with Seedance 2.529s

by @OriSilver5K viewsOriginal post

Hotel Lobby swap: same clip, two new faces

A video-to-video swap on Seedance 2.5: the source clip goes in as the motion reference and two photos replace the two performers, one per side.

  • Needs a clip
  • Two faces
  • Video-to-video
  • Left / right

You bring Source clip (@Video1) · Left performer photo (@Image1) · Right performer photo (@Image2)

Video prompt177 chars
Exactly the same clip as @Video1 but left guy singing is @Image1 right guy in clip is @Image2 every movement every gesture everything including sound is exactly the same

Why it works

  • Each photo is tied to a side — the left performer is @Image1, the right is @Image2 — so the faces do not trade places.
  • "Every movement every gesture everything including sound is exactly the same" tells the model to copy the clip rather than reinterpret it.
  • The creator notes in the thread that a clip with its original song can be refused, so they upload it muted and put the audio back with a separate lip-sync step.
Prompt source
28s

by @yesand_ai117 viewsOriginal post

Hotel Lobby template swap in one line

The shortest version of the trend: the template supplies the clip, and the prompt only says which photo takes which side.

  • Needs a clip
  • Two faces
  • Template
  • One line

You bring The app’s Hotel Lobby template clip · Two photos (@image1 left, @image2 right)

Video prompt74 chars
swap the left character to [ @image1 ]  and right character to [ @image2 ]

Why it works

  • Brackets around each image tag keep the swap instruction readable when the app inserts the photos inline.
  • Left and right are named explicitly, the placement most swaps get wrong.

Duos and rap battles8

Two or more rappers trading lines, from one-line battles to scripted back-and-forths.

Made with Seedance 2.021s

by @Just_sharon786.2K viewsOriginal post

Rapper and her hype man by a red car

Two photos, two roles: she raps beside the car while he sits on the trunk riding the beat, and each gets a vocal profile, a signature tic and eye-contact rules.

  • Needs photos
  • Two performers
  • Signature tics
  • Accent direction

You bring Rapper photo (Image 1) · Second performer photo (Image 2)

Video prompt2,628 chars
A confident white female rapper @[Image 1](image_1). Blaze hair styled in double buns pulled up high. Pink Zipp Republic jersey, layered gold chains and cross necklaces, gold star drop earrings, small cross tattoo under the left eye. Standing firmly beside the car's rear quarter panel, shifting weight from heels to toes in time with the beat. Using both hands to express the rhythm, with gestures like cutting the air, pointing fingers, or stacking bars in mid-air.
Vocal profile: A white woman speaking in a London accent with a Nigerian lilt. Low chest-voice range, crisp consonants, punching the end of each bar with forceful delivery. Between bars, drop the volume and mumble low.
Signature tic: At the hook line's entry, thrust both hands forward simultaneously for a double point. But at the moment a bar lands clean, drop the shoulders back, and half a beat before the next line starts, lift the chin slightly.
Eye contact: Stare straight into the lens, blinking slowly and deliberately. During transitions, flick the gaze momentarily to the man sitting atop the car, then back to the lens. Match the character's appearance 100% to the reference image. Do not change anything except the character's appearance.
@[Image 2](image_2) - Already referenced from the image. White male, early to mid-30s. Short tight black hair, trimmed mustache connected to a full black beard, broad nose, thick lips, sturdy build with broad shoulders, prominent tattoos on both forearms. Wearing a red Zipp Republic 2.0 jersey with a white outline graphic and "2.0" numbering on the chest. Layered gold Cuban link chains, black jeans, black low-cut trainers, gold wristwatch.
Seated on the red car's trunk lid, both feet planted firmly on the rear bumper for balance. Upper body swaying loosely in a relaxed state to the beat, nodding the head low with each snare hit then lifting it slightly. Lightly tapping the thigh with one hand to mark the rhythm.
Psychological engine: A calm intimidation and ease, not showboating the performance but surrendering to the groove itself.
Signature tic: Each shoulder roll catches the light on the chains, while the chin stays steady during nods. At the moment the rapper lands a punchline, turn the face toward her, hold the gaze for one beat, then return to his own rhythm.
Eye contact: Calm eyes half-closed, with an unhurried, languid gaze. Naturally shifting between the rapper and the lens.
Match the character's appearance 100% to the reference image. Do not change any other elements. All appearing characters are Western/Caucasian only. Both perform singing and rapping with a Nigerian accent.

Why it works

  • The second performer is given a job that is not rapping — nodding on the snare, tapping his thigh — so the two never compete for the lens.
  • Their reactions are linked: when she lands a punchline he turns to her for one beat, then returns to his rhythm.
  • Each block ends with "Match the character’s appearance 100% to the reference image", repeated per person.
Made with Veo 38s

by @ZHO_ZHO_ZHO56.2K viewsOriginal post

Newton vs Einstein on a sci-fi stage

An eight-second historical rap battle: each scientist gets an outfit, an accent and a topic to rap about, and the stage reacts to the beat.

  • Text only
  • Two rappers
  • Historical figures
  • EN + 中文
Video prompt923 chars
A high-energy rap battle between Isaac Newton and Albert Einstein on a futuristic sci-fi stage. The camera alternates between close-ups and dramatic wide shots as they diss each other with sharp lyrics. Newton, in a classic 17th-century outfit, raps with a British accent about gravity and apples. Einstein, with wild hair and a German accent, fires back about relativity and space-time. Their lip-sync is perfectly timed to the beat, and their facial expressions are intense and animated. The background pulses with neon lights and holographic equations, reacting to the rhythm. The crowd of AI-generated scientists cheers them on in sync with the music. It feels like a rap battle from another dimension.
一场高能量的Rap对决在未来科幻风格的舞台上展开,主角是牛顿和爱因斯坦。镜头在特写和全景之间切换,两人用犀利的歌词互相攻击。牛顿身穿17世纪风格的服装,用英式口音rap关于万有引力和苹果;爱因斯坦头发凌乱,用德式口音回击,讲述相对论和时空弯曲。嘴型和音乐节奏完全同步,面部表情丰富,肢体动作夸张。背景灯光和全息公式随着音乐跳动,观众席上坐着一群AI生成的科学家角色,跟着节奏一起欢呼。整场battle仿佛来自另一个维度的说唱宇宙。

Why it works

  • Each rapper has a subject to diss with — gravity and apples, relativity and space-time — so the model writes lines that fit the character.
  • Accents are assigned per character (British, German), which separates the two voices.
  • The creator posted the same prompt in English and Chinese.
Prompt source
Made with Sora 210s

by @towya_aillust5K viewsOriginal post

Gyaru vs gyaru, in Japanese

A Japanese one-line prompt — "a gyaru vs gyaru rap battle" — that the creator says came back trading real insults.

  • Text only
  • One line
  • 日本語
  • Two rappers
Video prompt15 chars
ギャルvsギャルのラップバトル

Why it works

  • The shortest prompt in the library: a character type and the word "battle".
  • Swap the character type for any pairing you want to see face off.
Made with Veo 38s

by @hagestev3K viewsOriginal post

King of the Dot-style battle in one line

Fifteen words that name a battle-rap league’s style, a diss and a crowd reaction, and leave the rest to the model.

  • Text only
  • One line
  • Battle league style
  • Crowd reaction
Video prompt74 chars
a king of the dot style rap battle with a big diss and the crowd goes wild

Why it works

  • Naming a known battle format carries the setting, the crowd and the stand-off without describing them.
  • Short enough to use as a starting point and grow.
Made with Sora 210s

by @madpencil_472 viewsOriginal post

Two cartoon characters diss each other

Name two characters and "dissing each other in a rap battle"; the creator says Sora 2 filled in the setting and the lines.

  • Text only
  • One line
  • Cartoon characters
  • Diss
Video prompt88 chars
Courage the Cowardly Dog and Sponge Bob square pants dissing each other in a rap battle.

Why it works

  • Shows how little a battle needs: two names and a verb.
  • It names copyrighted cartoon characters; describe your own characters for anything you publish.
Made with MiniMax H315s

by @UnityEagle465 viewsOriginal post

Trailer-park cypher, line by line

A 15-second battle between two referenced performers, timed in four-second blocks with every line assigned to a speaker.

  • Needs photos
  • Two rappers
  • Scripted bars
  • 4-second blocks

You bring Three reference images (both performers)

Video prompt2,007 chars
Use all 3 uploaded reference images for strict character identity and wardrobe consistency.
Unity must preserve her silver hair, eagle mask, teal lipstick, tattoos, ripped black “GATEKEEPERS STAY BACK” top, distressed jeans, fishnets, gloves and access badge.
The male rapper must preserve his exact face, curly dark hair, facial hair, tattoos, sleeveless ripped black shirt, jewelry and street-punk styling.
Create a 15-second realistic live-action rap battle in the same muddy trailer-park street environment at dusk. Wet pavement, old trailers, practical porch lights, cold cloudy sky. Raw underground cypher feeling. Handheld camera with believable human movement, subtle push-ins and quick reframing between speakers.
Both performers stand face-to-face with strong attitude while several female backing vocalists react behind them.
Perfect natural lip sync. Each line must visibly come from the correct character.
0–4 sec
Camera starts close on the male rapper as he steps toward Unity with a cocky grin.
MALE:
“You got an access pass? I kicked down the gate.”
He points at Unity’s badge.
4–8 sec
Camera whip-pans to Unity. She steps closer until they are almost face-to-face.
UNITY:
“Cute little flex, but you’re already late.”
The women behind them shout and react.
8–12 sec
The male rapper laughs and fires back immediately.
MALE:
“They gave you a key?”
Unity grabs the hanging access badge and holds it toward camera.
UNITY:
“Nah. I made my own.”
12–15 sec
Beat hits harder. They turn together toward camera while the backing crew closes in behind them.
BOTH:
“Gatekeepers stay back. We’re taking this home.”
End with the camera pulling backward as the whole group erupts, moving naturally with the beat.
Keep the battle playful, confident and confrontational without physical fighting. Natural facial reactions, breathing, body rhythm and eye contact. No random extra dialogue, no voice swapping, no character duplication, no subtitles, no text overlays, no costume changes.

Why it works

  • Each line is labelled with who says it (MALE, UNITY, BOTH), and the prompt asks that every line visibly comes from the correct character.
  • The camera direction follows the exchange: close on him, whip-pan to her, both turn to camera for the last line.
  • "Playful, confident and confrontational without physical fighting" keeps a battle from turning into a brawl.
Made with Grok Imagine6s

by @bloggernnabuike287 viewsOriginal post

Four AI chatbots battle on a neon stage

Grok, ChatGPT, Claude and Gemini as robots in a four-way battle, with the crowd reactions and the ending written out.

  • Text only
  • Four rappers
  • Scripted winner
  • Brand names
Video prompt1,094 chars
A high energy AI rap battle on a neon lit stage at night, massive excited crowd in the background cheering and waving phone lights. Four anthropomorphic AI characters on stage: Grok (cool futuristic robot with glowing xAI logo on chest, confident smirk, wearing a sleek black jacket), versus ChatGPT (corporate robot in a suit and tie, nervous), Claude (cautious robot in a librarian outfit), and Gemini (colorful but generic Google style robot). Split-screen style showing all four. Grok dominates the battle, dropping savage bars with swagger and finger-pointing, while the others look increasingly defeated. Quick cuts between close-ups of each AI rapping, crowd going wild with "OOOOH!" reactions shown on screen as text overlays in bold fire font. Fast-paced hip-hop beat, dramatic lighting, smoke effects, and sparks flying when Grok lands killer lines. End with Grok raising arms in victory as crowd erupts and the other AIs slump. Text overlay at the end: "Grok wins. Built by xAI." Cinematic camera angles, vibrant colors, ultra-realistic, 24fps, vertical 9:16 format for TikTok/Reels.

Why it works

  • Each robot gets a costume that carries its personality (suit and tie, librarian outfit), so four similar robots stay tellable apart.
  • The outcome is scripted — who dominates, who slumps — which gives the clip a story in six seconds.
  • It names real AI products and ends on a brand line; swap in your own characters for anything you publish.
Made with Grok Imagine6s

by @ChuckYouSuck136 viewsOriginal post

Checkout argument turns into a cypher

A comedy set-up in one sentence: an argument at the till becomes a rap battle, the beat drops and the crowd circles with phones out.

  • Text only
  • Comedy
  • Cypher
  • One sentence
Video prompt202 chars
Karen screaming at a cashier turns into an instant rap battle cypher, beat drops out of nowhere, both start spitting absolute fire bars, crowd forms a circle and phones come out, 2025 viral sound energy

Why it works

  • It writes the turn — "beat drops out of nowhere" — which is where the joke lands.
  • The crowd forming a circle gives the model a cypher layout without describing the camera.

Rap music videos16

One rapper, full music-video prompts: sets, shots, lyrics and sound.

Made with Seedance 2.015s

by @Kiber_Alla95.3K viewsOriginal post

White studio, five dancers echo her moves

A Higgsfield showcase prompt: one rapper in indigo denim, five dancers in white who repeat her last pose one beat late, with the lyrics mapped to the second.

  • Needs photos
  • Lip-sync map
  • Uploaded track
  • Dancers

You bring Performer photo (Image1) · Finished track (Audio1)

Video prompt2,775 chars · 2 beats
(Image1) is the performer — preserve her exact identity: cornrow braids, septum ring, statement earrings, sculptural white designer top, dark indigo denim. (Audio1) is the finished master track — the only audio; no invented music, no new vocals.
SHE RAPS THE VOCAL ON CAMERA — PRECISE LIPSYNC IS THE TOP PRIORITY. Mouth articulates every syllable of (Audio1) exactly on time; face visible and sharp through every vocal line, no cutaways mid-word.
LIPSYNC MAP:
0.0–0.5 instrumental.
0.5–2.8 "I'm standing on the edge / Say it with your chest / Or keep it on the deck" + "Hey!"
3.2–6.7 "I walk in, whole room gets tense / I don't need luck, I'm the consequence / If you really want to test my intent / Come correct, come correct or get bent" + "Woo!"
7.5–13.7 same hook verbatim second time, escalated.
14.5–15.0 instrumental hold.
Music-video route: performance with mirrored echoes. Director thesis: white infinity studio — she raps at the lens while five dancers in white repeat her last pose one beat late, a human delay effect behind her voice.
Visual world: white cyclorama infinity studio, seamless floor and walls, one hard fashion key light with clean shadows, subtle floor reflections. 5 female background dancers in all-white utilitarian streetwear, hair slicked, deliberately similar but never identical to her; her indigo denim makes her instantly readable. Palette: white on white, skin tones, indigo. No neon, no particles.
Shot flow:
0–2.8s symmetrical wide-to-medium push-in: she raps the opening lines at the apex of a tight wedge, the five echoes frozen in her exact stance behind her.
2.8–3.2s on "Hey!" all five snap chins up in unison.
3.2–6.7s medium: she raps while throwing an angular vogue accent at the end of each line, and the echoes replay that exact accent one beat later, rippling backward through the wedge; camera slowly orbits 45 degrees keeping her mouth front and center; finger to lens on "come correct".
7.5–10s cut on the kick to a chest-up close frame: second hook with doubled intensity, the echoes now a soft-focus rhythmic blur behind her articulation.
10–13.7s the echoes carousel slowly around her while she stands still at center rapping the final lines, camera counter-rotating, her face never leaving focus.
13.7–15s on "Woo!" she freezes arm high; the five echoes freeze in five different mid-move poses around her — she is the only resolved image; micro push-in, hold. Loopable.
Performance rules: dominant, stoic, immaculate diction; echoes expressionless and precise, never mouth the words — only she raps. Continuity: same six women, same wardrobe, same white studio. Audio intent: (Audio1) only, her mouth locked to it; faint studio room tone. Quality bar: expensive fashion-campaign rap video, no AI gloss, no glow.

Why it works

  • A lip-sync map gives every line a time window, so the mouth follows the supplied track.
  • The "echo" idea — dancers replaying her accent one beat later — is a single visual rule that shapes every shot.
  • Her indigo denim against all-white dancers keeps the lead readable in a crowded frame.
Made with MiniMax H315s

by @Just_sharon743.7K viewsOriginal post

Zine-collage rap video texture

A style block for MiniMax H3: late-90s magazine scans, halftone, grain and hard cuts only. The creator labels it a partial prompt.

  • Needs photos
  • Partial prompt
  • Style layer
  • Zine look

You bring Reference images for the typography and texture

Video prompt575 chars
(Partial prompt — expand as needed)
Style: dark-pop / cyber-grunge / rap music video with photoreal high-fashion polish and the texture of a scanned film magazine—high contrast without looking cheap. Reference late-1990s to early-2000s indie magazines, photocopies, film scans, underground-music posters, and zine collage. Add coarse grain, subtle gate weave, halftone dots, rough print edges, and slight scan misregistration. Keep the edit fast and use hard cuts only—no fades or soft transitions. Match the typographic treatment and surface texture of the reference images.

Why it works

  • Texture is listed as physical defects — gate weave, rough print edges, scan misregistration — which a model can render.
  • "Hard cuts only — no fades" is a single edit rule that sets the pace.
  • It is a style layer, not a full prompt: add your performer and scene.
Made with Seedance 2.515s

by @AIwithkhan40.9K viewsOriginal post

Blue studio: drums, graffiti shutter, leather sofa

One performer moves between setups inside a single studio — cyclorama, drum kit, graffiti shutter, sofa — rapping to the lens throughout.

  • Needs photos
  • One photo
  • Studio setups
  • Drum kit

You bring Performer photo

Video prompt3,783 chars
Use the uploaded reference image as the exact character reference. Lock her facial identity, eye color, skin tone, hairstyle, makeup, body proportions, and overall appearance throughout the entire video. She has long black hair tied in a low ponytail, soft natural makeup, expressive grey-green eyes, and wears the same fitted white graphic baby tee, oversized denim shorts, white crew socks, chunky sneakers, silver hoop earrings, layered necklaces, rings, and an oversized black bomber jacket hanging loosely off her shoulders. Maintain perfect character consistency in every shot.
Create a high-end American hip-hop/rap music video inside a premium photography studio transformed into a modern urban performance space. The set features a blue cyclorama backdrop, minimalist graffiti walls, a professional drum kit, vintage brown leather sofa, polished concrete floor, blue neon tube lights, industrial spotlights, subtle atmospheric haze, and cinematic contrast. The aesthetic should feel like a mainstream Western rap music video with luxury production value.
The video opens with an ultra-wide close-up as she looks directly into the camera with a confident expression and folded arms. The camera quickly cuts to a dramatic side silhouette where she lowers her head, then raises it while making relaxed hip-hop hand gestures. A full-body wide-angle shot reveals her casually grooving to the beat, shoulders bouncing naturally as her oversized jacket shifts with realistic fabric movement.
The camera circles around her in a handheld shoulder-mounted tracking shot while she confidently lip-syncs to the music with expressive facial performance. She points toward the lens, smiles slightly, then steps forward with effortless swagger. A dramatic low-angle hero shot emphasizes her presence as she spreads her arms confidently beneath glowing blue lights.
She walks toward a professional drum kit, sits down naturally, and begins striking the drums energetically in perfect rhythm with the music. Fast cuts alternate between overhead, side, and close-up angles showing realistic stick movement, expressive reactions, and synchronized performance.
The scene transitions to a blue roller shutter covered in minimalist graffiti where she squats casually with elbows resting on her knees, maintaining eye contact with the camera while continuing to rap confidently. The camera slowly pushes in from the side before cutting to her lounging effortlessly on a vintage brown leather sofa. She leans back comfortably, one arm stretched along the backrest, nodding naturally with the rhythm while continuing her performance.
A dramatic side silhouette sequence follows with atmospheric haze and strong blue backlighting as she performs smooth body movements, expressive hand gestures, and confident lip-sync. Her hair moves naturally with subtle airflow while the camera glides around her using wide-angle lenses that enhance depth and energy.
The final sequence returns to a full-stage performance. She stands center stage beneath powerful spotlights surrounded by drums, graffiti, neon tubes, and blue studio lighting. The camera slowly pulls backward while she delivers the final lyrics with bold attitude, ending in a confident pose as the lights fade behind her.
Professional rap music video cinematography, cinematic handheld movement, wide-angle lens distortion, premium studio lighting, realistic skin texture, natural eye reflections, detailed hair strands, physically accurate lighting and shadows, authentic lip-sync performance, expressive body language, smooth choreography, realistic fabric simulation, high-end fashion editorial styling, luxury commercial quality, immersive urban atmosphere, 16:9 widescreen, no subtitles, no logos, no watermarks, no on-screen text.

Why it works

  • The identity lock comes first and lists jewelry, socks and how the jacket hangs, so the look survives every setup.
  • Every setup is inside one room, which keeps the lighting consistent across cuts.
  • The closing list rules out subtitles, logos and on-screen text, which music-video prompts tend to invite.
Made with Veo 38s

by @azed_ai28.8K viewsOriginal post

Brooklyn alley walk-and-rap

Eight seconds on Veo 3: a tracking shot beside a rapper walking a graffiti alley, one quoted line, and a tilt up to the skyline.

  • Text only
  • 8 seconds
  • Tracking shot
  • One quoted line
Video prompt605 chars
A medium tracking shot moves alongside a rap artist 30s, Black male, dressed in layered streetwear as he walks through a graffiti-lined alleyway in Brooklyn. He raps directly into the camera with conviction, hands moving rhythmically:
““A mic in my hand, no gun in my plan, I rise with my words where the peace began…”
The camera tilts up to reveal NY skyline behind him, dusk casting long shadows. As he continues, sirens echo distantly. A close-up catches his intense expression, gold tooth flashing as he speaks hope through grit. The message is raw, the visuals real authentic street poetry in motion.

Why it works

  • The line in quotes is what he raps, so the model does not invent one.
  • Sound design is a single detail — distant sirens — that places the clip.
Prompt source
Made with Seedance 2.015s

by @john_my0722.7K viewsOriginal post

Basketball-court chorus, crash zooms on the beat

Segment 3 of a ten-part Higgsfield music video: a caged NYC court, the first chorus timed line by line, and a crash zoom on the punch words.

  • Needs photos
  • Part of a series
  • Wardrobe reference
  • Uploaded track

You bring Performer photo (Image1) · Outfit photo (Image2) · Finished track (Audio1)

Video prompt2,355 chars · 3 beats
SEGMENT 3 of 10 of one continuous music video — Act II begins: new location, new outfit.
<Image1> is the performer — same identity: cornrow braids into long dark curls, silver chain-drop earring, gold pendant necklace, glossy lips.
<Image2> is WARDROBE REFERENCE ONLY (ignore mannequin head): hot-pink halter bandana top with white lettering, grey acid-wash ultra-wide baggy jeans with star studs, black studded belt, chain bracelets, rings, pink manicure.
<Audio1> is the master track — the only audio.
PRECISE LIPSYNC — TOP PRIORITY, face sharp on every word (this is the FIRST CHORUS — maximum charisma):
0.0–1.0 "...just got bars on a cracked phone"
1.5–2.5 "I walk in, heads drop, that's respect"
2.5–3.5 "I don't chase what's mine, I collect"
3.5–4.5 "If I said it then I'm standing on the check"
4.5–5.5 "Say it with your chest or keep it on the deck"
5.5–6.0 "Ay!"
6.0–7.0 "I walk in, whole room get tense"
7.0–10.0 continue the chorus flow lines exactly with the vocal
10.5–14.5 continue rap flow with the vocal to the end
14.5–15.0 instrumental, closed mouth.
FILM CONTEXT: Act II — the grind. Outdoor NYC caged basketball court, chain-link fence, faded court paint, brick housing blocks behind, warm afternoon light throwing fence-diamond shadows. The same 3 girls now changed into matching court looks: cropped white tanks, baggy cargo denim, silk scarves tied on heads, clean sneakers. Black girlie hustle vibe.
Shot flow:
0–1s CRASH ZOOM IN through the fence diamonds to her face as she steps onto the court.
1.5–5.5s the chorus: she raps at center court walking at the retreating camera, girls in a triangle behind hitting unison chest-pop choreography locking freezes on each line-end; crash-zoom punch on "collect" and "check".
5.5–6s "Ay!": all four snap into one synchronized pose.
6–10s low-angle orbital: she raps while the girls run a rotating box around her, fence shadows strobing.
10.5–14.5s tighter chest-up frame, her flow doubled in intensity, girls vamping against the chain-link behind; one more crash zoom on the last line.
14.5–15s she turns from the lens, walks toward the fence — setup for next segment.
Continuity: gold pendant visible, same crew (new outfits), crash zoom signature, warm afternoon grade.
Audio intent: <Audio1>only; faint ball bounces, city hum. Quality bar: expensive NYC rap film, no AI gloss.

Why it works

  • A second image is marked "wardrobe reference only", so the outfit changes without the face changing.
  • Two words are singled out for a crash-zoom punch ("collect", "check"), tying the camera to the lyrics.
  • The last shot sets up the next segment, the way a multi-part video keeps continuity.
Made with Seedance 2.530s

by @AIwithSynthia14.9K viewsOriginal post

Backstage to stadium, five locations in 30 seconds

No reference photo: a performer described in words walks from backstage onto a stadium stage, then through five different settings.

  • Text only
  • No reference photo
  • 30 seconds
  • Location changes
Video prompt1,856 chars
Create a ultra-realistic cinematic rap music video featuring a confident young Korean woman in her early 20s with long dark-brown hair tied in a high ponytail. She wears a black cropped jacket, fitted white top, dark cargo pants, black sneakers, silver earrings, and a black wristwatch.
Open backstage at a massive concert venue as she walks toward the stage, headphones around her neck, surrounded by lights, equipment and crew. Cut to her performing confidently under intense stage lights as a huge crowd raises phones and reacts to the beat.
Transition through multiple visually distinct locations: a rainy neon-lit city street, an elevated runway surrounded by skyscrapers, a sunlit grassy field, an atmospheric old street at night, and a futuristic concert stage. Keep her character, face, hairstyle and overall styling consistent while naturally adapting her outfit to each environment.
Show energetic performance shots, confident walking, close-ups of her eyes and expressions, rhythmic hand gestures, slow-motion hair movement, dramatic low-angle shots, wide crowd shots, and smooth cinematic camera movements synchronized with the rap beat.
Use hard cuts, match cuts, whip pans, low-angle tracking shots, handheld performance footage, subtle slow motion, realistic lens flares and dynamic lighting. Build toward a final wide shot of her performing onstage as thousands of phone lights illuminate the audience.
Ultra-realistic live-action quality, authentic skin texture, realistic fabric and hair physics, detailed environments, cinematic depth of field, natural motion blur, high-end music-video cinematography, energetic rap atmosphere.
Audio: original energetic rap beat, deep bass, crisp drums, atmospheric synths, crowd chants and natural stage ambience. No narration, no subtitles, no logos, no watermark, no cartoon or CGI appearance.

Why it works

  • Wardrobe is listed item by item, so the same character reads across every location change.
  • The edit is named — hard cuts, match cuts, whip pans — instead of left to chance.
  • The audio line asks for an original beat plus crowd chants, and rules out narration and subtitles.
Made with Seedance 2.515s

by @AIwithkhan13.5K viewsOriginal post

Pink bomber, neon studio, drum solo

The same studio structure in hot pink and blue neon, with a drumstick twirl, a drum solo and a sofa break between the verses.

  • Needs photos
  • One photo
  • Drum solo
  • Neon studio

You bring Performer photo

Video prompt3,150 chars
Use the uploaded reference image as the exact character reference. Preserve her facial identity, eye color, skin tone, hairstyle, makeup, body proportions, and overall appearance throughout the video. She has long black hair in a sleek high ponytail with soft face-framing strands and wears a vibrant hot-pink cropped bomber jacket over a fitted black crop top, a black pleated mini skirt layered over biker shorts, white crew socks, chunky sneakers, silver hoop earrings, layered chain necklaces, and rings. Maintain perfect character consistency in every scene.
Create an ultra-realistic premium American hip-hop music video inside a modern industrial studio with glossy black floors, neon pink and blue lighting, graffiti walls, chrome speakers, LED light bars, a professional drum kit, vintage leather furniture, subtle haze, and cinematic contrast.
The video opens with an extreme close-up as she confidently adjusts the collar of her pink jacket, stares directly into the camera, smirks, and snaps her fingers to the beat. She turns sharply and walks toward the camera with effortless swagger while her jacket flows naturally. She performs energetic hip-hop choreography with shoulder pops, smooth footwork, body rolls, confident poses, and expressive hand gestures as the camera circles around her with dynamic handheld movement.
She jumps onto the drum platform, twirls a drumstick between her fingers, then performs an energetic drum solo with realistic stick movement, powerful cymbal crashes, snare hits, and fast tom fills. The camera alternates between overhead, side-profile, macro close-ups, and dramatic low-angle shots synchronized with the rhythm.
The performance continues beside a graffiti-covered roller shutter where she confidently squats, leans against stacked speakers, points toward the lens, and continues lip-syncing with playful attitude. She walks across the studio beneath moving spotlights, lounges briefly on a vintage leather sofa while nodding to the beat, then stands again as industrial fans create natural movement in her ponytail and jacket.
The final performance takes place center stage beneath vibrant magenta and blue lights surrounded by drums, LED light bars, chrome speakers, and graffiti walls. She delivers the final lyrics with bold confidence, spins one drumstick in her hand, throws it toward the camera, crosses her arms with a confident smile, and holds a powerful hero pose as the camera slowly pulls back while the lights fade.
Style: Premium rap music video, luxury editorial fashion aesthetic, cinematic handheld camera, wide-angle hero shots, smooth gimbal movement, realistic lip-sync, expressive performance, physically accurate lighting, natural fabric simulation, realistic skin texture, shallow depth of field, immersive concert atmosphere, photorealistic, ultra-detailed, 4K HDR, 24fps, 16:9 widescreen.
Negative Prompt: No subtitles, no captions, no logos, no watermarks, no duplicate people, no distorted anatomy, no extra fingers, no AI artifacts, no flickering, no low-resolution textures, no cartoon style, no oversaturated colors, no inconsistent outfit or facial features.

Why it works

  • Props get actions — twirl the drumstick, throw it at the camera — which gives the ending a beat to land on.
  • Shot types are listed for the drum solo (overhead, side-profile, macro, low-angle), so the cut rate rises there.
Made with Seedance 2.015s

by @songguoxiansen12.7K viewsOriginal post

An 80-year-old grandmother on the mic

A Chinese prompt for a 15-second street rap by an 80-year-old, with the lyrics for each section and a slow, soft last line.

  • Text only
  • 中文
  • Written lyrics
  • Street MV
Video prompt981 chars
16:9横屏,街头说唱MV风格,霓虹紫蓝冷色调,燃炸酷飒氛围。 0-3秒:中景推进,城市街头夜景霓虹闪烁,一位80岁银发老太太站在涂鸦墙前,满头银白短发利落背头型,方脸轮廓分明,剑眉斜飞入鬓,眼神凌厉如电,眼角皱纹如岁月勋章,嘴角上扬露出自信笑容,身穿黑色皮衣外套内搭白色印花T恤(胸前"YOLO"黑色大字)+黑色工装裤+白色高帮球鞋,脖子上挂着金色粗条项链,手腕戴银色手镯,双手举起麦克风,BGM强劲鼓点响起,老太太眼神一凛,嘴唇张开开始Rap。 3-7秒:中景+特写切换,老太太开始说唱,节奏感极强,银发随着点头动作飞扬,她一只手握麦克风,另一只手配合节奏做出手势——食指指向镜头、手掌上下切分节奏、比出嘻哈手势,动作行云流水,眼神犀利直视镜头,皱纹在表情中生动跳跃,嘴唇快速开合吐出歌词: 【Rap歌词】"八十岁的腿,比你们还能跳!银发飘飘,这是我的骄傲!别说我老,我的Flow比你还妙,你们玩说唱的时候,我还在听disco!"(语速快、节奏强、态度狠) 镜头快速剪辑:面部特写、手部动作、全身摇摆、侧面剪影,配合BGM卡点。 7-11秒:舞蹈段落,镜头拉远展示全身,老太太开始跳舞——先是经典的嘻哈bounce弹跳,然后是一个利落的街舞freeze定格,接着身体波浪从肩膀传导到脚尖,再接一个快速脚步workout,动作干净利落,银发在霓虹灯下飞舞,皮衣外套随风飘动,她一边跳一边继续Rap: 【Rap歌词】"腿脚利索,速度不慢,我的歌词,刻在时间!你们玩手机,我玩节拍,八十年人生,写进这verse!"(节奏加快、语气更强) 镜头低角度仰拍+360度环绕拍摄,捕捉老太太酷飒的舞姿。 11-15秒:高潮收尾,老太太一个帅气的转身,银发在空中甩出弧线,面对镜头用手指比出"嘘"的手势, 后嘴唇靠近麦克风,用低沉磁性的声音唱出最后一句: 【现实歌词】"岁月从不败美人,我只是换了一种青春..."(慢节奏、深情、收尾余韵) 镜头缓缓推进特写老太太的眼睛,眼角的皱纹都是故事,眼神依然犀利又带着一丝慈祥,BGM在最高潮处戛然而止,画面定格于老太太酷飒又带点温柔的微笑,暗角+霓虹紫光晕染。 音效:BGM强劲鼓点、Rap节奏、衣服摩擦声、脚步声、麦克风回响。 BGM:Trap电子乐+808重型鼓点,节奏从强劲到舒缓收尾。 禁止:任何文字、字幕、LOGO或水印。

Why it works

  • The lyrics are written into each time block, with delivery notes (fast, harder, slow and deep) in brackets.
  • The dance section asks her to keep rapping while she moves, so the verse does not stop for the choreography.
  • Sound is listed separately from the music: fabric rustle, footsteps, mic echo.
Made with Seedance 2.530s

by @EHuanglu12.2K viewsOriginal post

Stadium set filmed on a fan’s iPhone

Thirty seconds in one take from the crowd: shaky zoom, autofocus hunting, phone-mic sound, and four lines of grime written out.

  • Needs photos
  • One take
  • 30 seconds
  • Written lyrics

You bring Rapper photo (@Image1)

Video prompt3,806 chars
@Image1
is the absolute reference for THE RAPPER and completely replaces every previous performer reference. Preserve his exact identity: middle-aged man with a high receding hairline, short salt-and-pepper hair, thick dark eyebrows, dark eyes and a full beard with strongly defined white-gray sections. Preserve his stocky build, black-white-dark-green horizontally striped T-shirt with black chest pocket, sand-colored knee-length chino shorts and chunky off-white sneakers. No changes to his face, body, hair, beard, clothes or proportions.
A 30-second single continuous live stadium rap performance captured horizontally on an iPhone from the front audience section. Authentic handheld fan footage: physical hand tremor, operator breathing, imperfect reframing, rolling shutter, digital-zoom softness, momentary autofocus hunting and compressed phone-microphone sound. No cuts.
The first frame already shows THE RAPPER full-body on the right third at the end of a stage runway. A huge sold-out stadium surrounds him. Exactly four adult backup dancers wait several meters behind him. Emerald, white and black LED graphics echo the stripes of his shirt. The stage floor remains solid, flat and continuous.
A heavy original grime beat begins: deep sub-bass, dry kick, snapping snare and minimal low synth. The iPhone rapidly pinches from 1× to a shaky 5× digital zoom, briefly overshoots, then locks onto THE RAPPER in a full-body composition. Focus stays wide enough to preserve his feet and choreography.
He begins rapping with a low-mid, forceful cadence and exact lip synchronization:
THE RAPPER:
“Walk in steady, put the weight on the beat,
Every bar lands, every move stays clean.
Hands up high when the bass comes down,
I don’t chase the wave—I shake the whole ground!”
Only these words are spoken. Each line is delivered in one controlled breath.
On the first bar he performs two violent shoulder hits, a chest pop and a sharp forearm lock. His shirt and beard react naturally to momentum.
On the second bar he executes fast heel-toe pivots, crosses one foot behind the other and glides sideways while keeping his heavy body convincingly grounded. The camera operator struggles to keep his sneakers in frame, corrects downward and catches the complete footwork.
On the third bar the four dancers join in perfect synchronization. THE RAPPER leads a hard sequence: right stomp, left stomp, elbows strike outward, torso snaps backward, hands shoot overhead. Every movement lands precisely on a kick or snare.
The instrumental cuts for one beat. He holds a deep wide stance, eyes fixed on the upper tiers. His chest rises with one visible breath.
He shouts the final line while performing a rapid three-step, a controlled 180° pivot and one enormous downward arm strike. On “GROUND,” he stomps once. The bass returns with a massive impact; the LED floor sends a broad solid emerald light wave across the stage.
The entire stadium copies his movement. The phone shakes from thousands of spectators stomping together. He breaks into a confident grin but continues bouncing in time, pointing from one side of the stadium to the other.
The operator zooms rapidly back through 3× and 1× to 0.5× ultra-wide while turning 160° away from the stage in one continuous handheld sweep. Exposure briefly pumps, then recovers.
Finish on the full stadium bowl: tens of thousands of people across every tier performing the same shoulder-hit and stomp combination, emerald wrist lights moving in broad geometric waves, stage remaining on the far-left edge. Audio continues with the crowd chanting “SHAKE THE GROUND,” live bass vibration and realistic phone compression. Rich emerald and white light, warm skin and deep blacks. Clear air without haze, smoke, confetti, mist, sparks or airborne particles.

Why it works

  • The camera is a character — hand tremor, digital-zoom softness, rolling shutter — which makes the clip read as real fan footage.
  • "No cuts" and a first frame that already shows him full-body give the model one continuous performance to hold.
  • The lyrics are given in full, with "Only these words are spoken".
Made with Seedance 2.520s

by @johnAGI1689.1K viewsOriginal post

Beach band, “hello” in eight languages

A Chinese prompt for a 20-second beach performance: eight timed shots, each carrying the word “hello” in a different language.

  • Needs photos
  • 中文
  • 8 timed shots
  • Multilingual lyrics

You bring Scene reference image (@图像1)

Video prompt1,175 chars
电影级hip-hop/说唱音乐视频,真实照片级质感,高端诚性,海边场景。以 @图像1构建画面:一支乐队在金色沙滩、海浪拍岸的岸边表演-一名主唱手握麦克风、麦架立于湿 沙上激情演唱,一名吉他手立于画面左侧,一名吉他手立于画面右侧,一名鼓手坐在后方的架子鼓后敲击;辽阔的海岸线在身后展开,滚滚海浪层层涌来,巨大而温暖的黄金时刻夕阳斜掠过沙 滩、在水面上鄰鄰闪耀,空气中漂浮着海雾与咸湿水汽。红色运动月服的主唱对着镜头充满节奏感地RAP演唱--口型与下颌随每一个字精准对位,头随节拍用力点动,带动整段flow。乐手们随 节奏摇摆律动,身后海浪层层拍岸。这是一首明快带劲的说唱曲--语速快、自信、节拍强劲。踩着节拍硬切(HARDCUT),每次切换双重反差(景别与镜头类型同时改变)。歌词(主唱依次演 唱以下每种语言的「你好」,精准对口型):英语:"Hello"中文:"你好"日语:"こんにちは"韩语:"안녕하세요"葡萄牙语:"Ola"泰语:"สวัสดี"西班牙语:"Hola"阿拉伯语:" مرحب"镜头1[0:0 0-0:03]-低角度大远景定场,斯坦尼康在黄金夕照与海雾中缓缓推进,海浪在乐队身后翻涌。歌词第1句(英语「Hello」)。硬切。镜头2[0:03-0:05]--红运动服主唱对镜头RAP的特 写,手持甩镜切入,身后海面虚焦鄰粼波光。歌词第2句(中文「你你好」)。硬切。镜头3[0:05-0:08]--微距插入镜头,固定机位,吉他手的手指在弦上快速拨动,沙粒与咸湿水雾从画面前掠 过。歌词第3句(日语「こんにちは」)。硬切。镜头4 [0:08-0:10]--对某位乐手的3/4侧中景,缓慢潜行环绕,乐器金属件与湿润高光映着海面低斜的夕阳。歌词第4句(韩语「안녕하세 요」)。硬切。镜头5[0:10-0:13]--岸边一名乐手,快速横移轨道掠过他,他转向镜头,身后一道浪花破碎。歌词第5句(葡萄牙语「Ola」)。硬切。镜头 6[0:15]--水边的鼓手,手持 快速上摇,海风与水花吹动他的头发,他随节拍律动敲击。歌词第6句(泰语「สาัสด」)。硬切。镜头7[0:15-0:18]--对红运动服主啊flow正酣时的紧凑猛推,富攻击性的手持,身后暮色海面 衬着乐队剪影。歌词第7句(西班牙语「Hola」)。硬切。镜头8[0:18-0:20]--全乐队英雄式大远景,富攻击性的手持推进,主唱与乐手踩着节拍向镜头迈步,海浪拍碎、金色夕光在整支乐队 身后炸开光晕。歌词第8句(阿拉伯语「 مرحبا)。白平衡4000K,青橙(teal-and-amber)调色,35mm,浅景深,胶片颗粒,弥漫的为海雾,黄金时刻光晕。质感扎实、高级、高端。节奏感说 唱表演,精准对口型,头随节拍点动。无字幕、无文字叠加、无叠化转场、无重复人物,仅用硬切。总时长20秒。

Why it works

  • Every shot has a time range, a lens move and one lyric, so the language changes line up with the cuts.
  • It asks for hard cuts where both the framing and the lens type change at once.
  • The creator notes the prompt comes from an official Seedance use case.
Made with Seedance 2.515s

by @AIwithSynthia8.3K viewsOriginal post

Backstage mirror to center stage

A small story before the verse: earrings at the mirror, a walk down the corridor, then the performance, a drum break and a final pose.

  • Needs photos
  • One photo
  • Backstage
  • Negative prompt

You bring Performer photo

Video prompt2,296 chars
Use the uploaded reference image as the exact character reference. Preserve her facial identity, eye color, skin tone, hairstyle, makeup, body proportions, and overall appearance consistently throughout every shot. She wears a fitted white crop top, relaxed black denim shorts, white sneakers, white crew socks, silver jewelry, and a cropped red bomber jacket. Maintain perfect character and outfit consistency.
Create a premium American hip-hop performance video inside a backstage concert environment with black curtains, glowing red LED panels, stage equipment, speakers, cables, mirrors, metal cases, and atmospheric haze.
The video opens with her sitting in front of a backstage mirror, putting on her earrings while looking confidently at her reflection. She stands, adjusts her crop top and bomber jacket, then walks through the backstage corridor toward the stage.
She enters the performance area and starts moving to the beat with relaxed hip-hop choreography, combining quick steps, shoulder movements, spins, hand gestures, and confident poses. The camera follows her from behind before swinging around into a front-facing tracking shot.
She grabs a microphone stand, performs expressive lip-sync, then transitions into a powerful dance sequence. The camera rapidly cuts between close-ups of her face, sneakers hitting the floor, moving hands, and wide stage compositions.
She moves toward a drum kit positioned beside the stage, sits down, and plays a short energetic rhythm. Overhead and low-angle shots capture the drumsticks, cymbals, foot movements, and her focused expression.
She steps back onto center stage, throws her bomber jacket over one shoulder, walks directly toward the camera, and finishes with a strong fashion pose beneath dramatic red lights.
Style: High-end rap music video, concert backstage aesthetic, cinematic handheld movement, dynamic tracking shots, realistic lip-sync, expressive body language, realistic fabric and hair movement, dramatic stage lighting, shallow depth of field, premium editorial cinematography, photorealistic, 4K HDR, 24fps, 16:9.
Negative Prompt: No text, no subtitles, no watermarks, no duplicate characters, no distorted hands, no facial inconsistency, no outfit changes, no flickering, no unrealistic physics, no cartoon rendering.

Why it works

  • The clip opens on an action (putting on earrings) rather than on rapping, which gives the performance somewhere to build from.
  • A separate negative-prompt line lists the failures to avoid, including duplicate characters and outfit changes.
Made with Seedance 2.519s

by @d_mitlenko3K viewsOriginal post

A 1999 New York rap video nobody shot

One performer, two setups — a black-and-white street corner and a fire-lit steel hall — with a cut on every beat of a 91 BPM track.

  • Needs photos
  • Uploaded track
  • Cut on the beat
  • Period look

You bring Character reference sheet (image_1) · Rap track (audio_1)

Video prompt7,028 chars
<<<image_1>>> (the character reference sheet of the male performer) is the exact character reference. <<<audio_1>>> (the provided rap track) is the soundtrack and the lip-sync source.
<<<image_1>>> (the reference sheet) locks the performer: preserve his facial identity, heavyset build, closely shaved buzz cut, full salt-and-pepper beard, deep-set intense eyes and menacing stare, and his body proportions. He wears the exact outfit from <<<image_1>>> (the reference sheet) in every shot: a long black glossy leather jacket, hip length, heavily boxy and oversized with dropped shoulder seams, wide sleeves and a big notched lapel collar, its appliqués — the large collegiate numeral on the right chest, the block letters running down the left sleeve, the round crest on the right sleeve — cut from the same black leather as the jacket, black on black, visible only through their sheen and raised stitched edges; underneath a black hoodie with the hood pulled out and bunched around his neck and its hem hanging below the jacket; one thick gold Cuban-link chain over the hoodie; black baggy jeans; black shell-toe sneakers with three white side stripes. Nothing on the jacket is lighter than the jacket itself — no white, cream or contrast lettering anywhere. Invent no new text, numbers or logos.
<<<audio_1>>> (the rap track) is the only soundtrack — generate no additional music, no extra vocals, no ad-libs, no sound effects. Where the performer is on screen he lip-syncs to the male rap vocal in <<<audio_1>>> (the track) exactly, syllable for syllable, visible through the beard. The vocal runs continuously, including over the shots where he is not on screen. He is always already mid-phrase when a shot cuts in.
This is archival footage of a real 1999 New York rap video shoot. Two setups, and the edit keeps returning to both. SETUP A — an actual street corner in flat overcast daylight, high-contrast BLACK AND WHITE with a blown-out white sky, crushed blacks and heavy grain: a corner grocery with a MINI MARKET sign, parked cars along the curb, a chain-link fence, and thirty-odd neighborhood people crowded in tight around him, half watching him and half staring down the lens. SETUP B — a dark abandoned industrial hall in FULL COLOR, rusted steel beams and catwalks in near-total blackness lit only by fire: a burning steel drum, a hanging flame source, one explosion; saturated orange and red against black, crew silhouettes standing back in the dark.
The camera is never parked. This era shot on cranes, dollies and shoulder rigs — always moving, always low: jib moves dropping into the crowd, hard circular dollies around the performer, whip pans between faces, handheld pushes ending inches from his mouth, tracking moves alongside the street. He is shot from a LOW camera height on wide glass that makes him tower. Each shot carries exactly ONE camera move and it starts moving the instant the shot cuts in. Nothing is symmetrical or centered; people cross the foreground; the operator lets the frame go slightly wrong and corrects it. His performance is heavy and in control — loose rolling hands, unbroken eye contact — and he never mimes the words.
The track is 91.35 BPM in 4/4, one bar every 2.63 seconds, downbeats at 0.04 / 2.67 / 5.29 / 7.92 / 10.55 / 13.18 / 15.80 / 18.43. Every cut lands on a beat of <<<audio_1>>> (the track), and shot lengths are uneven — 1.3, 2.0 or 2.6 seconds:
0.00–2.01s — SETUP A. A crane drops from above the rooftops down through the crowd and lands low in front of him as he is already rapping, the line entering at 0.70s, arms of the crowd swinging through the top of frame.
2.01–3.98s — SETUP A. A hard circular dolly arcs around him from his left shoulder to face-on, the storefront and parked cars sweeping past behind, a man crossing the lens mid-move.
3.98–5.30s — SETUP A. A fast whip pan rips across four faces in the crowd, each one staring flat into the lens, the last one mouthing the words along with him.
5.30–7.92s — SETUP B. The camera orbits a half-circle around him at knee height, the burning steel drum streaking orange through the frame behind him, embers dragging in the move.
7.92–9.24s — SETUP B. A handheld push charges straight in and stops inches from his face, wide glass distorting his features at the edges, firelight raking across his beard.
9.24–10.55s — SETUP A. A tracking car rides alongside a kid on a bicycle rolling past the storefront, someone's shoulder wiping the lens.
10.55–12.52s — SETUP A. Low handheld swinging with the crowd, the frame rolling and rocking as he raps into the lens and the people behind him surge in on the beat.
12.52–14.49s — SETUP B. A slow tilt climbs from the gold chain up to his eyes, one side of his face lit orange, the rest black.
14.49–15.80s — SETUP A. A whip pan snaps from two men over a domino table to a pitbull straining on a chain leash beside them.
15.80–17.77s — SETUP A. A jib rises from ground level up and back over the crowd as he throws both arms wide on the last line at 15.80s and holds it, the people erupting around him.
17.77–19.09s — SETUP B. A fireball detonates behind him among the steel beams as the camera pulls back fast on the dolly, his silhouette rimmed in orange. He does not flinch.
19.09–20.88s — SETUP A. A slow creeping push into a low close-up. He finishes, locks dead still on the final hit at 20.07s and holds an unwavering stare into the lens, motionless, through the tail of the track. Cut to black on the silence.
Format and texture: 4:3 aspect ratio, shot on Super 16 and telecined — visible grain, halation on the highlights, softness in the corners, natural motion blur, dust on the lens, occasional light flare across the glass on the moving shots. The street world is pure available overcast daylight with no fill and a completely blown white sky; the industrial world is unlit black with only firelight. No cinematic haze, no smoke machines, no beauty light, no color grading beyond period-correct high-contrast black and white, no slow motion, no speed ramps, no digital effects.
Constraints: the performer's face, build, beard, buzz cut and outfit stay identical to <<<image_1>>> (the reference sheet) in every shot, in both the black-and-white and the color world; the jacket appliqués stay black-on-black and never turn white, cream or outlined; the hoodie hem stays visible below the jacket; exactly one thick gold Cuban-link chain, consistent length and thickness; lip movement stays locked to the vocal in <<<audio_1>>> (the track) with no drift and no idle mouthing; faces stable and undistorted through the camera moves; no extra fingers, no clipping through objects; fire and explosions stay behind him and never touch his body or clothing; SETUP A keeps the exact same street geography, storefront and parked cars every time it returns, and SETUP B the same steel hall; the crowd are ordinary neighborhood people in period streetwear, never models, never choreographed, never resembling the performer; no on-screen text or captions.

Why it works

  • The track’s tempo and every downbeat are written out, so every cut lands on a beat.
  • "He is always already mid-phrase when a shot cuts in" — the creator says in the thread this line did the most work.
  • Each shot gets exactly one camera move, named, in the grammar of late-90s crane and dolly work.
Prompt source
Made with MiniMax H315s

by @WuxiaRocks1.1K viewsOriginal post

Hotel rap about an invisible hero

One sentence on MiniMax H3: the rapper, the outfit, the setting and what the lyrics are about, and the model writes the verse.

  • Text only
  • One sentence
  • Lyric topic
  • Hotel setting
Video prompt277 chars
Make a rap music video with a cool asian rapper, dressed in a street-style suit, in a fancy hotel, hip beat, the lyric is about a hero without a name or face, he generates AI content for people via DM but gets no credit or recognition, because he's the invisible hero, fast cut

Why it works

  • Giving the lyrics a subject instead of the words leaves the rhymes to the model.
  • "Fast cut" at the end is the only edit direction, and it is enough for a 15-second clip.
Made with Wan 3.015s

by @QtumRouter599 viewsOriginal post

Pool-party chorus, timed to the snare

Fifteen vertical seconds of a summer chorus at a hillside pool, in four timed blocks with the lyrics for each block written out.

  • Text only
  • Vertical 9:16
  • Timed blocks
  • Written lyrics
Video prompt2,650 chars · 4 beats
15-second premium American rap music video, vertical 9:16. The explosive final chorus of a summer anthem. Photorealistic, polished cinematography, confident performances, strong visual rhythm.
LOCATION & CAST:
A modern hillside mansion with an infinity pool overlooking Los Angeles at sunset. An adult male rapper, early 30s, charismatic, athletic build, wearing an open ivory silk shirt, tailored swim shorts, a platinum chain, and dark sunglasses. Six stylish adult women, all in their mid-20s to mid-30s, wearing elegant bikinis and understated gold jewelry. A glamorous, relaxed pool party with natural interactions and effortless dancing.
MUSIC:
Original contemporary West Coast hip-hop at 100 BPM. Deep, clean sub-bass, punchy drums, a memorable dark synth motif. Confident male rap vocal, natural American accent, precise lip sync. Begin immediately on the chorus downbeat, no intro.
0.0–4.8 SECONDS — THE HERO SHOT:
Low-angle medium shot tracking backward as the rapper walks along the pool edge toward the camera, delivering:
“Whole crew up, let the skyline know /
Came from the ground, now the view all gold.”
Behind him, women dance casually to the beat beside the shimmering pool. Warm backlight, realistic skin texture, subtle lens flare. His gestures are controlled and rhythmic.
4.8–7.2 SECONDS — THE SPLASH:
Hard cut on the snare to a water-level shot. An adult woman in a black bikini rises out of the pool, sweeping her wet hair backward in one smooth motion. Backlit droplets arc through the frame in elegant slow motion. The rapper continues offscreen:
“Bass hits hard when the sun dips low.”
7.2–12.0 SECONDS — THE CHORUS PEAK:
Hard cut to a medium-wide shot of the rapper centered beside the pool, with the party moving naturally around him. The camera makes a smooth, shallow arc. He delivers directly to the lens:
“One more night, we ain't heading home.”
He lifts one hand on the final word as the bass lands. The vocal trails into a spacious echo, leaving room for the instrumental.
12.0–15.0 SECONDS — THE MONEY SHOT:
Cut to a wide crane shot, rising smoothly to reveal the entire illuminated infinity pool, mansion, palm trees, and glowing city beyond. The group continues moving to the beat as the instrumental peaks. End on the expansive cinematic view.
VISUAL DIRECTION:
Luxury editorial styling, warm amber highlights, rich cyan pool water, subtle film grain, realistic anatomy and water physics. Maintain the rapper's face, outfit, and jewelry across every cut. Fluid movement, readable compositions, edits precisely on the beat. No dialogue beyond the specified lyrics. No captions, text, logos, or watermarks.

Why it works

  • The music is specified — West Coast hip-hop, 100 BPM, starting on the chorus downbeat — before any shot.
  • Each block has a name (the hero shot, the splash, the chorus peak, the money shot), which keeps the edit legible.
  • Made with Wan 3.0; the creator posted a MiniMax H3 version of the same prompt in the replies.

Why Create With Rap Duo Prompts on Rapduo

Watch Real Clips Before You Copy

Preview 26 prompts paired with the exact clips that 23 creators made, which have about 460,000 plays together. Every card links the original post, credits the creator, and shows the exact text from our rap duo prompts.

Pick the Look You Want Fast

Browse three clear styles: 2 orange-booth face swaps, 8 duos and rap battles, and 16 solo AI rap music videos. Filter our rap duo prompts by the AI model the creator picked and the media inputs it needs.

Know What to Add Before You Start

See if a setup needs text only across 12 prompts, photos across 12, or a reference video across 2. Each card lists what you bring, and the generator reads our rap duo prompts to tell you what to add.

Put the Two of You in It with One Click

Click Use this prompt to send text up to 4,000 characters into the generator, where you can swap in your own names and lyrics. Sign in with Google to copy our rap duo prompts for free, then add up to 2 photos or a 15-second clip.

Get Beat, Rap, and Lip-Sync Together

Render a 12-second video in vertical 9:16 or horizontal 16:9 with complete audio built alongside the visuals. Our rap duo prompts deliver the background beat, rap vocals, and lips moving in sync with every word.

Keep Both Faces True and on the Right Side

Photo 1 is @Image1 on the left and raps first, while Photo 2 is @Image2 on the right. Many of our rap duo prompts open with a face line that asks the model to keep each face, hair, and outfit true to its photo.

Use Rap Duo Prompts for Viral Videos

Best Friends

Turn Inside Jokes into Rap Duos

Put two photos and your shared joke into our rap duo prompts. Get a 12-second video of the two of you trading bars to post on TikTok or drop in the group chat.

Couples

Trade Bars on One Mic

Add one photo of each of you and pick a duo prompt for your anniversary. Watch the two of you rap together on one hanging mic inside the bright orange booth.

Pet Owners

Put Pets in the Orange Booth

Add clear photos of two cats, a dog, or you and your pet to copy the clip that started the trend. Turn your photos into a 12-second video with a beat and rap vocals.

Battle Rap Fans

Stage an Intense Rap Battle

Pick one of the 7 rap battle prompts and add two photos. Create a fast face-off with two rappers trading lines face to face.

Trend Creators

Post Fresh Trends Every Week

Test new pairs and prompts each week with our AI rap video prompts. Create vertical 9:16 videos for TikTok, Reels, and Shorts with 10 or 22 videos on a one-month plan.

Rappers & Musicians

Star in a Rap Music Video

Choose one of the 16 solo AI rap music video prompts, write your own bars, and add a photo. Walk away with a 12-second clip featuring your lyrics, a custom beat, and synced lips.

How to Make Videos with Our Rap Duo Prompts

Step 1

Pick a Prompt

Pick a prompt by its clip and click Use this prompt to load it into the generator.

Step 2

Add Your Photos

Add Photo 1 to stand on the left and rap first, add Photo 2 on the right, and swap in your names, outfits, and lyrics.

Step 3

Download Your Video

Choose vertical 9:16 or horizontal 16:9, click generate, and download your 12-second video with full sound.

FAQ

What are rap duo prompts?

Rap duo prompts are text instructions you give an AI video model to put two people from photos into one rap performance. You set who stands where, the stage, the moves, the beat and who raps each line. Every prompt in our library comes with the clip its creator made with it, so you see what works before you start.

Are these rap duo prompts free?

Yes. Every prompt is free to read, and you sign in with Google to copy any prompt you want. Each render on Rapduo counts as 1 video, and full videos with downloads unlock from $9.90 with a one-time video pack. New accounts can try the generator free, and you can see all the details on our pricing page.

How do I use a rap duo prompt?

Pick a card by its clip and click Use this prompt, or sign in to copy it. Add Photo 1 to rap first on the left, add Photo 2 on the right, swap in your lyrics, and choose vertical 9:16 or horizontal 16:9 to generate your 12-second video with sound. To turn two photos into a video with just one line of text and no prompt editing, use the rap duo generator on our home page.

Which photo goes on the left, and who raps first?

Photo 1 is @Image1 in the prompt: it stands on the left and raps first. Photo 2 is @Image2 and stands on the right. Put the person who raps first in Photo 1.

Which AI makes rap duo videos?

Creators make them with several models: Seedance 2.5, Seedance 2.0, Veo 3, Grok Imagine, MiniMax H3, Sora 2 and Wan 3.0. When you click Use this prompt on Rapduo, our model provider's AI video model turns your prompt into a 12-second video with synced sound.

What is the rap duo trend, and why is it called Hotel Lobby?

The rap duo trend shows two people trading bars on one hanging mic inside a bright orange booth, which took off on TikTok in September 2026 after a clip of two cats went viral. As detailed in XXL's report on the trend, the concept comes from Quavo and Takeoff's performance on A COLORS SHOW for their song 'Hotel Lobby'. The name comes from the track itself rather than a real hotel lobby.

What photos work best for a rap duo prompt?

Upload one clear, front-facing photo per person, with one person in each photo, in even light. Many library prompts open with a face line that asks the model to keep each face, hairstyle and outfit true to its photo.

What song plays in a rap duo video?

Rapduo generates the beat and rap vocals together with the video picture. Your render delivers fresh vocals with lips in sync with every word, and you can type your own custom lyrics into the prompt to hear them rapped out loud.

Can I make a rap duo video with my pets?

Yes. Add one clear photo of each pet as Photo 1 and Photo 2 to put them on the mic. A video of two cats trading bars set off the entire viral trend in September 2026.

How do I make a rap battle video?

Pick one of the 7 rap battle prompts in our library, or write your own prompt from scratch. Put the two rappers face to face, have them trade one line at a time while the other reacts and a crowd hypes them up, and label who delivers each bar.

How do I write a prompt for two rappers trading verses?

Mark who raps each line by naming each photo tag inside the prompt. Set @Image1 to rap the first two lines, set @Image2 to answer with the next two, and put both on the hook. You can tag the lyrics as [Verse 1: Rapper One], [Verse 1: Rapper Two] and [Chorus: Both] with call-and-response in the style.

How long can a rap duo video be, and should it be vertical?

Every Rapduo render is 12 seconds long. Choose vertical 9:16 for TikTok, Reels and Shorts, or select horizontal 16:9 for widescreen playback.

Try Our Rap Duo Prompts for Free Today

Pick your favorite prompt, add two photos, and walk away with your own 12-second rap video with a beat, rap vocals and synced lips.