A teaser for your next release
Use the hook of a song you are about to put out, cast your bandmates, and post the clip the week before it drops.
Bring your own audio
Same orange booth, same mic, new words. Add two photos plus a few seconds of audio, and the lead lip-syncs your own song, a verse you wrote or a recorded message at the booth’s hanging mic while the partner dances.
Your two performers
One person each. Left raps, right hypes.
Stage
Video & sound
This video: 6 credits. Ready in about 5 to 15 minutes; failed generations return your credits. View pricing
The short answer
Updated
Think of a Hotel Lobby AI video as three inputs. The two photos supply the faces, the booth clip supplies the camera work and the choreography, and an audio track tells the lead what to mouth. Unless you change it, that track is the original recording from the Hotel Lobby trend. Upload your own song and it takes over that job alone; set, moves and price stay put.
Only the lead on the left performs the audio. Your partner on the right answers with nods, points and dance steps. Want two voices trading lines? An upload is the wrong tool: use AI rap with Both rap, which needs a Seedance model.
A Hotel Lobby AI video with your song is a short clip, not a full-length music video: 5, 8, 10, 12 or 15 seconds. Choose the hook, the punchline or the line people quote. Stick to recordings you own or are licensed to use. For more ways to turn a photo and a song into a clip, see the AI music video generator.
Want different choreography too? Under Soundtrack & moves, pick Your clip for the moves and Your song for the sound. A dance video of 2 to 15 seconds sets the moves and the length, your recording sets the lip-sync, and the clip’s own sound is removed before upload so they never compete.
Three steps
A sharp photo for each performer, or a single shared picture. The left spot is the voice of the video, so put whoever should perform the audio there.
In the Soundtrack & moves drawer, switch Sound to Your song and select the file. Longer than the video? Drag the start point to the part you want, and only that stretch is kept.
The sound never moves the price: 12 seconds of Wan 3.0 480p video cost 6 credits, ready in about 5 to 15 minutes. Play it back before you post; a failed render is refunded.
Ideas
Use the hook of a song you are about to put out, cast your bandmates, and post the clip the week before it drops.
Record ten seconds of a wedding, graduation or birthday toast on your phone and let the lead deliver the audio into the hanging mic for you.
The thirty-second rant someone left in the family chat, cut to its best ten seconds and delivered like a live set.
Write eight bars about the person in the right-hand photo, record them over a beat you own, and watch them hype their own roast.
Audio never changes the price. It comes from the quality and the length, exactly as with the Hotel Lobby track. Credits on Wan 3.0:
| Length | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | 3 | 4 | 8 |
| 8 seconds | 4 | 6 | 12 |
| 10 seconds | 5 | 8 | 16 |
| 12 seconds | 6 | 10 | 20 |
| 15 seconds | 7 | 12 | 24 |
Your song plays on Wan 3.0 only; Seedance writes its own AI rap. Packs start at 10 credits for $9.99.
Watch
Videos and posts from creators and other tools, each credited and linked to the original. Nothing loads from YouTube until you press play.
MP3, WAV, M4A, AAC or OGG. A phone voice memo is usually M4A, which works. If your browser cannot open a format, convert the file to MP3 and try again.
The left one. The lead carries the audio into the hanging mic, and your partner on the right only reacts, nodding and gesturing without singing your lines.
No. The original recording is already the default sound. The upload is for something of your own.
No. Every video is paid, the first one included, and uploaded audio costs the same as the default sound. Failed renders are refunded.
Yes. Switch to a Seedance model, set the sound to AI rap and describe the subject in up to 120 characters; the song is written for you, beat included. It cannot be mixed with an uploaded file.
Not with an upload: only the lead performs it. For a duet, use AI rap with Both rap so the two swap lines.
Upload it anyway. Your browser cuts out the part you want, as long as the video, and uploads only that.
One clear voice close to the mic, little background music, and a stretch that opens with a word, not with a pause. Spoken lines work too.
Close, not identical. The soundtrack is generated from your recording along with the picture, so the words and timing follow it, but small differences can appear. Watch the result before you post.
Ready when you are
Two photos and a few seconds of audio in the Hotel Lobby AI booth: 6 credits for 12 seconds in 480p, the same as the default sound.