Skip to content

Bring your own audio

Hotel Lobby AI, set to a song or voice note of your own

Same orange booth, same mic, new words. Add two photos plus a few seconds of audio, and the lead lip-syncs your own song, a verse you wrote or a recorded message at the booth’s hanging mic while the partner dances.

  • 2 to 15 seconds of your own song
  • MP3, WAV, M4A, AAC or OGG
  • No extra credits for audio
  • Works with your own dance clip
1

Your two performers

One person each. Left raps, right hypes.

2

Stage

3

Video & sound

This video: 6 credits. Ready in about 5 to 15 minutes; failed generations return your credits. View pricing

The short answer

Only the Hotel Lobby track changes when you upload

Updated

Think of a Hotel Lobby AI video as three inputs. The two photos supply the faces, the booth clip supplies the camera work and the choreography, and an audio track tells the lead what to mouth. Unless you change it, that track is the original recording from the Hotel Lobby trend. Upload your own song and it takes over that job alone; set, moves and price stay put.

Only the lead on the left performs the audio. Your partner on the right answers with nods, points and dance steps. Want two voices trading lines? An upload is the wrong tool: use AI rap with Both rap, which needs a Seedance model.

A Hotel Lobby AI video with your song is a short clip, not a full-length music video: 5, 8, 10, 12 or 15 seconds. Choose the hook, the punchline or the line people quote. Stick to recordings you own or are licensed to use. For more ways to turn a photo and a song into a clip, see the AI music video generator.

Pair the track with a dance clip

Want different choreography too? Under Soundtrack & moves, pick Your clip for the moves and Your song for the sound. A dance video of 2 to 15 seconds sets the moves and the length, your recording sets the lip-sync, and the clip’s own sound is removed before upload so they never compete.

Three steps

From song file to Hotel Lobby AI video in three steps

  1. Add the performers

    A sharp photo for each performer, or a single shared picture. The left spot is the voice of the video, so put whoever should perform the audio there.

  2. Upload the audio

    In the Soundtrack & moves drawer, switch Sound to Your song and select the file. Longer than the video? Drag the start point to the part you want, and only that stretch is kept.

  3. Render, listen, then post

    The sound never moves the price: 12 seconds of Wan 3.0 480p video cost 6 credits, ready in about 5 to 15 minutes. Play it back before you post; a failed render is refunded.

Ideas

Songs and voice notes worth putting on stage

01

A teaser for your next release

Use the hook of a song you are about to put out, cast your bandmates, and post the clip the week before it drops.

02

A toast you recorded

Record ten seconds of a wedding, graduation or birthday toast on your phone and let the lead deliver the audio into the hanging mic for you.

03

The voice memo everyone quotes

The thirty-second rant someone left in the family chat, cut to its best ten seconds and delivered like a live set.

04

Bars about your friend

Write eight bars about the person in the right-hand photo, record them over a beat you own, and watch them hype their own roast.

Inside the browser: how your file is trimmed

  1. The picker takes MP3, WAV, M4A, AAC or OGG, up to 200 MB. A short MP3 or WAV that already fits the video goes up as it is; anything longer, and every other format, is cut and converted to WAV in your browser.
  2. Drag the start point to the part you want. Everything outside a window as long as the video is cut away before upload; the window needs at least 2 seconds of sound.
  3. Your words come out of the lead under that hanging mic while the partner keeps the energy up. The soundtrack in the finished video is generated from your recording, so it follows it closely without being a bit-for-bit copy.

What your song can and cannot change

  • One voice. Only the lead lip-syncs; the partner hypes. Two voices need an AI rap duo instead.
  • Fifteen seconds at most. Enough for a hook or a toast, not a full music video.
  • One model and stage. Wan 3.0 in the Hotel Lobby booth takes your audio, up to 15 seconds; the Seedance models write their own AI rap instead.
  • Clean audio in, clean lip-sync out. Music over the voice or a noisy room blurs the mouth movement.
  • Not a voice changer. The lead performs your recording; it does not turn your voice into someone else’s.

What a video with your song costs

Audio never changes the price. It comes from the quality and the length, exactly as with the Hotel Lobby track. Credits on Wan 3.0:

Length480p720p1080p
5 seconds348
8 seconds4612
10 seconds5816
12 seconds61020
15 seconds71224

Your song plays on Wan 3.0 only; Seedance writes its own AI rap. Packs start at 10 credits for $9.99.

Watch

Own-song Hotel Lobby AI edits on YouTube and X

Videos and posts from creators and other tools, each credited and linked to the original. Nothing loads from YouTube until you press play.

Full AI Music Video Tutorial | Your own Face | Your own SongReference Example by Theos Stack on YouTube ↗

Questions

Questions about uploading audio

Still stuck? Contact us or see pricing.

What formats can my track be in?

MP3, WAV, M4A, AAC or OGG. A phone voice memo is usually M4A, which works. If your browser cannot open a format, convert the file to MP3 and try again.

Does the left or the right performer lip-sync?

The left one. The lead carries the audio into the hanging mic, and your partner on the right only reacts, nodding and gesturing without singing your lines.

Do I need to upload the Hotel Lobby song itself?

No. The original recording is already the default sound. The upload is for something of your own.

Can I try my own song without paying?

No. Every video is paid, the first one included, and uploaded audio costs the same as the default sound. Failed renders are refunded.

No recording yet: can it write one for me?

Yes. Switch to a Seedance model, set the sound to AI rap and describe the subject in up to 120 characters; the song is written for you, beat included. It cannot be mixed with an uploaded file.

Can both performers sing it as a duet?

Not with an upload: only the lead performs it. For a duet, use AI rap with Both rap so the two swap lines.

What if my track is longer than 15 seconds?

Upload it anyway. Your browser cuts out the part you want, as long as the video, and uploads only that.

Which recordings lip-sync most cleanly?

One clear voice close to the mic, little background music, and a stretch that opens with a word, not with a pause. Spoken lines work too.

Will the video sound exactly like my file?

Close, not identical. The soundtrack is generated from your recording along with the picture, so the words and timing follow it, but small differences can appear. Watch the result before you post.

Ready when you are

Hand the lead your own track

Two photos and a few seconds of audio in the Hotel Lobby AI booth: 6 credits for 12 seconds in 480p, the same as the default sound.