Dos personas, cinco ángulos, estudio blanco, micrófonos de brazo. Todo eso salió de un video de celular grabado en una recámara, en una sola toma, sin invitado y sin camarógrafo. Aquí está el prompt completo con el que se generó —sin recortar, tal cual corrió— y los cinco pasos para llegar a él: qué imagen necesitas, cómo se sacan las otras tres y qué exactamente se le entrega al modelo.
El prompt de esta página es el que corrió de verdad, el 8 de agosto de 2026, y las imágenes son las mismas que se le entregaron al modelo. Nada está reconstruido para la guía.
→ La materia prima es un screenshot de tu propia toma. Nada más.
→ Con ese screenshot, un editor de imágenes con IA te arma el plano general: mesa,
fondo blanco, entrevistador enfrente.
→ Del plano general —no de cero— salen los dos primeros planos. Ya tienes tus tres imágenes.
→ Al modelo le entras con tu video + las tres imágenes + el prompt, y devuelve el multicámara.
→ Lo que hace que se vea real no es el modelo: son el bloque de continuidad y la lista
de negativos, que ocupan casi un tercio del prompt.
Está escrito en inglés a propósito: el vocabulario de cámara —wide two-shot, locked tripod, 85mm— es el que el modelo entiende sin traducir, y los planos van numerados como en una escaleta de rodaje de verdad. Cambia sólo tres cosas: la descripción del anfitrión, la del entrevistador y las marcas de tiempo si tu toma no dura 17 segundos.
Lo que hace el trabajo pesado no es el estilo, son los seis bloques. Sobre todo dos: el que dice qué referencia manda sobre qué, y el que enumera lo que no puede pasar.
SEEDANCE 2.5 DURATION: 17 SECONDS ASPECT RATIO: 16:9 REFERENCE ASSIGNMENT: @video1 = the authoritative performance. Lock the host's exact speech, timing, mouth movement, head motion, hand gestures and body language to this video, frame for frame. Nothing about his delivery changes. @image1 = exact identity of the HOST: face, curly dark hair, build, black crew-neck t-shirt, forearm tattoo. @image2 = the authoritative SET and the exact identity of the INTERVIEWER, plus the authoritative wide two-shot composition for SHOT 01 and SHOT 05. SET (from @image2): White seamless studio, white walls meeting at a soft corner behind the table. Long light wood table running across the lower frame. Two black boom-arm broadcast microphones angled in from either side. A small black audio mixer with coloured pads at the centre of the table, a laptop lying flat beside it, black over-ear headphones resting on the table in front of each man, a cable running across the surface. Black mesh Eames-style office chairs. This set is identical in every single shot. INTERVIEWER (from @image2): Man in his early fifties, salt-and-pepper hair swept back, short greying beard, no glasses, charcoal grey knit sweater. He sits at the LEFT side of the table, in profile facing right toward the host. This is an invented character — he must not resemble any real or recognizable public figure. HE NEVER SPEAKS. His mouth stays closed for the entire clip. He is listening. HOST: The man from @image1 and @image2 — black t-shirt, curly dark hair. He sits at the RIGHT side of the table, facing left toward the interviewer. He is the only person who speaks. WORLD LOCK: Bright white podcast studio, soft even diffused key light with gentle falloff, clean natural skin tones, no colour cast, no cinematic grade. Photorealistic: real skin pores and texture, visible knit weave on the sweater, cotton texture on the t-shirt, realistic specular highlights on the microphone bodies and the wood grain of the table. Modern professional podcast production look. CAMERA LANGUAGE: This must read as a genuine multi-camera podcast edit, not as a moving AI shot. Every angle is a locked-off tripod camera with only the faintest breathing drift — no handheld shake, no push-ins, no orbits, no zooms, no dolly. All energy comes from HARD CUTTING between static cameras, like a real podcast vision switcher. 35mm on the wide, 85mm on the singles with shallow depth of field. SHOT 01 | 00:00 - 00:03 WIDE TWO-SHOT, exactly the composition, camera height and framing of @image2 — symmetrical side-on view, both men in profile, microphones and mixer between them, table across the lower frame. The host is ALREADY mid-sentence as the clip opens — no silence, no settling in, no establishing beat. The interviewer sits still, hands loosely clasped on the table, head angled toward the host. SHOT 02 | 00:03 - 00:08 HARD CUT. MEDIUM SINGLE on the HOST, 85mm, locked tripod, framed from mid-chest up, his microphone entering frame from the left, white wall softly out of focus behind him. Three-quarter angle, speaking toward the interviewer off-frame left. His hands move naturally as he talks, exactly as in @video1. SHOT 03 | 00:08 - 00:10.5 HARD CUT. MEDIUM SINGLE on the INTERVIEWER, 85mm, locked tripod, framed from mid-chest up, his microphone entering frame from the right. He listens: one slow blink, a single small nod, eyes steady on the host off-frame right. Mouth closed and still. Restrained, genuine attention — nothing theatrical. The host's voice continues uninterrupted over this shot. SHOT 04 | 00:10.5 - 00:15 HARD CUT. TIGHTER SINGLE on the HOST, 85mm — head and shoulders, closer than SHOT 02, microphone in the lower left of frame. This is the point of the clip. Hold here without cutting away until his final word is completely finished. Mouth articulation perfectly synchronized to the audio of @video1. SHOT 05 | 00:15 - 00:17 HARD CUT back to the WIDE TWO-SHOT — identical camera position and framing to SHOT 01 and @image2. The host has finished speaking and settles slightly back into his chair. The interviewer gives one small final nod. Two seconds of held silence, room tone only. No dialogue. CONTINUITY: The interviewer stays SCREEN LEFT and the host stays SCREEN RIGHT in every angle — their singles must respect this, the host always facing screen left, the interviewer always facing screen right. Both men keep the same clothes, same hair, same faces and the same seating positions throughout. Microphones, headphones, mixer, laptop, cable and chairs never move between cuts. Every shot is the same room, same light. AUDIO: Only the host's voice, taken exactly from @video1 — same voice, same language, same intonation, same pacing. No second voice. No voice-over. No narrator. No music, no score, no sound design. Only quiet studio room tone underneath. NEGATIVE: no camera movement, no handheld shake, no zoom, no push-in, no dolly, no orbit, no slow motion, no speed ramps, no subtitles, no text, no captions, no on-screen graphics, no logos, no watermarks, no split screen, no music, no second voice, no interviewer dialogue, no interviewer mouth movement, no extra people, no windows, no coloured lighting, no lens flare, no cinematic teal-orange grade, no beauty smoothing, no glasses on the interviewer, no identity drift between shots, no wardrobe changes, no crossing the line — screen directions never flip.
CONTINUITY y NEGATIVE. Sin ellos el modelo mueve la
cámara, cambia la ropa entre planos, voltea la línea de mirada y le pone música. Todo
eso, junto, es lo que delata que un video es generado.| Bloque | Qué amarra | Si falta |
|---|---|---|
| REFERENCE ASSIGNMENT | Qué referencia manda sobre qué: el video sobre la actuación, una imagen sobre tu cara, otra sobre el set | Cara distinta |
| SET / WORLD LOCK | El cuarto, los objetos sobre la mesa y la luz — descritos una vez, iguales en los cinco planos | Set que muta |
| CAMERA LANGUAGE | Todas las cámaras fijas en tripié; la energía sale del corte, no del movimiento | Look de IA |
| SHOT 01–05 | La escaleta con marcas de tiempo: qué plano entra en qué segundo | Un solo plano |
| CONTINUITY | Entrevistador siempre a la izquierda, anfitrión siempre a la derecha; nada se mueve entre cortes | Se voltea |
| AUDIO + NEGATIVE | Una sola voz, cero música — y la lista de todo lo que no puede aparecer | Segunda voz |
El orden importa más de lo que parece. Las tres imágenes que le vas a dar al modelo tienen que salir unas de otras, no generarse por separado — es lo único que garantiza que la ropa, la luz, la mesa y los micrófonos coincidan de un plano a otro.
Habla a cámara como si ya estuvieras en el podcast: sin cortes, sin moverte de sitio y con la cámara quieta. Cuando termines, pausa el video en un frame donde se te vea bien de frente y guarda el screenshot. Ese screenshot es la materia prima: de ahí sale tu cara, tu pelo, tu ropa y tu complexión.
ReglaEl frame que elijas manda sobre todo lo demás. Si sales con media cara tapada o mal iluminado, el modelo arrastra ese error a los cinco planos y no hay prompt que lo arregle.Súbelo y pídele que te siente en el set. La instrucción es tan simple como suena: «siéntalo en una mesa con fondo blanco, en un estudio de podcast, y ponle enfrente un entrevistador». Lo que regresa es el plano general de dos personas — y ese plano va a mandar después sobre el set completo.
OjoEl entrevistador tiene que ser un personaje inventado. Pedir que se parezca a alguien real o reconocible es justo lo que el prompt de la sección 01 prohíbe, y con razón.No los generes desde cero. Vuelve a subir el plano general y pide un recorte: primero «hazme un primer plano del de la derecha», y luego, otra vez desde la imagen grande, «hazme un primer plano del de la izquierda». Que las tres salgan de la misma imagen es lo que hace que todo coincida entre planos.
Ya estáCon esto tienes tus tres imágenes en fondo blanco: el plano general y los dos primeros planos.Al modelo le entras con cuatro cosas y cada una tiene un trabajo distinto: tu video crudo (manda sobre la voz, el ritmo y los gestos), tu primer plano (manda sobre tu identidad), el plano general (manda sobre el set y sobre el entrevistador) y el prompt completo. De ahí sale el multicámara de 17 segundos.
Aquí está el detalle que hace que el reel funcione: los dos videos arrancan en el segundo cero y corren en paralelo. Arriba el estudio, abajo tu recámara. Cuando el generado termina, tu toma sube al centro y ahí cae el llamado a la acción. La comparación es el contenido — sin los dos cuadros a la vez, no hay pieza.
Ninguna de las tres cosas de abajo tiene que ver con la calidad del modelo. Las tres se ganan o se pierden en el texto del prompt, y las tres son las que la gente nota sin saber por qué.
Esto funciona como contenido porque se enseña el antes y el después. El gancho del reel es la comparación, no el engaño: arriba el estudio que no existe, abajo la recámara donde de verdad se grabó. Presentarlo como material real es otra cosa, y ahí deja de ser un recurso creativo.
Sigue aprendiendo
Video en vivo · 12 sep 2026Una transmisión que nunca termina y que el público dirige
GPT Image · 22 ago 2026100 comandos para transformar cualquier foto con IA
Seedance 2.5 · 20 ago 2026Una toma fija, una toma de cine¿Lo quieres funcionando en tu negocio sin armarlo tú?
Escríbenos
Lo que acabas de leer es una receta suelta: sirve para una pieza. Lo que de verdad mueve un negocio es el sistema alrededor — que las ideas, los guiones, los videos y las respuestas a quien comenta se muevan solos, cada semana, sin que tú estés empujando. En Sotai armamos justo eso, a la medida de tu giro. Escríbenos, cuéntanos qué haces y en cinco minutos te decimos si te conviene o si con esta guía ya la hiciste.
Escríbenos →
Guía preparada por Sotai · 8 de agosto de 2026 · No estamos afiliados a ByteDance ni a Seedance