Breaking Down the AI Video Production Workflow

AI Video production social media

We built a repeatable AI video production workflow that does what seemed impossible: convincingly place real people into different eras of history.

The Rob in History series began as an internal research and development project inspired by a conversation with a potential client. The client wanted to explore using AI to transform a real person into a digital version of themselves for educational content, allowing them to interact with experts and historical subjects in a more engaging way.

That discussion sparked an idea within our team. Rather than simply explaining the technology, we decided to create our own proof of concept by placing Robert into different moments in history and building entertaining educational videos around him. The first episode took Robert back to the silent film era, while the second episode, First Flight Debate, explored the historical discussion surrounding the first powered flight by placing Robert and Eduardo in conversation with Alberto Santos Dumont.

Our goal was not simply to generate AI videos. It was to develop a repeatable production workflow capable of transforming real people into believable historical characters while maintaining visual consistency, storytelling, and cinematic quality.

Production Details

Script Development and Creative Direction

Every episode begins with the story.

For the first episode, we developed a script that placed Robert in the early days of cinema. The concept centered on the importance of video communication and how even silent films were capable of moving audiences emotionally. The story then connected that historical perspective with today’s world, showing how modern video combines image, sound, music, and storytelling to communicate even more effectively.

The creative direction focused heavily on the visual language of the 1910s and 1920s. The silent film era became the foundation for every artistic decision, from wardrobe and locations to camera composition, title cards, lighting, and pacing. The visual inspiration came from the classic black-and-white cinema associated with Charles Chaplin and the pioneers of silent filmmaking.

Once the story was defined, we broke the script into individual scenes, identified the most important storytelling moments, created a storyboard, and established the visual rhythm of each sequence before generating any images or animation.

For the second episode, First Flight Debate, we followed the same workflow while expanding the creative challenge. Instead of one main character, the story required Robert and Eduardo to interact naturally while debating aviation history with Alberto Santos Dumont. Writing believable dialogue between multiple AI-generated characters introduced a new level of planning and complexity during the storyboard stage.

Visual Style Development and AI Image Generation

After the storyboard was approved, we focused on creating the visual identity for each episode.

Several AI platforms were used together because no single application consistently produced the best results for every situation.

Primary tools included:

  • ChatGPT
  • Midjourney
  • Gemini
  • Adobe Photoshop, including its AI tools such as NanoBanana

One important difference from many AI projects was that we first built Robert as a reusable character. We photographed him from multiple angles and used those reference images to establish facial consistency before generating the scenes. This became the foundation for every historical version of Robert throughout the series.

For the second episode, the process expanded considerably with the introduction of Eduardo as a second recurring character. Building two consistent characters that could interact naturally in the same scene proved significantly more challenging than creating a single AI character. Maintaining facial consistency, body proportions, clothing, and interaction between both characters required continuous refinement throughout production.

As with every AI production, images often moved between different platforms before reaching their final form. Photoshop played a major role in refining compositions, correcting colors, cleaning imperfections, extending images, and preparing improved reference material for future generations.

Video Length and Production Complexity

Before generating animation, we built what we call the video skeleton.

Using Adobe Premiere Pro, we assembled the storyboard frames into a complete sequence, established temporary timing, added music, sound effects, narration, dialogue, and rough pacing so we could evaluate the story before investing significant time generating final animation.

This early edit became the production roadmap. It allowed us to determine how long each shot should be, where transitions belonged, which scenes required additional emphasis, and whether the pacing effectively supported the storytelling.

Because the series combines education with entertainment, balancing historical information and viewer engagement became one of the most important creative considerations throughout production.

Animation, Voice Over, and Audio Production

With the timing established, we began producing the animated shots.

For animation, we primarily used Runway and Artlist AI, testing different models depending on the specific needs of each shot. The models we relied on most were Kling 3.0, Seedance 2.0, and Runway Gen-4. Different scenes often benefited from different models, so we regularly compared multiple generations before selecting the version that best supported the story.

Like most AI productions, animation became a highly iterative process.

Reference images evolved continuously.

Prompts were rewritten repeatedly.

Scenes were regenerated whenever facial consistency, camera movement, expressions, or storytelling needed improvement.

Unlike some of our larger narrative productions, this series did not require original music. Instead of composing songs with Suno, we selected licensed music from our existing production music library that best matched the tone and historical period of each episode.

ElevenLabs was used to generate narration, create dialogue, clone voices when necessary, and refine delivery and timing so each performance matched the pacing and emotion of the story.

The second episode introduced an entirely new level of complexity because two recurring AI characters needed to speak, react, and interact naturally throughout the film. Achieving believable conversations required considerably more experimentation than producing scenes with only a single character.

Editing, Motion Graphics, and Post Production

Once enough animated shots had been generated, we assembled the final edit.

Adobe Premiere Pro became the central production environment.

The editing process involved selecting the strongest generated takes, trimming clips to capture only their best moments, replacing weaker animations with improved generations, refining pacing, building transitions, synchronizing dialogue and narration, and carefully mixing the music and sound effects into a cohesive final experience.

Adobe After Effects was used for compositing, transitions, and additional visual enhancements.

Photoshop continued to play an important role after animation generation by correcting colors, refining details, and preparing improved reference images whenever additional AI generations were needed.

Throughout production, several scenes required returning to the very beginning of the workflow. Rather than accepting weaker animation results, we often revisited the original reference images, refined them inside Photoshop, adjusted prompts, and regenerated the animation until the quality met our expectations.

Software Requirements, Revisions, and Production Experience

One of the biggest lessons learned during the Rob in History series was that AI filmmaking is rarely a linear process.

Very few shots are perfect on the first attempt.

Many scenes required multiple prompt revisions, new reference images, testing different AI models, regenerating animations several times, combining the best sections from multiple generations, and adapting the storytelling to the strengths and limitations of current AI technology.

The introduction of multiple recurring characters made this challenge even greater. Small inconsistencies that might go unnoticed with one character became much more visible once Robert and Eduardo appeared together, making careful iteration even more important.

An important part of professional AI production is recognizing when a scene has reached a publishable level of quality. While it is always possible to continue generating additional versions in search of an even better result, every production eventually reaches a point where the creative vision has been achieved well enough to move the project forward. Knowing when to continue refining a shot and when to commit to the strongest available version is an important part of the production experience.

This workflow closely resembles traditional filmmaking.

Final Thoughts

This series demonstrates how AI can be used not only to generate individual images or videos, but to build recurring characters capable of telling engaging educational stories across multiple episodes.

By combining traditional filmmaking techniques with modern AI tools, we developed a repeatable production workflow that transforms historical subjects into entertaining, visually engaging experiences while maintaining consistency across episodes.

The project reinforced an important lesson about AI filmmaking. The tools accelerate production, but they do not replace the creative process. Storytelling, planning, visual design, editing, sound, pacing, revision, and creative judgment remain the elements that ultimately determine the quality of the finished work.

Reach out to us for a free strategy consultation!

Contact

Let’s TALK!

1330 Fifth Ave, Ste 6e
New York, NY 10026
(646) 319-8609
rweiss@multivisiondigital.com

Second office address
City, State ZIP
(646) 319-8609
rweiss@multivisiondigital.com

Starting a new project? Contact Us!







     

    Copyright © 2026 MultiVisionDigital – Online Video Production Company for New Jersey & NYC, New York.
    All rights reserved.