Where does the demand for multiple versions of generative product videos come from?

When preparing a product video project, brands often need to cover multiple placement scenarios simultaneously. The same product may need to be presented on e-commerce detail pages, social media feeds, offline exhibition screens, and overseas advertising platforms. User attention, viewing environments, and content preferences vary greatly across different platforms, so directly reusing a single final video usually fails to meet expectations. The advantage of generative product videos is that, based on the same set of visual assets, different packaging versions can be quickly produced through an AI workflow, such as landscape and portrait formats, 15-second and 60-second cuts, Chinese voiceovers and English subtitles, or product feature demonstrations and brand emotional expressions.

AIGC commercial footage from the case study materials, observing the relationship between the camera, subject, and lighting.
Case study material stills, sourced from the research document "Ready Player One inside the oasis". This footage is used solely to observe camera work and production methods, and does not represent an ONCE client project. Source page Case study material page

The first step in project initiation focuses on defining the version matrix. The brand needs to list all planned placement platforms, the material specification requirements for each platform, target audience characteristics, and core communication messages. Only after receiving this list can the production team determine which versions can be mass-produced using generative technology and which require separate shooting or post-production. If a brand has not even finalized the placement platforms but demands the production team to deliver ten versions first, the project will most likely require rework, because the narrative pacing and visual composition of each version are completely different.

The standard for determining whether a version is suitable for AIGC generation is whether its visual elements are controllable. If a version needs to showcase the product's true materials, precise dimensions, or colors under specific lighting, the risk of AI generation is higher because generative models have limited ability to reproduce physical details. If a version focuses on abstract atmospheres, conceptual scenes, and stylized transitions, AIGC can significantly improve efficiency. Brands should break down the requirements of each version into a checklist of visual elements, marking item by item which must be real and which can accept generative representation.

Four types of materials brands must prepare during the project initiation phase

Insufficient material preparation is the most common cause of delays in generative product video projects. Brands must prepare at least four types of materials, none of which can be missing.

The first type is product assets, including high-definition product images, photos from different angles, 3D model files, and product demonstration videos. If the product has not yet been mass-produced, design drawings, renderings, or prototype photos must be provided. These materials are used for AI model training or reference image generation, and their quality directly affects the product's fidelity in the final video. The second type is brand assets, including the brand logo, standard color values, font specifications, links to past commercials, and brand visual reference images. Generative models need to understand the brand tone; otherwise, they may easily produce visuals that deviate from the intended style.

The third type is communication information, including the core selling points of this campaign, target audience profiles, competitor video references, and visual elements to avoid. For example, if the product focuses on safety performance, the visuals should not feature dangerous driving or exaggerated actions. The fourth type is technical constraints, including delivery platforms, delivery formats, resolution requirements, subtitle languages, whether multilingual voiceovers are needed, and copyright ownership requirements. This information must be written into the project brief so the production team can evaluate the workload and risks.

When preparing materials, brands must pay special attention to the selection of reference visuals. Reference visuals are primarily used to convey the stylistic direction. It is recommended to provide three to five video clips of different styles, with each clip annotated with the visual elements or pacing you wish to reference. At the same time, it must be clarified which reference visuals are copyright-restricted to prevent the generative model from accidentally imitating protected content during training.

How to leave room for generative post-production during the shooting phase

Even if the visuals rely primarily on AIGC generation, the shooting phase cannot be omitted. Live-action footage is used to provide product details, real lighting and shadows, actor expressions, and scene structures, which are difficult for AI to generate from scratch. During shooting, the production team must leave room for post-production generation. Specific practices include shooting product footage against solid-color backgrounds, multi-angle lighting separation footage, empty shots of scenes without people, and scene footage with tracking points.

Solid-color background footage is used for post-production keying and AI repainting; it is recommended to use a green or gray screen and avoid backgrounds with colors similar to the product. Multi-angle lighting separation footage refers to shooting the same product under different lighting conditions to facilitate the compositing of different atmospheres in post-production. Empty shots are used for AI-generated scene extensions, such as shooting a desk and later generating a product usage scene on it. Tracking point footage is used for moving shots, allowing AI-generated elements to attach to the real scene.

During shooting, brands must confirm whether the production team has a generative post-production workflow. Traditional shooting teams may only deliver raw footage and not be responsible for AI generation. Brands need to specify in the contract whether the format, codec, and color space of the shooting footage meet the requirements of AI tools. For example, some AI tools have minimum resolution requirements for footage or require specific frame rates. If these are not considered during shooting, the footage may not be directly usable in post-production, requiring reshoots or additional processing.

A common risk is that production teams use extensive color grading and filters to pursue visual aesthetics, resulting in the loss of color information in the footage. Generative models are sensitive to color, and if the footage's color is over-compressed, the AI-generated results may suffer from color shifts. Therefore, it is recommended to retain an ungraded raw file during shooting for AI processing. When reviewing the shot footage, brands should check for the presence of original color files and ensure they contain sufficient product details.

Generation strategies for different packaging versions in post-production

In the post-production phase, the core task of generative product videos is to transform a single set of visual assets into different versions. The production team must first establish a base version before deriving other versions. The base version is typically the most comprehensive and longest iteration, such as a 60-second brand promotional video. Other versions, such as 15-second social media shorts, 9-second bumper ads, and vertical formats, are extracted or reconstructed from the base version.

Generative tools can be used for scene extension, background replacement, style transfer, and dynamic interpolation. For example, if the base version features a product close-up with a live-action desktop background, generative tools can replace the background with abstract flowing particles to create a tech-focused version. Alternatively, live-action shots can be converted into a hand-drawn style to suit youth-oriented platforms. These operations require the production team to possess prompt engineering skills to precisely control the generated results.

Version differences extend beyond duration and resolution. Different versions require distinct narrative logic. A 15-second version might showcase only one core selling point, using a single voiceover line alongside dynamic product visuals. A 60-second version can tell the product development story, including user pain points, solutions, and usage scenarios. When reviewing versions, brands should focus on whether each version stands alone, rather than simply extracting clips from the base version. If a version is merely a cut, the audience will feel it is incomplete, and the campaign's effectiveness will be compromised.

In post-production, audio and subtitles are also part of version differentiation. Different language versions require re-dubbing and subtitling; generative voice tools can reduce dubbing costs, but brands must confirm that the voice style aligns with their brand image. Subtitle styling, placement, and timing must also be adjusted according to the platform; for instance, subtitles in vertical versions are usually centered, while those in horizontal versions can be placed at the bottom. The production team should provide subtitle style guidelines, and brands should inspect them item by item during acceptance.

Checklists and evaluation criteria for the acceptance phase

When accepting generative product videos, brands should not only look at the final output but also inspect intermediate files and metadata. Brands should establish a phased acceptance checklist that includes scripts, storyboards, shot footage, edited versions, color-graded versions, audio versions, subtitle files, masters, and source files. Each phase must have clear acceptance criteria to avoid discovering issues only during a final, unified review.

The acceptance criteria for scripts are accurate information, prominent selling points, and a language style that matches the brand's tone. The acceptance criteria for storyboards are clear visual descriptions, coherent shot logic, and each shot labeled with its generative or live-action method. The acceptance criteria for shot footage are that resolution, frame rate, and color space meet presets, product details are complete, and there are no continuity errors. The acceptance criteria for edited versions are that the pacing meets platform expectations, transitions are smooth, and the information hierarchy is clear. The acceptance criteria for color-graded versions are unified colors, accurate product colors, and coordinated brightness across different scenes.

Audio acceptance includes voiceovers, background music, and sound effects. Voiceovers must be checked for accurate pronunciation, appropriate pacing, and matching emotion. Background music must be checked for copyright licensing scope to ensure it covers all distribution platforms. Sound effects must be checked for synchronization with on-screen actions. Subtitle acceptance requires checking for text accuracy, timeline synchronization, and consistent styling. Master and source file acceptance requires checking format, codec, resolution, color depth, and whether editable project files are included.

During acceptance, brands must pay special attention to the copyright issues of AI-generated content. The training data for generative models may contain copyrighted material, and the generated results may produce similar content. Brands should require the production team to provide the versions and setting records of the generative tools, as well as copyright statements. If the brand plans to use the video for commercial advertising, it must confirm that the generated content does not infringe on third-party rights. It is recommended to clearly state in the contract that the production team will bear responsibility for any copyright disputes arising from the generated content.

Situations where generative product videos are not suitable

Not all product videos are suitable for AIGC generation. Before project initiation, brands should evaluate the following unsuitable conditions to avoid investing resources without achieving expectations.

The first situation is when the product appearance requires extremely high precision, such as medical devices, precision instruments, and automotive exteriors. The dimensions, colors, and materials of these products must be presented accurately, and AI generation is prone to slight deviations that could lead to consumer misunderstanding. The second situation is when the product involves safety certifications or legal compliance, such as food, pharmaceuticals, and children's products. These videos need to show real testing processes or certification marks, and AI-generated content may not pass the review.

The third situation is when the brand requires highly consistent visual assets, such as globally unified commercials. If the brand requires all markets to use exactly the same visuals and audio, the generative version may cause visual inconsistencies due to platform differences, thereby increasing management costs. The fourth situation is when the budget and schedule are extremely tight, but the number of versions is massive. Although the generative workflow can produce content in batches, each version still requires manual review and adjustment, making it difficult to guarantee quality if time is insufficient.

The fifth situation is when the brand lacks internal review capabilities. Every frame of a generative video may contain AI hallucinations, requiring a professional team to check them frame by frame. If the brand does not have enough internal staff or external consultants to conduct the review, it is recommended not to use AIGC, or to choose a hybrid mode where key shots are live-action and supplementary shots are generated.

Next steps

If a brand is considering an AIGC commercial or AI product video project, it is recommended to conduct a minimum viable test first. Select one product, determine two versions—one landscape and one portrait—and have the production team use a generative workflow to produce a sample video. Keep the testing cycle within two weeks and the budget within 30% of a standard shoot. The purpose of the test is to verify the generative tools' accuracy in reproducing product details, the efficiency of version switching, and whether the internal review process is smooth.

After the test is completed, the brand should evaluate whether the results meet the communication goals. If the sample video meets expectations, then plan the full version matrix. If the sample video has flaws, determine whether it is a tool limitation or a workflow issue before deciding whether to adjust the plan. Generative product video is a rapidly iterating field, and brands do not need to pursue perfection in one step, but rather accumulate experience through small-scale trial and error. The services publicly listed on the ONCE website include AIGC commercials, AI product videos, generative imagery, and brand AI video workflows, which can serve as a reference starting point. The final decision should be based on the brand's actual needs, budget, and team capabilities, rather than chasing technological concepts.

If you are preparing an AIGC commercial project, you can first organize the brief, reference visuals, product or company materials, delivery platforms, and copyright scope, and then viewAIGC Video Services Page, bringing communication from abstract preferences down to executable production boundaries.