Skip to main navigation Skip to search Skip to main content

Domain-Adaptive Panoramic Video Generation using Image Diffusers

Research output: Contribution to journalArticlepeer-review

Abstract

Creating panoramic videos is still a challenging task, mainly due to the lack of large, well-annotated training datasets and the heavy computing power required for realistic, high-fidelity synthesis. These problems become even more severe if the proposed method is dealing with specialized domains such as automotive or indoor environments. Most of the time, the spatial or temporal aspect of precision gets lost when traditional methods are applied to wide-angle frames. As a result, distortions or inconsistent motion between views happen. To overcome these problems, the proposed method came up with a hierarchical two-stage framework called DAPanVidGen (Domain-Adaptive Panoramic Video Generation using Image Diffusers). The first stage is a text-to-image diffusion model that is fine-tuned on panoramic datasets so that it can better grasp the geometry and visual priors of wide-angle imagery. The second stage builds upon this basis and trains the image latents to become video-level latents that describe motion across frames. During execution, the adapted image model initially generates a rough latent representation that is further transformed into a series of video latents by a motion-aware recurrent temporal-attention diffusion operation. This operation allows the system to keep the temporal coherence of the videos and not to flicker in different parts of the panoramic scenes. To further support the wrap-around continuity, the framework is equipped with a special video-latent branch along with a multi-diffusion fusion component. An adaptive decoding module is capable of ensuring that the changes from one frame to the next are not visible, even in the case of long video sequences. Visually stable and semantically rich results in outdoor as well as indoor panoramic scenarios confirmed by the proposed method when assessed on such benchmarks as KITTI-360, Sun360, and Cityscapes.

Original languageEnglish
Pages (from-to)30984-30996
Number of pages13
JournalIEEE Access
Volume14
DOIs
Publication statusAccepted/In press - 2026

All Science Journal Classification (ASJC) codes

  • General Computer Science
  • General Materials Science
  • General Engineering

Fingerprint

Dive into the research topics of 'Domain-Adaptive Panoramic Video Generation using Image Diffusers'. Together they form a unique fingerprint.

Cite this