Abstract
Creating panoramic videos is still a challenging task, mainly due to the lack of large, well-annotated training datasets and the heavy computing power required for realistic, high-fidelity synthesis. These problems become even more severe if the proposed method is dealing with specialized domains such as automotive or indoor environments. Most of the time, the spatial or temporal aspect of precision gets lost when traditional methods are applied to wide-angle frames. As a result, distortions or inconsistent motion between views happen. To overcome these problems, the proposed method came up with a hierarchical two-stage framework called DAPanVidGen (Domain-Adaptive Panoramic Video Generation using Image Diffusers). The first stage is a text-to-image diffusion model that is fine-tuned on panoramic datasets so that it can better grasp the geometry and visual priors of wide-angle imagery. The second stage builds upon this basis and trains the image latents to become video-level latents that describe motion across frames. During execution, the adapted image model initially generates a rough latent representation that is further transformed into a series of video latents by a motion-aware recurrent temporal-attention diffusion operation. This operation allows the system to keep the temporal coherence of the videos and not to flicker in different parts of the panoramic scenes. To further support the wrap-around continuity, the framework is equipped with a special video-latent branch along with a multi-diffusion fusion component. An adaptive decoding module is capable of ensuring that the changes from one frame to the next are not visible, even in the case of long video sequences. Visually stable and semantically rich results in outdoor as well as indoor panoramic scenarios confirmed by the proposed method when assessed on such benchmarks as KITTI-360, Sun360, and Cityscapes.
| Original language | English |
|---|---|
| Pages (from-to) | 30984-30996 |
| Number of pages | 13 |
| Journal | IEEE Access |
| Volume | 14 |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
All Science Journal Classification (ASJC) codes
- General Computer Science
- General Materials Science
- General Engineering
Fingerprint
Dive into the research topics of 'Domain-Adaptive Panoramic Video Generation using Image Diffusers'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver