Yu, Jongmin, Sun, Zhongtian, Panagiotis Alexandridis, Konstantinos, Aviles-Rivero, Angelica I., Yang, Jinhong (2026) A latent diffusion for stable frame interpolation. IEEE Transactions on Neural Networks and Learning Systems, . pp. 1-14. ISSN 2162-237X. E-ISSN 2162-2388. (doi:10.1109/TNNLS.2026.3713915) (The full text of this publication is not currently available from this repository. You may be able to access a copy if URLs are provided) (KAR id:116729)
| The full text of this publication is not currently available from this repository. You may be able to access a copy if URLs are provided. | |
| Contact us about this publication | |
| Official URL: https://doi.org/10.1109/TNNLS.2026.3713915 |
|
Abstract
This article presents a novel approach to video frame interpolation (VFI), called latent diffusion for stable frame interpolation (LD4SFI). LD4SFI leverages a latent diffusion model (LDM) enhanced by a vector-quantized spatiotemporal variational autoencoder (VQ-STVAE). Our method captures intrinsic orthogonal relationships in high-dimensional spatiotemporal data and seamlessly integrates complementary information across diverse video frame sequences. Given the robust sampling capabilities of LDMs, LD4SFI is conditioned on a disparity map that describes the motion dynamics between two neighboring frames. The disparity map is applied to feed explicit spatiotemporal differences into a diffusion model (DM), thereby improving the spatiotemporal smoothness of the interpolation results. With this, LD4SFI efficiently generates interpolated frames, significantly improving the continuity and visual quality of the video content. On UCF-101, densely annotated video segmentation (DAVIS), and SNU-FILM, LD4SFI achieves competitive or improved accuracy relative to existing state-of-the-art (SOTA) methods. LD4SFI produces 0.016 learned perceptual image patch similarity (LPIPS), 36.219 peak signal-to-noise ratio (PSNR), 0.974 structural similarity index (SSIM), 0.031 FloLPIPS, and 20.105 Fréchet inception distance (FID) for the UCF-101 dataset, and 0.072 LPIPS, 30.261 PSNR, 0.912 SSIM, 0.108 FloLPIPS, and 8.037 FID for the DAVIS dataset. Additionally, LD4SFI achieves the highest measured throughput among DMs on a single V100. Experimental results show that LD4SFI outperforms existing SOTA methods, demonstrating highly competitive SOTA performance across standard VFI benchmarks.
| Item Type: | Article |
|---|---|
| DOI/Identification number: | 10.1109/TNNLS.2026.3713915 |
| Additional information: | For the purpose of open access, the author(s) has applied a Creative Commons Attribution (CC BY) licence to any Author Accepted Manuscript version arising. |
| Subjects: | Q Science > QA Mathematics (inc Computing science) |
| Institutional Unit: | Schools > School of Computing |
| Former Institutional Unit: |
There are no former institutional units.
|
| Depositing User: | Zhongtian Sun |
| Date Deposited: | 06 Oct 2026 08:49 UTC |
| Last Modified: | 06 Oct 2026 08:49 UTC |
| Resource URI: | https://kar.kent.ac.uk/id/eprint/116729 (The current URI for this page, for reference purposes) |
- Export to:
- RefWorks
- EPrints3 XML
- BibTeX
- CSV
- Depositors only (login required):

https://orcid.org/0000-0003-0489-5203
Altmetric
Altmetric