AI-Driven 3D/4D Content Generation, Editing, and Animation Using Diffusion Models and Gaussian Splatting: A Narrative Review

Main Article Content

Yanning Liu

Keywords

3D Gaussian Splatting, diffusion models, text-to-3D generation, 4D generation, 3D editing, neural rendering, score distillation sampling, motion generation, narrative review

Abstract

The convergence of diffusion models and 3D Gaussian Splatting (3D -GS) has catalyzed a paradigm shift in AI-driven 3D/4D content creation, enabling high-quality generation, editing, and animation from natural language instructions. This narrative review syn thesizes representative works published between 2023 and 2025 across four interconnected sub-areas: text-to-3D/4D generation, 3D scene editing, dynamic scene animation, and motion and video generation. Through comparative analysis of representation types, supervision strategies, and temporal modeling approaches, we identify convergent design principles—coarse- to-fine optimization in distillation-based methods, explicit representations, regularization against score - distillation variance, and decomposition for controllability—that recur across sub-areas. We catalog six critical limitations and propose concrete directions for improvement for each. This review serves as a foundational reference for researchers entering this rapidly evolving field.

Abstract 9 | PDF Downloads 3

References

  • [1] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. NeurIPS, 2020.
  • [2] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in Proc. ICLR, 2021.
  • [3] B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “DreamFusion: Text-to-3D using 2D diffusion,” in Proc. ICLR, 2023.
  • [4] B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3D Gaussian splatting for real-time radiance field rendering,” ACM Trans. Graph., vol. 42, no. 4, pp. 1–14, 2023.
  • [5] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” in Proc. ECCV, 2020.
  • [6] C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M. -Y. Liu, and T. -Y. Lin, “Magic3D: High-resolution text-to-3D content creation,” in Proc. IEEE/CVF CVPR, 2023, pp. 300– 310.
  • [7] H. Ling, S. W. Kim, A. Torralba, S. Fidler, and K. Kreis, “Align your Gaussians: Text-to-4D with dynamic 3D Gaussians and composed diffusion models,” in Proc. IEEE/CVF CVPR, 2024, pp. 8576–8588.
  • [8] Y. Jiang, C. Yu, C. Cao, F. Wang, W. Hu, and J. Gao, “Animate3D: Animating any 3D model with multi- view video diffusion,” arXiv:2407.11398v2, 2024.
  • [9] A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa, “Instruct-NeRF2NeRF: Editing 3D scenes with instructions,” in Proc. IEEE/CVF ICCV, 2023, pp. 19683–19693.
  • [10] Y. Chen, Z. Chen, C. Zhang, F. Wang, X. Yang, Y. Wang, Z. Cai, L. Yang, H. Liu, and G. Lin, “GaussianEditor: Swift and controllable 3D editing with Gaussian splatting,” in Proc. IEEE/CVF CVPR, 2024, pp. 21476–21485.
  • [11] Y.-H. Huang, Y. -T. Sun, Z. Yang, X. Lyu, Y. -P. Cao, and X. Qi, “SC-GS: Sparse-controlled Gaussian splatting for editable dynamic scenes,” in Proc. IEEE/CVF CVPR, 2024, pp. 4220–4230.
  • [12] T. Wimmer, M. Oechsle, M. Niemeyer, and F. Tombari, “Gaussians-to-Life: Text-driven animation of 3D Gaussian splatting scenes,” arXiv:2411.19233v2, Mar. 2025.
  • [13] M. Zhang, Z. Cai, L. Pan, F. Hong, X. Guo, L. Yang, and Z. Liu, “MotionDiffuse: Text-driven human motion generation with diffusion model,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 6, pp. 4115–4128, Jun. 2024.
  • [14] J. Xing, M. Xia, Y. Liu, Y. Zhang, Y. Zhang, Y. He, H. Liu, H. Chen, X. Cun, X. Wang, Y. Shan, and T.- T. Wong, “Make-Your-Video: Customized video generation using textual and structural guidance,” IEEE Trans. Vis. Comput. Graph., vol. 31, no. 2, pp. 1526–1541, Feb. 2025.
  • [15] Y. Men, Y. Yao, M. Cui, and L. Bo, “MIMO: Controllable character video synthesis with spatial decomposed modeling,” in Proc. IEEE/CVF CVPR, 2025, pp. 21181–21191.
  • [16] B. Fei, J. Xu, R. Zhang, Q. Zhou, W. Yang, and Y. He, “3D Gaussian splatting as a new era: A survey,” IEEE Trans. Vis. Comput. Graph., vol. 31, no. 8, pp. 4429–4449, 2024.
  • [17] T. Wu, G. Yang, Z. Li, K. Zhang, Z. Liu, L. Guibas, D. Lin, and G. Wetzstein, “GPT -4V(ision) is a human-aligned evaluator for text-to-3D generation,” in Proc. IEEE/CVF CVPR, 2024, pp. 22227–22238.
  • [18] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proc. IEEE/CVF CVPR, 2022.
  • [19] C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, et al., “Photorealistic text-to-image diffusion models with deep language understanding,” in Proc. NeurIPS, 2022.
  • [20] Y. Huang, J. Wang, A. Zeng, H. Cao, X. Qi, Y. Shi, Z.-J. Zha, and L. Zhang, “DreamWaltz: Make a scene with complex 3D animatable avatars,” in Proc. NeurIPS, 2023.
  • [21] Z. Wang, C. Lu, Y. Wang, F. Bao, Z. Li, H. Su, and J. Zhu, “ProlificDreamer: High-fidelity and diverse text-to-3D generation with variational score distillation,” in Proc. NeurIPS, 2023.
  • [22] J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng, “DreamGaussian: Generative Gaussian splatting for efficient 3D content creation,” in Proc. IEEE/CVF CVPR, 2024, pp. 872–882.
  • [23] J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “LGM: Large multi-view Gaussian model for high-resolution 3D content creation,” in Proc. ECCV, 2024.
  • [24] Y. Shi, P. Wang, J. Ye, L. Mai, K. Li, and X. Yang, “MVDream: Multi-view diffusion for 3D generation,” in Proc. ICLR, 2024.
  • [25] J. Chung, S. Lee, H. Nam, J. Lee, and K. M. Lee, “LucidDreamer: Domain-free generation of 3D Gaussian splatting scenes,” arXiv:2311.13384, 2024.
  • [26] T. Xie, Z. Zong, Y. Qiu, X. Li, Y. Feng, Y. Yang, and C. Jiang, “PhysGaussian: Physics-integrated 3D Gaussians for generative content creation,” in Proc. IEEE/CVF CVPR, 2024, pp. 6766–6775.
  • [27] Z. Yang, X. Gao, W. Zhou, S. Jiao, Y. Zhang, and X. Jin, “Deformable 3D Gaussians for high-fidelity monocular dynamic scene reconstruction,” in Proc. IEEE/CVF CVPR, 2024, pp. 20341–20351.