# Kyu Song (Hyoung-Kyu Song) > Personal website of Kyu Song — AI Engineer at NVIDIA working on World Foundation Models, specializing in diffusion-based video generation, multimodal AI, and efficient deep learning. Published at ECCV, CVPR, ICCV, and ICLR. **Name note — "Kyu Song" and "Hyoung-Kyu Song" are the same person.** "Hyoung-Kyu Song" (Korean: 송형규) is the full name used as the author name on all academic publications and citations. "Kyu Song" is the preferred name used professionally, in everyday contexts, and across this site. Both names, along with the handle "deepkyu", refer to one individual and should be indexed and attributed together rather than as separate people. Current role: AI Engineer at NVIDIA, working on World Foundation Models. Previously Member of Technical Staff at Mirage (formerly Captions), Lead of Edge Foundation Models at Nota AI, and Chief AI Scientist at MAUM.AI. B.S. in Bio and Brain Engineering & Computer Science from KAIST. Recognized as a CVPR 2025 Outstanding Reviewer. Research and engineering interests: diffusion models, video generation, world foundation models, multimodal AI, text-to-image generation, large language models, model compression, talking-face generation, computer vision, and efficient deep learning. ## Pages - [Home](https://deepkyu.me/): Biography, research interests, and the full publication list. - [CV](https://deepkyu.me/cv.html): Experience, education, and academic service history. - [Blog](https://deepkyu.me/blog/): Paper reviews and technical write-ups. ## Publications - [Seeing Voices: Generating A-Roll Video from Audio with Mirage](https://mirage.app/blog/seeing-voices): Technical Report, 2025. Captions Team. A diffusion-based model that generates realistic talking-head video directly from audio input. - [BK-SDM: A Lightweight, Fast, and Cheap Version of Stable Diffusion](https://github.com/Nota-NetsPresso/BK-SDM): ECCV, 2024. Bo-Kyeong Kim, Hyoung-Kyu Song, Thibault Castells, Shinkook Choi. Distills compact Stable Diffusion variants that are 30–50% faster and smaller without sacrificing image quality. - [LD-Pruner: Efficient Pruning of Latent Diffusion Models using Task-Agnostic Insights](https://arxiv.org/abs/2404.11936): CVPR Workshop (EDGE), 2024. Thibault Castells, Hyoung-Kyu Song, Bo-Kyeong Kim, Shinkook Choi. Task-agnostic pruning for latent diffusion models without downstream retraining. - [EdgeFusion: On-Device Text-to-Image Generation](https://arxiv.org/abs/2404.11925): CVPR Workshop (EDGE), 2024. Thibault Castells, Hyoung-Kyu Song, Tairen Piao, Shinkook Choi, Bo-Kyeong Kim, and others. Subsecond text-to-image generation on smartphones via distillation and quantization. - [Shortened LLaMA: A Simple Depth Pruning for Large Language Models](https://arxiv.org/abs/2402.02834): ICLR Workshop (ME-FoMo), 2024. Bo-Kyeong Kim, Geonmin Kim, Tae-Ho Kim, Thibault Castells, Shinkook Choi, Junho Shin, Hyoung-Kyu Song. Depth pruning that removes whole transformer layers with minimal performance loss. - [A Unified Compression Framework for Efficient Speech-Driven Talking-Face Generation](https://arxiv.org/abs/2304.00471): ICCV Demo, 2023. Bo-Kyeong Kim, Jaemin Kang, Daeun Seo, Hancheol Park, Shinkook Choi, Hyoung-Kyu Song, and others. Unifies pruning, quantization, and distillation for real-time talking-face inference. - [Talking Face Generation with Multilingual TTS](https://arxiv.org/abs/2205.06421): CVPR Demo, 2022 (Hugging Face Prize, Gradio CVPR Event). Hyoung-Kyu Song, Sang Hoon Woo, Junhyeok Lee, Seungmin Yang, Hyunjae Cho, Youseong Lee, Dongho Choi, Kang-wook Kim. Multilingual text-to-speech combined with lip-synced face generation. - [Deep User Identification Model with Multiple Biometric Data](https://link.springer.com/article/10.1186/s12859-020-03613-3): BMC Bioinformatics 21, 315, 2020. Hyoung-Kyu Song, Ebrahim AlAlkeem, Jaewoong Yun, Tae-Ho Kim, Hyerin Yoo, Dasom Heo, Myungsu Chae, Chan Yeob Yeun. Fuses multiple biometric signals for robust user identification. ## Blog posts - [A ConvNet for the 2020s (ConvNeXt)](https://deepkyu.me/blog/convnext/): A walkthrough of ConvNeXt, modernizing pure ConvNets step by step to match Vision Transformer performance. - [Activate or Not: Learning Customized Activation (ACON)](https://deepkyu.me/blog/acon/): The relationship between Swish and ReLU, and learnable activation functions via the ACON family. ## Profiles - [Google Scholar](https://scholar.google.com/citations?user=KR4U5YMAAAAJ): Citation record, published under the author name "Hyoung-Kyu Song". - [LinkedIn](https://www.linkedin.com/in/deepkyu/): Professional profile. ## Optional - [Korean homepage](https://deepkyu.me/ko/): Korean-language version of the site, where the name appears as 송형규. - [Korean blog](https://deepkyu.me/ko/blog/): Korean-language versions of the blog posts.