2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Paper • 2501.00958 • Published 6 days ago • 87
PERSE: Personalized 3D Generative Avatars from A Single Portrait Paper • 2412.21206 • Published 8 days ago • 15
Byte Latent Transformer: Patches Scale Better Than Tokens Paper • 2412.09871 • Published 26 days ago • 85
ColorFlow: Retrieval-Augmented Image Sequence Colorization Paper • 2412.11815 • Published 23 days ago • 26
MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation Paper • 2412.04448 • Published Dec 5, 2024 • 9
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions Paper • 2411.07461 • Published Nov 12, 2024 • 22
Acoustic Volume Rendering for Neural Impulse Response Fields Paper • 2411.06307 • Published Nov 9, 2024 • 5
Learning Video Representations without Natural Videos Paper • 2410.24213 • Published Oct 31, 2024 • 15
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation Paper • 2410.13726 • Published Oct 17, 2024 • 11
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis Paper • 2410.08261 • Published Oct 10, 2024 • 50
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation Paper • 2409.04410 • Published Sep 6, 2024 • 24
The VoxCeleb Speaker Recognition Challenge: A Retrospective Paper • 2408.14886 • Published Aug 27, 2024 • 10
CSGO: Content-Style Composition in Text-to-Image Generation Paper • 2408.16766 • Published Aug 29, 2024 • 17
Grokfast: Accelerated Grokking by Amplifying Slow Gradients Paper • 2405.20233 • Published May 30, 2024 • 6
EMO: Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions Paper • 2402.17485 • Published Feb 27, 2024 • 190