2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Paper • 2501.00958 • Published 6 days ago • 87
PERSE: Personalized 3D Generative Avatars from A Single Portrait Paper • 2412.21206 • Published 8 days ago • 15
Byte Latent Transformer: Patches Scale Better Than Tokens Paper • 2412.09871 • Published 26 days ago • 85
ColorFlow: Retrieval-Augmented Image Sequence Colorization Paper • 2412.11815 • Published 23 days ago • 26
MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation Paper • 2412.04448 • Published Dec 5, 2024 • 9
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions Paper • 2411.07461 • Published Nov 12, 2024 • 22