Han-Bit Kang

hbkang

AI & ML interests

Recent Activity

liked a dataset 2 days ago

DAMO-NLP-SG/multimodal_textbook

updated a collection 2 days ago

cool-papers

upvoted a paper 2 days ago

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

View all activity

Organizations

None yet

hbkang's activity

upvoted a paper 2 days ago

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Paper • 2501.00958 • Published 6 days ago • 87

upvoted a paper 6 days ago

PERSE: Personalized 3D Generative Avatars from A Single Portrait

Paper • 2412.21206 • Published 8 days ago • 15

upvoted a paper 21 days ago

Byte Latent Transformer: Patches Scale Better Than Tokens

Paper • 2412.09871 • Published 26 days ago • 85

upvoted 2 papers 22 days ago

Wonderland: Navigating 3D Scenes from a Single Image

Paper • 2412.12091 • Published 22 days ago • 15

ColorFlow: Retrieval-Augmented Image Sequence Colorization

Paper • 2412.11815 • Published 23 days ago • 26

upvoted a paper 28 days ago

APOLLO: SGD-like Memory, AdamW-level Performance

Paper • 2412.05270 • Published Dec 6, 2024 • 38

upvoted a paper about 1 month ago

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

Paper • 2412.04448 • Published Dec 5, 2024 • 9

upvoted 3 papers about 2 months ago

upvoted a paper 2 months ago

Learning Video Representations without Natural Videos

Paper • 2410.24213 • Published Oct 31, 2024 • 15

upvoted 3 papers 3 months ago

DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation

Paper • 2410.13726 • Published Oct 17, 2024 • 11

Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

Paper • 2410.08261 • Published Oct 10, 2024 • 50

FAN: Fourier Analysis Networks

Paper • 2410.02675 • Published Oct 3, 2024 • 25

upvoted 3 papers 4 months ago

Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Paper • 2409.04410 • Published Sep 6, 2024 • 24

The VoxCeleb Speaker Recognition Challenge: A Retrospective

Paper • 2408.14886 • Published Aug 27, 2024 • 10

CSGO: Content-Style Composition in Text-to-Image Generation

Paper • 2408.16766 • Published Aug 29, 2024 • 17

upvoted a paper 7 months ago

Grokfast: Accelerated Grokking by Amplifying Slow Gradients

Paper • 2405.20233 • Published May 30, 2024 • 6

upvoted a paper 9 months ago

COCONut: Modernizing COCO Segmentation

Paper • 2404.08639 • Published Apr 12, 2024 • 28

upvoted a paper 10 months ago

EMO: Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Paper • 2402.17485 • Published Feb 27, 2024 • 190