Mayank Mishra's picture

Mayank Mishra

mayank-mishra

·

https://mayank31398.github.io/

AI & ML interests

Large Language Models, Distributed Training and Inference

Recent Activity

upvoted a collection 18 days ago

Granite 3.1 Language Models

new activity 18 days ago

ibm-granite/granite-3.1-8b-instruct:Exceptional creative writer

authored a paper 18 days ago

Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models

View all activity

Articles

Improving Hugging Face Training Efficiency Through Packing with Flash Attention

Saving Memory Using Padding-Free Transformer Layers during Finetuning

Aurora-M: The First Open Source Biden-Harris Executive Order Red teamed Multilingual Language Model

Organizations

mayank-mishra's activity

upvoted a collection 18 days ago

Granite 3.1 Language Models

A series of language models with 128K context length trained by IBM licensed under Apache 2.0 license. • 8 items • Updated 19 days ago • 47

upvoted a paper 2 months ago

SelfCodeAlign: Self-Alignment for Code Generation

Paper • 2410.24198 • Published Oct 31, 2024 • 23

upvoted a collection 2 months ago

SmolLM2

State-of-the-art compact LLMs for on-device applications: 1.7B, 360M, 135M • 15 items • Updated 15 days ago • 197

upvoted a collection 3 months ago

Granite 3.0 Language Models

A series of language models trained by IBM licensed under Apache 2.0 license. We release both the base pretrained and instruct models. • 8 items • Updated 19 days ago • 96

upvoted 2 papers 4 months ago

The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Paper • 2408.15237 • Published Aug 27, 2024 • 38

Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler

Paper • 2408.13359 • Published Aug 23, 2024 • 23

upvoted an article 5 months ago

Article

Improving Hugging Face Training Efficiency Through Packing with Flash Attention

Aug 21, 2024

• 25

upvoted a collection 5 months ago

Power-LM

Dense & MoE LLMs trained with power learning rate scheduler. • 4 items • Updated Oct 17, 2024 • 15

upvoted a paper 5 months ago

Transformer Explainer: Interactive Learning of Text-Generative Models

Paper • 2408.04619 • Published Aug 8, 2024 • 156

upvoted 5 papers 6 months ago

OpenDevin: An Open Platform for AI Software Developers as Generalist Agents

Paper • 2407.16741 • Published Jul 23, 2024 • 69

Enhancing Training Efficiency Using Packing with Flash Attention

Paper • 2407.09105 • Published Jul 12, 2024 • 14

Scaling Granite Code Models to 128K Context

Paper • 2407.13739 • Published Jul 18, 2024 • 19

The infrastructure powering IBM's Gen AI model development

Paper • 2407.05467 • Published Jul 7, 2024 • 2

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Paper • 2407.02490 • Published Jul 2, 2024 • 23

upvoted a collection 6 months ago

Experimental Checkpoints

As name suggests • 1 item • Updated Jun 30, 2024 • 1

upvoted an article 7 months ago

Article

Aligning Large Language Models with BRAIn

By

•

Jun 11, 2024

• 8

upvoted a collection 7 months ago

Dolomite Engine Sample

This collections contains a sample dataset and model trained via dolomite-engine. Repo: https://github.com/ibm-granite/dolomite-engine/ • 2 items • Updated Jun 30, 2024 • 1

upvoted 3 papers 8 months ago

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Paper • 2405.12981 • Published May 21, 2024 • 28

Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization

Paper • 2404.03605 • Published Apr 4, 2024 • 1

Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Paper • 2405.04324 • Published May 7, 2024 • 22