School Information System

A Visual Guide to Quantization

Maarten Grootendorst

Demystifying the Compression of Large Language Models

As their name suggests, Large Language Models (LLMs) are often too large to run on consumer hardware. These models may exceed billions of parameters and generally need GPUs with large amounts of VRAM to speed up inference.

As such, more and more research has been focused on making these models smaller through improved training, adapters, etc. One major technique in this field is called quantization.

Share

Fast Lane Literacy™ by sedso

does

Explore teaching tips and learn more about the word does.