Explore indexBack to Terms

Quantized Low-Rank Adaptation

Quantized Low-Rank Adaptation (QLoRA) quantizes frozen base model weights down to 4-bit precision (such as NormalFloat4) while computing gradients through 16-bit LoRA adapter matrices. Combined with double quantization and paged optimizers, QLoRA enables fine-tuning large models on consumer-grade GPUs with minimal accuracy loss.

No public content is connected to this entity yet.