Explore indexBack to Terms
Quantized Low-Rank Adaptation
Quantized Low-Rank Adaptation (QLoRA) quantizes frozen base model weights down to 4-bit precision (such as NormalFloat4) while computing gradients through 16-bit LoRA adapter matrices. Combined with double quantization and paged optimizers, QLoRA enables fine-tuning large models on consumer-grade GPUs with minimal accuracy loss.
No public content is connected to this entity yet.