Today for AI

Hacker News AI · 2026/10/8 21:29:46

LittleBit-2 实现亚 1-bit LLM 压缩:通过潜在几何对齐优化低秩分解量化

原标题:Sub-1-Bit LLM Compression via Latent Factorization
68AI 研判分
核心综述

LittleBit-2 提出了一种将大语言模型压缩至亚 1-bit 级别的新方法,核心在于对权重矩阵进行低秩潜在因子分解并二值化。该研究通过引入内部潜在旋转与联合迭代量化(Joint-ITQ)技术,解决了初始化阶段的潜在几何错位问题,从而在保持原始推理架构且无额外开销的前提下,实现了低至 0.1 bits/weight 的极致压缩。

报道全文原始报道全文

本文目录10 个章节

论文

LittleBit-2:通过潜在几何对齐最大化亚 1-bit LLM 中的频谱能量增益 (ICML 2026)<br> Banseok Lee, Youngmin Kim<br>

LittleBit:通过潜在因子分解实现超低比特量化 (NeurIPS 2025)<br> Banseok Lee*, Dongkyu Kim*, Youngcheon You, Youngmin Kim<br>


摘要

LittleBit 通过将每个稠密权重矩阵分解为低秩潜在因子,对这些因子进行二值化,并通过轻量级可学习缩放恢复幅度信息,从而将大语言模型压缩至亚 1-bit 范围。这使得在推理时保留原始模型架构的同时,能够实现包括每权重 0.1 bits 在内的极致压缩。

LittleBit-2 通过解决初始化阶段的潜在几何错位问题改进了这一方案。它在量化感知训练(QAT)之前应用内部潜在旋转与联合迭代量化(Joint-ITQ),使 SVD 导出的潜在因子与二进制超立方体对齐。LittleBit-2 的初始化方法作为可选功能提供(--use_itq),且不会带来额外的推理开销。


亮点

  • 亚 1-bit 压缩: 专为每权重 1.0 到 0.1 bits 设计。
  • LittleBit-2 可选启用: 使用 --use_itq 启用 Joint-ITQ 初始化,以改善潜在几何对齐。
  • 无推理时变更: LittleBit-2 仅修改初始化过程;部署后的因子分解层保持不变。
  • 支持 QAT: 支持结合 SmoothSign 和可选残差因子分解的量化感知训练。

支持的模型

当前代码库支持以下模型:

  • OPT
  • Llama 及 Llama 2/3
  • Phi-4
  • Qwen2.5 及 QwQ
  • Gemma 2 及 Gemma 3
  • Qwen3

安装

我们推荐使用 Python 3.12。

BASH
conda create -n littlebit python=3.12
conda activate littlebit

# Install CUDA toolkit. Adjust the CUDA version if needed.
conda install nvidia/label/cuda-12.4.1::cuda-toolkit -c nvidia/label/cuda-12.4.1

# Install PyTorch.
pip install torch==2.8.0+cu124 torchvision==0.23.0+cu124 torchaudio==2.8.0+cu124 --index-url https://download.pytorch.org/whl/cu124

# Install dependencies.
pip install -r requirements.txt

[!IMPORTANT] 为了复现论文结果,请使用 transformers 4.51.x 版本。较新的 transformers 版本可能会改变模型内部结构或评估行为。

BASH
pip install "transformers==4.51.*"

用法

训练

使用量化感知训练(Quantization-Aware Training)对模型进行训练。默认情况下,LittleBitLinear 使用原始的仅 SVD 初始化方法。若要启用 LittleBit-2(Joint-ITQ),请传入 --use_itq True。

单 GPU

BASH
CUDA_VISIBLE_DEVICES=0 python -m main \
    --model_id meta-llama/Llama-2-7b-hf \
    --dataset c4_wiki \
    --save_dir ./outputs/Llama-2-7b-LittleBit-2 \
    --num_train_epochs 5.0 \
    --per_device_train_batch_size 4 \
    --lr 4e-05 \
    --warmup_ratio 0.02 \
    --report wandb \
    --quant_func SmoothSign \
    --quant_mod LittleBitLinear \
    --residual True \
    --eff_bit 1.0 \
    --kv_factor 1.0 \
    --min_split_dim 8 \
    --l2l_loss_scale 10.0

# Opt-in to LittleBit-2 initialization
# --use_itq True

多 GPU 配合 DeepSpeed

BASH
deepspeed --num_gpus=4 main.py \
    --model_id meta-llama/Llama-2-7b-hf \
    --dataset c4_wiki \
    --save_dir ./outputs/Llama-2-7b-LittleBit-2 \
    --ds_config_path configs/zero3.json \
    --num_train_epochs 5.0 \
    --per_device_train_batch_size 4 \
    --lr 4e-05 \
    --report wandb \
    --quant_func SmoothSign \
    --quant_mod LittleBitLinear \
    --residual True \
    --eff_bit 1.0 \
    --kv_factor 1.0 \
    --min_split_dim 8

评估

评估本地检查点或托管在 Hugging Face Hub 上的模型。

BASH
# From a local directory
CUDA_VISIBLE_DEVICES=0 python eval.py \
    --model_id ./outputs/Llama-2-7b-LittleBit-2 \
    --seqlen 2048 \
    --ppl_task wikitext2,c4 \
    --zeroshot_task boolq,piqa,hellaswag,winogrande,arc_easy,arc_challenge,openbookqa

# From the Hugging Face Hub
CUDA_VISIBLE_DEVICES=0 python eval.py \
    --model_id username/littlebit-llama-7b-0.1bpw \
    --seqlen 2048 \
    --ppl_task wikitext2

旧版检查点

较旧的检查点可能不包含 littlebit_config.json。在这种情况下,请显式传入量化参数:

BASH
CUDA_VISIBLE_DEVICES=0 python eval.py \
    --model_id ./outputs/Legacy-Llama-2-7b \
    --quant_func SmoothSign \
    --quant_mod LittleBitLinear \
    --split_dim 1024

参数加载优先级:

  1. 显式的 CLI 参数
  2. 模型目录中的 littlebit_config.json
  3. 针对较旧检查点的 config.json 回退机制

引用

如果您觉得这项工作有用,请引用:

BIBTEX
@inproceedings{lee2026littlebit2,
  title={LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment},
  author={Lee, Banseok and Kim, Youngmin},
  booktitle={Proceedings of the 43rd International Conference on Machine Learning},
  year={2026}
}
BIBTEX
@inproceedings{lee2025littlebit,
  title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
  author={Lee, Banseok and Kim, Dongkyu and You, Youngcheon and Kim, Youngmin},
  booktitle={Advances in Neural Information Processing Systems},
  year={2025}
}