What it is
AirLLM is a Python library that cuts inference memory so 70B models can run on a single 4GB GPU without quantization, distillation, or pruning. It keeps only one layer on the GPU at a time, so VRAM needs depend on layer size rather than total model size. Models are loaded through a single AutoModel interface using a Hugging Face repo ID or a local path. It also supports LoRA-style training of large models on small VRAM.
Who it's for
- Developers who want to run 70B+ open LLMs on low-VRAM consumer GPUs
- Users who want to fine-tune very large models (e.g. Qwen3.8-Flash-Next, Qwen3.8-27B) on small GPUs
- Mac users with Apple silicon who want to run large models locally
Requirements
Requirements
- Python with the airllm pip package
- Sufficient disk space in the Hugging Face cache directory, since the original model is split and saved layer-wise
- bitsandbytes and airllm later than 2.0.0 for 4bit/8bit compression
- On macOS: mlx and torch installed, and Apple silicon only
- A Hugging Face token (hf_token) for gated models such as meta-llama/Llama-2-7b-hf
- For Qwen3.8-Flash-Next: a transformers build with in-tree qwen4_exp and ~360GB of checkpoint disk
- For Qwen3.8-27B: transformers 5.8+
- For Kimi K3: compressed-tensors and flash-attn, a CUDA 12 build of torch, and transformers 4.56.x
Setup
Install the package
Install the airllm pip package.
bashpip install airllmEnable compression (optional)
Install bitsandbytes and make sure airllm is later than 2.0.0, then pass compression='4bit' or '8bit' when initializing the model.
bashpip install -U bitsandbytes pip install -U airllm
Examples
Basic inference with AutoModel
pythonfrom airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")
input_text = [
'What is the capital of United States?',
]
input_tokens = model.tokenizer(input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH,
padding=False)
generation_output = model.generate(
input_tokens['input_ids'].cuda(),
max_new_tokens=20,
use_cache=True,
return_dict_in_generate=True)
output = model.tokenizer.decode(generation_output.sequences[0])
print(output)What it does: Loads a model by Hugging Face repo ID, tokenizes a prompt, generates 20 new tokens on the GPU and decodes the result.
Enable 4-bit model compression
pythonmodel = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
compression='4bit' # specify '8bit' for 8-bit block-wise quantization
)What it does: Uses block-wise quantization of the weights to shrink loading size, which the README says can speed up inference by up to 3x.
Run a gated model with a Hugging Face token
pythonmodel = AutoModel.from_pretrained("meta-llama/Llama-2-7b-hf", #hf_token='HF_API_TOKEN')What it does: Shows where to supply hf_token for gated models, as described in the FAQ (the token argument is commented out in the source).
Train a LoRA adapter via the Python API
pythonfrom airllm import AirLLMLoRAQwen4Exp
trainer = AirLLMLoRAQwen4Exp(
"Qwen/Qwen3.8-Flash-Next",
max_seq_len=512,
lora_r=16,
delete_original=True,
)
tok = trainer.tokenizer
if tok.pad_token_id is None:
tok.pad_token = tok.eos_token
encoded = tok(
"Your training text here.",
return_tensors="pt",
truncation=True,
max_length=512,
)
loss = trainer.train_step(
encoded["input_ids"].cuda(),
attention_mask=encoded.get("attention_mask"),
)
print(loss)
trainer.save_adapter("qwen38-flash-next-lora.pt")What it does: Streams frozen weights one decoder layer at a time while keeping the adapters on the GPU, runs one training step and saves the adapter.
Train from the command line
bashpython air_llm/examples/train_qwen38_flash_next_lora.py \
--data my_data.jsonl \
--seq-len 512 \
--epochs 1 \
--save-adapter qwen38-flash-next-lora.ptWhat it does: Runs the example training script from the repo root on a .jsonl dataset.
Pros & cons
Pros
- Pro:Runs 70B models on a ~4GB GPU without quantization, distillation, or pruning, and scales up to very large models such as DeepSeek-V3 (671B, ~12GB)
- Pro:Single AutoModel interface that works with a Hugging Face ID or a local path across many model families
- Pro:Optional 4bit/8bit block-wise compression that can speed up inference by up to 3x
- Pro:Also supports training/fine-tuning of huge models on small VRAM, plus macOS (Apple silicon) support
Cons
- Con:Splitting the model layer-wise is very disk-consuming; running out of disk space causes errors like MetadataIncompleteBuffer
- Con:Some newer models have strict requirements (specific transformers versions, CUDA 12 torch, flash-attn, ~360GB of checkpoint disk for Flash-Next)
- Con:Prefetching is only supported by AirLLMLlama2 for now, and macOS support is limited to Apple silicon
Images
