Glyd Compresses AI Models Without Losing a Single Bit

An open source project called Glyd claims to cut the GPU memory required by AI models by roughly a third while preserving every weight exactly. The tool compresses model parameters from 16-bit floating point to about 11 bits and decodes them inside the GPU's matrix multiplication unit rather than in system memory.

The Problem It Solves

Running large language models requires substantial GPU memory. A 16-bit floating point weight uses 8 bits for its exponent, but in a trained model those exponent bits carry only about 2.6 bits of actual information. That leaves significant unused capacity. Glyd identifies and eliminates that redundancy without discarding any data.

The practical effect is that models which previously required two GPUs can now run on one, and models that barely fit on a card can now operate with room to spare for larger context windows.

How It Works

Glyd operates in three stages. First, it identifies the waste in each bf16 weight's exponent. Second, it encodes common exponents with short bit sequences and stores rare exponents in a side list. Sign and mantissa bits remain unchanged. The result averages about 11 bits per weight with nothing rounded.

Third, during inference, the GPU reads the compressed bytes and decodes them directly in registers inside the tensor cores. No bf16 copy is ever written back to memory. The decoding happens in the same operation that performs the matrix multiply.

No, It Is Not Quantization

Glyd differs fundamentally from 4-bit quantization methods like GGUF Q4, AWQ, or GPTQ. Quantization rounds weights to fewer levels, which changes the model and alters its outputs. Glyd stores the same values in fewer bits and returns every one of them exactly. The project describes itself as closer to a zip file than to quantization.

The accuracy difference is negligible. Across measured models, Glyd's MMLU answers match bf16 on 98.0% to 100% of questions, and perplexity stays within 0.06% of bf16. The matrix products compute terms in a slightly different order than cuBLAS, as any two GPU kernels will, so edge cases can go either way.

Performance Benchmarks

Because generating a token reads every weight once, reading a third fewer bytes saves time at many sequence lengths. On an RTX 4080 SUPER running Qwen2.5-7B, Glyd used 25% less time at a single sequence and 28% less at 32 sequences. An A10 showed 28% and 18% improvements respectively.

The picture is not uniformly positive. At higher sequence counts on A100 and H100 hardware, Glyd was slightly slower, at 6% and 8% respectively. Decode steps were measured for query, key, value, and gate operations merged as vLLM runs them.

What Fits Where

The compression makes a meaningful difference for mid-size models on consumer hardware. A Qwen3.8 27B model requires 51.7 GB in bf16 but 41.0 GB with Glyd, fitting under a 48 GB GPU card. Models like Gemma 3 12B, Phi-4 14B, and R1 Distill Qwen 14B fit only with Glyd on a 24 GB GPU.

Smaller models like Mistral 7B v0.3 and Llama 3.1 8B fit on a 24 GB GPU either way, though Glyd still reduces their memory footprint.

How It Compares

Lossless weight compression is an active research area. Glyd compresses to 10.8 or 12.0 bits per weight and decodes inside the matrix multiply in registers. The DFloat11 method (NeurIPS 2025) achieves about 11 bits but decompresses each block to bf16 in memory before running. ZipServ (ASPLOS 2026) reaches 11.35 bits inside the matrix multiply, with kernels within a few percent of Glyd's performance.

All three methods preserve the model bit for bit. Four-bit quantization methods achieve the smallest footprint by far but produce a different model whose answers change.

Availability and Plans

Glyd is open source and available on GitHub. Currently, GPU inference runs from a PyTorch harness: clone the repository, run the setup script, and execute the benchmark tool against any model released in bf16. A pip package with native vLLM integration is in progress, along with a command-line tool for fitting, packing, and serving models.

The codec underneath has already been built and installs through Homebrew, Cargo, or a prebuilt wheel. Benchmarks have been measured on an RTX 4080 SUPER, RTX A6000, A10, A100, and H100. Blackwell GPUs have not been tested yet. The project was announced September 27, 2026.