I have all the details from the paper. Writing the article now. /var/www/simpleprog-website/articles/odin-fhe-llama.mdcontent

Running a large language model under fully homomorphic encryption has always been a question of trade-offs: privacy comes at the cost of enormous computational overhead. A new open-source system called Odin pushes past that bottleneck, achieving a 4.51 times speedup over the previous state of the art when running Llama-3-8B inference on encrypted data.

The Privacy Problem

When users send prompts to cloud LLM services, the text passes through a model provider who can read it. That creates a privacy risk for anyone submitting sensitive material, from medical queries to proprietary business data. Fully homomorphic encryption offers a theoretical solution by letting a server perform inference on ciphertexts without ever decrypting the input. But representing data as encrypted data adds enormous storage and computational overhead that has made practical FHE inference prohibitively slow.

The core difficulty is not encryption itself but layout. In CKKS-based inference, a packing scheme maps logical tensors to ciphertexts and slots. That mapping determines how many ciphertexts are needed, how expensive linear layers become, and how data moves between layers, attention mechanisms, and nonlinear operations.

What Odin Changes

Odin takes a co-design approach, building the ciphertext packing strategy around the model's execution pattern rather than treating them as separate problems. It starts from a THOR-style baseline, where the primary bottleneck is the encoding of model weights into plaintext. The system replaces this with a feature-major cross-layer layout that unifies residual connections and layer interfaces, building transient intra-operator layouts for linear projections and attention.

This eliminates redundant plaintext encoding of weights across wide projection layers, where the same values get encoded multiple times. Inside the attention mechanism, the QK^T matrix multiplication produces scores that Softmax can consume directly without intermediate repacking. The probability output from Softmax feeds directly into the value projection, skipping the layout conversion that previously slowed the pipeline.

Handling Nonlinear Operations

Attention is not the only bottleneck. Nonlinear activation functions are typically implemented through polynomial approximation under FHE, which introduces its own overhead. Odin uses minimax polynomial approximation with input-range control and joint error allocation guided by actual model quality requirements. By carefully distributing approximation error across the computation graph, the system reduces the required polynomial degree and multiplicative depth without noticeably degrading output quality.

Lower polynomial degree means fewer sequential multiplications, which under homomorphic encryption translates directly to shorter evaluation circuits and faster inference.

Results on Llama-3-8B

Evaluated on Llama-3-8B weights with a 128-token input, Odin runs all 32 Transformer layers on a single NVIDIA H100 80 GB GPU. End-to-end server-side FHE evaluation completes in 366.4 seconds with 58.9 GiB of peak device memory. The comparable THOR baseline takes 1651.9 seconds under identical model, input, CKKS parameters, and hardware conditions.

The team claims Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3. The code is publicly available, which matters because prior implementations were either proprietary or limited to smaller models that did not exercise the full complexity of a 32-layer transformer architecture.

What This Means for Deployments

A four-fold speedup does not make FHE inference fast by ordinary standards. Three hundred and sixty-six seconds for a single prompt is nowhere near the sub-second latency that production chatbots require. But the gap between theoretical privacy and practical privacy just narrowed considerably, and the open-source nature of the implementation means other researchers can build on top of it.

The paper points to a path where privacy and utility are not mutually exclusive in cloud AI services. That path is still long, but Odin removes one of the biggest obstacles standing in the way.