DeepL has published a detailed account of how it moved its language model training and inference from 16-bit to 8-bit floating point precision, a shift that doubled inference throughput while maintaining translation quality and accelerating training by roughly 50 percent. The company's engineering team built its next-generation LLMs using NVIDIA H100 GPUs, which natively support FP8 data types through a new generation of Tensor Cores.
Why FP8 Matters
Modern LLM computation is dominated by matrix multiplications. Reducing the precision of those multiplications from 16-bit BFloat16 to 8-bit FP8 halves the memory required to store weights and activations, and it doubles the number of operations a given GPU can perform per unit of time. The trade-off is a narrower range of representable numbers and lower precision.
DeepL's engineering argument is that LLM training does not require the full precision that BF16 offers. The framework uses two FP8 formats in tandem: E4M3, which devotes more bits to the mantissa and is used during the forward pass for predicting the next token distribution, and E5M2, which devotes more bits to the exponent and is used during the backward pass for computing gradients where precision matters less than range. This mixed-precision approach lets the system use each format for the task it is best suited to.
The Training Journey
DeepL transitioned its existing training codebase to FP8 using NVIDIA's Transformer Engine, a library designed specifically to accelerate transformer models with FP8 support. The library manages the conversion between formats and handles scaling factors that prevent overflow and underflow when multiplying low-precision tensors. Each weight tensor carries an associated 32-bit scaling factor, and the multiplication formula adjusts accordingly.
Measured by Model FLOPS Utilization, the efficiency of DeepL's training pipeline rose from 44.6 percent to 67 percent with FP8. Working with NVIDIA to optimize Transformer Engine usage, the team incrementally improved performance by an additional 25 percent over 15 months, reaching 80 percent MFU. On the quality front, a 1.5-billion-parameter model trained on three trillion tokens showed only minimal degradation in training loss compared to BF16, and validation perplexity for English-German translation showed no measurable decline.
Inference at Scale
For production inference, DeepL uses NVIDIA TensorRT-LLM, which takes trained model weights and builds an optimized inference engine using kernel fusion, optimized CUDA code, KV caching, and continuous in-flight batching. Running FP8 in this pipeline changes the relationship between throughput and latency in a significant way.
At most batch sizes, FP8 delivers double the throughput of BF16 at the same latency level. For DeepL, which must serve millions of translations per day, this effectively doubles the capacity of its LLM fleet without requiring additional hardware. The company's next-generation models outperform their predecessors by 1.4 times for European language pairs and 1.7 times for more complex pairs like English and Japanese, all while maintaining the same response latency.
What Comes Next
DeepL has deployed a new NVIDIA DGX SuperPOD based on GB200 systems, which introduces Tensor Cores capable of native FP4 tensor operations. The company's engineering team views this as the beginning of the next phase of precision reduction. If a single byte delivered the gains described above, the team is now exploring what half a byte can do.