Why Post-Training Data Selection Matters for LLM Adaptation

The choice of post-training data substantially affects large language model downstream performance. This makes data curation a high-leverage activity for teams that fine-tune or align models after initial training.

Full-Gradient Ranking and Its Computational Cost

Gradient-based data selection ranks training examples by how well their parameter gradients align with those of a small validation set. Ranking with full-parameter gradients requires an expensive backward pass on every sample. For large candidate pools this backward-pass cost makes exhaustive computation intractable on practical timescales.

Output-Layer Gradients as a Proxy for Data Relevance

The paper investigates whether output-layer gradients can serve as a cheaper approximation. Output-layer gradients compute the partial derivative of the loss with respect to the final linear layer's weights. This operation requires only the forward pass through the model, avoiding the backward pass through all hidden layers. The authors find that even though output-layer and full gradients may rank individual samples differently, the batches they select exhibit aligned gradient directions.

LESSER: A Drop-In Wrapper for Output-Layer Feature Extraction

The authors implement LESSER as a drop-in wrapper around existing selection methods. LESSER replaces the feature-extraction component with output-layer gradient computation. The wrapper reduces the feature-extraction FLOP cost by 9.7× for supervised fine-tuning benchmarks and by 3.0× for reinforcement learning benchmarks. Downstream task performance is tracked against full-gradient baselines to ensure selection quality is preserved.

FLOP Savings and Batch-Level Gradient Alignment

Empirical results show that LESSER achieves the reported FLOP reductions while maintaining downstream task accuracy close to full-gradient selection. The paper reports that individual sample rankings may diverge between output-layer and full gradients, but the selected batches produce comparable gradient alignment. This suggests that the selection criterion operates at the batch level rather than depending on precise per-sample ranking order.

When Output-Layer Gradients Offer Less Precision

The approximation relies on the assumption that output-layer gradients capture sufficient signal for data relevance. For models where early layers dominate the decision boundary or for tasks requiring deep feature alignment, the output-layer proxy may be less effective. The abstract does not provide a per-architecture breakdown, but the empirical results indicate the method works across the evaluated SFT and RL suites.

Integration Path for LLM Development Teams

Teams performing post-training data selection can integrate LESSER with minimal code changes. By replacing the full-gradient computation step with output-layer gradient extraction they gain roughly an order of magnitude FLOP reduction for SFT pipelines. The drop-in nature means existing selection workflows continue to function with only the feature-extraction component swapped. Monitoring downstream task metrics provides verification that selection quality has not degraded after the switch.

Read the paper on arXiv