Why Post-Training Data Selection Matters for LLM Adaptation
The choice of post-training data substantially affects large language model downstream performance. This makes data curation a high-leverage activity for teams that fine-tune or align models after initial training.
Full-Gradient Ranking and Its Computational Cost
Gradient-based data selection ranks training examples by how well their parameter gradients align with those of a small validation set. Ranking with full-parameter gradients requires an expensive backward pass on every sample. For large candidate pools this backward-pass cost makes exhaustive computation intractable on practical timescales.
Output-Layer Gradients as a Proxy for Data Relevance
The paper investigates whether output-layer gradients can serve as a cheaper approximation. Output-layer gradients compute the partial derivative of the loss with respect to the final linear layer's weights. This operation requires only the forward pass through the model, avoiding the backward pass through all hidden layers. The authors find that even though output-layer and full gradients may rank individual samples differently, the batches they select exhibit aligned gradient directions.
LESSER: A Drop-In Wrapper for Output-Layer Feature Extraction
The authors implement LESSER as a drop-in wrapper around existing selection methods. LESSER replaces the feature-extraction component with output-layer gradient computation. The wrapper reduces the feature-extraction FLOP cost by 9.7× for supervised fine-tuning benchmarks and by 3.0× for reinforcement learning benchmarks. Downstream task performance is tracked against full-gradient baselines to ensure selection quality is preserved.
FLOP Savings and Batch-Level Gradient Alignment
Empirical results show that LESSER achieves the reported FLOP reductions while maintaining downstream task accuracy close to full-gradient selection. The paper reports that individual sample rankings may diverge between output-layer and full gradients, but the selected batches produce comparable gradient alignment. This suggests that the selection criterion operates at the batch level rather than depending on precise per-sample ranking order.
When Output-Layer Gradients Offer Less Precision
The approximation relies on the assumption that output-layer gradients capture sufficient signal for data relevance. For models where early layers dominate the decision boundary or for tasks requiring deep feature alignment, the output-layer proxy may be less effective. The abstract does not provide a per-architecture breakdown, but the empirical results indicate the method works across the evaluated SFT and RL suites.
Integration Path for LLM Development Teams
Teams performing post-training data selection can integrate LESSER with minimal code changes. By replacing the full-gradient computation step with output-layer gradient extraction they gain roughly an order of magnitude FLOP reduction for SFT pipelines. The drop-in nature means existing selection workflows continue to function with only the feature-extraction component swapped. Monitoring downstream task metrics provides verification that selection quality has not degraded after the switch.