GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G²MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G²MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
Why uniform distillation fails under curvature heterogeneity
Graph Neural Networks are the standard tool for learning on graphs, but their reliance on message passing makes inference slow in latency-sensitive applications. A growing line of work distils GNN knowledge into a structure-free Multi-Layer Perceptron, enabling fast graph-free inference. Most methods follow GLNN and match teacher logits with a KL loss, sometimes injecting structural priors through positional encodings, codebooks, or neighborhood alignment cues. These methods improve the student's information source, but they usually treat graph structure as uniformly useful: every node receives the same type of distillation pressure regardless of its topological role.
We argue that this uniform view misses an important source of teacher-student mismatch. Message passing changes representations differently across graph regions. Boundary nodes often contain graph-induced directions that are hard to infer from features alone, while dense interior nodes are often smoothed toward class-level prototypes. This produces two spectral failure modes in GNN-to-MLP distillation. On sparse graphs, the student has lower effective rank than the teacher, which we call spectral underfit. On dense graphs, the teacher may collapse to a low-rank geometry, while the feature-only student retains unsupported variation, which we call spectral overfit.
We characterize the student's coverage of the teacher's geometry in the teacher's eigenbasis. For each teacher eigen-direction with positive eigenvalue, we compute the ratio of student energy to teacher energy along that direction. When this ratio is below one, the student has insufficient energy along a teacher direction, which we call spectral underfit; when above one, the student retains energy in directions that the teacher kernel weakly supports, which we call spectral overfit. We summarize the diagonal mismatch by a weighted sum of squared deviations, where high-energy teacher directions receive larger weight.
Proposition 3.2 links Ollivier-Ricci curvature to local aggregation disagreement. ORC measures the Wasserstein-1 distance between neighborhood walks of adjacent nodes. Low-curvature edges mark boundary regions where random-walk neighborhoods disagree, while high-curvature edges mark dense regions where aggregation strongly smoothes representations. ORC is a locally computable proxy for where graph-dependent teacher information is likely to matter, without requiring construction or diagonalization of the teacher-student kernel mismatch.
Standard pointwise KL distillation sums node-wise divergences that depend on the teacher's marginal predictive distribution. Two students may match the same teacher logits while inducing different neighborhood geometries in hidden space. Pointwise KL does not directly control the graph-induced representation geometry. Reducing spectral distortion therefore requires an additional objective sensitive to neighborhood-level teacher-student alignment.
Two failure modes and the ideal distillation objective
We characterize the student's coverage of the teacher's geometry in the teacher's eigenbasis. Let the teacher kernel be decomposed as a sum of eigencomponents with eigenvalues and eigenvectors. For each direction with positive eigenvalue, we compute the ratio of student energy to teacher energy along that direction. When this ratio is below one, the student has insufficient energy along a teacher direction, which we call spectral underfit; when above one, the student retains energy in directions that the teacher kernel weakly supports, which we call spectral overfit.
We summarize the diagonal mismatch by a weighted sum of squared deviations, where high-energy teacher directions receive larger weight. This quantity measures how much MLP energy differs from teacher energy along teacher eigendirections, with high-energy directions receiving larger weight.
Ollivier-Ricci curvature provides a locally computable proxy for where graph-dependent teacher information is likely to matter. Proposition 3.2 indicates that curvature separates locally incoherent boundary neighborhoods from coherent interior neighborhoods. Low-curvature edges correspond to pairs whose local random-walk neighborhoods are difficult to align, while high-curvature edges correspond to locally coherent neighborhoods.
The two modes concentrate at different parts of the curvature spectrum: spectral underfit appears mainly in low-curvature boundary regions, whereas spectral overfit appears mainly in high-curvature interior subgraphs. A uniformly weighted loss does not explicitly distinguish these regimes. Increasing pressure on boundary-discriminative directions also increases pressure on already-smoothed interior regions, and the converse is also true. This motivates a curvature-stratified objective in which prediction-side and representation-side alignment receive opposite curvature-dependent emphasis.
G²MLP: curvature-guided graph Wasserstein alignment
We approximate the energy-weighted alignment objective through three practical relaxations: global kernel alignment is replaced by local neighborhood distribution matching, discrete optimal transport is replaced by a diagonal Gaussian Wasserstein-2 surrogate, and uniform node weighting is replaced by Ollivier-Ricci curvature-dependent weighting. These steps are tractable approximations designed to preserve the graph-dependent teacher geometry most relevant to spectral underfit and spectral overfit.
From global alignment to local neighborhood measures
The graph-dependent part of the teacher kernel is generated through message-passing neighborhoods. We approximate the energy-weighted alignment by matching local teacher and student representation distributions. For each node, we define local representation measures as weighted sums of Dirac measures at neighbor representations, using the lazy random walk distribution over the closed neighborhood. Matching these local measures in Wasserstein-2 distance encourages agreement in both local means and local dispersion. Aggregating over nodes gives a graph Wasserstein objective.
Efficient neighborhood alignment via moment matching
- Each Wasserstein-2 distance over neighborhood atoms costs cubic time in the number of atoms per node.
- To obtain a scalable surrogate, we match the first two marginal moments of the teacher and student neighborhood measures.
- We use a diagonal Gaussian approximation for the neighborhood measures, where the mean is the weighted average of neighbor representations and the covariance is the coordinate-wise variance.
- The resulting closed-form distance splits into a feature-side term and a logit-side term, enabling curvature-dependent weighting.
- This reduces the per-node cost to linear in the hidden dimension.
Curvature-adaptive weighting
We use Ollivier-Ricci curvature as a local proxy for allocating the alignment budget, rather than as an exact estimator of kernel eigenvalues. We instantiate this proxy with opposite monotone weights for the logit and feature components: low-curvature nodes receive stronger prediction-side alignment to correct spectral underfit, while high-curvature nodes receive stronger representation-side alignment to reduce spectral overfit. The normalization enforces that curvature redistributes emphasis without changing the overall loss scale. Hyperparameters control the strength of stratification.
The curvature-adaptive Wasserstein objective combines per-node terms with curvature weights, computable in linear time per iteration in the number of nodes and edges times hidden dimension. This construction keeps the approximation aligned with the two regimes: low-curvature nodes receive stronger prediction-level correction, while high-curvature nodes receive stronger neighborhood representation matching where the Gaussian moment surrogate is most appropriate.
Full training objective
The full G²MLP training objective combines cross-entropy on labeled nodes, the standard KL distillation term, and the curvature-adaptive Wasserstein term. The stop-gradient normalizes the curvature-adaptive Wasserstein gradient by a batch-level constant without changing its direction, making its magnitude comparable to entropy-scaled losses. The KL term preserves probability-simplex calibration, while the curvature-adaptive Wasserstein term aligns local teacher-student geometry.
Experimental results across six benchmarks
We evaluate on six node-classification benchmarks: three citation networks (Cora, Citeseer, Pubmed), two Amazon co-purchase graphs, and the large-scale OGB graph ogbn-Arxiv. The first five follow the CPF setting; ogbn-Arxiv follows the standard OGB split. We compare against vanilla MLP, GLNN, KRD, FF-G2M, and other baselines. G²MLP achieves the best graph-free accuracy on all six datasets, improving over the strongest baseline by 0.16 to 1.66 percentage points. G²MLP further surpasses the GraphSAGE teacher on five of the six datasets, with gains up to 4.02 percentage points on Citeseer; only on ogbn-Arxiv does the student fall short of the teacher, reflecting the known difficulty of feature-only inference on this large-scale graph.
Under the production setting that mixes inductive and transductive evaluation, G²MLP achieves the best graph-free production accuracy on 5 of 6 datasets, with gains of 0.16–1.60 percentage points over the strongest baseline. The advantage is most pronounced in the inductive regime, where G²MLP improves by 2.07–2.69 percentage points on Cora, Citeseer, Pubmed, and A-computer—substantially larger than under transductive evaluation, suggesting curvature-stratified alignment transfers better to unseen nodes than nodewise distillation.
Extension to Graph Transformer teachers and link prediction
Graph Transformers incur quadratic attention cost at inference, making distillation to MLP students particularly attractive. We evaluate three Graph Transformer teachers: GraphGPS with hybrid MPNN-plus-global attention, NAGphormer with per-node hop-token attention, and a baseline Graph Transformer with local masked attention. G²MLP transfers without architectural changes to Graph Transformer teachers and link prediction. Under pure-distillation regimes, G²MLP achieves competitive link prediction scores, and the edge-BCE term can be ablated while retaining performance. G²MLP also reduces the cross-regime gap between sparse and dense attention mechanisms in GraphGPS.
Inference efficiency benefits
The deployed model remains a standard MLP and requires no graph access at inference. Graph-dependent quantities are used only during training. Compared to GNN inference, the MLP student evaluates orders of magnitude faster. The curvature-adaptive Wasserstein objective adds minimal overhead since curvature can be precomputed once per training epoch, and the diagonal Gaussian Wasserstein-2 surrogate reduces per-node cost to linear in the hidden dimension.
Conclusion
We identify spectral underfit and spectral overfit as two failure modes of GNN-to-MLP distillation, and connect them to boundary and interior graph regions. We derive an energy-weighted alignment objective that characterizes teacher-student geometric mismatch and motivates curvature-dependent allocation of distillation supervision. We propose G²MLP, a curvature-guided distillation framework that combines prediction- and representation-level alignment while preserving graph-free MLP inference. Across node-classification benchmarks, G²MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both sparse and dense regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.