Conformal prediction has become a standard tool for uncertainty quantification because it produces prediction sets with finite-sample coverage guarantees under minimal assumptions. Practitioners routinely treat the size of these sets as a proxy for uncertainty—smaller sets mean higher confidence, larger sets mean more ambiguity. But the information-theoretic justification for this heuristic has remained incomplete. This paper closes that gap by building a decision-theoretic generalization of entropy tailored to set-valued prediction and proving that conformal set-size reduction satisfies the core properties of an information measure.

Why Conformal Prediction Needs an Information-Theoretic Foundation

Split conformal prediction takes a pretrained model and a calibration dataset, computes nonconformity scores, and thresholds them to produce prediction sets that cover the true label with probability at least 1 minus alpha. The procedure is distribution-free and requires only exchangeability. In classification, two common score functions are the probability score (negative predicted probability of each class) and the adaptive prediction sets (APS) score (cumulative probability of classes at least as likely as the true label, with randomization at the boundary). Both yield sets whose expected size shrinks as the model becomes more confident.

Researchers have long used expected set size as a measure of inefficiency, and instance-level set size as a heuristic for uncertainty. Prior work established one-sided bounds linking set size to conditional entropy using list-decoding arguments. But these bounds do not establish whether set-size reduction itself behaves as a principled information measure. An information measure should satisfy certain axioms: non-negativity, a chain rule or integral representation connecting it to Shannon information, and a data processing inequality. None of these had been shown for conformal set-size reduction.

Generalized Entropy for Set-Valued Prediction

The authors start from DeGroot's decision-theoretic entropy. Fix an action space and a loss function. The generalized entropy of Y given X is the expected Bayes risk—the minimum expected loss achievable by choosing an action after observing X. Generalized mutual information between X and Y given Z is the reduction in Bayes risk when X is added to Z. This framework recovers Shannon entropy when the action space is the probability simplex and the loss is log loss.

To adapt this to set-valued prediction, the authors define a family of loss functions indexed by lambda in (0, 1]. For a predicted set Gamma, the lambda-loss is the cardinality of Gamma plus one over lambda times the indicator that the true label falls outside Gamma. The action space is the power set of the label space. Minimizing expected lambda-loss trades off set size against coverage. The Bayes-optimal predictor is a threshold set: include all labels whose predicted probability exceeds lambda. This is exactly the conformal prediction set produced by an oracle model at threshold lambda.

Conformal entropy H-lambda(Y given X) is the expected minimum lambda-loss. Conformal mutual information I-lambda(X; Y given Z) is the reduction in conformal entropy when X is observed. These measures inherit non-negativity and a data processing inequality from the general framework. A conformal information divergence D-lambda(p parallel q) captures the excess risk of acting under a predicted distribution q when the true distribution is p.

Conformal Mutual Information and Its Integral Representation

The central theoretical result is an exact integral representation of Shannon mutual information in terms of the conformal family. For any random variables X, Y, Z, the Shannon mutual information I(X; Y given Z) in nats equals the integral from 0 to 1 of I-lambda(X; Y given Z) d-lambda. This means Shannon information is the uniform average of the entire conformal information family. Each lambda captures a different aspect of predictive power: small lambda corresponds to high-coverage regimes where the predictor must include many labels, while large lambda corresponds to low-coverage regimes where only the most probable labels matter.

This representation is not merely formal. In split conformal prediction with the probability score, the calibrated threshold tau-hat-alpha corresponds to lambda-hat-alpha = -tau-hat-alpha. The conformal mutual information evaluated at this data-dependent lambda-hat-alpha therefore approximates the slice of Shannon information relevant to the chosen coverage level. The authors also extend the construction to the APS score by introducing a random index Lambda and a randomization variable U, showing the same integral representation holds in expectation.

The Sandwich Bounds: Set-Size Reduction Between Two Measures

Consider a base feature set Z and an augmented feature set (X, Z). Let X-1 = Z and X-2 = (X, Z). The expected set-size reduction Delta-alpha is the difference in expected conformal set cardinality between the base and augmented predictors. The authors prove that, conditional on the calibration data, Delta-alpha is sandwiched between two calibration-dependent conformal mutual information terms plus correction terms:

I-lambda-hat-alpha-2(X; Y given Z) minus D-lambda-hat-alpha-2 plus epsilon-Delta over lambda-hat-alpha-2 less-than-or-equal Delta-alpha less-than-or-equal I-lambda-hat-alpha-1(X; Y given Z) plus D-lambda-hat-alpha-1 plus epsilon-Delta over lambda-hat-alpha-1

Here lambda-hat-alpha-i are the calibrated thresholds for each feature set, D terms are conformal information divergences measuring model misspecification, and epsilon-Delta is the difference in conditional miscoverage probabilities. The lower bound uses the augmented feature set's calibrated lambda, the upper bound uses the base feature set's calibrated lambda. Both bounds hold for any fixed calibration dataset.

Taking expectations over the calibration data yields a cleaner sandwich bound in expectation, with additional O(1 over sqrt-n) terms from the variance of the inverse calibrated thresholds. When the model is well-specified (D terms vanish) and the calibration set is large, the expected set-size reduction is approximately the conformal mutual information at the relevant lambda. This justifies using set-size reduction as an empirical proxy for information gain.

Data Processing Inequality for Conformal Sets

An information measure should satisfy a data processing inequality: processing data cannot increase information. The authors show that conformal set-size reduction obeys an approximate DPI. Suppose a degraded feature tilde-X is obtained from X through a noisy channel, forming a Markov chain Y to (X, Z) to (tilde-X, Z). Let tilde-Delta-alpha be the set-size reduction using tilde-X instead of X. Then:

tilde-Delta-alpha less-than-or-equal Delta-alpha plus D-lambda-hat-alpha-2 plus (tilde-epsilon-2 minus epsilon-2) over lambda-hat-alpha-2

In expectation, the DPI holds up to the model divergence and an O(1 over sqrt-n) calibration term. This means that, on average, adding noise to a feature cannot increase the apparent information gain measured by set-size reduction, modulo finite-sample and misspecification effects. The result also extends to the APS score using the random-index construction.

Empirical Validation Across Eleven Classification Tasks

The authors validate their theory on eleven datasets spanning tabular and image classification: three synthetic datasets (S1, S2, S3), Wine Quality, Human Activity Recognition, Letter Recognition, Covertype, Fashion-MNIST, CIFAR-10, CIFAR-100, and ImageNet-1k. They train neural networks and gradient-boosted trees, apply split conformal prediction with both probability and APS scores, and compute the sandwich bounds and DPI violations across a grid of nominal coverage levels from 0.01 to 0.99.

The sandwich bounds hold empirically. For the probability score, the observed Delta-alpha falls between the lower and upper bounds in over 99 percent of dataset-coverage combinations. The gap between bounds narrows as the calibration set grows, consistent with the O(1 over sqrt-n) theory. The DPI violations are rare and small in magnitude—typically less than 0.05 in set-size units—and diminish with larger calibration sets. The APS score results mirror the probability score results, confirming the random-index extension.

The authors also verify the integral representation by numerically integrating the estimated I-lambda curves and comparing to estimated Shannon mutual information. The two agree within estimation error across all datasets, providing empirical support for the exact theoretical identity.

Feature Selection: When Set-Size Reduction Disagrees With Mutual Information

The paper's most practical finding emerges from a greedy feature selection experiment. Starting from a base feature set, the authors iteratively add the feature that maximizes either the expected set-size reduction Delta-alpha (averaged over several coverage levels) or the Shannon mutual information with the label. They compare the resulting feature rankings across the eleven datasets.

The rankings often disagree. On datasets like Covertype and CIFAR-100, the Spearman correlation between the two rankings is below 0.4. Features that score low on Shannon mutual information can produce large set-size reductions, and vice versa. The divergence occurs because Shannon mutual information averages over all lambda uniformly, while Delta-alpha at a specific coverage level weights the lambda-slice corresponding to that coverage. A feature that strongly affects the tails of the predictive distribution—shifting mass just enough to include or exclude labels at the calibrated threshold—can have outsized impact on conformal set size but modest impact on overall mutual information.

This has concrete implications. If the goal is to reduce prediction set sizes at a target coverage level—for instance, to minimize human review workload in a high-recall setting—selecting features by set-size reduction is more effective than selecting by Shannon mutual information. Conversely, if the goal is general predictive power, Shannon mutual information remains appropriate. The two metrics capture different aspects of feature utility.

What This Means for Practitioners

The paper provides a theoretical license for a common practice. Developers who use conformal prediction and track set sizes as a measure of uncertainty can now cite a formal justification: set-size reduction approximates a well-defined information measure, satisfies a sandwich bound linking it to conformal mutual information, and obeys a data processing inequality. The theory also clarifies the limitations. The sandwich bounds depend on calibration-dependent lambdas and model misspecification terms. With small calibration sets or poorly calibrated models, the correspondence between set-size reduction and information gain degrades. The DPI holds only approximately, with slack from finite-sample effects.

For feature selection, the choice between set-size reduction and Shannon mutual information should be driven by the downstream objective. If the pipeline uses conformal prediction at a fixed coverage level, set-size reduction is the more direct metric. The paper's code and experiments offer a template for computing these quantities in practice. More broadly, the work establishes a bridge between the conformal prediction literature and information theory, opening the door to further cross-pollination—for example, using conformal information measures for active learning, model compression, or fairness auditing where set-valued outputs are natural.