Animal fur reconstruction from multi-view images has long been stuck behind a data problem. Human hair benefits from several large-scale capture datasets, but no comparable resource exists for animal fur. The variation across species, the density of strands, and the difficulty of ground-truth capture have kept the field reliant on synthetic approximations or artist-created grooms. FurE, a new method from researchers at the University of Washington and NVIDIA, bypasses the dataset gap entirely by transferring a strand decoder trained on human hair to animal fur reconstruction. The result is a per-strand, editable groom recovered in a fraction of the time required by prior dense optimization approaches.
The dataset void and why it matters
Hair and fur share structural similarities: both are dense collections of thin, curved strands anchored to a skin surface. But animal fur introduces variability that human hair datasets do not capture. A single animal can exhibit different fur lengths, thicknesses, and curl patterns across its body. Inter-species differences compound this — compare the short, dense coat of a rabbit to the long, coarse mane of a lion. Existing methods for human hair reconstruction assume access to thousands of captured strands for training. When those methods are applied to animals, they either fail to generalize or require per-instance optimization that takes hours.
The core difficulty is that fur covers most of the animal's body. Self-occlusion is extreme. A single view reveals only a fraction of the strands. Multi-view setups help, but the correspondence problem across views becomes intractable at strand-level density. Without a dataset to learn a prior, prior work resorted to either hand-crafted procedural models or per-instance optimization from scratch. Both approaches have clear limits: procedural models lack photorealism, and per-instance optimization is too slow for practical pipelines.
How the root-conditioned latent field works
FurE represents fur as a collection of strands, each defined by a root position on the animal's surface and a latent code that determines its 3D geometry. The key design decision is to optimize a continuous latent field over the surface rather than optimizing each strand independently. For a given surface point, the field outputs a latent vector. This vector is fed into a PCA-based decoder that produces the full strand geometry.
The decoder is the transfer component. It is trained exclusively on human hair strands from existing datasets. Principal component analysis on aligned human hair strands yields a low-dimensional basis that captures the dominant modes of strand shape variation — overall curvature, twist, tapering, and higher-frequency undulations. The decoder maps a latent code (typically 8-16 dimensions) to coefficients in this PCA basis, which are then linearly combined with the mean strand shape to produce the final 3D curve.
Why does a human hair decoder work for animal fur? The PCA basis captures geometric primitives of strand-like structures: smooth curves with consistent cross-section, anchored at one end, tapering toward the tip. These primitives are shared across mammals. The latent codes that FurE optimizes for each surface point essentially select and weight these primitives to match the observed fur appearance. The human hair prior constrains the solution space to physically plausible strands, while the per-instance optimization adapts the latent field to the specific animal's fur characteristics.
Surface-constrained Gaussian Frosting for the defurred body
Strand reconstruction requires knowing where the skin surface lies beneath the fur. FurE estimates this defurred body using a surface-constrained Gaussian Frosting representation. The method starts with a coarse mesh of the furred animal obtained from multi-view stereo or a neural surface reconstruction method. It then models the fur layer as a volumetric density field defined by Gaussian kernels anchored to the mesh surface.
Each Gaussian kernel has a position (on the surface), a scale (fur thickness), and an opacity. The collection of kernels forms a "frosting" layer over the mesh. By rendering this representation and comparing to input images, the method optimizes both the underlying mesh vertices and the kernel parameters. The surface constraint — kernels stay attached to the mesh — prevents the frosting from drifting away from the body. Part-based priors further regularize the solution: the body is segmented into semantic regions (head, torso, legs, tail), and each region has a learned prior on fur thickness and density.
The output is a clean defurred mesh with per-vertex fur thickness estimates. This mesh provides the root positions for the strand latent field. The thickness estimates also serve as a loss term during strand optimization: strands in regions with higher estimated thickness should be longer and denser.
Optimization pipeline and the 10x speedup
The full optimization alternates between three stages. First, the Gaussian Frosting representation is optimized to recover the defurred mesh and thickness map. Second, the root-conditioned latent field is initialized — typically by interpolating from a small set of manually placed or automatically detected guide strands. Third, the latent field is refined by differentiable rendering.
The rendering loss compares multi-view images of the reconstructed strands against the input photographs. Strands are rendered as thin ribbons or tubes with a simple shading model. The loss includes photometric terms, silhouette alignment, and the thickness consistency term from the frosting stage. Crucially, the PCA decoder is differentiable, so gradients flow from pixel-space losses back to the latent codes.
The 10x speedup over prior dense per-strand optimization comes from three factors. First, the latent field reduces the number of optimization variables dramatically: instead of optimizing thousands of strand control points independently, the method optimizes a continuous field parameterized by a small neural network or a grid of latent codes. Second, the human hair decoder provides a strong prior that keeps optimization in a plausible region of the solution space, reducing the number of iterations needed. Third, the surface-constrained frosting stage provides a good initialization for both root positions and thickness, avoiding the slow convergence of joint shape-and-strand optimization from random initialization.
Quantitative results on the FurBench synthetic benchmark show FurE matches the strand-level accuracy of the previous state-of-the-art method (Neural Fur, SIGGRAPH 2024) while reducing optimization time from 2.3 hours to 14 minutes on a single RTX 4090. On real-world captures of a rabbit, a fox, and a long-haired cat, FurE produces grooms that are visually indistinguishable from Neural Fur's outputs in qualitative evaluation, with PSNR and LPIPS metrics within 2% of the slower method.
Limitations and trade-offs
The human hair decoder imposes a structural bias. Strands that deviate significantly from the human hair PCA manifold — for example, the flattened, ribbon-like strands of some aquatic mammals, or the hollow, medullated guard hairs of cold-climate species — cannot be perfectly represented. The decoder can approximate them, but the approximation error increases with structural divergence.
The Gaussian Frosting representation assumes fur thickness varies smoothly across the surface. This breaks down at sharp boundaries like the edge of a mane or the transition from body fur to bare skin on a muzzle. The part-based priors help but require manual or learned segmentation, which can fail on unusual poses or heavy self-occlusion.
The method also assumes the input multi-view images have consistent lighting and exposure. Strong shadows cast by fur onto the body surface can confuse the frosting stage, leading to thickness overestimation in shadowed regions. The paper notes this as a failure mode in outdoor captures with directional sunlight.
Finally, the 10x speedup is measured against dense per-strand optimization. Methods that use neural radiance fields or 3D Gaussian splatting for the entire fur volume (without explicit strands) can be faster still, but they do not produce editable strand grooms. FurE targets the specific use case where an explicit, editable strand representation is required — for simulation, grooming tools, or downstream rendering with production hair shaders.
Practical implications for graphics pipelines
For a developer integrating fur reconstruction into a 3D pipeline, FurE offers a clear path: capture multi-view images of an animal, run the frosting stage to get a defurred mesh and thickness map, then run the latent field optimization to get a strand groom. The output is a set of strand curves in a standard format (Alembic or USD curves) compatible with Houdini, Maya, Blender, and production renderers like RenderMan or Arnold.
The human hair decoder is a fixed asset — train it once on a human hair dataset (the paper uses the HairSalon dataset, 50k strands from 200 subjects) and reuse it across all animal species. The per-instance cost is the optimization time, now low enough for overnight or even interactive use with a good initialization.
If a project requires species-specific strand structures not covered by the human hair manifold, the decoder can be retrained on a small synthetic dataset of that fur type. The paper shows that even 500 synthetic strands generated from a procedural model for a target species can adapt the decoder effectively, because the PCA basis only needs to capture the residual variation not already explained by the human hair basis.
The open-source release includes the trained decoder, the frosting optimization code, and a PyTorch implementation of the latent field with differentiable rendering. Integration into a neural rendering pipeline requires only a camera calibration step and a multi-view image set. No animal-fur dataset collection is needed.
Read the paper on arXiv