A developer has released a new open source tool that converts tens of millions of geographic point records into smooth, visually usable rasters in seconds — a task that traditionally could take hours or days with standard GIS software.
The Problem with Visualizing Massive Point Datasets
Working with large spatial datasets is a persistent pain point for data professionals. A single CSV file might contain 20 million rows of modeled risk data with geographic coordinates, and the task is to turn that into a continuous map. Traditional tools like ArcGIS, QGIS, and GRASS can choke on this volume, sometimes running for days to produce a single raster. The alternative approaches have their own problems.
A grid-bin cumulative mean — dividing a cumulative sum raster by a cumulative count raster — is computationally cheap but produces patchy, aesthetically unpleasing results, especially across large countries where most of the land area has few data points. The method is also locked to a single resolution: small grids leave regional areas empty, while large grids blur dense urban regions. Inverse distance weighting, or IDW, produces more visually appealing results but carries a computational cost that scales as the product of points multiplied by output cells. At a million points and a million output cells, a naive implementation would require up to a trillion comparisons.
How ppgrid Works
The new tool, ppgrid, takes a different approach by approximating inverse distance weighting through a push-pull mipmap pyramid. The method draws on a technique described in a 1996 paper by Gortler et al. called The Lumigraph, originally designed for reconstructing missing samples across a multiresolution image pyramid — a problem structurally similar to the patchy raster problem.
The process begins by assigning each point into two grids: one accumulating values and another counting points. This pre-binning step operates in linear time relative to the number of points. A mipmap pyramid is then built from those grids, creating increasingly coarse layers through repeated two-by-two block sums. Where data is dense, the fine grid provides precise local estimates. Where data is sparse, coarser layers supply decreasingly confident values, eliminating the ugly gaps that plague simpler approaches.
The push phase runs backward from the coarsest level, upsampling estimates into finer layers and blending them with local information. A confidence value — calculated as the minimum of the cell count divided by a saturation threshold and one — determines how much weight a cell gives to its parent layer's estimate. Cells with abundant local data trust themselves; sparse cells lean on broader estimates.
Critically, the sum and count grids remain separate throughout the pyramid. Averaging at every level would give a cell containing a single point the same weight as a cell containing a thousand, which would distort the weighted mean. Carrying them separately preserves accuracy.
Performance and Benchmarking
Benchmarks were conducted on an M4 MacBook Pro with 24GB of memory using a synthetic dataset of 100,000 points in the Melbourne area. Against the traditional gdal_grid tool, ppgrid was approximately 20 times faster. The author's previous cumulative mean approach was roughly four times slower than ppgrid and produced noticeably worse visual results.
The benchmark was limited to 100-meter resolution because attempts to run gdal_grid at 10 meters caused the test machine to struggle severely. At comparable computational cost to a basic rasterization operation, ppgrid produces a higher fidelity raster without the gaps and artifacts typical of IDW outputs at boundaries.
Design Decisions for Practical Use
Beyond the core algorithm, ppgrid includes several practical features. A fill distance limit prevents the coarsest pyramid level from inventing values across the entire raster using points hundreds of kilometers away. A support raster is included in the output, recording the effective spatial scale behind each pixel's estimate — smaller values mean local evidence strongly supports the value, while larger values indicate the estimate was borrowed from a coarser layer.
The tool also performs preflight calibration on raw input values, testing identity, logarithmic, square-root, and percentile transforms before selecting the one where location best explains the variation. Fill distance is calibrated through spatially blocked cross-validation rather than random sampling, since random validation on clustered data can produce misleadingly optimistic results. The output is returned as a percentile on a 0-to-100 scale, encoded as an int16 GeoTIFF, which makes vectorization straightforward and keeps file sizes manageable.
The tool is distributed under the MIT License, with source code on GitHub and a package on PyPI. The repository includes 13,580 Melbourne property sales records with dense inner-city observations and sparser outer areas to demonstrate the pyramid's adaptive behavior. The author notes that the tool was built with help from local AI coding agents and that future work will focus on improving spatial correctness and statistical soundness for cases where the rasters need to serve purposes beyond visualization.