geomermaids · H3 as a GeoParquet Pyramid Builder

My talk at the H3 Conference 2026, New York, 23 September, hosted by Fused and Uber. 36 minutes, questions included.

H3 as a GeoParquet Pyramid Builder, Guillaume Sueur, H3 Conference 2026 Play the video

Slides: download the deck (PDF, 10 MB). Video: watch on YouTube. The measurements were made with GeoPQ Workbench, free and open source.

The problem

GeoParquet is the best format we have for geo-analytics: columnar, compressed, read straight from object storage with HTTP range requests. It is bad at one thing: showing the whole dataset. A Cloud Optimized GeoTIFF carries overviews, so a viewer reads a little when zoomed out and a lot when zoomed in. GeoParquet has no equivalent. Ask for the state and you read the state, every row, at a scale where most parcels are a fraction of a pixel.

And nothing in the file says what a geometry will cost before you open it: the same column holds 4-vertex squares and 300,000-vertex multipolygons. Paste the URL of a 500 MB file into a GIS, and you wait for all of it before the first feature is drawn.

The idea: H3 as the pyramid

Partition the full-detail data by H3 cell, then build coarser levels on top, one file per cell, named by the cell id. For the 2,557,398 MassGIS parcels that means 697 files at resolution 6, the parcels themselves, untouched and Hilbert-sorted. Each level above is built from the one below: one file per parent cell, one row per child cell.

A reader picks the level from the scale, computes the cells under the viewport, and opens those files. Zoom in and it opens the children. No index file, no tile scheme, no lookup service. And nobody looking at the map ever sees a hexagon: the cells organize the files, not the picture.

Three ways to climb a level

The method is written in each overview file's metadata, so a file opened alone says what it is.

Massachusetts parcels, simplify method: 2,557,398 rows, 390 MB
Simplify. Same rows, fewer vertices. Fine for lines, or to keep every feature. 390 MB.
Massachusetts parcels, prune method: 159,861 rows, 78 MB
Prune. Keep the biggest, drop the rest. Fine for points, leaves holes in polygons. 159,861 rows, 78 MB.
Massachusetts parcels, dissolve method: 123 rows, 10 MB
Dissolve. Merge everything inside each child cell, then simplify. Keeps count and area. 123 rows, 10 MB.
304 MB → 10 MBwhole state, same reader
123 rowson screen, still a table
+6%on disk for the dissolve overviews

The dissolve levels are not pictures: each row is an H3 aggregate with count and area_sum, so the map at state scale and a GROUP BY return the same answer.

-- what the map shows at state scale
SELECT count, area_sum FROM read_parquet('parcels/r4/*.parquet');
-- what it shows on a Boston block
SELECT * FROM read_parquet('parcels/r6/862a30667ffffff.parquet');

Why H3

What it is not

It is not a replacement for vector tiles. If the map is the product, PMTiles are smaller, faster and served from anywhere: use them. But a tile is a picture of the data, with the attributes you chose to bake in. When the map is a workbench and every zoom level has to stay a table you can query and join, the pyramid keeps the data and only removes what you cannot see.

It is a convention, not a standard, with one dataset and one reader so far: mine. The open questions from the talk:

If you have an opinion on any of them, I want to hear it.

Chapters

  1. 0:00 Introduction
  2. 3:40 Why geospatial is special: we want to see the map
  3. 5:45 We don't know what's inside the box
  4. 8:59 Raster data has an easier path
  5. 10:08 GeoParquet in short
  6. 14:26 What GeoParquet lacks: a spatial index, overviews
  7. 15:09 Overview options: a second geometry column, a sidecar file
  8. 17:25 Overview options: bounding boxes
  9. 18:51 Overview options: level of detail inside the file
  10. 20:15 The H3 pyramid
  11. 22:55 Three ways to climb: simplify, prune, dissolve
  12. 26:46 What H3 really gave us
  13. 28:29 What this is not: a replacement for vector tiles
  14. 29:12 Questions I don't have the answer to
  15. 31:00 Q&A: very large polygons, baked-in levels, why not PMTiles

The level-of-detail-inside-the-file option builds on Kanahiro Iguchi's cloud-optimized GeoParquet work.

Links

Building something similar, or want a dataset turned into a pyramid like this one? Get in touch.