geomermaids · H3 as a GeoParquet Pyramid Builder
My talk at the H3 Conference 2026, New York, 23 September, hosted by Fused and Uber. 36 minutes, questions included.
Slides: download the deck (PDF, 10 MB). Video: watch on YouTube. The measurements were made with GeoPQ Workbench, free and open source.
The problem
GeoParquet is the best format we have for geo-analytics: columnar, compressed, read straight from object storage with HTTP range requests. It is bad at one thing: showing the whole dataset. A Cloud Optimized GeoTIFF carries overviews, so a viewer reads a little when zoomed out and a lot when zoomed in. GeoParquet has no equivalent. Ask for the state and you read the state, every row, at a scale where most parcels are a fraction of a pixel.
And nothing in the file says what a geometry will cost before you open it: the same column holds 4-vertex squares and 300,000-vertex multipolygons. Paste the URL of a 500 MB file into a GIS, and you wait for all of it before the first feature is drawn.
The idea: H3 as the pyramid
Partition the full-detail data by H3 cell, then build coarser levels on top, one file per cell, named by the cell id. For the 2,557,398 MassGIS parcels that means 697 files at resolution 6, the parcels themselves, untouched and Hilbert-sorted. Each level above is built from the one below: one file per parent cell, one row per child cell.
A reader picks the level from the scale, computes the cells under the viewport, and opens those files. Zoom in and it opens the children. No index file, no tile scheme, no lookup service. And nobody looking at the map ever sees a hexagon: the cells organize the files, not the picture.
Three ways to climb a level
The method is written in each overview file's metadata, so a file opened alone says what it is.
The dissolve levels are not pictures: each row is an H3 aggregate with
count and area_sum, so the map at state scale and a
GROUP BY return the same answer.
-- what the map shows at state scale
SELECT count, area_sum FROM read_parquet('parcels/r4/*.parquet');
-- what it shows on a Boston block
SELECT * FROM read_parquet('parcels/r6/862a30667ffffff.parquet'); Why H3
- The cell id is the file name. No manifest to download, no directory to list.
- It knows its parent and its children. Moving between levels is arithmetic on the name: that is the pyramid.
- It is a
GROUP BYkey. Building a level is an aggregation, not a spatial join.
What it is not
It is not a replacement for vector tiles. If the map is the product, PMTiles are smaller, faster and served from anywhere: use them. But a tile is a picture of the data, with the attributes you chose to bake in. When the map is a workbench and every zoom level has to stay a table you can query and join, the pyramid keeps the data and only removes what you cannot see.
It is a convention, not a standard, with one dataset and one reader so far: mine. The open questions from the talk:
- Should overviews live in the GeoParquet spec, or stay a convention on top of it?
- Hexagons, or any grid with parents and children, such as A5?
- How to aggregate many columns when each needs its own rule: sum, mean, majority?
- Who writes the second reader?
If you have an opinion on any of them, I want to hear it.
Chapters
- 0:00 Introduction
- 3:40 Why geospatial is special: we want to see the map
- 5:45 We don't know what's inside the box
- 8:59 Raster data has an easier path
- 10:08 GeoParquet in short
- 14:26 What GeoParquet lacks: a spatial index, overviews
- 15:09 Overview options: a second geometry column, a sidecar file
- 17:25 Overview options: bounding boxes
- 18:51 Overview options: level of detail inside the file
- 20:15 The H3 pyramid
- 22:55 Three ways to climb: simplify, prune, dissolve
- 26:46 What H3 really gave us
- 28:29 What this is not: a replacement for vector tiles
- 29:12 Questions I don't have the answer to
- 31:00 Q&A: very large polygons, baked-in levels, why not PMTiles
The level-of-detail-inside-the-file option builds on Kanahiro Iguchi's cloud-optimized GeoParquet work.
Links
- Slides (PDF, 36 slides).
- GeoPQ Workbench: opens, grades, rewrites and queries GeoParquet, H3- and A5-partitioned datasets included. Source on GitHub.
- GeoParquet writing cookbook: partitioning by attribute and by H3 cell, in DuckDB.
- Great datasets, in GeoParquet: the open datasets mentioned in the introduction.
- H3, and the hosts of the conference, Fused.
Building something similar, or want a dataset turned into a pyramid like this one? Get in touch.