Data Formats#
gedih3 produces and consumes several data formats throughout the pipeline. This page describes each format, its structure, and when to use it.
H3 Database (Internal Format)#
Created by gh3_build. Optimized for repeated queries with Dask and DuckDB.
h3_database/
├── h3_03=838041fffffffff/
│ ├── 838041fffffffff.metadata.json
│ ├── year=2019/
│ │ ├── 838041fffffffff.2019.0.parquet
│ │ └── 838041fffffffff.2019.0.metadata.json
│ ├── year=2020/
│ │ ├── 838041fffffffff.2020.0.parquet
│ │ └── 838041fffffffff.2020.0.metadata.json
│ │ ...
├── h3_03=83804cfffffffff/
│ ├── 83804cfffffffff.metadata.json
│ ├── year=2019/
│ │ ...
│ ...
├── gedih3_build_log.json
└── _manifest.txt
Nested hive-partitioned — first by H3 cell at the partition level (default: level 3, ~12,000 km²), then by year
Each H3 directory contains a cell-level metadata file and yearly sub-directories; each year holds a
.parquetdata file and a companion.metadata.jsonThis two-level scheme caps file size and makes it easy to append new data without touching existing files
gedih3_build_log.jsonrecords build metadata (products, variables, region, resolution levels);_manifest.txtlists all partition pathsNot designed for direct use with external tools — use
gh3_extractto produce user-friendly flat files. If you must read the partition files directly, select them withgedih3.gh3_select_partitions(source, region)rather than intersecting the cell polygons yourself (see the selection-safety note below)
Partition metadata sidecar (*.metadata.json)#
Each partition’s sidecar carries two spatial fields that are not interchangeable:
Field |
Meaning |
Safe for partition selection? |
|---|---|---|
|
The exact H3 cell polygon (GeoJSON) |
No — H3 children overhang their parent, so a region that clips this polygon silently misses boundary shots |
|
|
Yes — intersect ROIs against this |
The same overhang-padded bbox is also embedded in each parquet file’s GeoParquet footer (columns.geometry.bbox). To compute the padded extent for any cell ID directly, use gedih3.h3_partition_bbox(cell_id, partition_level).
Build Log Keys#
Key |
Description |
|---|---|
|
Fine H3 resolution used for shot-level indexing |
|
Coarse H3 resolution used for partitioning (directory names) |
|
GEDI products included |
|
Column schema |
|
Spatial extent |
H3 Dual-Level Structure#
The H3 database uses two H3 resolution levels simultaneously:
Partition level (default: 3) — determines the directory structure. A query for a specific region only reads tiles that overlap that region.
Index level (default: 12) — the H3 cell ID assigned to each individual GEDI shot, stored as a column/index in every parquet file.
Parent/child caveat: H3 parent hexagons are not perfectly geometrically inclusive of their children. When aggregating across resolution levels,
gh3_aggregateusesh3.cell_to_parent()which assigns each shot to its closest parent, which is consistent and fast but not a strict geometric containment. The same property means a partition’s stored shots can fall slightly outside the partition’s own cell polygon — so selecting partition files by exact polygon intersection silently drops boundary shots. Usegh3_select_partitions(or the paddedbbox) for direct access. See H3 Indexing for details.
Simplified Dataset (User-Friendly Format)#
Created by gh3_extract and gh3_aggregate. Flat Parquet files for use with any tool.
output/
├── 838041fffffffff.parquet
├── 83804cfffffffff.parquet
├── 83804efffffffff.parquet
└── gedih3_dataset.json
Files named by H3 or EGI partition ID
gedih3_dataset.jsondescribes the whole dataset (index type, columns, aggregation, etc.)Readable with pandas, R, QGIS, DuckDB, and any other Parquet-compatible tool
Used as input for
gh3_aggregate,gh3_rasterize,gh3_from_img,gh3_from_polygon,gh3_update
# Read with pandas
import pandas as pd
df = pd.read_parquet('/path/to/output/838041fffffffff.parquet')
# Load all files with gedih3
import gedih3.gh3driver as gh3
gdf = gh3.gh3_load(source='/path/to/output/').compute()
Dataset Metadata (gedih3_dataset.json)#
Key |
Description |
|---|---|
|
|
|
Spatial resolution level |
|
Partition tile size |
|
Data columns included |
|
Aggregation method (if from |
GeoTIFF (Raster Output)#
Created by gh3_rasterize or the -R flag in gh3_aggregate. Standard GeoTIFF files compatible with GDAL, QGIS, R (terra), Python (rasterio,rioxarray), and virtually any GIS tool.
# Tiled output (one file per partition)
gh3_rasterize -d aggregated/ -o rasters/ --compress LZW
# Single merged raster
gh3_rasterize -d aggregated/ -m -o output.tif --compress LZW
# Select specific variables
gh3_rasterize -d aggregated/ -l agbd_l4a_mean -o rasters/
Key properties:
Tiled by default — output is split by spatial partition for efficient access
Compression support —
LZW,DEFLATE,ZSTD,NONEBIGTIFF support — for files exceeding 4 GB
Time-series naming — when produced from time-windowed data, files are named after their temporal windows
# Load GeoTIFF output in Python
import rioxarray
xds = rioxarray.open_rasterio('agbd_mean.tif')
xds.plot()
Other Supported Formats#
gh3_export (Python API) and gh3_extract support additional output formats beyond Parquet:
Format |
Extension |
Notes |
|---|---|---|
GeoParquet |
|
Default; includes geometry for spatial tools |
Feather |
|
Fast in-memory format |
GeoPackage |
|
OGC standard vector format; QGIS native |
HDF5 |
|
For compatibility with scientific workflows; no geometry |
Shapefile |
|
Legacy vector format; column name length limited |
CSV |
|
Tabular export; no geometry |
Parquet Schema#
Each simplified dataset Parquet file contains:
Index column:
h3_XX(H3 cell ID, string) oregiXX(EGI hash, uint64)Data columns: product variables (e.g.,
agbd_l4a,rh_098_l2a)Geometry (optional):
geometrycolumn (WKB Point geometries in EPSG:4326)Metadata: stored in Parquet file metadata (accessible via
pyarrow)
Inspecting Files#
# Inspect schema from CLI
gh3_read_schema /path/to/output/abc123.parquet
gh3_read_schema /path/to/database/
Choosing Between H3 and EGI#
Consideration |
H3 |
EGI |
|---|---|---|
Grid shape |
Hexagonal |
Square |
Coordinate system |
WGS84 (EPSG:4326) |
EASE-Grid 2.0 (EPSG:6933) |
Rasterization |
Requires hex-to-pixel interpolation |
Direct 1:1 mapping |
GEDI L4B compatible |
No |
Yes |
Parent/child nesting |
Approximate (see above) |
Exact |
Default in gedih3 |
Yes |
No |
Support by external software |
Yes |
No |
EGI is the right choice when you need perfect pixel alignment or when producing gridded datasets for interoperability with raster-native workflows. For general analysis and exploratory work, H3 is simpler and faster. See EGI Indexing for a detailed comparison.