Data Formats#

gedih3 produces and consumes several data formats throughout the pipeline. This page describes each format, its structure, and when to use it.


H3 Database (Internal Format)#

Created by gh3_build. Optimized for repeated queries with Dask and DuckDB.

h3_database/
├── h3_03=838041fffffffff/
│   ├── 838041fffffffff.metadata.json
│   ├── year=2019/
│   │   ├── 838041fffffffff.2019.0.parquet
│   │   └── 838041fffffffff.2019.0.metadata.json
│   ├── year=2020/
│   │   ├── 838041fffffffff.2020.0.parquet
│   │   └── 838041fffffffff.2020.0.metadata.json
│   │   ...
├── h3_03=83804cfffffffff/
│   ├── 83804cfffffffff.metadata.json
│   ├── year=2019/
│   │   ...
│   ...
├── gedih3_build_log.json
└── _manifest.txt
  • Nested hive-partitioned — first by H3 cell at the partition level (default: level 3, ~12,000 km²), then by year

  • Each H3 directory contains a cell-level metadata file and yearly sub-directories; each year holds a .parquet data file and a companion .metadata.json

  • This two-level scheme caps file size and makes it easy to append new data without touching existing files

  • gedih3_build_log.json records build metadata (products, variables, region, resolution levels); _manifest.txt lists all partition paths

  • Not designed for direct use with external tools — use gh3_extract to produce user-friendly flat files. If you must read the partition files directly, select them with gedih3.gh3_select_partitions(source, region) rather than intersecting the cell polygons yourself (see the selection-safety note below)

Partition metadata sidecar (*.metadata.json)#

Each partition’s sidecar carries two spatial fields that are not interchangeable:

Field

Meaning

Safe for partition selection?

h3_geometry

The exact H3 cell polygon (GeoJSON)

No — H3 children overhang their parent, so a region that clips this polygon silently misses boundary shots

bbox

[minlon, minlat, maxlon, maxlat], overhang-padded to enclose every shot actually stored in the partition

Yes — intersect ROIs against this

The same overhang-padded bbox is also embedded in each parquet file’s GeoParquet footer (columns.geometry.bbox). To compute the padded extent for any cell ID directly, use gedih3.h3_partition_bbox(cell_id, partition_level).

Build Log Keys#

Key

Description

h3_index_level

Fine H3 resolution used for shot-level indexing

h3_partition_level

Coarse H3 resolution used for partitioning (directory names)

products

GEDI products included

columns

Column schema

region

Spatial extent

H3 Dual-Level Structure#

The H3 database uses two H3 resolution levels simultaneously:

  • Partition level (default: 3) — determines the directory structure. A query for a specific region only reads tiles that overlap that region.

  • Index level (default: 12) — the H3 cell ID assigned to each individual GEDI shot, stored as a column/index in every parquet file.

Parent/child caveat: H3 parent hexagons are not perfectly geometrically inclusive of their children. When aggregating across resolution levels, gh3_aggregate uses h3.cell_to_parent() which assigns each shot to its closest parent, which is consistent and fast but not a strict geometric containment. The same property means a partition’s stored shots can fall slightly outside the partition’s own cell polygon — so selecting partition files by exact polygon intersection silently drops boundary shots. Use gh3_select_partitions (or the padded bbox) for direct access. See H3 Indexing for details.


Simplified Dataset (User-Friendly Format)#

Created by gh3_extract and gh3_aggregate. Flat Parquet files for use with any tool.

output/
├── 838041fffffffff.parquet
├── 83804cfffffffff.parquet
├── 83804efffffffff.parquet
└── gedih3_dataset.json
  • Files named by H3 or EGI partition ID

  • gedih3_dataset.json describes the whole dataset (index type, columns, aggregation, etc.)

  • Readable with pandas, R, QGIS, DuckDB, and any other Parquet-compatible tool

  • Used as input for gh3_aggregate, gh3_rasterize, gh3_from_img, gh3_from_polygon, gh3_update

# Read with pandas
import pandas as pd
df = pd.read_parquet('/path/to/output/838041fffffffff.parquet')

# Load all files with gedih3
import gedih3.gh3driver as gh3
gdf = gh3.gh3_load(source='/path/to/output/').compute()

Dataset Metadata (gedih3_dataset.json)#

Key

Description

index_type

"h3" or "egi"

index_level

Spatial resolution level

partition_level

Partition tile size

columns

Data columns included

agg

Aggregation method (if from gh3_aggregate)


GeoTIFF (Raster Output)#

Created by gh3_rasterize or the -R flag in gh3_aggregate. Standard GeoTIFF files compatible with GDAL, QGIS, R (terra), Python (rasterio,rioxarray), and virtually any GIS tool.

# Tiled output (one file per partition)
gh3_rasterize -d aggregated/ -o rasters/ --compress LZW

# Single merged raster
gh3_rasterize -d aggregated/ -m -o output.tif --compress LZW

# Select specific variables
gh3_rasterize -d aggregated/ -l agbd_l4a_mean -o rasters/

Key properties:

  • Tiled by default — output is split by spatial partition for efficient access

  • Compression supportLZW, DEFLATE, ZSTD, NONE

  • BIGTIFF support — for files exceeding 4 GB

  • Time-series naming — when produced from time-windowed data, files are named after their temporal windows

# Load GeoTIFF output in Python
import rioxarray
xds = rioxarray.open_rasterio('agbd_mean.tif')
xds.plot()

Other Supported Formats#

gh3_export (Python API) and gh3_extract support additional output formats beyond Parquet:

Format

Extension

Notes

GeoParquet

.parquet

Default; includes geometry for spatial tools

Feather

.feather

Fast in-memory format

GeoPackage

.gpkg

OGC standard vector format; QGIS native

HDF5

.h5

For compatibility with scientific workflows; no geometry

Shapefile

.shp

Legacy vector format; column name length limited

CSV

.csv

Tabular export; no geometry


Parquet Schema#

Each simplified dataset Parquet file contains:

  • Index column: h3_XX (H3 cell ID, string) or egiXX (EGI hash, uint64)

  • Data columns: product variables (e.g., agbd_l4a, rh_098_l2a)

  • Geometry (optional): geometry column (WKB Point geometries in EPSG:4326)

  • Metadata: stored in Parquet file metadata (accessible via pyarrow)

Inspecting Files#

# Inspect schema from CLI
gh3_read_schema /path/to/output/abc123.parquet
gh3_read_schema /path/to/database/

Choosing Between H3 and EGI#

Consideration

H3

EGI

Grid shape

Hexagonal

Square

Coordinate system

WGS84 (EPSG:4326)

EASE-Grid 2.0 (EPSG:6933)

Rasterization

Requires hex-to-pixel interpolation

Direct 1:1 mapping

GEDI L4B compatible

No

Yes

Parent/child nesting

Approximate (see above)

Exact

Default in gedih3

Yes

No

Support by external software

Yes

No

EGI is the right choice when you need perfect pixel alignment or when producing gridded datasets for interoperability with raster-native workflows. For general analysis and exploratory work, H3 is simpler and faster. See EGI Indexing for a detailed comparison.