Isochrone

Region Data Pipeline

This document describes the full pipeline for turning a place in OpenStreetMap into web-loadable assets for this app.

The pipeline now has one external source-of-truth config file and one Python entry point:

Naming/source-of-truth rule:

What Exists Today

Implemented:

Still manual:

Stage 1: Fetch Raw OSM JSON

Run:

./data_pipeline/region-data.py fetch

For the current CLI options and examples, run ./data_pipeline/region-data.py --help or ./data_pipeline/region-data.py <subcommand> --help.

region-data.py prefers the repository .venv/bin/python when that virtualenv exists, so direct execution keeps using the project dependencies even if your shell is currently on another interpreter.

The location list is loaded from data_pipeline/regions.json.

Optional naming fields in that file:

Boundary discovery for each region is also configured there. The optional subdivisionDiscoveryModes array controls how the boundary query finds child administrative relations:

If omitted, both modes are used. Regions whose parent area scans are expensive can disable "area" and use only "subarea", which is how London is configured.

Outputs go to data_pipeline/input/ and are named:

Examples:

This stage only downloads raw Overpass API responses.

Operational debugging behavior:

Query templates used by this stage:

Boundary extracts are written in a download-friendly shape:

The build step reconstructs boundary polylines from those refs, so fetch does not depend on inline way geometry being present in the Overpass response. The boundary query supports both area containment and explicit subarea membership so it can adapt to regions whose administrative relations are modeled differently. It always includes the selected place relation itself in the output, so the build step can fall back to an outer-boundary basemap when no matching child subdivisions exist for that region.

To avoid fetching every configured region, filter by id:

./data_pipeline/region-data.py fetch --only paris

To fetch only one raw input class:

./data_pipeline/region-data.py fetch --only luxembourg-country --components ways
./data_pipeline/region-data.py fetch --only luxembourg-country --components boundaries

Stage 2-4: Build Renderable And Deployable Artifacts

One command now performs:

Run:

./data_pipeline/region-data.py build > web/src/data/locations.json

Notes:

To build only one artifact class:

./data_pipeline/region-data.py build --only luxembourg-country --components graph
./data_pipeline/region-data.py build --only luxembourg-country --components boundary

Optional coast/water context:

The low-level boundary simplifier can also attach a clipped water-polygon layer to the same output JSON as the administrative boundaries:

./data_pipeline/scripts/simplify_boundary_json.py \
  --input data_pipeline/input/rhode-island-district-boundaries.osm.json \
  --output data_pipeline/output/rhode-island-district-boundaries-canvas.json \
  --resolution 25 \
  --units meters \
  --include-coast

Notes:

Forest, inland-water, waterway, and airport context:

Unlike coastal water, this context is always fetched and rendered — there is no opt-in flag. It comes from the same Overpass request as admin boundaries (see docs/overpass_boundary_query.sh’s .naturalArea block), not an external download, so every region gets it automatically once the boundary input is (re)fetched with the current query.

If the boundary input file exists but contains zero Overpass elements, the build step fails explicitly and tells you to rerun fetch for that region. That usually means an older fetch silently produced an empty payload before the stricter fetch validation was added.

The combined command also supports partial selection:

./data_pipeline/region-data.py all --only luxembourg-country --fetch-components ways --build-components graph

Stage 5: Register The Region In The UI

The build and all commands already emit the correct manifest JSON for web/src/data/locations.json. That manifest carries through optional localizedNames from data_pipeline/regions.json so the web app can localize the location menu without introducing a second naming source of truth.

Example:

./data_pipeline/region-data.py build > web/src/data/locations.json

The top-bar location menu reads that file and loads the matching graph and boundary assets.

Stage 6: Publish The New Assets

If the region should be available on GitHub Pages, update .github/workflows/pages.yml so it copies the new files into the site artifact.

Current workflow only publishes Berlin:

For a new region such as Paris, add copies for:

Paris Example

Assuming you only want Paris:

./data_pipeline/region-data.py fetch --only paris
./data_pipeline/region-data.py build --only paris > web/src/data/locations.json

This produces:

And web/src/data/locations.json receives:

{
  "locations": [
    {
      "id": "paris",
      "name": "Paris",
      "graphFileName": "paris-graph.bin.gz",
      "boundaryFileName": "paris-district-boundaries-canvas.json"
    }
  ]
}

Full Process Checklist

  1. Edit data_pipeline/regions.json if the configured region list or per-region metadata should change
  2. Fetch raw Overpass JSON with ./data_pipeline/region-data.py fetch
  3. Build canvas basemaps, binary graphs, gzip artifacts, and stdout manifest with ./data_pipeline/region-data.py build
  4. Redirect stdout to web/src/data/locations.json when the UI should load those regions
  5. Update GitHub Pages workflow if the region should ship in the deployed site

That is the full process as the repository currently stands.