This document describes the full pipeline for turning a place in OpenStreetMap into web-loadable assets for this app.
The pipeline now has one external source-of-truth config file and one Python entry point:
data_pipeline/regions.jsondata_pipeline/region-data.pyNaming/source-of-truth rule:
data_pipeline/regions.json is authoritative for each region’s canonical name and optional localizedNamesweb/src/data/locations.json is generated output for the web app and should inherit those names from the pipeline rather than being edited separatelynameImplemented:
data_pipeline/regions.jsondata_pipeline/region-data.py fetchdata_pipeline/region-data.py builddata_pipeline/region-data.py allweb/src/data/locations.jsonStill manual:
epsg) per regionweb/src/data/locations.json.github/workflows/pages.ymlRun:
./data_pipeline/region-data.py fetch
For the current CLI options and examples, run ./data_pipeline/region-data.py --help or ./data_pipeline/region-data.py <subcommand> --help.
region-data.py prefers the repository .venv/bin/python when that virtualenv exists, so direct execution keeps using the project dependencies even if your shell is currently on another interpreter.
The location list is loaded from data_pipeline/regions.json.
Optional naming fields in that file:
name: canonical fallback/display name used when no locale-specific override existslocalizedNames: optional object keyed by locale code such as de or frBoundary discovery for each region is also configured there. The optional
subdivisionDiscoveryModes array controls how the boundary query finds child
administrative relations:
"area": scan administrative relations inside the selected place area"subarea": follow explicit subarea membership from the selected place relationIf omitted, both modes are used. Regions whose parent area scans are expensive can
disable "area" and use only "subarea", which is how London is configured.
Outputs go to data_pipeline/input/ and are named:
<slug>-routing.osm.json<slug>-district-boundaries.osm.jsonExamples:
data_pipeline/input/paris-routing.osm.jsondata_pipeline/input/paris-district-boundaries.osm.jsonThis stage only downloads raw Overpass API responses.
Operational debugging behavior:
elements list, the pipeline treats that as a failed fetch and writes sidecar debug files next to the intended output:
<output>.failed-query.ql<output>.failed-curl-stderr.txt<output>.failed-response-body.txt<output>.failed-response-headers.txt<output>.failed-curl-stdout.txt when curl produced stdoutQuery templates used by this stage:
docs/overpass_routing_query.shdocs/overpass_boundary_query.shBoundary extracts are written in a download-friendly shape:
The build step reconstructs boundary polylines from those refs, so fetch does not depend on inline way geometry being present in the Overpass response.
The boundary query supports both area containment and explicit subarea membership so it can adapt to regions whose administrative relations are modeled differently.
It always includes the selected place relation itself in the output, so the build step can fall back to an outer-boundary basemap when no matching child subdivisions exist for that region.
To avoid fetching every configured region, filter by id:
./data_pipeline/region-data.py fetch --only paris
To fetch only one raw input class:
./data_pipeline/region-data.py fetch --only luxembourg-country --components ways
./data_pipeline/region-data.py fetch --only luxembourg-country --components boundaries
One command now performs:
Run:
./data_pipeline/region-data.py build > web/src/data/locations.json
Notes:
build reads raw inputs from data_pipeline/input/build writes generated artifacts to data_pipeline/output/build prints only the UI-ready locations manifest JSON to stdoutTo build only one artifact class:
./data_pipeline/region-data.py build --only luxembourg-country --components graph
./data_pipeline/region-data.py build --only luxembourg-country --components boundary
Optional coast/water context:
The low-level boundary simplifier can also attach a clipped water-polygon layer to the same output JSON as the administrative boundaries:
./data_pipeline/scripts/simplify_boundary_json.py \
--input data_pipeline/input/rhode-island-district-boundaries.osm.json \
--output data_pipeline/output/rhode-island-district-boundaries-canvas.json \
--resolution 25 \
--units meters \
--include-coast
Notes:
--include-coast is opt-in; normal boundary builds do not fetch or attach water polygonswater_features--coast-source <local.zip|local.shp> overrides the default download source for offline or debug runsForest, inland-water, waterway, and airport context:
Unlike coastal water, this context is always fetched and rendered — there is no opt-in flag. It comes from the same Overpass request as admin boundaries (see docs/overpass_boundary_query.sh’s .naturalArea block), not an external download, so every region gets it automatically once the boundary input is (re)fetched with the current query.
natural=wood/landuse=forest → forest_features, natural=water → inland_water_features, aeroway=aerodrome → airport_features (all written into the boundary canvas JSON; polygons under ~1 hectare are dropped as digitizing noise)waterway=river|canal|stream → waterway_features, each with a navigable flag derived from the boat tag (boat=yes/permissive/designated → navigable, boat=no → not navigable, otherwise canals default navigable and rivers/streams default not)natural=water and aeroway=aerodrome multipolygon relations are also included (unlike forest/landuse, which stay ways-only) — large real-world features like a tidal river or a major airport are frequently mapped as relations, not a single way. Relation member ways are stitched into closed ring(s) (_stitch_ways_into_rings in boundary_canvas.py), with outer/inner member roles wound in opposite directions so islands/holes render correctly under the nonzero-winding fill rule. This is deliberately narrower than a full generic multipolygon implementation: only these two specific tags get relation support, because a relation-typed area scan over a broad tag (boundary=administrative) is the exact query shape that was previously found too expensive for some regions (see the subdivisionDiscoveryModes note above) — natural=water/aeroway=aerodrome relations are rare/localized enough by comparison that the scan is cheap (measured ~7s for all of Greater London).fetch --components boundary for a region to pick them upIf the boundary input file exists but contains zero Overpass elements, the build step fails explicitly and tells you to rerun fetch for that region. That usually means an older fetch silently produced an empty payload before the stricter fetch validation was added.
The combined command also supports partial selection:
./data_pipeline/region-data.py all --only luxembourg-country --fetch-components ways --build-components graph
epsg, subdivisionAdminLevel, output filenames, and relation selectors come from data_pipeline/regions.jsongraph-walk.bin / graph-walk.bin.gz filenames because that is what the web runtime currently references by defaultThe build and all commands already emit the correct manifest JSON for web/src/data/locations.json.
That manifest carries through optional localizedNames from data_pipeline/regions.json so the web app can localize the location menu without introducing a second naming source of truth.
Example:
./data_pipeline/region-data.py build > web/src/data/locations.json
The top-bar location menu reads that file and loads the matching graph and boundary assets.
If the region should be available on GitHub Pages, update .github/workflows/pages.yml so it copies the new files into the site artifact.
Current workflow only publishes Berlin:
data_pipeline/output/berlin-district-boundaries-canvas.jsondata_pipeline/output/graph-walk.bin.gzFor a new region such as Paris, add copies for:
data_pipeline/output/paris-district-boundaries-canvas.jsondata_pipeline/output/paris-graph.bin.gzAssuming you only want Paris:
./data_pipeline/region-data.py fetch --only paris
./data_pipeline/region-data.py build --only paris > web/src/data/locations.json
This produces:
data_pipeline/input/paris-routing.osm.jsondata_pipeline/input/paris-district-boundaries.osm.jsondata_pipeline/output/paris-district-boundaries-canvas.jsondata_pipeline/output/paris-graph.bindata_pipeline/output/paris-graph.bin.gzdata_pipeline/output/paris-graph-summary.jsonAnd web/src/data/locations.json receives:
{
"locations": [
{
"id": "paris",
"name": "Paris",
"graphFileName": "paris-graph.bin.gz",
"boundaryFileName": "paris-district-boundaries-canvas.json"
}
]
}
data_pipeline/regions.json if the configured region list or per-region metadata should change./data_pipeline/region-data.py fetch./data_pipeline/region-data.py buildweb/src/data/locations.json when the UI should load those regionsThat is the full process as the repository currently stands.