ActiveInferenceInstitute

geo-infer-data

Data connectors, ETL pipelines, and data management for geospatial datasets. Use when loading spatial data from databases, APIs, files (GeoJSON, Shapefile, GeoParquet), or building data transformation pipelines.

ActiveInferenceInstitute 15 2 Updated 6mo ago

Resources

10
GitHub

Install

npx skillscat add activeinferenceinstitute/geo-infer/geo-infer-data

Install via the SkillsCat registry.

About this skill

GEO-INFER-DATA provides data connectors, ETL pipelines, and data management tools for geospatial datasets, supporting sources like PostgreSQL/PostGIS, REST APIs, and file formats including GeoJSON, Shapefile, and GeoParquet. It addresses the need to ingest, transform, validate, and cache spatial data with built-in coordinate bounds checking and schema validation. Use it when loading spatial data from various sources or building data transformation pipelines for geospatial workflows.

SKILL.md

GEO-INFER-DATA

Instructions

Core Capabilities

  • Connectors: PostgreSQL/PostGIS, SQLite/SpatiaLite, REST APIs, file I/O
  • Formats: GeoJSON, Shapefile, GeoParquet, GeoTIFF, CSV with coordinates
  • ETL pipelines: Extract → Transform → Load with spatial awareness
  • Caching: Spatial tile caching, query result caching
  • Validation: Schema validation, coordinate bounds checking

Key Imports

from geo_infer_data.connectors.database import DatabaseConnector
from geo_infer_data.core.pipeline import ETLPipeline
from geo_infer_data.formats.geojson import GeoJSONLoader

Examples

from geo_infer_data.formats.geojson import GeoJSONLoader
from geo_infer_data.core.validation import CoordinateValidator

loader = GeoJSONLoader()
features = loader.load("buildings.geojson")
print(f"Loaded {len(features)} features")

# Validate coordinates against WGS84 bounds
validator = CoordinateValidator()
valid, invalid = validator.validate(features)
print(f"Valid: {len(valid)}, Out-of-bounds: {len(invalid)}")
from geo_infer_data.core.pipeline import ETLPipeline

pipeline = ETLPipeline(name="census_ingest")
pipeline.extract(source="postgresql://db/census", query="SELECT * FROM tracts")
pipeline.transform(operations=["reproject_to_4326", "validate_bounds"])
pipeline.load(target="geoparquet", path="output/census.parquet")
pipeline.run()

Guidelines

  • SQL uses parameterized queries (:param placeholders) — never string interpolation
  • All coordinate data validated against WGS84 bounds
  • Test: uv run python -m pytest GEO-INFER-DATA/tests/ -v

Integrations

  • SPACE → Spatial indexing of loaded datasets
  • GIT → Version control for spatial data
  • API → Data source for spatial query endpoints
  • IOT → Sensor data ingestion pipelines
  • EXAMPLES → Example ETL workflows