01 / Flagship / Verified

NASA Earth Observation Event Intelligence Platform

A local geospatial event pipeline built to test identity, lineage, deterministic replay, distributed processing, recovery, spatial serving, and measurable system boundaries.

Version 1.0 Local completeMeasured locallyAWS not deployed

System / Designed and implemented locally

From governed source to spatial evidence.

NASA FIRMS observations enter a controlled Python ingestion path, retain source identity and lineage, and become deterministic replay events for Kafka and Spark processing. Silver and Gold Parquet feed PostgreSQL/PostGIS; a read-only FastAPI layer serves a Streamlit dashboard. Airflow coordinates the batch workflow and its recovery contract.

  1. InputNASA FIRMS · Python ingestion · checksums
  2. TransformationDeterministic replay · Kafka · Spark batch and Structured Streaming
  3. IntelligenceSilver/Gold Parquet · lineage · validation · PostGIS
  4. OutputRead-only FastAPI · bounded cache · Streamlit

Python · Apache Kafka · Apache Spark · Airflow · PostgreSQL/PostGIS · FastAPI · Streamlit · Docker

Evidence / Verified and measured

One million events through the complete local pipeline.

Underlying detections
10,000
NASA FIRMS detections selected for controlled replay
Complete local pipeline
1,000,000
Replay events, not one million original observations
Spark batch
149.502 s
4 CPUs · 4 GiB container · 32 shuffle partitions
Measured throughput
6,688.88/s
Spark batch events per second
PostgreSQL/PostGIS
1M passed
Counts, geometry, identity, lineage, and aggregates reconciled
Portable suite
100 passed
108 collected · 8 environment-dependent skips · 0 failed

Reliability / Passed

Stable identities make reruns inspectable.

Kafka offset reconciliation, Spark checkpoints, staged database loads, stable run identities, and conflict detection prevent silent duplication or overwrite. A checkpoint restart consumed no new Kafka records; an idempotent PostgreSQL rerun inserted zero rows and recognized all one million existing events.

Kafka and PostgreSQL restart tests preserved committed offsets and serving truth. A deliberate database identity conflict rolled back without partial mutation. Airflow verified ordering, bounded retries, failure propagation, and safe rerun behavior on its integration profile.

Capacity boundary / Passed + failed

The 10M result diverges.

Passed

Generation and read-back

Exactly 10,000,000 replay events were generated from 10,000 underlying detections and independently reconciled. Two deterministic outputs produced the same SHA-256 checksum.

Failed

Spark processing

The 10M Spark attempt did not complete. It reached a verified local Java heap exhaustion boundary after approximately 629 seconds; no Spark output or admitted manifest was produced.

Meaning: 10M deterministic generation and verification passed. 10M Spark, Kafka, streaming, PostgreSQL, Airflow, API, and dashboard execution are not claimed.

Cloud boundary / Designed · locally validated · not deployed

AWS exists as architecture, not execution.

CloudFormation defines bounded S3, EMR Serverless, CloudWatch, IAM, KMS, budget, alarm, tagging, and teardown controls. Thirty-two resources passed infrastructure tests and local template validation.

AWS resources created
0
AWS workloads executed
0
Actual AWS cost
$0.00

Limitations / Explicit

What this evidence does not claim.

  • Complete local validation stops at one million events.
  • 10M generation is not 10M Spark or full-platform validation.
  • Local Kafka uses one KRaft broker and does not demonstrate broker failover or high availability.
  • Recorded timings are local sequential measurements, not multi-user load tests.
  • AWS infrastructure was not deployed, and EMR Serverless was not executed.
  • No live users, production telemetry, production workload, or universal capacity is claimed.
Inspect the repositoryOpen résuméContact Nitheesh