01 / Flagship / Verified
NASA Earth Observation Event Intelligence Platform
A local geospatial event pipeline built to test identity, lineage, deterministic replay, distributed processing, recovery, spatial serving, and measurable system boundaries.
System / Designed and implemented locally
From governed source to spatial evidence.
NASA FIRMS observations enter a controlled Python ingestion path, retain source identity and lineage, and become deterministic replay events for Kafka and Spark processing. Silver and Gold Parquet feed PostgreSQL/PostGIS; a read-only FastAPI layer serves a Streamlit dashboard. Airflow coordinates the batch workflow and its recovery contract.
- InputNASA FIRMS · Python ingestion · checksums
- TransformationDeterministic replay · Kafka · Spark batch and Structured Streaming
- IntelligenceSilver/Gold Parquet · lineage · validation · PostGIS
- OutputRead-only FastAPI · bounded cache · Streamlit
Python · Apache Kafka · Apache Spark · Airflow · PostgreSQL/PostGIS · FastAPI · Streamlit · Docker
Evidence / Verified and measured
One million events through the complete local pipeline.
- Underlying detections
- 10,000 NASA FIRMS detections selected for controlled replay
- Complete local pipeline
- 1,000,000 Replay events, not one million original observations
- Spark batch
- 149.502 s 4 CPUs · 4 GiB container · 32 shuffle partitions
- Measured throughput
- 6,688.88/s Spark batch events per second
- PostgreSQL/PostGIS
- 1M passed Counts, geometry, identity, lineage, and aggregates reconciled
- Portable suite
- 100 passed 108 collected · 8 environment-dependent skips · 0 failed
Reliability / Passed
Stable identities make reruns inspectable.
Kafka offset reconciliation, Spark checkpoints, staged database loads, stable run identities, and conflict detection prevent silent duplication or overwrite. A checkpoint restart consumed no new Kafka records; an idempotent PostgreSQL rerun inserted zero rows and recognized all one million existing events.
Kafka and PostgreSQL restart tests preserved committed offsets and serving truth. A deliberate database identity conflict rolled back without partial mutation. Airflow verified ordering, bounded retries, failure propagation, and safe rerun behavior on its integration profile.
Capacity boundary / Passed + failed
The 10M result diverges.
Generation and read-back
Exactly 10,000,000 replay events were generated from 10,000 underlying detections and independently reconciled. Two deterministic outputs produced the same SHA-256 checksum.
Spark processing
The 10M Spark attempt did not complete. It reached a verified local Java heap exhaustion boundary after approximately 629 seconds; no Spark output or admitted manifest was produced.
Meaning: 10M deterministic generation and verification passed. 10M Spark, Kafka, streaming, PostgreSQL, Airflow, API, and dashboard execution are not claimed.
Cloud boundary / Designed · locally validated · not deployed
AWS exists as architecture, not execution.
CloudFormation defines bounded S3, EMR Serverless, CloudWatch, IAM, KMS, budget, alarm, tagging, and teardown controls. Thirty-two resources passed infrastructure tests and local template validation.
- AWS resources created
- 0
- AWS workloads executed
- 0
- Actual AWS cost
- $0.00
Limitations / Explicit
What this evidence does not claim.
- Complete local validation stops at one million events.
- 10M generation is not 10M Spark or full-platform validation.
- Local Kafka uses one KRaft broker and does not demonstrate broker failover or high availability.
- Recorded timings are local sequential measurements, not multi-user load tests.
- AWS infrastructure was not deployed, and EMR Serverless was not executed.
- No live users, production telemetry, production workload, or universal capacity is claimed.