Apache Spark contributor
SPARK-59146 fixes the SQL analyzer so qualified access to source columns survives a pipe SET, with planner changes and SQL test coverage. Open upstream.
> I build
Software Engineer on HPE’s Distributed Data Platform. I run the batch and streaming pipelines that move 1+ TB a day across Spark, Kafka, Debezium and Trino. I contribute upstream to Apache Spark and PyDeequ, and I want to know what it would take to make data quality a first-class property of open datasets.
A day in the platform
Orchestrated by Airflow, running on EKS, watched by Prometheus and Grafana.
Most of the interesting failures in a data platform never happen in the model. They happen three systems upstream, in a schema nobody owned, in a CDC stream that quietly dropped a column, in a join that was fine until the partition skewed. That is the layer I work at.
At Hewlett Packard Enterprise I design and operate the batch and streaming pipelines behind the Distributed Data Platform: Spark and Scala for processing, Kafka and Debezium for change capture, Trino for query, all of it on Kubernetes and AWS. It moves 1+ TB a day and feeds 15+ BI data models across Customer Success, FinOps and Engineering, so reliability is not an abstraction here. It is a pager.
The rest of my time goes to understanding the systems underneath. I wrote RaftDB, a distributed database built from scratch with Raft consensus, an LSM storage engine and a Calcite query layer. I did not want consensus, recovery and query optimization to stay black boxes, so I built them. I contribute upstream to Apache Spark and AWS PyDeequ, and I am working out what a cloud-native data quality framework for open-source datasets would take.
Production data infrastructure, and the degree that led into it.
Changes sent to the projects I depend on every day.
Filter by domain. Everything public links straight to the source.
Projects I run, fork and follow closely enough to patch.
Working toward a cloud-native data quality framework for open-source datasets. My notes on the papers behind it go here as I write them.
SPARK-59146 fixes the SQL analyzer so qualified access to source columns survives a pipe SET, with planner changes and SQL test coverage. Open upstream.
DQDL support via EvaluateDataQuality landed in AWS PyDeequ, with tests and docs. It now ships to everyone using the library.
Batch and streaming pipelines on Spark, Kafka, Debezium and Trino, with reliable CDC, low-latency sync and a 95%+ incident SLA record behind them.
Raft consensus, leader election, log replication, WAL, snapshots, SSTables, Bloom filters and compaction, with a cost-based SQL layer on Apache Calcite over the top.
EKS and Kubernetes, Terraform, Argo CD and CI/CD, plus production observability through Prometheus, Grafana, OpenTelemetry, CloudWatch and Humio.
An ongoing review of validation systems, error detection, label noise and LLM-driven cleaning, working toward a cloud-native data quality framework for open datasets.
Conferences, teams and the occasional whiteboard.
Open to conversations about distributed systems, data platform work, database internals and research collaborations. The fastest route is email.