Anmol Mishra
0%
About Experience Projects Research Highlights Gallery Contact Résumé ↓
Available for conversations

Anmol Mishra

> I build

Software Engineer on HPE’s Distributed Data Platform. I run the batch and streaming pipelines that move 1+ TB a day across Spark, Kafka, Debezium and Trino. I contribute upstream to Apache Spark and PyDeequ, and I want to know what it would take to make data quality a first-class property of open datasets.

Anmol Mishra
HPE · Distributed Data Platform
1+ TB / day

A day in the platform

Postgressources
Debeziumchange capture
Kafkastreaming
Sparktransform
Iceberglake
Trinoquery
BI models15+ delivered

Orchestrated by Airflow, running on EKS, watched by Prometheus and Grafana.

01 / About

I care about what happens to data
before anyone models it.

Most of the interesting failures in a data platform never happen in the model. They happen three systems upstream, in a schema nobody owned, in a CDC stream that quietly dropped a column, in a join that was fine until the partition skewed. That is the layer I work at.

At Hewlett Packard Enterprise I design and operate the batch and streaming pipelines behind the Distributed Data Platform: Spark and Scala for processing, Kafka and Debezium for change capture, Trino for query, all of it on Kubernetes and AWS. It moves 1+ TB a day and feeds 15+ BI data models across Customer Success, FinOps and Engineering, so reliability is not an abstraction here. It is a pager.

The rest of my time goes to understanding the systems underneath. I wrote RaftDB, a distributed database built from scratch with Raft consensus, an LSM storage engine and a Calcite query layer. I did not want consensus, recovery and query optimization to stay black boxes, so I built them. I contribute upstream to Apache Spark and AWS PyDeequ, and I am working out what a cloud-native data quality framework for open-source datasets would take.

Based in
Bangalore, India
Currently
SWE, Distributed Data Platform · HPE
Education
B.Tech CSE, VIT Chennai · 8.81 CGPA
Focus
Distributed systems, database internals, data quality
LinkedIn
1+
TB processed / day
15+
BI data models
95%+
Incident SLA met
3+
Yrs in production data
2
Upstream OSS PRs
02 / Experience

Where the work happened

Production data infrastructure, and the degree that led into it.

02b / Open source

Contributions upstream

Changes sent to the projects I depend on every day.

03 / Projects

Things I built to find out how they work

Filter by domain. Everything public links straight to the source.

03b / Tracked upstream

The stack I work inside

Projects I run, fork and follow closely enough to patch.

04 / Research

Data quality, read properly

Working toward a cloud-native data quality framework for open-source datasets. My notes on the papers behind it go here as I write them.

05 / Highlights

The short version

Apache Spark contributor

SPARK-59146 fixes the SQL analyzer so qualified access to source columns survives a pipe SET, with planner changes and SQL test coverage. Open upstream.

PyDeequ, merged

DQDL support via EvaluateDataQuality landed in AWS PyDeequ, with tests and docs. It now ships to everyone using the library.

1+ TB/day in production

Batch and streaming pipelines on Spark, Kafka, Debezium and Trino, with reliable CDC, low-latency sync and a 95%+ incident SLA record behind them.

RaftDB, from scratch

Raft consensus, leader election, log replication, WAL, snapshots, SSTables, Bloom filters and compaction, with a cost-based SQL layer on Apache Calcite over the top.

Cloud-native by default

EKS and Kubernetes, Terraform, Argo CD and CI/CD, plus production observability through Prometheus, Grafana, OpenTelemetry, CloudWatch and Humio.

Data quality research

An ongoing review of validation systems, error detection, label noise and LLM-driven cleaning, working toward a cloud-native data quality framework for open datasets.

07 / Contact

Let’s talk about hard data problems.

Open to conversations about distributed systems, data platform work, database internals and research collaborations. The fastest route is email.