Software Engineer II @ American Express
Distributed Systems β’ Databases β’ Caching β’ AI Infrastructure
Portfolio β’ LinkedIn β’ Medium
I'm a software engineer focused on the infrastructure behind modern software systems β distributed systems, databases, caching, and AI infrastructure.
I like understanding systems from first principles:
- How do systems behave under load?
- Where do they stop degrading gracefully?
- Why do failures propagate?
- How do databases and caches behave internally?
- What trade-offs exist between performance, correctness, and reliability?
I explore these questions through open-source contributions, reproducible experiments, systems projects, and technical writing.
I contribute to production infrastructure projects, primarily around correctness, crash-safety, reliability, and performance.
dragonflydb/dragonfly β high-performance in-memory datastore
My merged contributions include fixes across the core C++ engine involving:
- Cluster slot-migration correctness
- Production crash / SIGABRT root-cause analysis
- Cross-shard replication and geo-query crashes
- Sorted-set flag handling
- Search schema validation
- Scripting-path crashes
Several of these contributions are shipped in the v2.0.0 release.
valkey-io/valkey-search
Contributed fixes involving:
- Index-count correctness after coordinator RDB loads
- Thread-pool failure handling
- Shutdown and worker-state correctness
BetterDB-inc/monitor
Contributed reliability and security fixes around:
- MCP server HTTP error handling
- Internal error information exposure
- Docker networking / DNS behavior
I also have ongoing contributions across:
Valkey β’ FoundationDB β’ RocksDB β’ other distributed-systems projects
covering areas such as replication, persistence, indexing, pattern matching, and native memory management.
β All Pull Requests
β Dragonfly Commits
β BetterDB Monitor Commits
Overload experimentation for distributed systems.
SlimyBug is a fault-injection and experimentation platform for studying how distributed systems behave under overload.
Instead of learning failure modes only from production incidents, SlimyBug creates controlled, reproducible experiments around:
- Retry amplification
- Cascading failures
- Circuit breakers
- Admission control
- Connection-pool saturation
- Overload onset
- Dependency latency propagation
- Retry jitter
- Admission deferral
- Connection-pool self-locking
Where exactly does a system stop degrading gracefully β and why?
I've run 12+ causal experiments measuring collapse boundaries and evaluating mitigation strategies. For example, a minimal circuit breaker reduced load reaching a saturated dependency by up to 44% in one experiment, followed by dedicated multi-run validation.
Stack: Python β’ FastAPI β’ PostgreSQL β’ Docker β’ Toxiproxy β’ Prometheus β’ Grafana β’ OpenTelemetry
Learned eviction policies for semantic LLM caches.
SmartEvict investigates whether lightweight learned policies can outperform traditional cache eviction heuristics for semantic LLM workloads.
The project benchmarks a Dueling DQN policy against:
- LRU
- FIFO
- GDSF
using real conversational workloads including LMSYS-Chat-1M and WildChat-1M.
π¦ Repository
π Research Artifact / DOI
At American Express, I work on GenAI infrastructure and backend systems supporting internal financial-data workflows.
Some of the systems I've worked on include:
- Custom MCP infrastructure for contextual financial data access
- Graph-based financial data modeling using PostgreSQL + Apache AGE
- Two-layer semantic caching using pgvector + Redis
- LLM observability and tracing
- Agent memory infrastructure
- Sandboxed Python/Docker execution for data analysis
- FastAPI microservices and production backend infrastructure
Previously at EY, I worked on enterprise NLP/LLM systems, metadata intelligence, vector retrieval, and safety infrastructure for LLM-facing systems.
Across these systems, my focus has consistently been on performance, reliability, correctness, and productionization.
Distributed Systems
Databases & Storage Engines
Caching Systems & Eviction
Replication & Consensus
Performance Engineering
Reliability Engineering
AI / LLM Infrastructure
Developer Infrastructure
Real-Time Cryptocurrency Arbitrage Detection System
A real-time arbitrage detection system built around a high-performance C++ engine processing high-frequency price updates across exchanges.
Stack: C++ β’ WebSockets β’ CUDA β’ React
Dynamic inference optimization using Deep Q-Learning
Implemented early-exit strategies for CNN inference using Deep Q-Learning to dynamically determine when computation can stop while maintaining model accuracy.
Stack: Python β’ TensorFlow β’ Deep Q-Learning
I write about the experiments, investigations, and engineering lessons behind my systems work.
Topics include:
Distributed Systems β’ Databases β’ Caching β’ AI Infrastructure β’ Performance β’ Reliability β’ Software Engineering
π Read on Medium
Languages & Systems
Python β’ C++ β’ Go β’ Distributed Systems β’ System Design
Databases & Messaging
PostgreSQL β’ Redis β’ MongoDB β’ Kafka β’ RabbitMQ β’ Elasticsearch
AI / LLM Infrastructure
OpenAI API β’ LangChain β’ MCP β’ RAG β’ Vector Search β’ pgvector
Backend & Infrastructure
FastAPI β’ Flask β’ Microservices β’ Docker β’ AWS β’ GCP β’ CI/CD β’ Prometheus β’ Grafana β’ OpenTelemetry
- Cache internals & eviction algorithms
- Database internals
- Replication & consensus protocols
- Distributed-systems failure modes
- AI serving infrastructure
- LLM caching
- Performance engineering
- Reproducible systems experimentation
Build systems. Run experiments. Produce evidence. Share findings.
I believe the best way to understand complex systems is to build them, break them, measure them, and investigate why they behave the way they do.
π Portfolio β’ πΌ LinkedIn β’ π Medium β’ π§ Email
Build systems. Run experiments. Produce evidence. Share findings.