Skip to content

Latest commit

Β 

History

97 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 

Repository files navigation

Hi πŸ‘‹ I am Shikha Pandey

Software Engineer II @ American Express
Distributed Systems β€’ Databases β€’ Caching β€’ AI Infrastructure

Portfolio β€’ LinkedIn β€’ Medium


About

I'm a software engineer focused on the infrastructure behind modern software systems β€” distributed systems, databases, caching, and AI infrastructure.

I like understanding systems from first principles:

  • How do systems behave under load?
  • Where do they stop degrading gracefully?
  • Why do failures propagate?
  • How do databases and caches behave internally?
  • What trade-offs exist between performance, correctness, and reliability?

I explore these questions through open-source contributions, reproducible experiments, systems projects, and technical writing.


πŸ”§ Systems & Infrastructure Contributions

I contribute to production infrastructure projects, primarily around correctness, crash-safety, reliability, and performance.

Dragonfly

dragonflydb/dragonfly β€” high-performance in-memory datastore

My merged contributions include fixes across the core C++ engine involving:

  • Cluster slot-migration correctness
  • Production crash / SIGABRT root-cause analysis
  • Cross-shard replication and geo-query crashes
  • Sorted-set flag handling
  • Search schema validation
  • Scripting-path crashes

Several of these contributions are shipped in the v2.0.0 release.

Valkey Search

valkey-io/valkey-search

Contributed fixes involving:

  • Index-count correctness after coordinator RDB loads
  • Thread-pool failure handling
  • Shutdown and worker-state correctness

BetterDB

BetterDB-inc/monitor

Contributed reliability and security fixes around:

  • MCP server HTTP error handling
  • Internal error information exposure
  • Docker networking / DNS behavior

Additional OSS Work

I also have ongoing contributions across:

Valkey β€’ FoundationDB β€’ RocksDB β€’ other distributed-systems projects

covering areas such as replication, persistence, indexing, pattern matching, and native memory management.

β†’ All Pull Requests
β†’ Dragonfly Commits
β†’ BetterDB Monitor Commits


πŸš€ Current Work

πŸ› SlimyBug

Overload experimentation for distributed systems.

SlimyBug is a fault-injection and experimentation platform for studying how distributed systems behave under overload.

Instead of learning failure modes only from production incidents, SlimyBug creates controlled, reproducible experiments around:

  • Retry amplification
  • Cascading failures
  • Circuit breakers
  • Admission control
  • Connection-pool saturation
  • Overload onset
  • Dependency latency propagation
  • Retry jitter
  • Admission deferral
  • Connection-pool self-locking

What I'm investigating

Where exactly does a system stop degrading gracefully β€” and why?

I've run 12+ causal experiments measuring collapse boundaries and evaluating mitigation strategies. For example, a minimal circuit breaker reduced load reaching a saturated dependency by up to 44% in one experiment, followed by dedicated multi-run validation.

Stack: Python β€’ FastAPI β€’ PostgreSQL β€’ Docker β€’ Toxiproxy β€’ Prometheus β€’ Grafana β€’ OpenTelemetry


🧠 SmartEvict

Learned eviction policies for semantic LLM caches.

SmartEvict investigates whether lightweight learned policies can outperform traditional cache eviction heuristics for semantic LLM workloads.

The project benchmarks a Dueling DQN policy against:

  • LRU
  • FIFO
  • GDSF

using real conversational workloads including LMSYS-Chat-1M and WildChat-1M.

πŸ“¦ Repository
πŸ“– Research Artifact / DOI


πŸ’Ό Production Engineering

At American Express, I work on GenAI infrastructure and backend systems supporting internal financial-data workflows.

Some of the systems I've worked on include:

  • Custom MCP infrastructure for contextual financial data access
  • Graph-based financial data modeling using PostgreSQL + Apache AGE
  • Two-layer semantic caching using pgvector + Redis
  • LLM observability and tracing
  • Agent memory infrastructure
  • Sandboxed Python/Docker execution for data analysis
  • FastAPI microservices and production backend infrastructure

Previously at EY, I worked on enterprise NLP/LLM systems, metadata intelligence, vector retrieval, and safety infrastructure for LLM-facing systems.

Across these systems, my focus has consistently been on performance, reliability, correctness, and productionization.


πŸ”¬ Areas of Interest

Distributed Systems
Databases & Storage Engines
Caching Systems & Eviction
Replication & Consensus
Performance Engineering
Reliability Engineering
AI / LLM Infrastructure
Developer Infrastructure


πŸ§ͺ Systems Projects

⚑ ArbiSim

Real-Time Cryptocurrency Arbitrage Detection System

A real-time arbitrage detection system built around a high-performance C++ engine processing high-frequency price updates across exchanges.

Stack: C++ β€’ WebSockets β€’ CUDA β€’ React


🧠 Early-Exit CNN

Dynamic inference optimization using Deep Q-Learning

Implemented early-exit strategies for CNN inference using Deep Q-Learning to dynamically determine when computation can stop while maintaining model accuracy.

Stack: Python β€’ TensorFlow β€’ Deep Q-Learning


✍️ Writing

I write about the experiments, investigations, and engineering lessons behind my systems work.

Topics include:

Distributed Systems β€’ Databases β€’ Caching β€’ AI Infrastructure β€’ Performance β€’ Reliability β€’ Software Engineering

πŸ“š Read on Medium


πŸ› οΈ Technologies

Languages & Systems

Python β€’ C++ β€’ Go β€’ Distributed Systems β€’ System Design

Databases & Messaging

PostgreSQL β€’ Redis β€’ MongoDB β€’ Kafka β€’ RabbitMQ β€’ Elasticsearch

AI / LLM Infrastructure

OpenAI API β€’ LangChain β€’ MCP β€’ RAG β€’ Vector Search β€’ pgvector

Backend & Infrastructure

FastAPI β€’ Flask β€’ Microservices β€’ Docker β€’ AWS β€’ GCP β€’ CI/CD β€’ Prometheus β€’ Grafana β€’ OpenTelemetry


🌱 Currently Exploring

  • Cache internals & eviction algorithms
  • Database internals
  • Replication & consensus protocols
  • Distributed-systems failure modes
  • AI serving infrastructure
  • LLM caching
  • Performance engineering
  • Reproducible systems experimentation

πŸ“Š Engineering Philosophy

Build systems. Run experiments. Produce evidence. Share findings.

I believe the best way to understand complex systems is to build them, break them, measure them, and investigate why they behave the way they do.


πŸ“« Connect

🌐 Portfolio β€’ πŸ’Ό LinkedIn β€’ πŸ“ Medium β€’ πŸ“§ Email


Build systems. Run experiments. Produce evidence. Share findings.