AI & Full-Stack Engineering Leader · Founder · 15+ years shipping production systems
I design and ship AI systems that hold up in production: LLM agents, retrieval pipelines, evaluation frameworks, and the data platforms underneath them. I have taken AI products from a blank repo to live customers as a founder, led engineering teams in enterprise settings, and defined the AI standards that other engineers build on.
- Lead AI initiatives from 0 → 1. I identify where AI solves a real customer or operational problem, align stakeholders across departments, and own delivery end to end: architecture, model serving, evaluation, and production support.
- Build agentic and retrieval systems over real enterprise data. Multi-agent workflows, hybrid retrieval with reranking, citation-grounded generation, and streaming LLM orchestration across OpenAI, Anthropic, and open-source models.
- Make AI measurable. I establish evaluation practice on Braintrust: regression datasets, scoring rubrics, prompt and retrieval A/B tests, and model comparisons that gate every change before it reaches users.
- Engineer for cost and latency. Edge-optimized open-source model stacks, Redis-backed caching, model-aware routing, and on-device inference. One deployment cut AI operating costs by 40% with no loss in response quality.
- Grow engineers. I have managed a team of four, mentored 8+ engineers, and set the prompting, review, and integration standards that raised team throughput and shortened PR cycle time.
- Founder & Principal Engineer, Dotori.ai — Built CosReg, a live multi-agent compliance analyst that screens cosmetic formulas against 86,000+ ingredients across US, EU, and global regulations, and Signal, an ML price-prediction platform delivered through web and API.
- AI / Senior Full-Stack Engineer at a national membership organization serving 3M+ members — Drive the organization's AI program: a production RAG platform over 10+ years of institutional data, a voice-enabled AI advisor agent, agentic workflows over a 6.9M-contact CRM, and the enterprise BI platform they reason over.
- Graduate Researcher, Columbia University — Open-source AI ecosystems and enterprise adoption: data sovereignty, model licensing, and open-vs-proprietary tradeoffs.
- ollama_exo_proxy_server — Secure, high-performance LLM inference gateway for distributed Exo and Ollama clusters. Model-aware routing and load balancing, API-key security, rate limiting, built-in RAG and vector search, Redis-backed KV caching, and application-level cluster management without Kubernetes. FastAPI · Gunicorn · Redis · MongoDB · ChromaDB · Nginx · Docker
- gen-ai — Full-stack Generative AI playground implementing CNNs, GANs, diffusion models, energy-based models, and LLMs behind a FastAPI backend, with Dockerized deployment and Jupyter support for experimentation. PyTorch · FastAPI · Docker · Jupyter
- AI-Companion — Real-time AI companion providing virtual family interaction for dementia patients. React · TypeScript · Convex
- AML Detection Platform (private, under NDA) — End-to-end anti-money-laundering system combining unsupervised anomaly detection, PCA, and ensemble risk scoring with analyst-facing dashboards. Dockerized and version-controlled for high-stakes automated decisioning. Python · HBOS · Isolation Forest · ECOD · XGBoost · Docker
- Deep-Learning-LLM-Wounded-Treatment — Clinical decision-support platform: CNN multi-class wound classification (90%+ accuracy), vector similarity search over historical cases, and a local LLM for explainable case summaries. PyTorch · CNNs · Vector Search · Docker
- NLP-LLM-Job-Recommendation — NLP platform over LinkedIn job postings using NER, topic modeling, and transformer embeddings to produce explainable compatibility scores and skill-gap analysis. Python · Transformers · NER · Topic Modeling
- Skin-Lesion-Generation-Diffusion — Synthetic dermoscopic image generation with four generative models trained on the HAM10000 and ISIC datasets. PyTorch · Diffusion Models · GANs
| Area | Expertise |
|---|---|
| LLM & Agentic Systems | Autonomous agents, multi-step tool-using workflows, function calling, structured-output pipelines, streaming, real-time voice agents, multi-provider orchestration (OpenAI, Anthropic, Ollama, Exo) |
| Retrieval | RAG architecture, hybrid retrieval and reranking, embeddings, Qdrant, ChromaDB, citation-grounded generation over messy enterprise data |
| Evaluation & Reliability | Braintrust, regression datasets, scoring rubrics, prompt and retrieval A/B testing, Git-style model versioning, reproducible AI behavior in production |
| Deep Learning | PyTorch, TensorFlow, CNNs, GANs, diffusion models, energy-based models, transformers |
| Classical ML & Anomaly Detection | XGBoost, HBOS, Isolation Forest, ECOD, PCA, scikit-learn, pandas |
| Backend & Data | Python, TypeScript/Node.js, PHP/Symfony, SQL · FastAPI, Flask, Express, REST, microservices, event-driven architectures, Celery · PostgreSQL, MySQL/MariaDB, MSSQL, MongoDB, Redis, AWS Redshift |
| Frontend & Mobile | React, Next.js, TypeScript, React Native, Swift |
| Cloud & Delivery | AWS (Lambda, Redshift), Docker, Kubernetes, Nginx, GitHub Actions, Jenkins, serverless, edge and on-device inference |
- Ship, measure, iterate. Fast to a working system, rigorous about measuring it, disciplined about what changes next.
- Build vs. buy on evidence. Hosted APIs where they win, open-source models where they win, and the numbers to show which is which.
- Standards over heroics. Reusable patterns, review discipline, and documentation so the team moves faster than any one person.
- AI-first delivery. I use Claude, Cursor, and similar tools daily across design, scaffolding, refactoring, and testing, and I teach teams to do the same well.
- M.S., Applied Analytics — Columbia University
- B.E. — The City College of New York
I am always open to conversations about LLM infrastructure, agentic systems, evaluation, and building AI teams. Ask me about Python, LLM system design, retrieval, and taking AI products from concept to production.
- GitHub: @hyper07
- Email: kk3789@columbia.edu



