AI systems performance engineer · Product leader · Founder · Advisor · Three-time O'Reilly author
I build and explain AI systems, from GPU kernels and distributed training to inference serving and the products that put those systems to work. I founded PipelineAI and have worked at AWS, Databricks, and Netflix. Today I build open source performance tools and advise teams bringing AI systems into production.
- AI Performance Engineering: The code and labs behind my latest book. Topics include CUDA, PyTorch, distributed training, profiling, and inference.
- GPU Performance Tuning: Agent workflows and MCP tools for profiling, benchmarking, and improving GPU inference.
- Agent Harness Optimization: A workbench for evaluating agent prompts, tools, transcripts, and traces with the evidence attached.
- Claude Founder Kit: A runnable path from the first Claude API call through evals, cost controls, and product activation.
- AI Systems Performance Engineering: A field guide to performance across AI hardware and software. Browse the code.
- Generative AI on AWS: Co-authored with Antje Barth. Browse the code.
- Data Science on AWS: Co-authored with Antje Barth. Browse the code.
- Generative AI with Large Language Models: The DeepLearning.AI course I co-teach. See the course on fregly.com.
I cohost AI Performance Engineering with Antje Barth. Our technical sessions reach a worldwide community of more than 100,000 people. Watch the YouTube channel or start with these conversations:
- AI Systems Performance Engineering on SuperDataScience
- Software and Hardware Codesign on MLOps Community
- More talks and community sessions on fregly.com
For recent projects, books, talks, and contact details, visit fregly.com.





