
Engineering
Feb 24, 2025 · 5 min
Docker's Dynamic MCP and the Context Window Problem
05 min read

Inference
Jan 30, 2025 · 3 min
LPUs and What They Do That GPUs Do Not
03 min read

Experiments
Dec 18, 2024 · 4 min
Fine-Tuning Large Language Models: PEFT, LoRA, and QLoRA in Practice
04 min read

Inference
Dec 10, 2024 · 4 min
Deploying AI Models with NVIDIA NIM: Production LLMs as Microservices
04 min read

Evaluation
Oct 28, 2024 · 4 min
How Retrieval-Augmented Generation Works
04 min read

Experiments
Oct 14, 2024 · 1 min
SpaceX Caught a Rocket With the Launch Tower
01 min read

Systems
Sep 20, 2024 · 2 min
Windows to Linux: What Changed in How I Work
02 min read

Systems
Sep 05, 2024 · 4 min
Why C++ Still Matters: The Backbone of High-Performance AI and Systems
04 min read