The reading list
Essays, tech notes, and the in-between.
- AI·3 min read
CUDA, ROCm, Vulkan, and Metal: Why the GPU Software Stack Matters for Local AI
CUDA is neither a gaming requirement nor an empty AI slogan. Understand GPU backends, AMD's ROCm, Vulkan compute, and Apple's Metal before comparing local inference hardware.

- AI·3 min read
llama.cpp, vLLM, or LM Studio: Which Local LLM Tool Does What?
Separate the desktop app, inference engine, model file, and GPU backend before choosing how to run a local language model.

- AI·3 min read
Beyond Transformers: Five Model Families and the Names You Know
A practical guide to model architectures, with a map of where Qwen3.8, MiniMax-M3, GLM-5.3, Gemma 4, RWKV, and Mamba fit—and how dense and MoE differ.

- AI·3 min read
From Prompts to Harnesses: How AI Coding Became a System
Prompt engineering shapes a request. Context engineering supplies the evidence. Harness engineering makes the coding assistant's work loop useful, bounded, and verifiable.

- AI·3 min read
One Halo, Several Halos, or a Desktop Full of GPUs?
Choose a serving shape for latency, concurrency, and context capacity without confusing replicas, model splitting, or remote cache transfer.

- AI·3 min read
Prefix Caching, Chunked Prefill, and the Trouble with Chunks
Trace shared and changed prompts through cache blocks and scheduling slices, and resolve what people mean by chunked prefixing.

- AI·3 min read
Stop Reprocessing the Same Prompt: Caching for Coding Agents
Arrange stable repository context for reuse, compare provider cache controls, and distinguish API savings from coding-assistant subscription behavior.

- AI·3 min read
Two GPUs, One Narrow Cable: Where Does the Traffic Go?
Follow model traffic through an upstream link and a PCIe switch, and learn what the Franken Strix experiment does and does not prove.

- AI·3 min read
Your Model Fits. Why Does Your Conversation Run Out of Memory?
Build a memory budget for a 128 GB Halo: weights, attention state, runtime buffers, and several people asking questions at once.

- AI·3 min read
What Actually Happens When You Send an AI Prompt?
Follow a repository question through tokenization, prefill, and decode to see why memory capacity, bandwidth, and compute solve different problems.
