Category
AI
Reading paths
Read end to end.
Local AI Toolkit: Runtimes, GPU Backends, and Model Files
Three practical primers on choosing llama.cpp, vLLM, or LM Studio; understanding CUDA, ROCm, Vulkan, and Metal; and reading GGUF, quantization, dense-model, and MoE labels. A companion to the standalone model-architecture guide and Running AI Yourself.
Start →Running AI Yourself: Where the Time and Memory Go
Six practical explainers following a coding assistant from its first repository question to a shared local service: inference, memory budgets, PCIe topology, prompt caching, scheduling, and hardware placement.
Start →The Parallel Developer
An eight-part series on running multiple features in flight with git worktrees, OpenSpec, Beads, AI agents, and OpenGSD — from the first disciplined phase loop to safely operating concurrent workstreams.
Start →Greenfield
Engineering, AI, and the businesses we build from scratch. Solo-narrated essay episodes on the rules, layers, and judgment calls behind working with coding agents in real codebases.
Start →Browse all series →
See every reading path on the site.
- ·3 min read
Agent Config, Installed Tool, and Live Behavior Are Three Different Things
A model or workflow can be correct in source and still be absent from the installed runtime. Verify each layer before claiming it works.

- ·3 min read
Running Independent AI Research Trials
Two isolated research environments can work from the same source material while keeping their analysis independent and reviewable.

- ·3 min read
Two Codex Accounts, One Session History
Use CODEX_HOME for a second Codex login, then share only the session files you need to resume by ID. Here is the setup I use, what stays separate, and what I actually verified.

- ·3 min read
I Tuned My Local AI Server. Then I Tried Ollama.
Vulkan, Flash Attention, speculative decoding, thread counts, and a custom service. After all that tuning, Ollama matched or slightly beat my Qwen3.8 setup on the same mini PC.

- ·3 min read
GGUF, Quantization, Dense, and MoE: How to Read a Local Model Download
Decode model filenames and parameter counts without confusing a file format, numerical precision, expert routing, and the memory needed to run a conversation.

- ·3 min read
CUDA, ROCm, Vulkan, and Metal: Why the GPU Software Stack Matters for Local AI
CUDA is neither a gaming requirement nor an empty AI slogan. Understand GPU backends, AMD's ROCm, Vulkan compute, and Apple's Metal before comparing local inference hardware.

- ·3 min read
llama.cpp, vLLM, or LM Studio: Which Local LLM Tool Does What?
Separate the desktop app, inference engine, model file, and GPU backend before choosing how to run a local language model.

- ·3 min read
Beyond Transformers: Five Model Families and the Names You Know
A practical guide to model architectures, with a map of where Qwen3.8, MiniMax-M3, GLM-5.3, Gemma 4, RWKV, and Mamba fit—and how dense and MoE differ.

- ·3 min read
From Prompts to Harnesses: How AI Coding Became a System
Prompt engineering shapes a request. Context engineering supplies the evidence. Harness engineering makes the coding assistant's work loop useful, bounded, and verifiable.

- ·3 min read
One Halo, Several Halos, or a Desktop Full of GPUs?
Choose a serving shape for latency, concurrency, and context capacity without confusing replicas, model splitting, or remote cache transfer.
