Series

Local AI Toolkit: Runtimes, GPU Backends, and Model Files

Three practical primers on choosing llama.cpp, vLLM, or LM Studio; understanding CUDA, ROCm, Vulkan, and Metal; and reading GGUF, quantization, dense-model, and MoE labels. A companion to the standalone model-architecture guide and Running AI Yourself.

3 parts · first published

Local AI Toolkit: Runtimes, GPU Backends, and Model Files
0:000:00