Local LLM
A 27B model on two 4GB Maxwell GPUs: inference tuning
What worked, what failed, and why quality gates mattered while running Bonsai-27B on two 4GB Quadro K2200s.
From tokens to value
What worked, what failed, and why quality gates mattered while running Bonsai-27B on two 4GB Quadro K2200s.
A practical engineering write-up on running qwen3.6:35b with Ollama on dual Tesla P40s: throughput, context length, prompt cache, MTP, and coding benchmarks.
A practical total-cost and performance look at used Tesla P40s versus newer 24GB and 32GB GPUs for local LLM inference.
I believe my iPhone has a specific kind of AGI - I have Codex running locally on an old iPhone that can build apps, deploy, test and iterate locally - all without a Mac. After my last weekend's