Gemma 4 uses per-layer embeddings, a shared KV cache to cut memory, and dual RoPE for mixed local-global attention. Activates 3.8B of 26B params.
Hands-on experience with AI, self-hosting, Linux, and the developer tools Botmonster actually uses
Home Assistant smart irrigation: local control, $25-89 hardware
Build a smart garden irrigation system with Home Assistant, a rain sensor, and automations. No subscriptions, full local control, $25-89 in parts.
PCIe bifurcation: add 4 NVMe drives for $25-50 per adapter
PCIe bifurcation splits one x16 slot into four x4 lanes, so you can add up to four NVMe drives on a $25 adapter at full per-drive bandwidth.
Python Memory Optimization: 50-80% Reduction with memray
Profile and optimize Python memory with memray, tracemalloc, and objgraph. Covers leak detection, generators, __slots__, and cutting peak use.
Running Gemma 4 26B MoE on 8GB VRAM: three strategies that work
Run Google Gemma 4 26B MoE on a budget 8GB VRAM GPU using aggressive quantization, GPU-CPU layer offloading, and tensor parallelism. Full setup guide.
Self-host Plausible Analytics: 1 KB script, no cookies
Self-host Plausible on a cheap VPS with Docker Compose: a 1 KB gzipped privacy-first tracking script, zero cookies, a real Google Analytics swap.






