Google's Gemma 4 spans 2.3B to 31B parameters. Compare hardware needs, speed, and quality across all variants when running locally through Ollama.
AI
Hands-on guides to LLMs, agents, prompt engineering, and the AI tools Botmonster runs every day for real work, not demos.
Self-hosted AI search: combine SearXNG and a local RAG pipeline
Build a private AI search engine with SearXNG and RAG. Runs on single machine with 12 GB VRAM, delivering cited answers without external queries.
Three tiers of AI pair programming: from autocomplete to autonomous overnight agents
Three tiers of AI development: inline completions for flow, agent sprints for features, and overnight batch runs. Route each task to the right one.
Fine-tuning Gemma 4 with Unsloth on a single GPU: a practical guide
Fine-tune Gemma 4 with Unsloth on a single GPU using QLoRA: dataset prep, the full training workflow, and GGUF export for RTX 4090 or Colab.
Gemma 4 vs Qwen 3.5 vs Llama 4: which open model should you actually use? (2026)
Gemma 4, Qwen 3.5, and Llama 4 compared on benchmarks, licensing, speed, and hardware so you can pick the right open model fast.
Local meeting transcriber: Whisper, Ollama, structured notes
Transcribe and summarize meetings locally using Whisper and Llama. Privacy-first pipeline with speaker diarization and automated note generation.
Botmonster Tech




