Build a private AI search engine with SearXNG and RAG. Runs on single machine with 12 GB VRAM, delivering cited answers without external queries.
AI
Hands-on guides to LLMs, agents, prompt engineering, and the AI tools Botmonster runs every day for real work, not demos.
Three tiers of AI pair programming: from autocomplete to autonomous overnight agents
Three tiers of AI development: inline completions for flow, agent sprints for features, and overnight batch runs. Route each task to the right one.
Fine-tuning Gemma 4 with Unsloth on a single GPU: a practical guide
Fine-tune Gemma 4 with Unsloth on a single GPU using QLoRA: dataset prep, the full training workflow, and GGUF export for RTX 4090 or Colab.
Gemma 4 vs Qwen 3.5 vs Llama 4: which open model should you actually use? (2026)
Gemma 4, Qwen 3.5, and Llama 4 compared on benchmarks, licensing, speed, and hardware so you can pick the right open model fast.
Local meeting transcriber: Whisper, Ollama, structured notes
Transcribe and summarize meetings locally using Whisper and Llama. Privacy-first pipeline with speaker diarization and automated note generation.
Route Ollama, vLLM, OpenAI through one LiteLLM API
Unify access to Ollama, vLLM, OpenAI, Anthropic, and Google models behind one endpoint. Routing, load balancing, and rate limiting with LiteLLM.
Botmonster Tech




