Three tiers of AI development: inline completions for flow, agent sprints for features, and overnight batch runs. Route each task to the right one.
AI
Hands-on guides to LLMs, agents, prompt engineering, and the AI tools Botmonster runs every day for real work, not demos.
Self-hosted AI search: combine SearXNG and a local RAG pipeline
Build a private AI search engine with SearXNG and RAG. Runs on single machine with 12 GB VRAM, delivering cited answers without external queries.
Claude Code skills ecosystem: 1,340+ installable agent skills for AI coding assistants
Explore Claude Code's 1,340+ agentic skills ecosystem: major repositories, top skills by category, building custom skills, and future directions.
Running multiple AI coding agents in parallel: patterns that actually work
Coordinate multiple AI coding agents in parallel using file ownership, iteration caps, and review gates to beat single-agent throughput on real work.
Route Ollama, vLLM, OpenAI through one LiteLLM API
Unify access to Ollama, vLLM, OpenAI, Anthropic, and Google models behind one endpoint. Routing, load balancing, and rate limiting with LiteLLM.
Gemma 4 vs Qwen 3.5 vs Llama 4: which open model should you actually use? (2026)
Gemma 4, Qwen 3.5, and Llama 4 compared on benchmarks, licensing, speed, and hardware needs so you can pick the right open model for your work fast.
Botmonster Tech




