<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Llm on rHomelab</title><link>https://rhomelab.com/tags/llm/</link><description>Recent content in Llm on rHomelab</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Tue, 22 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://rhomelab.com/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>Self-Hosted AI: Running Ollama and Open WebUI on Your Homelab</title><link>https://rhomelab.com/self-hosted/self-hosted-ai-ollama-open-webui/</link><pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate><guid>https://rhomelab.com/self-hosted/self-hosted-ai-ollama-open-webui/</guid><description>Running your own local LLM stack with Ollama and Open WebUI, what hardware you actually need, which models make sense at homelab scale, and where a self-hosted model beats (and loses to) a cloud API.</description><content:encoded><![CDATA[<p>Every other self-hosted app on this site replaces a subscription with something you run yourself. AI is a little different, the frontier cloud models are genuinely ahead of anything you can run on homelab-grade hardware, and pretending otherwise sets you up for disappointment. What self-hosting AI actually buys you is privacy for anything you don&rsquo;t want leaving your network, zero per-token cost for high-volume low-stakes tasks, and a model that keeps working when your internet doesn&rsquo;t. Ollama and Open WebUI are the two pieces most homelabbers land on to get there, and they&rsquo;re worth understanding separately before you stand up either one.</p>
<h2 id="the-hardware-reality-before-anything-else">The hardware reality, before anything else</h2>
<p>The single number that determines what you can run is VRAM, not CPU, not system RAM, not disk speed. Large language models are a pile of weights that has to sit in GPU memory for inference to run at a usable speed, and the model&rsquo;s file size on disk is roughly what it needs in VRAM to load. A model quantized to 4-bit (the common default for local use, trading a small amount of accuracy for a large drop in size) that&rsquo;s listed as needing &ldquo;8GB&rdquo; means an 8GB card is close to the minimum, not a comfortable margin, once you account for context window and any other GPU load on the box.</p>
<p>Realistic tiers, assuming 4-bit quantized models and a dedicated GPU:</p>
<ul>
<li><strong>8GB VRAM (e.g. an older mid-range card):</strong> small models in the 7-8B parameter range, usable for chat, summarization, and simple drafting. Fine, but you&rsquo;ll notice the ceiling on anything requiring real reasoning depth.</li>
<li><strong>12-16GB VRAM:</strong> mid-size models (roughly 13-14B), a meaningfully better default for general use without stepping up to serious hardware.</li>
<li><strong>24GB+ VRAM:</strong> larger models (30B and up, depending on quantization) start to feel closer to a genuinely capable general-purpose assistant, still short of frontier cloud models but a real step up.</li>
</ul>
<p>CPU-only inference works and is how most people start, it&rsquo;s just slow, expect real time-to-first-token delays and low tokens-per-second, tolerable for an occasional query, painful for anything conversational or high-volume. If a GPU is available on the box (even an older consumer card you weren&rsquo;t otherwise using), pass it through, the difference in usability is not subtle.</p>
<h2 id="ollama-the-model-runtime">Ollama: the model runtime</h2>
<p>Ollama is the piece that actually loads and runs models. It wraps <code>llama.cpp</code> under the hood, handles model downloading and quantization variants through a simple pull-based interface (<code>ollama pull &lt;model&gt;</code>), and exposes an OpenAI-compatible-ish local API on port 11434 that other tools, including Open WebUI, talk to. You don&rsquo;t hand-manage GGUF files or compile anything yourself unless you want to, Ollama&rsquo;s model library covers the popular open-weight families and updates as new releases land.</p>
<p>Running it in a homelab context, a few things matter more than the quickstart guides mention:</p>
<ul>
<li><strong>Run it as its own service, in its own LXC or VM</strong>, not bolted onto whatever box happens to have a GPU. Model downloads are large (multi-gigabyte per model, and you&rsquo;ll likely keep more than one around) and inference is CPU/GPU-intensive in bursts, isolating it means a heavy inference run doesn&rsquo;t compete with anything else you&rsquo;re running.</li>
<li><strong>GPU passthrough is the real setup cost.</strong> If Ollama is landing in a VM, that means proper PCIe passthrough with IOMMU groups and vfio-pci binding. If it&rsquo;s landing in an LXC instead, that&rsquo;s the simpler device-passthrough path (<code>/dev/nvidia*</code> or <code>/dev/dri</code> bound into the container), covered in more depth in this site&rsquo;s article on Proxmox GPU passthrough for VMs vs LXCs, worth reading before you commit to either approach here.</li>
<li><strong>Set <code>OLLAMA_HOST=0.0.0.0</code></strong> deliberately, not by accident, if you want other machines on your network reaching it. The default binds to localhost only, which is the safer starting point until you&rsquo;ve decided who else should be able to hit the API.</li>
<li><strong>Model storage adds up fast.</strong> A handful of model variants at a few gigabytes each turns into real disk usage quickly, point Ollama&rsquo;s model directory at storage you&rsquo;ve actually planned for, not whatever&rsquo;s left on the boot volume.</li>
</ul>
<h2 id="open-webui-the-interface-people-actually-want">Open WebUI: the interface people actually want</h2>
<p>Ollama by itself is a command-line tool and a local API, functional but not something you&rsquo;re going to hand a family member. Open WebUI sits in front of it (or in front of any OpenAI-compatible backend, including actual cloud APIs if you want one interface for both) and gives you a ChatGPT-style web interface: conversation history, multiple users with their own logins, model switching per conversation, and document upload for basic retrieval-augmented generation against your own files.</p>
<p>Deploy it as its own container, pointed at Ollama&rsquo;s API address, and put it behind the same reverse proxy and TLS setup you&rsquo;re already running for everything else self-hosted on your network, this is not a service you want to expose without at least basic auth in front of it if there&rsquo;s any chance of it reaching beyond your LAN. The multi-user support is genuinely useful the moment more than one person in the house wants to use it, each account gets its own history and doesn&rsquo;t see anyone else&rsquo;s conversations by default.</p>
<p>Worth knowing going in: the document upload / RAG feature is convenient for quick lookups against a handful of files, but it&rsquo;s not a substitute for a real vector database and retrieval pipeline if you&rsquo;re trying to build something more serious against a large document set. Treat it as &ldquo;ask questions about a PDF I just uploaded,&rdquo; not as the foundation of a knowledge base.</p>
<h2 id="picking-a-model-without-overpromising">Picking a model without overpromising</h2>
<p>This is the part where honesty matters most. A locally-run open-weight model, even a good one, is not a drop-in replacement for a frontier cloud model like Claude or GPT for hard reasoning, long-context work, or anything where being wrong has a real cost. What it&rsquo;s genuinely good for at homelab scale:</p>
<ul>
<li>Summarizing your own notes, emails, or logs without sending them anywhere</li>
<li>Drafting text you&rsquo;re going to edit anyway</li>
<li>Simple classification and extraction tasks run in bulk, where per-token cloud pricing would add up</li>
<li>A privacy-sensitive assistant for anything you specifically don&rsquo;t want touching a third party&rsquo;s servers</li>
<li>Learning how these systems actually work, which is reason enough on its own for a lot of homelabbers</li>
</ul>
<p>If you&rsquo;re weighing this against paying for a cloud subscription, be clear-eyed about the actual tradeoff: you&rsquo;re not getting equivalent capability for free, you&rsquo;re trading capability for privacy, offline availability, and no per-request cost on the tasks where a smaller model is genuinely good enough. For anything that actually needs the strongest available reasoning, that&rsquo;s still a cloud API call, and pretending your homelab box replaces that is the kind of overclaim that gets homelabbers a bad reputation with people who tried it once and got a worse answer than they expected.</p>
<h2 id="where-this-fits-in-a-real-homelab">Where this fits in a real homelab</h2>
<p>If you already have a GPU-capable Proxmox host, or you&rsquo;re weighing whether to reinstall an older card headless for exactly this purpose, this is one of the more satisfying uses for spare GPU hardware that would otherwise sit idle. Size your model choice to the VRAM you actually have rather than chasing the largest model that technically fits, a smaller model that responds in a couple of seconds beats a larger one that makes you wait, for most day-to-day use. Start with Ollama and a single mid-size model, add Open WebUI once you&rsquo;ve confirmed the API responds the way you expect, and resist the urge to load five different models &ldquo;just in case,&rdquo; disk and load time both add up faster than it looks like they will.</p>
<h2 id="bottom-line">Bottom line</h2>
<p>Ollama gives you a model runtime that pulls, quantizes, and serves open-weight models with minimal ceremony. Open WebUI gives you the chat interface people actually want to use, with multi-user support and basic document Q&amp;A on top. Together they&rsquo;re a genuinely useful addition to a homelab, as long as you&rsquo;re running them for what local models are actually good at, privacy, cost control on bulk tasks, and offline availability, rather than expecting them to match a frontier cloud model on hard problems. Size the hardware to the model tier you want, isolate the service in its own container, and treat the result as a capable assistant for the right jobs, not a full replacement for the AI you&rsquo;re already paying for.</p>
]]></content:encoded></item></channel></rss>