A self-hosted browser interface with models running locally on your GPU
Open WebUI is a self-hosted chat application that you can run entirely on your hardware platform. This gives you privacy and security because everything stays on your system. The Open WebUI application operates offline, and your queries go to a model running locally.
Open WebUI is browser based, so you can use it to chat with a model from your laptop, or directly on the hardware platform as a desktop.
This playbook shows you two paths to run Open WebUI:
Both paths lead to the same outcome: chatting with a model running on your hardware platform through Open WebUI.
You'll download an Open WebUI container image with Ollama onto your hardware platform, run the container, then use the Open WebUI browser interface to download and run a model. Then you'll chat with it.
The container setup includes integrated Ollama for model management, persistent data storage, and GPU acceleration for model inference.
Required:
Optional:
Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.
| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
|---|---|---|---|---|
| DGX Spark | DGX OS (Linux) | 128 GB Unified Memory | Open WebUI + Ollama container (ghcr.io/open-webui/open-webui:ollama); desktop on port 8080, NVIDIA Sync custom app on port 12000 | — |
Hardware requirements
gpt-oss:20b or about 25 GB for qwen3.6:latest)Software requirements
docker ps)8080 locally, or port 12000 via NVIDIA Sync)Browse models you can pull and run with the integrated Ollama stack in the Ollama library. Use the tags and sizes that fit your hardware platform’s memory and storage.
| Hardware platform | More recipes |
|---|---|
| DGX Spark | Ollama library |
Use the Open WebUI Remotely or Open WebUI on Desktop tab for the base workflow. For agentic workloads, see the Agent-ready Models tab.