---
title: "Chat with LLMs Using Open WebUI and Ollama"
publisher: "nvidia"
type: "playbook"
updated: "2026-09-22T17:56:15.206Z"
description: "A self-hosted browser interface with models running locally on your GPU"
canonical: "https://build.nvidia.com/playbooks/open-webui.md"
---

# Basic idea

Open WebUI is a self-hosted chat application that you can run entirely on your hardware platform. This gives you privacy and security because everything stays on your system. The Open WebUI application operates offline, and your queries go to a model running locally.

Open WebUI is browser based, so you can use it to chat with a model from your laptop, or directly on the hardware platform as a desktop.

This playbook shows you two paths to run Open WebUI:

* **Remotely with NVIDIA Sync:** Run Open WebUI on a remote hardware platform from your laptop.
* **Manually on a desktop:** Run Open WebUI on a local desktop session, or if you want to get under the hood.

Both paths lead to the same outcome: chatting with a model running on your hardware platform through Open WebUI.

# What you'll accomplish

You'll download an Open WebUI container image with Ollama onto your hardware platform, run the container, then use the Open WebUI browser interface to download and run a model. Then you'll chat with it.

The container setup includes integrated Ollama for model management, persistent data storage, and GPU acceleration for model inference.

# What to know before starting

**Required:**

- Basic familiarity with terminal commands
- For the remote path: cutting and pasting terminal commands, and how to use NVIDIA Sync ([documentation](https://docs.nvidia.com/sync/latest/direct-connections.html#nvidia-sync-direct-connections))
- For the desktop path: familiarity with Docker commands

**Optional:**

- Experience browsing models on the [Ollama models page](https://ollama.com/search) (for example, [qwen3.6](https://ollama.com/library/qwen3.6) and [gpt-oss](https://ollama.com/library/gpt-oss))

# Supported hardware platforms

Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.

| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
| :---- | :---- | :---- | :---- | :---- |
| **DGX Spark** | DGX OS (Linux) | 128 GB Unified Memory | Open WebUI + Ollama container (`ghcr.io/open-webui/open-webui:ollama`); desktop on port `8080`, NVIDIA Sync custom app on port `12000` | — |

# Prerequisites

**Hardware requirements**

- Supported hardware platform — see Supported hardware platforms matrix above
- Hardware platform powered on, networked, and reachable for local desktop or NVIDIA Sync access
- Enough disk space for the container image and models (about 7 GB for the container image; about 15 GB for `gpt-oss:20b` or about 25 GB for `qwen3.6:latest`)

**Software requirements**

- Docker available on the hardware platform (`docker ps`)
- Web browser access to Open WebUI (port `8080` locally, or port `12000` via NVIDIA Sync)
- Network access from the hardware platform to download the container image and models
- Remotely with NVIDIA Sync path: NVIDIA Sync installed on your laptop and connected to your hardware platform

# Find model recipes

Browse models you can pull and run with the integrated Ollama stack in the [Ollama library](https://ollama.com/library). Use the tags and sizes that fit your hardware platform’s memory and storage.

| Hardware platform | More recipes |
| ----------------- | ------------ |
| **DGX Spark** | [Ollama library](https://ollama.com/library) |

Use the **Open WebUI Remotely** or **Open WebUI on Desktop** tab for the base workflow. For agentic workloads, see the **Agent-ready Models** tab.

# Time & risk

- **Estimated time:** 15–20 MIN for setup, including the Open WebUI container download and model download (time varies with your internet speed)
- **Risk level:** Low
- Docker permission issues may require user group changes and a session restart
- Large model downloads may take significant time depending on network speed
- **Rollback:** Stop and remove the container, image, and volumes (see Cleanup in each instructions tab). Cleanup is optional and destructive to chat history and downloaded models.
- **Last Updated:** 07/31/2026
- Run Open WebUI with integrated Ollama on your hardware platform, chat from a browser, and pull models locally

## More

- [Open WebUI Remotely](/playbooks/open-webui/sync.md)
- [Open WebUI on Desktop](/playbooks/open-webui/instructions.md)
- [Agent-ready Models](/playbooks/open-webui/agent-ready-models.md)
- [Troubleshooting](/playbooks/open-webui/troubleshooting.md)