A self-hosted browser interface with models running locally on your GPU
Agent-ready models are tuned for agentic workloads — tool calling, reasoning traces, and long multi-turn sessions. Use this tab to pick a recommended model for your hardware platform, then pull it through Open WebUI’s integrated Ollama stack using either instructions tab.
| Hardware platform | Recommended agent-ready model | Ollama library tag |
|---|---|---|
| DGX Spark | Agent-ready Qwen3.6-35B-A3B (Ollama) | qwen3.6:35b-a3b |
Complete the setup through Step 5 in the Open WebUI Remotely tab (through creating the administrator account), or through Step 4 in the Open WebUI on Desktop tab, so Open WebUI is running and you are signed in.
Then pull and select the recommended model in the Open WebUI interface:
qwen3.6:35b-a3b in the search field.In the chat text area, enter a prompt such as Write me a haiku about GPUs and AI, then press Enter and wait for the response.
For agentic use, prefer a context window of at least 32K tokens. If your hardware platform has additional memory headroom, 64K or higher is recommended so multi-turn sessions and tools have room to work. Adjust context settings in Open WebUI or the model controls for your session when available.