Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Run NemoClaw with a Local LLM

    30 MIN

    Build a local AI assistant in an OpenShell sandbox with vLLM inference and optional Telegram

    • Agentic Workflow
    • DGX Spark
    • DGX Station
    • NemoClaw
    • OpenClaw
    • OpenShell
    • Telegram
    • vLLM
    NemoClaw on GitHub
    OverviewOverviewInstructionsInstructionsMulti-nodeMulti-nodeTroubleshootingTroubleshooting
    SymptomCauseFix
    nemoclaw: command not found after installShell PATH not updatedRun source ~/.bashrc (or source ~/.zshrc for zsh), or open a new terminal window.
    Installer fails with Node.js version errorNode.js version below 22.16Install Node.js 22.16+: curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash - && sudo apt-get install -y nodejs then re-run the installer.
    npm install fails with EACCES permission errornpm global directory not writablemkdir -p ~/.npm-global && npm config set prefix ~/.npm-global && export PATH=~/.npm-global/bin:$PATH then re-run the installer. Add the export line to ~/.bashrc to make it permanent.
    Docker permission deniedUser not in docker groupsudo usermod -aG docker $USER, then log out and back in.
    Gateway fails with cgroup / "Failed to start ContainerManager" errorsOlder OpenShell or Docker still using a private cgroup namespace for the gateway so kubelet cannot see cgroup v2 controllersFirst upgrade OpenShell (re-run the Phase 1 nemoclaw.sh install so you get a build that sets host cgroupns on the gateway container). If it still fails, force Docker's default to host mode by running the daemon.json cgroup fix below, then run sudo systemctl restart docker.
    Gateway fails with "port 8080 is held by container..."Another OpenShell gateway or container is using port 8080Run nemoclaw onboard (or nemoclaw onboard --resume) again. NemoClaw probes the existing managed gateway, reuses it if healthy, and recreates stale gateway state when it can do so safely. See the NemoClaw commands documentation.
    Sandbox creation failsStale gateway state or DNS not propagatedRun nemoclaw onboard (or nemoclaw onboard --resume) again. NemoClaw probes the existing managed gateway, reuses it if healthy, and recreates stale gateway state when it can do so safely. See the NemoClaw commands documentation.
    CoreDNS crash loopKnown issue on some hardware platform configurationsRe-run the NemoClaw installer (curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash) which includes the CoreDNS fix. If the issue persists, see NemoClaw troubleshooting.
    "No GPU detected" during onboardSome hardware platforms report unified memory differentlyExpected on some hardware platforms. The wizard still works and uses vLLM for inference.
    Inference timeout or hangsvLLM not running, the model is still loading, or the server is not reachableCheck the vLLM server: curl http://127.0.0.1:8000/v1/models should list the model you selected during onboarding or Express Install. If the request hangs, wait for Application startup complete, then check nemoclaw my-assistant status for the Inference health line.
    Agent gives no response or is very slowFirst response can be slow, especially with larger modelsResponse time depends on model size (30B: a few seconds; larger models may take longer). Verify inference route: nemoclaw my-assistant status.
    Port 18789 already in useAnother process is bound to the portlsof -i :18789 then kill <PID>. If needed, kill -9 <PID> to force-terminate.
    Web UI port forward dies or dashboard unreachablePort forward not activeopenshell forward stop 18789 my-assistant then openshell forward start 18789 my-assistant --background.
    Web UI shows origin not allowedAccessing via localhost instead of 127.0.0.1Use http://127.0.0.1:18789/#token=... in the browser. The gateway origin check requires 127.0.0.1 exactly.

    daemon.json cgroup fix

    Use this script as the fallback for the cgroup / "Failed to start ContainerManager" row above. It validates any existing /etc/docker/daemon.json, writes a .bak backup, sets default-cgroupns-mode to host, and atomically replaces the file. It exits non-zero with an error on stderr if anything fails, leaving the original daemon.json untouched.

    sudo python3 - <<'PY'
    import json, os, shutil, sys, tempfile
    
    path = '/etc/docker/daemon.json'
    try:
        if os.path.exists(path):
            with open(path) as f:
                data = json.load(f)
            if not isinstance(data, dict):
                raise ValueError(f'{path} is not a JSON object')
        else:
            data = {}
    except (json.JSONDecodeError, ValueError, OSError) as e:
        print(f'error: failed to read {path}: {e}', file=sys.stderr)
        sys.exit(1)
    
    if os.path.exists(path):
        try:
            shutil.copy2(path, path + '.bak')
        except OSError as e:
            print(f'error: failed to back up {path}: {e}', file=sys.stderr)
            sys.exit(1)
    
    data['default-cgroupns-mode'] = 'host'
    
    target_dir = os.path.dirname(path) or '/'
    fd, tmp = tempfile.mkstemp(prefix='daemon.json.', dir=target_dir)
    try:
        with os.fdopen(fd, 'w') as f:
            json.dump(data, f, indent=2)
            f.write('\n')
        os.chmod(tmp, 0o644)
        os.replace(tmp, path)
    except OSError as e:
        if os.path.exists(tmp):
            try:
                os.unlink(tmp)
            except OSError:
                pass
        print(f'error: failed to write {path}: {e}', file=sys.stderr)
        sys.exit(1)
    PY
    

    NOTE

    Some hardware platforms use Unified Memory Architecture (UMA), which enables dynamic memory sharing between the GPU and CPU. With many applications still updating to take advantage of UMA, you may encounter memory issues even when within the memory capacity of the hardware platform. If that happens, manually flush the buffer cache with:

    sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
    

    Optional — Messaging (Telegram)

    Skip this section if you are not using Telegram or another messaging channel.

    SymptomCauseFix
    Telegram bridge does not startTelegram channel is not configured, the sandbox gateway is unhealthy, or Telegram startup/config failedRun nemoclaw <name> status and nemoclaw <name> logs to confirm the failure. If the sandbox gateway is unhealthy, run nemoclaw <name> recover. If Telegram is not configured, add it from the optional Telegram section at the bottom of Instructions, or rerun nemoclaw onboard and enable Telegram. See the NemoClaw Troubleshooting guide.
    Telegram stops responding after sandbox rebuildDuplicate bot-token consumer, missing DM allowlist, BotFather group privacy mode, inference failure, policy denial, or rebuilt channel config issueRun nemoclaw <name> status and nemoclaw <name> logs. Look for Telegram 409 Conflict, allowlist warnings, privacy-mode issues, inference errors, or policy denials. If configuration needs to change, rerun nemoclaw onboard. See the NemoClaw Troubleshooting guide.
    Telegram bot receives messages but does not replyInbound Telegram delivery works, but the agent turn, inference call, policy check, allowlist/mention gate, or outbound reply failedRun nemoclaw <name> status and nemoclaw <name> logs. Check for inbound Telegram update, outbound send, inference, and policy-denial messages. Fix the logged cause; run nemoclaw <name> recover only if the sandbox gateway is unhealthy. For Telegram configuration changes, rerun nemoclaw onboard. See the NemoClaw Troubleshooting guide.

    Resources

    • NemoClaw
    • NemoClaw Documentation
    • OpenClaw Documentation
    • vLLM Recipes
    • DGX Spark Documentation
    • DGX Spark Forum
    • DGX Station Support
    • NVIDIA Developer Forums
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation