Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

45 results for

Filters

  • Free Endpoint
    18
  • Partner Endpoint
    16
  • Download Available
    15
  • Image-to-Text
    3
  • Code Generation
    1
  • OpenRouter
    15
  • Deepinfra
    14
  • GMI Cloud
    10
  • Together AI
    7
  • Bitdeer
    5
  • NVIDIA
    29
  • Mistral AI
    3
  • Qwen
    3
  • DeepSeek AI
    2
  • OpenAI
    2
  • Developer
    21
  • AI Engineer
    17
  • Ml Engineer
    16
  • Hpc Developer
    9
  • Application Developer
    7
  • AI And Machine Learning
    18
  • Physical AI
    4
  • Infrastructure
    1
  • B200
    1
  • H100 80GB HBM3
    1
  • H200
    1
  • NeMo Megatron Bridge
    9
  • Jetson
    5
  • TAO Toolkit
    5
  • Video Search and Summarization (VSS)
    2
  • DALI
    1
  • Systematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper. Covers the Three Walls framework, parallel folding, recompute strategy, dispatcher choice, and CUDA-graph bring-up.
    Skill
    Developer
    2K
    1mo

    MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
    Skill
    Developer
    2K
    1mo

    Representative MoE training playbooks by hardware platform and model family. Summarizes rounded throughput bands, parallelism patterns, and common tuning stacks.
    Skill
    Developer
    2K
    1mo

    Long-context MoE training guidance for Megatron Bridge. Covers CP sizing, selective recompute, dispatcher choices, and practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.
    Skill
    Developer
    2K
    1mo
    Items per page
    of 2 pages

    Practical guidance for training MoE VLMs in Megatron Bridge. Compares FSDP and 3D-parallel approaches, using rounded lessons from Qwen3-VL, Qwen3-Next, and other multimodal experiments.
    Skill
    Developer
    2K
    1mo

    Choose the right MoE token dispatcher (`alltoall`, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage. Summarizes patterns from DSV3, Qwen3, Qwen3-Next, and VLM bring-up work.
    Skill
    Developer
    2K
    1mo
    DeepSeek AI
    DownloadableFree Endpoint

    deepseek-v4-pro

    DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
    Model
    Moe
    8M
    2mo
    NVIDIA
    DownloadableFree Endpoint

    nemotron-3-nano-30b-a3b

    Open, efficient MoE model with 1M context, excelling in coding, reasoning, instruction following, tool calling, and more
    Model
    MoE
    12M
    7mo
    NVIDIA
    DownloadableFree Endpoint

    nemotron-3-super-120b-a12b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    Model
    MoE
    60M
    4mo
    Qwen
    DownloadableFree Endpoint

    qwen3.5-397b-a17b

    Next-gen Qwen 3.5 VLM (400B MoE) brings advanced vision, chat, RAG, and agentic capabilities.
    Model
    MoE
    16M
    5mo
    DeepSeek AI
    DownloadableFree Endpoint

    deepseek-v4-flash

    DeepSeek V4 Flash is a 284B MoE model with 1M-token context optimized for fast coding and agents.
    Model
    MoE
    15M
    2mo
    NVIDIA
    DownloadableFree Endpoint

    nemotron-3-ultra-550b-a55b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    Model
    Agent
    52M
    1mo

    DALI imperative dynamic mode (`nvidia.dali.experimental.dynamic`, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks.
    Skill
    Developer
    2K
    1mo

    Plan and apply safe Jetson headless-mode changes to reclaim GUI and daemon memory.
    Skill
    Developer
    891
    27d
    OpenAI
    DownloadableFree Endpoint

    gpt-oss-120b

    Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
    Model
    reasoning
    45M
    11mo
    OpenAI
    DownloadableFree Endpoint

    gpt-oss-20b

    Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
    Model
    reasoning
    18M
    11mo
    Moonshotai
    DownloadableFree Endpoint

    kimi-k2.6

    1T multimodal MoE for long-horizon coding, agentic tool use, and image/video understanding.
    Model
    Multimodal
    16M
    2mo
    Poolside
    Free Endpoint

    laguna-xs-2.1

    Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
    Model
    Agentic AI
    3d
    Meta
    Deprecation in 9dFree Endpoint

    llama-4-maverick-17b-128e-instruct

    A general purpose multimodal, multilingual 128 MoE model with 17B parameters.
    Model
    language generation
    16M
    1y
    Mistral AI
    DownloadableFree Endpoint

    mistral-small-4-119b-2603

    Hybrid MoE model unifying instruct, reasoning, and coding with multimodal input and 256k context
    Model
    code generation
    11M
    4mo
    Mistral AI
    DownloadableFree Endpoint

    mixtral-8x7b-instruct-v0.1

    An MOE LLM that follows instructions, completes requests, and generates creative text.
    Model
    Advanced Reasoning
    1M
    1y
    Qwen
    DownloadableFree Endpoint

    qwen3-next-80b-a3b-instruct

    Qwen3-Next Instruct blends hybrid attention, sparse MoE, and stability boosts for ultra-long context AI.
    Model
    text-generation
    25M
    10mo
    Qwen
    DownloadableFree Endpoint

    qwen3.5-122b-a10b

    122B MoE LLM (10B active) for coding, reasoning, multimodal chat. Agent-ready.
    Model
    tool calling
    15M
    4mo
    Stepfun-ai
    Deprecation in 9dFree Endpoint

    step-3.5-flash

    200B open-source reasoning engine with sparse MoE powering frontier agentic AI.
    Model
    Agentic
    10M
    5mo