Search results
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text appsDownloadableFree Endpointdiffusiongemma-26b-a4b-it
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and mathDownloadableFree Endpointgpt-oss-20b
Items per page
of 1 pages