Search results
High-throughput serving with RadixAttention, structured output, and an OpenAI-compatible API High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API Cut memory ~3.5× vs FP16 while keeping accuracy close to FP8, then validate with an OpenAI-compatible endpoint
Items per page
of 1 pages