SGLang Omni: Serving Omni & Multimodal Models | Nebius Science Paper Club

Join a session with Jiaxin Deng of the SGLang Omni team.

SGLang-Omni is a high-performance, multi-stage serving runtime for omni, speech, and TTS models, extending SGLang beyond single-loop autoregressive decoding to pipelines that mix text, image, audio, and video in and out.

The talk will cover SGLang Omni’s design along with broader SGLang LLM inference work, followed by Q&A and open discussion with Nebius researchers and the community.

Read the paper before →

About Nebius Science Paper Club

Nebius Science Paper Club is a webinar series led by Nebius researchers. Each session brings together paper authors and practitioners to discuss new ideas, research, and discoveries in AI.

The session is open to researchers, engineers, students, and anyone curious about the paper.

Follow how AI and science co-evolve →

Key takeaways

SGLang Omni coordinates multiple inference stages to serve voice and omni models efficiently.

It reuses SGLang’s autoregressive inference capabilities and adds scheduling across components such as encoders, language models and audio decoders.

Each stage can use an execution strategy suited to its workload.

Try Nebius AI Cloud console today

Get immediate access to NVIDIA GPUs, along with CPU resources, storage and additional services through our user-friendly self-service console.