| | SGLang and Miles Add Day-0 Support for DeepSeek-v4.1 (lmsys.org) |
| 1 point by aray07 6 days ago | past | discuss |
|
| | Pushing the Limits of Serving DeepSeek-V4-Pro (lmsys.org) |
| 1 point by gmays 19 days ago | past |
|
| | Miles v0.1: Production-level Post-training (lmsys.org) |
| 3 points by gmays 21 days ago | past |
|
| | Pushing the Limits of Serving DeepSeek-V4-Pro (lmsys.org) |
| 2 points by vzhou842 28 days ago | past |
|
| | Miles v0.1: Production-level Post-training (lmsys.org) |
| 2 points by nblintao 35 days ago | past |
|
| | The next generation of speculative decoding: DFlash and Spec V2 (lmsys.org) |
| 1 point by ronfriedhaber 53 days ago | past |
|
| | Agent-Assisted SGLang Development: An Initial Exploration (lmsys.org) |
| 1 point by gmays 77 days ago | past |
|
| | The next generation of speculative decoding: DFlash and Spec V2 (lmsys.org) |
| 4 points by gmays 3 months ago | past |
|
| | No Token Left Behind: Demystifying Token-in-Token-Out in Miles (lmsys.org) |
| 2 points by kkm 3 months ago | past |
|
| | DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles (lmsys.org) |
| 80 points by mji 4 months ago | past | 10 comments |
|
| | Pipeline Parallelism in SGLang: Scaling to Million-Token Contexts (lmsys.org) |
| 3 points by roody_wurlitzer 6 months ago | past |
|
| | Pipeline Parallelism in SGLang: Scaling to Million-Token Contexts and Beyond (lmsys.org) |
| 1 point by gmays 8 months ago | past |
|
| | Production-Ready Speculative Decoding Models and Framework (lmsys.org) |
| 1 point by gmays 9 months ago | past |
|
| | Mini-SGLang: Efficient Inference Engine in a Nutshell (lmsys.org) |
| 2 points by matt_d 9 months ago | past |
|
| | Power Up FSDP2 as a Flexible Training Back End for Miles (lmsys.org) |
| 1 point by gmays 9 months ago | past |
|
| | NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference (lmsys.org) |
| 115 points by yvbbrjdr 11 months ago | past | 93 comments |
|
| | Deploying DeepSeek on 96 H100 GPUs (lmsys.org) |
| 285 points by GabrielBianconi on Aug 29, 2025 | past | 80 comments |
|
| | Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP: 2.7x Throughput (lmsys.org) |
| 1 point by gmays on June 17, 2025 | past |
|
| | Match DeepSeek's inference system performance with SGLang (lmsys.org) |
| 1 point by echaozh on May 6, 2025 | past |
|
| | Does style matter? Disentangling style and substance in Chatbot Arena (lmsys.org) |
| 2 points by ZeljkoS on Feb 9, 2025 | past |
|
| | Faster JSON Decoding for LLMs (lmsys.org) |
| 1 point by gaocegege on Dec 18, 2024 | past |
|
| | Does style matter? Disentangling style and substance in Chatbot Arena (lmsys.org) |
| 1 point by scottfr on Aug 29, 2024 | past |
|
| | LLM Lookahead Decoding (lmsys.org) |
| 2 points by mr-ai on Aug 20, 2024 | past |
|
| | From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline – Lmsys Org (lmsys.org) |
| 2 points by swyx on Aug 3, 2024 | past |
|
| | Faster Open-Source Llama3 Serving with SGLang Runtime (vs. TensorRT-LLM, VLLM) (lmsys.org) |
| 4 points by yvbbrjdr on July 25, 2024 | past |
|
| | RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing (lmsys.org) |
| 4 points by adr1an on July 5, 2024 | past |
|
| | RouteLLM: An Open-Source Framework for Cost-Effective LLM Routing (lmsys.org) |
| 4 points by not-chatgpt on July 1, 2024 | past | 1 comment |
|
| | Introducing Hard Prompts Category in Chatbot Arena (lmsys.org) |
| 1 point by JumpCrisscross on June 21, 2024 | past |
|
| | Introducing Hard Prompts Category in Chatbot Arena (lmsys.org) |
| 1 point by CharlesW on May 20, 2024 | past |
|
| | Hard Prompts Category in Chatbot Arena (lmsys.org) |
| 1 point by imjonse on May 17, 2024 | past |
|
|
| More |