Support multi-thread model weight loading #7277

xianzhiT · 2025-06-17T13:39:18Z

Motivation

Support multi-thread model weight loading when using the --load-format "auto", which can speedup weight load process for users with distributed storage backend or NVME ssd disk.

Tests have found that storing DeepSeek-R1-0528 model weights on a single NVME drive for cold start of services can speed up model loading time by 6.26 times

Modifications

base on PR #4822

Multi-threaded Weight Loading: Implemented multi-threaded loading for safetensors and PyTorch (.pt/.bin) model weights, significantly accelerating the model loading process, especially for large sharded models.
Introduced a new --model-loader-extra-config command-line argument: Allowing users to enable multi-threaded loading (enable_multithread_load) and specify the number of threads (num_threads) via a JSON string.

Cmd Example:

echo "1" > /proc/sys/vm/drop_caches

python3 -m sglang.launch_server \
 --model-path /data/DeepSeek-R1-0528 \
 --tp 8 \
 --trust-remote-code \
 --mem-fraction-static 0.9 \
 --host 0.0.0.0 \
 --port 8001 \
 --enable-metric \
 --log-requests \
 --log-requests-level 0 \
 --allow-auto-truncate  \
 --quantization fp8 \
 --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 8}'

Test DeepSeek-R1-0528 safetensor loading in NVME ssd PCIE Gen4 (Speed 16GT/s, Width x4)

	Time consumption before optimization	Time consumption after optimization	Load Weight Speed up
weight in single NVME ssd	702s	112s	6.26x
weight in raid0 with four NVME ssd	656s	62s	10.58x

Checklist

Format your code according to the Code Formatting with Pre-Commit.
Add unit tests as outlined in the Running Unit Tests.
Update documentation / docstrings / example tutorials as needed, according to Writing Documentation.
Provide throughput / latency benchmark results and accuracy evaluation results as needed, according to Benchmark and Profiling and Accuracy Results.
For reviewers: If you haven't made any contributions to this PR and are only assisting with merging the main branch, please remove yourself as a co-author when merging the PR.
Please feel free to join our Slack channel at https://slack.sglang.ai to discuss your PR.- [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ] - [ ]

gemini-code-assist

Summary of Changes

Hello @xianzhiT, I'm Gemini Code Assist¹! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a significant performance improvement for model weight loading by enabling multi-threaded operations for safetensors and PyTorch model formats. It provides a new command-line interface for users to control this feature, leading to faster cold starts for services, particularly beneficial for large models stored on fast I/O systems like NVMe SSDs.

Highlights

Multi-threaded Model Weight Loading: Implemented multi-threaded loading for safetensors and PyTorch (.pt/.bin) model weights. This significantly accelerates the model loading process, especially for large sharded models, with reported speedups of over 6x on NVMe SSDs and over 10x on RAID0 NVMe setups.
Configurable Multi-threading via CLI: Introduced a new command-line argument, --model-loader-extra-config, which allows users to enable multi-threaded loading (enable_multithread_load: true) and specify the number of threads (num_threads) using a JSON string. A default of 8 threads is used if num_threads is not specified when multi-threading is enabled.
Integration into Model Loading Pipeline: The new multi-threading logic is integrated into the DefaultModelLoader to conditionally use the new multi-threaded iterators (multi_thread_safetensors_weights_iterator and multi_thread_pt_weights_iterator) based on the provided model_loader_extra_config.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point in your pull request via creating an issue comment (i.e. comment on the pull request page) using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands.

Feature	Command	Description
Code Review	`/gemini review`	Performs a code review for the current pull request in its current state.
Pull Request Summary	`/gemini summary`	Provides a summary of the current pull request in its current state.
Comment	@gemini-code-assist	Responds in comments when explicitly tagged, both in issue comments and review comments.
Help	`/gemini help`	Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist is currently in preview and may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments to provide feedback.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

gemini-code-assist

Code Review

The pull request introduces multi-threaded weight loading for safetensors and PyTorch models, controlled by a new command-line argument. The changes include adding necessary imports, new weight iterator functions, and conditional logic to enable multi-threading. I have suggested some minor improvements related to error handling and clarity.

python/sglang/srt/model_loader/loader.py

python/sglang/srt/model_loader/weight_utils.py

garrett4wade

Disclaimer: I'm not a core contributor - my review was requested by @zhaochenyang20

Thanks for the contribution! The implementation looks good overall, though there are a few minor style points that could be addressed.

One question: does the author report the total launch time, or specifically the model loading time (disk to memory)? I'm curious about the multi-threading performance gains for end-to-end launch time from disk to GPU, even though PCIe bandwidth should theoretically be much higher.

python/sglang/srt/model_loader/loader.py

python/sglang/srt/model_loader/weight_utils.py

python/sglang/srt/model_loader/loader.py

BraveY · 2025-06-18T03:26:24Z

@xianzhiT , @garrett4wade — Great work! I submitted a similar PR a few days ago, aiming to mmap the weight file into pagecache to mitigate cold start issues in inference. Would either of you be willing to review my PR and share your feedback?

xianzhiT · 2025-06-18T03:38:56Z

Disclaimer: I'm not a core contributor - my review was requested by @zhaochenyang20

Thanks for the contribution! The implementation looks good overall, though there are a few minor style points that could be addressed.

One question: does the author report the total launch time, or specifically the model loading time (disk to memory)? I'm curious about the multi-threading performance gains for end-to-end launch time from disk to GPU, even though PCIe bandwidth should theoretically be much higher.

Thanks for the review, appreciate your feedback. I'll address the minor style points you mentioned and update the code accordingly.

The report is the time consume for model weight loading from disk to memory. I calculated the time by counting the time interval between "Load weight begin" and "Load weight end" in the log.

xianzhiT · 2025-06-18T04:00:19Z

Disclaimer: I'm not a core contributor - my review was requested by @zhaochenyang20

Thanks for the contribution! The implementation looks good overall, though there are a few minor style points that could be addressed.

One question: does the author report the total launch time, or specifically the model loading time (disk to memory)? I'm curious about the multi-threading performance gains for end-to-end launch time from disk to GPU, even though PCIe bandwidth should theoretically be much higher.

Disclaimer: I'm not a core contributor - my review was requested by @zhaochenyang20
Thanks for the contribution! The implementation looks good overall, though there are a few minor style points that could be addressed.
One question: does the author report the total launch time, or specifically the model loading time (disk to memory)? I'm curious about the multi-threading performance gains for end-to-end launch time from disk to GPU, even though PCIe bandwidth should theoretically be much higher.

Thanks for the review, appreciate your feedback. I'll address the minor style points you mentioned and update the code accordingly.

The report is the time consume for model weight loading from disk to memory. I calculated the time by counting the time interval between "Load weight begin" and "Load weight end" in the log.

The nvme ssd disk I use to store model weights is PCIE GEN4 width 4, the theoretical upper limit for read bandwidth is 8GB/s. In the optimized test, I observed through iotop that the read bandwidth can reach around 6.3GB, which is close to the theoretical bandwidth limit.

I have also test the e2e launch time speed, we can get server cold start in about 160s.

	E2E Launch Time consumption before optimization (DeepGEMM JIT cached)	E2E Launch Time consumption after optimization ((DeepGEMM JIT cached))	Speedup
weight in single NVME ssd	756s	160s	4.725x

python/sglang/srt/model_loader/weight_utils.py

BraveY · 2025-06-18T10:07:56Z

Update the test result between this PR and my PR:
Weight in raid0 with three NVME ssd. The test model is ds-r1 fp8.

Multi-thread with the config: --tp-size 8 --quantization fp8 --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 8}'

$grep -C 5 "Load weight" sglang-test-2025-06-18_15:18.log
[2025-06-18 15:19:16 TP0] MLA optimization is turned on. Use flashinfer backend.
[2025-06-18 15:19:16 TP0] Chunked prefix cache is turned on.
[2025-06-18 15:19:16 TP0] Init torch distributed begin.
[2025-06-18 15:19:18 TP0] sglang is using nccl==2.26.2
[2025-06-18 15:19:22 TP0] Init torch distributed ends. mem usage=1.12 GB
[2025-06-18 15:19:24 TP0] Load weight begin. avail mem=93.52 GB
[2025-06-18 15:19:24 TP0] Detected fp8 checkpoint.
Multi-thread loading shards:   0% Completed | 0/163 [00:00<?, ?it/s]
Multi-thread loading shards:   1% Completed | 1/163 [01:22<3:42:45, 82.50s/it]
Multi-thread loading shards: 100% Completed | 163/163 [01:22<00:00,  1.98it/s]

[2025-06-18 15:20:48 TP0] Load weight end. type=DeepseekV3ForCausalLM, dtype=torch.bfloat16, avail mem=13.97 GB, mem usage=79.54 GB.
[2025-06-18 15:20:51 TP4] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP7] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP3] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP2] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP0] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB

The loading time is 1m24s: 15:19:24->15:20:48

My PR with barrier commit id 4e31c43. Config: --tp-size 8 --quantization fp8 --load-format prefetch_auto

$grep -C 5 "Load weight" sglang-test-2025-06-18_15:27.log
[2025-06-18 15:28:26 TP0] MLA optimization is turned on. Use flashinfer backend.
[2025-06-18 15:28:26 TP0] Chunked prefix cache is turned on.
[2025-06-18 15:28:26 TP0] Init torch distributed begin.
[2025-06-18 15:28:28 TP0] sglang is using nccl==2.26.2
[2025-06-18 15:28:32 TP0] Init torch distributed ends. mem usage=1.12 GB
[2025-06-18 15:28:34 TP0] Load weight begin. avail mem=93.52 GB
[2025-06-18 15:28:34 TP0] Detected fp8 checkpoint.
[2025-06-18 15:28:35 TP6] Mmaping 20 files concurrently
[2025-06-18 15:28:35 TP5] Mmaping 20 files concurrently
[2025-06-18 15:28:35 TP3] Mmaping 20 files concurrently
[2025-06-18 15:28:35 TP4] Mmaping 20 files concurrently
--
Loading safetensors checkpoint shards:  99% Completed | 161/163 [00:51<00:00,  3.16it/s]
Loading safetensors checkpoint shards:  99% Completed | 162/163 [00:52<00:00,  2.84it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:52<00:00,  2.76it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:52<00:00,  3.11it/s]

[2025-06-18 15:30:09 TP0] Load weight end. type=DeepseekV3ForCausalLM, dtype=torch.bfloat16, avail mem=13.07 GB, mem usage=80.44 GB.
[2025-06-18 15:30:09 TP5] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB
[2025-06-18 15:30:09 TP2] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB
[2025-06-18 15:30:09 TP0] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB
[2025-06-18 15:30:09 TP0] Memory pool end. avail mem=9.38 GB
[2025-06-18 15:30:09 TP3] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB

The loading time is 1m35s: 15:28:34->15:30:09

My PR without barrier commit id 1ae9d55.

$grep -C 5 "Load weight" sglang-test-2025-06-18_16:24.log
[2025-06-18 16:24:55 TP0] MLA optimization is turned on. Use flashinfer backend.
[2025-06-18 16:24:55 TP0] Chunked prefix cache is turned on.
[2025-06-18 16:24:55 TP0] Init torch distributed begin.
[2025-06-18 16:24:58 TP0] sglang is using nccl==2.26.2
[2025-06-18 16:25:01 TP0] Init torch distributed ends. mem usage=1.12 GB
[2025-06-18 16:25:03 TP0] Load weight begin. avail mem=93.52 GB
[2025-06-18 16:25:03 TP0] Detected fp8 checkpoint.
[2025-06-18 16:25:03 TP5] Mmaping 20 files concurrently
[2025-06-18 16:25:04 TP1] Mmaping 21 files concurrently
[2025-06-18 16:25:04 TP0] Mmaping 21 files concurrently
[2025-06-18 16:25:04 TP3] Mmaping 20 files concurrently
--
Loading safetensors checkpoint shards:  99% Completed | 161/163 [00:43<00:00,  3.78it/s]
Loading safetensors checkpoint shards:  99% Completed | 162/163 [00:43<00:00,  3.51it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:43<00:00,  3.48it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:43<00:00,  3.72it/s]

[2025-06-18 16:26:28 TP0] Load weight end. type=DeepseekV3ForCausalLM, dtype=torch.bfloat16, avail mem=13.84 GB, mem usage=79.68 GB.
[2025-06-18 16:26:38 TP5] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP1] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP6] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP0] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP7] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB

The loading time is 1m25s:16:25:03->16:26:28.

We obtained the same performance boost. Excellent work!

xianzhiT · 2025-06-18T12:24:41Z

Update the test result between this PR and my PR: Weight in raid0 with three NVME ssd. The test model is ds-r1 fp8.

Multi-thread with the config: --tp-size 8 --quantization fp8 --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 8}'

$grep -C 5 "Load weight" sglang-test-2025-06-18_15:18.log
[2025-06-18 15:19:16 TP0] MLA optimization is turned on. Use flashinfer backend.
[2025-06-18 15:19:16 TP0] Chunked prefix cache is turned on.
[2025-06-18 15:19:16 TP0] Init torch distributed begin.
[2025-06-18 15:19:18 TP0] sglang is using nccl==2.26.2
[2025-06-18 15:19:22 TP0] Init torch distributed ends. mem usage=1.12 GB
[2025-06-18 15:19:24 TP0] Load weight begin. avail mem=93.52 GB
[2025-06-18 15:19:24 TP0] Detected fp8 checkpoint.
Multi-thread loading shards:   0% Completed | 0/163 [00:00<?, ?it/s]
Multi-thread loading shards:   1% Completed | 1/163 [01:22<3:42:45, 82.50s/it]
Multi-thread loading shards: 100% Completed | 163/163 [01:22<00:00,  1.98it/s]

[2025-06-18 15:20:48 TP0] Load weight end. type=DeepseekV3ForCausalLM, dtype=torch.bfloat16, avail mem=13.97 GB, mem usage=79.54 GB.
[2025-06-18 15:20:51 TP4] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP7] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP3] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP2] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB
[2025-06-18 15:20:51 TP0] KV Cache is allocated. #tokens: 69931, KV size: 4.58 GB

The loading time is 1m24s: 15:19:24->15:20:48

My PR with barrier commit id 4e31c43. Config: --tp-size 8 --quantization fp8 --load-format prefetch_auto

$grep -C 5 "Load weight" sglang-test-2025-06-18_15:27.log
[2025-06-18 15:28:26 TP0] MLA optimization is turned on. Use flashinfer backend.
[2025-06-18 15:28:26 TP0] Chunked prefix cache is turned on.
[2025-06-18 15:28:26 TP0] Init torch distributed begin.
[2025-06-18 15:28:28 TP0] sglang is using nccl==2.26.2
[2025-06-18 15:28:32 TP0] Init torch distributed ends. mem usage=1.12 GB
[2025-06-18 15:28:34 TP0] Load weight begin. avail mem=93.52 GB
[2025-06-18 15:28:34 TP0] Detected fp8 checkpoint.
[2025-06-18 15:28:35 TP6] Mmaping 20 files concurrently
[2025-06-18 15:28:35 TP5] Mmaping 20 files concurrently
[2025-06-18 15:28:35 TP3] Mmaping 20 files concurrently
[2025-06-18 15:28:35 TP4] Mmaping 20 files concurrently
--
Loading safetensors checkpoint shards:  99% Completed | 161/163 [00:51<00:00,  3.16it/s]
Loading safetensors checkpoint shards:  99% Completed | 162/163 [00:52<00:00,  2.84it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:52<00:00,  2.76it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:52<00:00,  3.11it/s]

[2025-06-18 15:30:09 TP0] Load weight end. type=DeepseekV3ForCausalLM, dtype=torch.bfloat16, avail mem=13.07 GB, mem usage=80.44 GB.
[2025-06-18 15:30:09 TP5] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB
[2025-06-18 15:30:09 TP2] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB
[2025-06-18 15:30:09 TP0] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB
[2025-06-18 15:30:09 TP0] Memory pool end. avail mem=9.38 GB
[2025-06-18 15:30:09 TP3] KV Cache is allocated. #tokens: 55457, KV size: 3.63 GB

The loading time is 1m35s: 15:28:34->15:30:09

My PR without barrier commit id 1ae9d55.

$grep -C 5 "Load weight" sglang-test-2025-06-18_16:24.log
[2025-06-18 16:24:55 TP0] MLA optimization is turned on. Use flashinfer backend.
[2025-06-18 16:24:55 TP0] Chunked prefix cache is turned on.
[2025-06-18 16:24:55 TP0] Init torch distributed begin.
[2025-06-18 16:24:58 TP0] sglang is using nccl==2.26.2
[2025-06-18 16:25:01 TP0] Init torch distributed ends. mem usage=1.12 GB
[2025-06-18 16:25:03 TP0] Load weight begin. avail mem=93.52 GB
[2025-06-18 16:25:03 TP0] Detected fp8 checkpoint.
[2025-06-18 16:25:03 TP5] Mmaping 20 files concurrently
[2025-06-18 16:25:04 TP1] Mmaping 21 files concurrently
[2025-06-18 16:25:04 TP0] Mmaping 21 files concurrently
[2025-06-18 16:25:04 TP3] Mmaping 20 files concurrently
--
Loading safetensors checkpoint shards:  99% Completed | 161/163 [00:43<00:00,  3.78it/s]
Loading safetensors checkpoint shards:  99% Completed | 162/163 [00:43<00:00,  3.51it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:43<00:00,  3.48it/s]
Loading safetensors checkpoint shards: 100% Completed | 163/163 [00:43<00:00,  3.72it/s]

[2025-06-18 16:26:28 TP0] Load weight end. type=DeepseekV3ForCausalLM, dtype=torch.bfloat16, avail mem=13.84 GB, mem usage=79.68 GB.
[2025-06-18 16:26:38 TP5] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP1] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP6] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP0] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB
[2025-06-18 16:26:38 TP7] KV Cache is allocated. #tokens: 67872, KV size: 4.44 GB

The loading time is 1m25s:16:25:03->16:26:28.

We obtained the same performance boost. Excellent work!

Great. Looks we both achieve great performance.

xianzhiT · 2025-06-18T12:30:08Z

@garrett4wade The addressed code problem has been fixed. Also test in pytorch.bin models, works fine.
Test longchat-13b-16k(https://huggingface.co/lmsys/longchat-13b-16k) pytorch.bin shard loading in NVME ssd PCIE Gen4 (Speed 16GT/s, Width x4)

	Time consumption before optimization	Time consumption after optimization	Load Weight Speed up
weight in single NVME ssd	11s	3s	3.6x

guoyuhong · 2025-06-24T17:14:17Z

LGTM. I verified some of the combinations. It worked fine.
CC @zhyncs

zhaochenyang20 · 2025-06-24T18:28:20Z

I think this is important for RL, especially for update_weights_from_disk.

Python's `argparse does` have automatic prefix matching (also called abbreviation matching), but it only works when there's no ambiguity. After #7277, a new arg `--model-loader-extra-config` was introduced and it caused a lot CI failure: ``` launch_server.py: error: ambiguous option: --model could match --model-path, --model-loader-extra-config ``` This PR adds an alias `--model` for `--model-path`.

@mickqian

* Use seq_len_fill_value in the cuda graph runners (sgl-project#7233) * support custom weight loader for model runner (sgl-project#7122) Co-authored-by: kavioyu <kavioyu@tencent.com> * Fix AMD speculative decoding (sgl-project#7252) * [Refactor] OAI Server components (sgl-project#7167) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * OAI Server Skeleton & Core Utility Endpoints (sgl-project#7179) * [amd] Opt dsv3 moe (sgl-project#7160) Co-authored-by: wunhuang <wunhuang@amd.com> * update ci node for xeon (sgl-project#7265) * feat: mtp support dp-attention (sgl-project#6081) Co-authored-by: austindeng <austindeng@tencent.com> Co-authored-by: tianqilin.99 <tianqilin.99@bytedance.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: ch-wan <cwan39@gatech.edu> * support qwen2 running on ascend npu device (sgl-project#7022) Co-authored-by: 刁莹煜 <diaoyingyu1@hisilicon.com> * Fix Deepseek R1 0528 FP4 tensor name mismatch issue during weights loading. (sgl-project#7164) * bugfix(tool call ebnf): Fix EBNF generation for optional function parameters (sgl-project#7283) * Fix AWQ Dequant and Weight Loading of deepseek v2 (sgl-project#6842) * fix: resolve b200 dsv3 mtp issue (sgl-project#7286) * ci: Fix test_ebnf_generate_all_optional_function_params (sgl-project#7288) * fix: only enable flash_attn test on sm80 sm90 (sgl-project#7289) * [PD] Support get local ip from NIC for PD disaggregation (sgl-project#7237) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * [PD] Add custom memory pool option to support Mooncake PD with NVLink (sgl-project#7264) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * Upstreaming hicache bug fixes (sgl-project#7267) * Update python API of activation, topk, norm and rope and remove vllm dependency (sgl-project#6614) Co-authored-by: Wu, Chunyuan <chunyuan.wu@intel.com> Co-authored-by: jianan-gu <jianan.gu@intel.com> Co-authored-by: sdp <sdp@gnr799219.jf.intel.com> * Fix hicache benchmark script bug - some sampled input_request is [] (sgl-project#7300) * chore: change logs from`INFO` to `DEBUG` for dp and add force quit for tokenizer manager (sgl-project#7251) * update invalid link in doc (sgl-project#7297) * Fix mini_lb for PD with long output: limit chunk size of decode response (sgl-project#7301) Signed-off-by: ch-tiger1 <xyz@ch-tech.ip-ddns.com> Co-authored-by: ch-tiger1 <xyz@ch-tech.ip-ddns.com> * Fix profiler error when there are idle passes (sgl-project#7003) * [pd] optimize dockerfile for pd disaggregation (sgl-project#7319) Co-authored-by: zhyncs <me@zhyncs.com> * Merge PDLB (Prefill-Decode Load Balancer) into SGLang Router (sgl-project#7096) * Add more refactored openai test & in CI (sgl-project#7284) * fix: resolve blackwell deepep image issue (sgl-project#7331) * add seed in CPU UTs to avoid flaky failure (sgl-project#7333) * Multi-Stage Awake: Support Resume and Pause KV Cache and Weights separately (sgl-project#7099) * Reintroduce tiny fix sampler error when prob is not contiguous (sgl-project#7354) * [Refactor] Clean up radix cache related API (sgl-project#7303) Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * Put `_normalize_rid` before other normalization in `io_struct` (sgl-project#7363) * [PD] Transfer hidden states for mtp when disaggregation (sgl-project#7242) * [Bugfix][PD] Set conclude state before clear when failure happens (sgl-project#7362) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * docs: update installation (sgl-project#7366) * [Docker] optimize dockerfile remove deepep and blackwell merge it to… (sgl-project#7343) Co-authored-by: Yineng Zhang <me@zhyncs.com> * Clean unused import for mimo mtp model (sgl-project#7370) * [Bugfix]Fix hang bug using dp attention with HiRadixCache (sgl-project#7159) Signed-off-by: huanglong <huanglong@linux.alibaba.com> * [Doc] add embedding rerank doc (sgl-project#7364) * Fix judgment condition for enabling Deepseek V3/R1 shared expert fusion optimization (sgl-project#7371) * Feat/refactor embedding server (sgl-project#7322) * Purge VerlEngine (sgl-project#7326) Signed-off-by: Ata Fatahi <immrata@gmail.com> * support return logprobs for pipeline (sgl-project#7356) Co-authored-by: Zhang Kaihong <zhangkaihong.zkh@alibaba-inc.com> * [PD] Optimize custom mem pool usage and bump mooncake version (sgl-project#7393) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * Support THUDM/GLM-4-0414 (GLM-Z1) Glm4ForCausalLM architecture. (sgl-project#5485) * Refine OpenAI serving entrypoint to remove batch requests (sgl-project#7372) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> Co-authored-by: Chang Su <csu272@usc.edu> * [Feature] Comprehensive Hybrid Parallelism Support (sgl-project#6389) * [DeepSeekNextN] fix: residual of head norm can be None (sgl-project#7398) * [OAI refactor] Add rerank and score serving (sgl-project#7399) Co-authored-by: Chang Su <chang.s.su@oracle.com> * [OAI Server Refactor] [ChatCompletions & Completions] Implement UsageInfo Processor (sgl-project#7360) Co-authored-by: Chang Su <chang.s.su@oracle.com> * Fix All-Gather under world size one (sgl-project#7219) * Optimize DP attn scheduling for speculative decoding (sgl-project#7285) * Update usage_processor.py (sgl-project#7402) * Fix 7285 Merge Conflicts (sgl-project#7403) * chore: upgrade mooncake-transfer-engine 0.3.4 (sgl-project#7401) * [OAI Server Refactor] [ChatCompletions & Completions] Support Return Hidden State (sgl-project#7329) Signed-off-by: keru <rukeyang@gmail.com> * Remove batches api in docs & example (sgl-project#7400) * [BugFix]: fix EmbeddingReqInput single input error (sgl-project#7396) * [BugFix]fix qwen25 invoke function call streaming responses with curly braces as the starting indicator (sgl-project#7394) * fix overlap pagecount (sgl-project#6984) Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> * fix: Fix CI test_function_call_parser.py (sgl-project#7425) * Fix CPU offloading for MLA memory pool (sgl-project#7409) * [fix] PD disaggregation when enable mtp and tp!=dp (sgl-project#7420) * feat(oai refactor): Replace `openai_api` with `entrypoints/openai` (sgl-project#7351) Co-authored-by: Jin Pan <jpan236@wisc.edu> * Refactor LoRAManager and LoRAMemoryPool state management logic for dynamic LoRA loading support (sgl-project#7412) * refactor(test): reorganize OpenAI test file structure (sgl-project#7408) * [minor] simplify the `TokenToKVPoolAllocator` (sgl-project#7414) * Tiny add logging for GC (sgl-project#7406) * FlashInfer NVFP4 MoE with EP & 2-stream shared expert (sgl-project#7327) Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com> Co-authored-by: alcanderian <alcanderian@gmail.com> * Remove copy after bmm (sgl-project#7441) * Fix torch compile run (sgl-project#7391) Co-authored-by: wunhuang <wunhuang@amd.com> Co-authored-by: Sai Enduri <saimanas.enduri@amd.com> * [misc] Add PD service discovery support in router (sgl-project#7361) * add fused moe config for qwen3 in triton3.3.1 (sgl-project#7445) * Fix CUDA Graph Check under Deepep with DP FFN (sgl-project#7451) * Update hyperparameter_tuning.md (sgl-project#7454) * feat: integrate deepgemm into EPMoE (sgl-project#6821) Co-authored-by: tianqilin.99 <tianqilin.99@bytedance.com> Co-authored-by: TianQiLin666666 <1834987979@qq.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * Solve docker build failed in the virtual machine (sgl-project#7290) Co-authored-by: wunhuang <wunhuang@amd.com> Co-authored-by: Sai Enduri <saimanas.enduri@amd.com> Co-authored-by: HAI <hixiao@gmail.com> * Fix a bug in BatchTokenIDOut & Misc style and dependency updates (sgl-project#7457) * [CI] Upgrade mooncake to 0.3.4.post1 to fix 8 gpu tests (sgl-project#7472) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * Fix prefill OOM due to wrong token calculation when page > 1 (sgl-project#7397) * feat(func_call): Add more check in `BaseFormatDetector.parse_streaming_increment` (sgl-project#7479) * Fix dtype for idle input in spec decoding (sgl-project#7456) * update mooncake in dockerfile (sgl-project#7480) * kvcache io kernels and test case (sgl-project#7382) * [perf] slightly imporve DeepSeek-R1-FP4 TP8 (sgl-project#7481) * Quick fix for DeepGemm requant to also cover MTP. (sgl-project#7378) * Support weight loading without mmap (sgl-project#7469) * ci: Revert openai_server related tests in AMD suites (sgl-project#7449) * Perormance: Enable cuda graph for dp idle batch (sgl-project#7269) Co-authored-by: austindeng <austindeng@tencent.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: ch-wan <cwan39@gatech.edu> * bugfix: Prevent global mutation of conv.stop_str across requests (sgl-project#7347) Co-authored-by: Chang Su <chang.s.su@oracle.com> * Fix RequestValidationError response format (sgl-project#7487) * Fix MTP with Deepseek R1 Fp4 (sgl-project#7376) * chore: bump sgl-kernel v0.2.0 (sgl-project#7490) * chore: bump v0.4.8 (sgl-project#7493) * [AMD] add aiter fused moe in DeepEP path (sgl-project#7268) * enable aiter_biased_grouped_topk kernel (sgl-project#7423) * [PD Disaggregation] replace transfer with batch transfer for better performance (sgl-project#7236) * Remove cumsum_buffer initilization (sgl-project#7439) * [benchmark] fbgemm benchmark support bandwidth report and support fbgemm_cutlass_gmm (sgl-project#7422) * Support multi-thread model weight loading (sgl-project#7277) * [PD] NIXL: Register kv args in advance and cleanup finished requests (sgl-project#6717) * fix: Add `--model` as an alias for `--model-path` in server_args (sgl-project#7505) * misc: Improvement to serving_chat.py and add more ut (sgl-project#7489) * Fuse sorted_token_ids padding to moe_align_block_size kernel (sgl-project#7437) * [OAI] patch origin request_id logic (sgl-project#7508) * [PD][Spec] Fix hidden state transfer for spec decode (sgl-project#7516) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * EPLB support for MTP (sgl-project#7510) * clean duplicate code (sgl-project#7512) * [ci] add router benchmark script and CI (sgl-project#7498) * fix: force synchronization between TP workers when update_weights (sgl-project#6626) Co-authored-by: dangkai.dk <dangkai.dk@alibaba-inc.com> * [CPU] [BF16] Call fused_experts_cpu, weight_packed_linear and bmm_cpu kernel in DeepSeek model (sgl-project#6641) Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg> * [CI] Upgrade mooncake to v0.3.4.post2 to fix potential slice failed bug (sgl-project#7522) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * npu fused op (sgl-project#7386) Co-authored-by: Li Junwen <lijunwen13@hisilicon.com> * feat: send kvmetrics from sglang scheduler (sgl-project#6721) * [PD] Add different TP sizes support for no-MLA models (sgl-project#6793) Co-authored-by: shangmingc <csmthu@gmail.com> Co-authored-by: Shangming Cai <caishangming@linux.alibaba.com> * enable aiter fp8 blockscale quant (sgl-project#7520) * take aiter get_rope back (sgl-project#7521) * Fix typo of flash_cache (sgl-project#7513) * feat: add return hidden_states at async generation (sgl-project#7507) * minor: 'role' must be system/assistant/tool, but case insensitive for now (sgl-project#7499) * Fix FP8 KV Cache Support in FA3 Backend (sgl-project#7148) * Fix gathered_buffer issues in tbo (sgl-project#7531) * [PD] Raise error for incompatible mooncake version and some minor fixes (sgl-project#7527) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * [CMake] Fix sgl-kernel CMakeLists for Blackwell (sgl-project#7543) * Add Tencent HunYuanMoEV1 model support (sgl-project#7549) * Update seed in CPU UTs to avoid flaky failure with single test (sgl-project#7544) * chore: improve ci bug reporting (sgl-project#7542) * chore: remove vlm unnecessary import (sgl-project#7541) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> Co-authored-by: yhyang201 <yhyang201@gmail.com> Co-authored-by: Mick <mickjagger19@icloud.com> * chore: bump v0.4.8.post1 (sgl-project#7559) * [PD][NIXL] Set is_sorted=False to fix NIXL_ERR_NOT_FOUND (sgl-project#7330) * [Fix] incorrect assert in EPLB (sgl-project#7575) * Updates Gemma3n MLP layer to adapt latest transformers version (sgl-project#7573) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * Fix MTP error when enabling two-batch overlap (sgl-project#7569) * Add e2e test for multi instance multi stage memory release/resume occupuation (sgl-project#7208) Signed-off-by: Ata Fatahi <immrata@gmail.com> * [CI] Add CI Testing for Prefill-Decode Disaggregation with Router (sgl-project#7540) * Updates transformers and timm dependencies (sgl-project#7577) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * feat: support compatibility between MTP and two-batch-overlap (sgl-project#7225) Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> * Move multimodal processors into a separate folder (sgl-project#7581) * Fix broken CI TestVILAServer (sgl-project#7610) * [router] add centralized configuration module for sgl-router (sgl-project#7588) * Fix: Minicpm (sgl-project#7612) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * Hybrid kv cache for LLaMA4 (sgl-project#6563) Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: tarinkk <rt572@physics.rutger.edu> Co-authored-by: tarinkk <rt572@rutgers.physics.edu> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> * [CPU] add optimizations for INT8 and FP8 DeepSeek (sgl-project#6769) Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com> * Tiny add logs for expert location updater (sgl-project#7308) * Fix flakiness in LoRA batch test. (sgl-project#7552) * [BUG] fix local_rank in initialize_dp_attention (sgl-project#7584) * Support dynamic LoRA loading / unloading in engine/server API (sgl-project#7446) * [PD] Respect sampling_params.max_new_tokens when PD disaggregation is activated (sgl-project#7598) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * fix unit tests (sgl-project#7618) * Let ep_scatter support arbitrary strides / ue8m0 format (sgl-project#7309) * Let EP prefill support new DeepGEMM (sgl-project#7310) * docs: add gb200 nvl72 and a16z grant (sgl-project#7620) * oai: Adds support for OpenAI chat completions API in bench_serving (sgl-project#7036) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com> Co-authored-by: Mick <mickjagger19@icloud.com> * [bugfix] Remove PR comment posting from Rust benchmark workflow (sgl-project#7625) * [Minor] clean up multimodal processor and tokenizer manager (sgl-project#7624) * Add dsv3 fused a gemm to sgl-kernel (sgl-project#7630) * Add @mickqian as the CODEOWNERS of multimodal (sgl-project#7636) * Fix stream reasoning parser and Adds Kimi reasoning parser (sgl-project#7432) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * Fix sgl-router startup crash (sgl-project#7619) * [bugfix] fix runtime dropping panic in editable (sgl-project#7628) * Move files related to EPLB (sgl-project#7580) * [misc] reduce weird rope_scaling_factor warning (sgl-project#7176) * [AMD] Add unit-test-sgl-kernel-amd to AMD CI (sgl-project#7539) * Update CODEOWNERS (sgl-project#7640) * [EAGLE] remove a wrong adjustment for page_size > 1 & topk > 1 in server_args.py (sgl-project#7643) * [CPU] add c++ kernel to bind CPU cores and memory node (sgl-project#7524) * Improve streaming, log_level, memory report, weight loading, and benchmark script (sgl-project#7632) Co-authored-by: Kan Wu <wukanustc@gmail.com> * Add dsv3 router gemm kernel (sgl-project#7627) * chore: upgrade flashinfer v0.2.7 jit (sgl-project#7663) * [doc] update lws doc for pd (sgl-project#7318) * Fix: sync prepare_fp8_layer_for_marlin with latest vllm changes (sgl-project#7648) * Add small requirements for benchmark/parse_result tools (sgl-project#7671) * [CPU] remove process_group from inputs of shm_allreduce and shm_allgather (sgl-project#7486) * chore: bump sgl-kernel v0.2.1 (sgl-project#7675) * support llama4 eagle3 (sgl-project#6985) Co-authored-by: shuaills <shishuaiuoe@gmail.com> Co-authored-by: Shenggui Li <somerlee.9@gmail.com> Co-authored-by: Yingyi Huang <yingyihuang2000@outlook.com> Co-authored-by: yizhang2077 <1109276519@qq.com> * Refactor mm processors and Enable mixed modality processing (sgl-project#7629) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * upgrade sgl kernel to 0.2.1 for main (sgl-project#7676) * add description for llama4 eagle3 (sgl-project#7688) * fix(model loader): use safe_open to prevent file handle leaks. (sgl-project#7684) * chore: upgrade flashinfer v0.2.7.post1 (sgl-project#7698) * Improve error handling for requests with unloaded LoRA path(s) (sgl-project#7642) * Apply dsv3_fused_a_gemm kernel (sgl-project#7635) * Fix GPTQMarlinMoE (sgl-project#7697) * [1/n] apply wna16marlin kernel in moe weight only quantization (sgl-project#7683) Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com> Co-authored-by: yych0745 <1398089567@qq.com> Co-authored-by: HandH1998 <1335248067@qq.com> Co-authored-by: 弋云 <yiyun.wyt@antgroup.com> Co-authored-by: walker-ai <2398833647@qq.com> * Apply dsv3 router gemm kernel for deepseek-r1 fp4 (sgl-project#7677) * [AMD] Temporarily disable test_no_overlap_scheduler and test_vision_chunked_prefill (sgl-project#7717) * [RL] add --skip-warmup (sgl-project#7416) * [RL] support update_weights_from_distributed with different group and multiple weights (sgl-project#7292) * [router] add --log-level to sgl-router (sgl-project#6512) * [b200] support trt-llm allreduce fuse rms_norm_add kernel (sgl-project#7621) * [CPU] Bind threads and numa node for each TP rank (sgl-project#6549) Co-authored-by: srinarayan-srikanthan <srinarayan.srikanthan@intel.com> * Support non-contiguous query input for extend/decode attention (sgl-project#7462) * Support updating weights at once by stopping all requests (sgl-project#6698) Signed-off-by: Tianyu Zhou <albert.zty@antgroup.com> Co-authored-by: Zilin Zhu <zhuzilinallen@gmail.com> * Fix num_tokens_pre_allocated in disaggregation log (sgl-project#7714) * [CPU] [sgl-kernel] set dispatch key of initialize to CatchAll (sgl-project#7734) * [CPU] fix all_reduce and all_gather (sgl-project#6770) Co-authored-by: blzheng <beilei.zheng@intel.com> * fix awq and dsv3 fused gemm compatible (sgl-project#7735) * [CI][Router] Fix bench_one_batch_server for pd router test (sgl-project#7731) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> * Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture (sgl-project#7278) Co-authored-by: HydraQYH <QYH820@Outlook.com> Co-authored-by: TianQiLin666666 <1834987979@qq.com> * fix dsv3 fused proj check (sgl-project#7738) * Ascend attention backend(PA&MLA) (sgl-project#7722) Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: VDV1985 <vladdv85@mail.ru> * [fix] fix dsv3_router_gemm filter (sgl-project#7750) * [CPU] refine CPU integration code (sgl-project#7647) * [CPU] support the case where num_attention_heads or intermediate_size is not divisible by the TP size (sgl-project#6771) * support qwen3 dense model dp attention (sgl-project#7681) * [optimize] add two stream norm for qwen3 (sgl-project#7740) Co-authored-by: ispobock <ispobaoke@gmail.com> * feat: use D2D instead of H2H in pp (sgl-project#7673) Co-authored-by: alpha-baby <fujianhao1997@qq.com> * [Bug] add flashinfer bool check for fusedmoe in Qwen moe models (sgl-project#7723) * [fix] put cpu in the first priority in get_device() (sgl-project#7752) * [optimize] fuse renormalize into moe_topk_softmax (sgl-project#7744) Co-authored-by: ispobock <ispobaoke@gmail.com> * chore: bump sgl-kernel 0.2.2 (sgl-project#7755) * fix CI: update native api ipynb (sgl-project#7754) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> * fuse renormal into moe topk softmax kernel python code (sgl-project#7751) Co-authored-by: ispobock <ispobaoke@gmail.com> Co-authored-by: zhyncs <me@zhyncs.com> * Remove type conversion and fix id map in topk (sgl-project#7759) * Add V2-lite model test (sgl-project#7390) Co-authored-by: DiweiSun <105627594+DiweiSun@users.noreply.github.com> * refactor llama4 dp attention logic (sgl-project#7729) * fix(docs): fix the broken link in `docs/references/production_metrics.md` (sgl-project#7741) Signed-off-by: rudeigerc <rudeigerc@gmail.com> * [fix] update bench_speculative.py for compatibility (sgl-project#7764) Signed-off-by: Kay Yan <kay.yan@daocloud.io> * Move mem_fraction_static adjustment for multimodal models to `server_args.py` & Fix session control & Other cleanups (sgl-project#7748) * [RL] Add --nccl-port to prevent port conflict (sgl-project#7418) * [RL] add pause and continue generation for async rl training (sgl-project#7419) * [Fix] Alloc return type error (sgl-project#7778) Signed-off-by: Capronir <839972205@qq.com> * [feat] Support EAGLE3 for Qwen (sgl-project#7745) Co-authored-by: 纬杭 <ximing.wxm@antgroup.com> Co-authored-by: zyksir <zyksir@outlook.com> * saving hidden_states.clone() (sgl-project#7705) * [1/n]: add cutlass W4A8 moe kernel for hopper architecture (sgl-project#7772) Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com> Co-authored-by: yicwang <yichen.wang@bytedance.com> * add model: qwen2-audio (sgl-project#7596) * Optimize Hopper CUTLASS FP8 Blockwise Grouped GEMM Kernel in Small K Scenario (sgl-project#7782) * Embedding parallel by attn_tp (sgl-project#7623) * fix: fix apply_shuffle_mul_sum (sgl-project#7444) * chore: bump sgl-kernel v0.2.3 (sgl-project#7784) * fix: use nvidia-nccl-cu12 2.27.5 (sgl-project#7787) * DP Attention with Auto DeepEP Dispatch (sgl-project#7222) * chore: upgrade sgl-kernel v0.2.3 (sgl-project#7786) * Fix incorrect spec_num_draft_tokens in draft_extend (sgl-project#7757) * [fix] fix misusing of is_cuda (sgl-project#7790) * Add treemask mode to build_eagle_tree & release sgl-kernel 0.2.3 (sgl-project#7756) Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com> * chore: bump sgl-kernel v0.2.4 (sgl-project#7800) * ci: fix port args (sgl-project#7792) * Fix CI test OOM issue. (sgl-project#7799) * chore: upgrade sgl-kernel v0.2.4 (sgl-project#7801) * chore: bump v0.4.9 (sgl-project#7802) * fix merge conflict issue * fix hpu attention nonetyep issue * fix alignment * fix alignment2 * Ci failure fixes * fix attention-backend choices --------- Signed-off-by: Xinyuan Tong <justinning0323@outlook.com> Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> Signed-off-by: ch-tiger1 <xyz@ch-tech.ip-ddns.com> Signed-off-by: huanglong <huanglong@linux.alibaba.com> Signed-off-by: Ata Fatahi <immrata@gmail.com> Signed-off-by: keru <rukeyang@gmail.com> Signed-off-by: Tianyu Zhou <albert.zty@antgroup.com> Signed-off-by: rudeigerc <rudeigerc@gmail.com> Signed-off-by: Kay Yan <kay.yan@daocloud.io> Signed-off-by: Capronir <839972205@qq.com> Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com> Signed-off-by: Mohit Sinha <msinha@habana.ai> Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com> Co-authored-by: KavioYu <67678385+yukavio@users.noreply.github.com> Co-authored-by: kavioyu <kavioyu@tencent.com> Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com> Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com> Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com> Co-authored-by: wunhuang <wunhuang@amd.com> Co-authored-by: DiweiSun <105627594+DiweiSun@users.noreply.github.com> Co-authored-by: u4lr451 <u4lr451@gmail.com> Co-authored-by: austindeng <austindeng@tencent.com> Co-authored-by: tianqilin.99 <tianqilin.99@bytedance.com> Co-authored-by: Qiaolin Yu <liin1211@outlook.com> Co-authored-by: ch-wan <cwan39@gatech.edu> Co-authored-by: Yijie Zhu <762412795@qq.com> Co-authored-by: 刁莹煜 <diaoyingyu1@hisilicon.com> Co-authored-by: Charles Chen <pychen96@gmail.com> Co-authored-by: Chang Su <chang.s.su@oracle.com> Co-authored-by: AniZpZ <zhuangsen.zp@antgroup.com> Co-authored-by: Yineng Zhang <me@zhyncs.com> Co-authored-by: shangmingc <caishangming@linux.alibaba.com> Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu> Co-authored-by: YanbingJiang <yanbing.jiang@intel.com> Co-authored-by: Wu, Chunyuan <chunyuan.wu@intel.com> Co-authored-by: jianan-gu <jianan.gu@intel.com> Co-authored-by: sdp <sdp@gnr799219.jf.intel.com> Co-authored-by: Binyao Jiang <byjiang1996@gmail.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: linzhuo <15313137931lz@gmail.com> Co-authored-by: ch-tiger1 <tiger@ch-tech.ip-ddns.com> Co-authored-by: ch-tiger1 <xyz@ch-tech.ip-ddns.com> Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com> Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com> Co-authored-by: Simo Lin <linsimo.mark@gmail.com> Co-authored-by: Jinn <47354855+jhinpan@users.noreply.github.com> Co-authored-by: Stefan He <hebiaobuaa@gmail.com> Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com> Co-authored-by: Atream <80757050+Atream@users.noreply.github.com> Co-authored-by: Li Hui <lambert80.ios@gmail.com> Co-authored-by: Huang Long <121648372+LLLL114@users.noreply.github.com> Co-authored-by: woodx <124784234+woodx9@users.noreply.github.com> Co-authored-by: Ata Fatahi <immrata@gmail.com> Co-authored-by: strgrb <zhangkaihong.zkh@antgroup.com> Co-authored-by: Zhang Kaihong <zhangkaihong.zkh@alibaba-inc.com> Co-authored-by: Wenbo Yang <solrex@users.noreply.github.com> Co-authored-by: Chang Su <csu272@usc.edu> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com> Co-authored-by: Keyang Ru <rukeyang@gmail.com> Co-authored-by: ehuaa <ehuamail@163.com> Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com> Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com> Co-authored-by: Jin Pan <jpan236@wisc.edu> Co-authored-by: Lifu Huang <lifu.hlf@gmail.com> Co-authored-by: Trevor Morris <tmorris@nvidia.com> Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com> Co-authored-by: alcanderian <alcanderian@gmail.com> Co-authored-by: Ke Bao <ISPObaoke@163.com> Co-authored-by: Sai Enduri <saimanas.enduri@amd.com> Co-authored-by: Yi Zhang <1109276519@qq.com> Co-authored-by: xutizhou <xutingz@nvidia.com> Co-authored-by: TianQiLin666666 <1834987979@qq.com> Co-authored-by: HAI <hixiao@gmail.com> Co-authored-by: Yuhong Guo <guoyuhong1985@outlook.com> Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com> Co-authored-by: Alex Sun <alex.s@amd.com> Co-authored-by: valarLip <103567126+valarLip@users.noreply.github.com> Co-authored-by: Francis <38564764+ssssnow@users.noreply.github.com> Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com> Co-authored-by: xianzhiT <xianzhitang@tencent.com> Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com> Co-authored-by: DangKai <dangkai4u@outlook.com> Co-authored-by: dangkai.dk <dangkai.dk@alibaba-inc.com> Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg> Co-authored-by: ll819214 <18801269230@163.com> Co-authored-by: Li Junwen <lijunwen13@hisilicon.com> Co-authored-by: zixuanzhang226 <zixuanzhang@bytedance.com> Co-authored-by: Hongbo Xu <1320612015@qq.com> Co-authored-by: shangmingc <csmthu@gmail.com> Co-authored-by: eigen <52445717+yyihuang@users.noreply.github.com> Co-authored-by: mlmz <54172054+minleminzui@users.noreply.github.com> Co-authored-by: Ruihang Lai <ruihangl@cs.cmu.edu> Co-authored-by: Meng, Peng <pengmeng@tencent.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: yhyang201 <yhyang201@gmail.com> Co-authored-by: tarinkk <129432511+tarinkk@users.noreply.github.com> Co-authored-by: tarinkk <rt572@physics.rutger.edu> Co-authored-by: tarinkk <rt572@rutgers.physics.edu> Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com> Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com> Co-authored-by: Sheng Qi <shengqi2018@pku.edu.cn> Co-authored-by: finetune <82650881+finetunej@users.noreply.github.com> Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com> Co-authored-by: Kan Wu <wukanustc@gmail.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: narutolhy <582909902@qq.com> Co-authored-by: lukec <118525388+sleepcoo@users.noreply.github.com> Co-authored-by: shuaills <shishuaiuoe@gmail.com> Co-authored-by: Shenggui Li <somerlee.9@gmail.com> Co-authored-by: Yingyi Huang <yingyihuang2000@outlook.com> Co-authored-by: Simon_CQK <cqk0100@gmail.com> Co-authored-by: Kyungmin Lee <30465912+lkm2835@users.noreply.github.com> Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com> Co-authored-by: yych0745 <1398089567@qq.com> Co-authored-by: HandH1998 <1335248067@qq.com> Co-authored-by: 弋云 <yiyun.wyt@antgroup.com> Co-authored-by: walker-ai <2398833647@qq.com> Co-authored-by: Zilin Zhu <zhuzilinallen@gmail.com> Co-authored-by: srinarayan-srikanthan <srinarayan.srikanthan@intel.com> Co-authored-by: Albert <albert.zty@antgroup.com> Co-authored-by: Ziming Huang <1520787127@qq.com> Co-authored-by: ayrnb <70835312+ayrnb@users.noreply.github.com> Co-authored-by: HydraQYH <QYH820@Outlook.com> Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: VDV1985 <vladdv85@mail.ru> Co-authored-by: ispobock <ispobaoke@gmail.com> Co-authored-by: TianyuZhang1214 <tianyuzhang1214@163.com> Co-authored-by: alpha-baby <fujianhao1997@qq.com> Co-authored-by: Yuchen Cheng <rudeigerc@gmail.com> Co-authored-by: Kay Yan <kay.yan@daocloud.io> Co-authored-by: Caproni <40862361+Capronir@users.noreply.github.com> Co-authored-by: Ximingwang-09 <72070413+Ximingwang-09@users.noreply.github.com> Co-authored-by: 纬杭 <ximing.wxm@antgroup.com> Co-authored-by: zyksir <zyksir@outlook.com> Co-authored-by: SijiaYang <yangsijia.614@bytedance.com> Co-authored-by: yicwang <yichen.wang@bytedance.com> Co-authored-by: Leng Yue <lengyue@lengyue.me> Co-authored-by: Qi Yuhang <45795032+HydraQYH@users.noreply.github.com> Co-authored-by: Gang Chen <13298548+MoonBall@users.noreply.github.com> Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com> Co-authored-by: jay <jthakur@habana.ai>

xianzhiT requested review from merrymercy, Ying1123, hnyls2002, zhyncs, ispobock and ByronHsu as code owners June 17, 2025 13:39

gemini-code-assist bot reviewed Jun 17, 2025

View reviewed changes

xianzhiT marked this pull request as draft June 17, 2025 15:23

xianzhiT marked this pull request as ready for review June 17, 2025 15:26

garrett4wade reviewed Jun 18, 2025

View reviewed changes

xianzhiT closed this Jun 18, 2025

xianzhiT reopened this Jun 18, 2025

BraveY reviewed Jun 18, 2025

View reviewed changes

python/sglang/srt/model_loader/weight_utils.py Outdated Show resolved Hide resolved

xianzhiT force-pushed the feature/support_multithread_safetenosr_load branch from 6b1312f to 4801661 Compare June 18, 2025 06:32

BraveY mentioned this pull request Jun 18, 2025

feat: add load format 'prefetch_auto' for parallel mmap prefetching #7209

Open

6 tasks

xianzhiT requested a review from garrett4wade June 19, 2025 12:39

Support multi-thread model weight loading

fd697cd

xianzhiT force-pushed the feature/support_multithread_safetenosr_load branch from 4b9bd2b to fd697cd Compare June 24, 2025 09:46

Merge branch 'main' into feature/support_multithread_safetenosr_load

49e4bf8

zhyncs merged commit 9f1787f into sgl-project:main Jun 24, 2025
63 of 73 checks passed

CatherineSue mentioned this pull request Jun 24, 2025

fix: Add --model as an alias for --model-path in server_args #7505

Merged

6 tasks

yudian0504 mentioned this pull request Jun 25, 2025

fix: incorrect dtype when load_model #7517

Open

6 tasks

Edenzzzz mentioned this pull request Jun 27, 2025

[Feature] Load weights from distributed hao-ai-lab/FastVideo#470

Merged

3 tasks

tianyuzhou95 mentioned this pull request Jul 11, 2025

concurrently load weights of DeepseekV2ForCausalLM #7943

Merged

6 tasks

chenxijun1029 added a commit to chenxijun1029/sglang that referenced this pull request Jul 17, 2025

Support multi-thread model weight loading (sgl-project#7277)

47ed307

chenxijun1029 pushed a commit to chenxijun1029/sglang that referenced this pull request Jul 17, 2025

Support multi-thread model weight loading (sgl-project#7277)

b0f2759

shuaills pushed a commit to shuaills/sglang that referenced this pull request Jul 21, 2025

Support multi-thread model weight loading (sgl-project#7277)

74613c3

Support multi-thread model weight loading #7277

Support multi-thread model weight loading #7277

Conversation

xianzhiT commented Jun 17, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Motivation

Modifications

Checklist

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Summary of Changes

Highlights

Footnotes

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

garrett4wade left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

BraveY commented Jun 18, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

xianzhiT commented Jun 18, 2025

Uh oh!

xianzhiT commented Jun 18, 2025

Uh oh!

Uh oh!

BraveY commented Jun 18, 2025

Uh oh!

xianzhiT commented Jun 18, 2025

Uh oh!

xianzhiT commented Jun 18, 2025

Uh oh!

guoyuhong commented Jun 24, 2025

Uh oh!

Uh oh!

zhaochenyang20 commented Jun 24, 2025

Uh oh!

Uh oh!

xianzhiT commented Jun 17, 2025 •

edited

Loading

BraveY commented Jun 18, 2025 •

edited

Loading