-
-
Notifications
You must be signed in to change notification settings - Fork 20k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Parser] Migrate Kimi K3 to Parser Engine
k3
kimi
tool-calling
#50229
opened Jul 29, 2026 by
chaunceyjiang
Collaborator
•
Draft
4 tasks
[Bugfix][Kimi K3] Fix message-level tools, response_format passthroug…
bug
Something isn't working
frontend
k3
kimi
tool-calling
#50228
opened Jul 29, 2026 by
wangln19
Contributor
Loading…
4 tasks
[Bugfix] Avoid non-contiguous CPU tensors in Llama-4 ModelOpt weight …
bug
Something isn't working
llama
Related to Llama models
#50227
opened Jul 29, 2026 by
Dev-with-Mouzan
Loading…
[Kernel] Optimize SM103 GDN causal-conv prefills
ci/build
performance
Performance-related issues
#50226
opened Jul 29, 2026 by
Alizen-1009
Loading…
[Cleanup] Remove orphaned output_text_buffer_length and stale spec-decode startup warning after V0 removal
v1
#50225
opened Jul 29, 2026 by
xiaoyaoqilan
Loading…
[Docs] Update Quickstart with Apple Silicon/Metal compatibility note and safetensors-compatible default model
documentation
Improvements or additions to documentation
#50224
opened Jul 29, 2026 by
xiaoyaoqilan
Loading…
[Docs] Add SharedStorageConnector→ExampleConnector migration note and fix connector inventory
documentation
Improvements or additions to documentation
#50223
opened Jul 29, 2026 by
xiaoyaoqilan
Loading…
[CI] Fix MXFP8 MOE backend selection tests on gfx942
ready
ONLY add when PR is ready to merge/full CI is needed
rocm
Related to AMD ROCm
#50222
opened Jul 29, 2026 by
fxmarty-amd
Contributor
Loading…
fix(security): enforce audio decode duration limit in NanoNemotronVL
#50221
opened Jul 29, 2026 by
jperezdealgaba
Contributor
Loading…
[CPU][s390x] Optimize inference perf and add oneDNN INT8 GEMM for s390x
ci/build
cpu
Related to CPU backends
documentation
Improvements or additions to documentation
#50219
opened Jul 29, 2026 by
R3hankhan123
Contributor
Loading…
4 tasks
[Cleanup] Remove orphaned output_text_buffer_length and stale spec-decode warning after V0 removal- #50225
v1
#50218
opened Jul 29, 2026 by
AdaAibaby
Loading…
[Bugfix][CI/Build] Honor VLLM_PRECOMPILED_WHEEL_COMMIT=nightly
bug
Something isn't working
ci/build
#50216
opened Jul 29, 2026 by
BIT-Orange
Loading…
4 tasks done
[Bugfix][Quantization] Fix fused AutoGPTQ overrides and mixed-group MoE loading
bug
Something isn't working
quantization
#50214
opened Jul 29, 2026 by
ZX-ModelCloud
Loading…
4 tasks done
[Doc]: Update Quickstart documentation with Apple Silicon / Metal execution
documentation
Improvements or additions to documentation
#50213
opened Jul 29, 2026 by
xiaoyaoqilan
Loading…
4 tasks
[ROCm][Perf][Model] Fuse Qwen3-VL attention prologue into single AITER kernel
qwen
Related to Qwen models
rocm
Related to AMD ROCm
#50212
opened Jul 29, 2026 by
vorapolsiloai
Loading…
4 tasks
[XPU] Fix inc int4 model
intel-gpu
Related to Intel GPU
quantization
#50209
opened Jul 29, 2026 by
mayuyuace
Contributor
Loading…
[Bugfix][KV Connector][Mooncake] Preserve HMA region strides
bug
Something isn't working
kv-connector
v1
#50208
opened Jul 29, 2026 by
Dao007forever
Contributor
Loading…
[XPU][CI]Add back tests/v1/e2e/general/test_correctness_sliding_window.py::test_sliding_window_retrieval[True-1-5-google/gemma-3-1b-it]
ci/build
intel-gpu
Related to Intel GPU
#50207
opened Jul 29, 2026 by
zxd1997066
Contributor
Loading…
4 tasks done
trtllm fp8 moe sm100 compatibility
nvidia
#50205
opened Jul 29, 2026 by
JaredforReal
Contributor
Loading…
4 tasks
[Kernel] Harden top_k_per_row against NaN and under-filled output
#50201
opened Jul 29, 2026 by
Smallfu666
Loading…
[Bugfix][Rust Frontend] Select earliest-completing stop string
bug
Something isn't working
rust
#50200
opened Jul 29, 2026 by
samlaf
Loading…
[XPU][CI]Adjust Samplers test ENV for Intel GPU
ci/build
intel-gpu
Related to Intel GPU
#50199
opened Jul 29, 2026 by
zxd1997066
Contributor
Loading…
4 tasks done
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.