Skip to content

Pull requests: vllm-project/vllm

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Bugfix][Kimi K3] Fix message-level tools, response_format passthroug… bug Something isn't working frontend k3 kimi tool-calling
#50228 opened Jul 29, 2026 by wangln19 Contributor Loading…
4 tasks
[Bugfix] Avoid non-contiguous CPU tensors in Llama-4 ModelOpt weight … bug Something isn't working llama Related to Llama models
#50227 opened Jul 29, 2026 by Dev-with-Mouzan Loading…
[Kernel] Optimize SM103 GDN causal-conv prefills ci/build performance Performance-related issues
#50226 opened Jul 29, 2026 by Alizen-1009 Loading…
[CI] Fix MXFP8 MOE backend selection tests on gfx942 ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm
#50222 opened Jul 29, 2026 by fxmarty-amd Contributor Loading…
fix(security): enforce audio decode duration limit in NanoNemotronVL
#50221 opened Jul 29, 2026 by jperezdealgaba Contributor Loading…
Fix MoE fused sum row offsets
#50220 opened Jul 29, 2026 by happyyzy Loading…
[CPU][s390x] Optimize inference perf and add oneDNN INT8 GEMM for s390x ci/build cpu Related to CPU backends documentation Improvements or additions to documentation
#50219 opened Jul 29, 2026 by R3hankhan123 Contributor Loading…
4 tasks
[Bugfix][CI/Build] Honor VLLM_PRECOMPILED_WHEEL_COMMIT=nightly bug Something isn't working ci/build
#50216 opened Jul 29, 2026 by BIT-Orange Loading…
4 tasks done
[Bugfix][Quantization] Fix fused AutoGPTQ overrides and mixed-group MoE loading bug Something isn't working quantization
#50214 opened Jul 29, 2026 by ZX-ModelCloud Loading…
4 tasks done
[Doc]: Update Quickstart documentation with Apple Silicon / Metal execution documentation Improvements or additions to documentation
#50213 opened Jul 29, 2026 by xiaoyaoqilan Loading…
4 tasks
[ROCm][Perf][Model] Fuse Qwen3-VL attention prologue into single AITER kernel qwen Related to Qwen models rocm Related to AMD ROCm
#50212 opened Jul 29, 2026 by vorapolsiloai Loading…
4 tasks
[XPU] Fix inc int4 model intel-gpu Related to Intel GPU quantization
#50209 opened Jul 29, 2026 by mayuyuace Contributor Loading…
[Bugfix][KV Connector][Mooncake] Preserve HMA region strides bug Something isn't working kv-connector v1
#50208 opened Jul 29, 2026 by Dao007forever Contributor Loading…
trtllm fp8 moe sm100 compatibility nvidia
#50205 opened Jul 29, 2026 by JaredforReal Contributor Loading…
4 tasks
[Bugfix][Rust Frontend] Select earliest-completing stop string bug Something isn't working rust
#50200 opened Jul 29, 2026 by samlaf Loading…
[XPU][CI]Adjust Samplers test ENV for Intel GPU ci/build intel-gpu Related to Intel GPU
#50199 opened Jul 29, 2026 by zxd1997066 Contributor Loading…
4 tasks done
ProTip! Type g p on any issue or pull request to go back to the pull request listing page.