-
Notifications
You must be signed in to change notification settings - Fork 2.6k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
Hybrid GDN (Qwen3.5-class) serving appears host/sync-bound at high concurrency — GPU only ~35% busy; expected?
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentInference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16976 In NVIDIA/TensorRT-LLM;[Bug]: Structured Output with DSpark Speculative Drafter
bugSomething isn't workingSomething isn't workingPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16967 In NVIDIA/TensorRT-LLM;Support NumPy token IDs in executor APIs
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#16910 In NVIDIA/TensorRT-LLM;[Bug]: Llama-3-70B FP16/BF16 inference fails while FP8 works on TensorRT-LLM 0.21.c-rc0 with 8x RTX 6000D
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Low PrecisionLower-precision formats (INT8/INT4/FP8) for TRTLLM quantization (AWQ, GPTQ).Lower-precision formats (INT8/INT4/FP8) for TRTLLM quantization (AWQ, GPTQ).Scale-out<NV>Multi-GPU and distributed inference scaling issues, tensor/pipeline/data parallelism<NV>Multi-GPU and distributed inference scaling issues, tensor/pipeline/data parallelismStatus: Open.#16899 In NVIDIA/TensorRT-LLM;Gemma4 attention_k_eq_v: NVFP4 v_proj scale tensors not duplicated from k_proj → KeyError: 'v' in fused-QKV loader
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Model optimization<NV>Model-specific performance optimizations and tuning<NV>Model-specific performance optimizations and tuningPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16867 In NVIDIA/TensorRT-LLM;[Performance]: MAX_UTILIZATION causes a 40.6% output-throughput drop and 9.1x TPOT p99 under KV-cache oversubscription (Llama-3.3-70B-FP8, H200, v1.2.1)
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentKV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16827 In NVIDIA/TensorRT-LLM;[Bug] Vanilla attention uses incorrect KV indices for layer-specific cache layouts
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceStatus: Open.#16801 In NVIDIA/TensorRT-LLM;TensorRT-LLM PyTorch backend Qwen3.5-VL parity issue: HF/vLLM greedy output matches, TRT-LLM repeats to max tokens
bugSomething isn't workingSomething isn't workingMultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16792 In NVIDIA/TensorRT-LLM;[Bug] DSpark speculative decoding: accept length collapses to ~1 at generation batch size > 1 in disaggregated serving
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16767 In NVIDIA/TensorRT-LLM;gRPC GenerateRequest.max_tokens is required — should be optional with an engine-side default (parity with LLM API)
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#16549 In NVIDIA/TensorRT-LLM;[New Model]: thinkingmachines/Inkling
new modelRequest to add a new modelRequest to add a new modelStatus: Open.#16507 In NVIDIA/TensorRT-LLM;[Bug]: find_input_mm_embeds is order-dependent for mixed full-prefill and partial VLM requests
MultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsStatus: Open.#16460 In NVIDIA/TensorRT-LLM;