-
Notifications
You must be signed in to change notification settings - Fork 2.6k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
Gemma4 attention_k_eq_v: NVFP4 v_proj scale tensors not duplicated from k_proj → KeyError: 'v' in fused-QKV loader
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Model optimization<NV>Model-specific performance optimizations and tuning<NV>Model-specific performance optimizations and tuningPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16867 In NVIDIA/TensorRT-LLM;Add REFUTE scientific critique + calibration benchmark
Doc<NV>TRTLLM's textual/illustrative materials: API refs, guides, tutorials. Improvement & clarity.<NV>TRTLLM's textual/illustrative materials: API refs, guides, tutorials. Improvement & clarity.Status: Open.#16855 In NVIDIA/TensorRT-LLM;[Performance]: MAX_UTILIZATION causes a 40.6% output-throughput drop and 9.1x TPOT p99 under KV-cache oversubscription (Llama-3.3-70B-FP8, H200, v1.2.1)
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentKV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16827 In NVIDIA/TensorRT-LLM;[Bug]: SA (Suffix Automaton) speculative decoding crashes the executor event loop when a prompt is near
max_seq_len—default_max_tokensis deduced against the draft-inflatedmax_seq_lenbugSomething isn't workingSomething isn't workingInference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Speculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16826 In NVIDIA/TensorRT-LLM;[Bug] Vanilla attention uses incorrect KV indices for layer-specific cache layouts
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceStatus: Open.#16801 In NVIDIA/TensorRT-LLM;TensorRT-LLM PyTorch backend Qwen3.5-VL parity issue: HF/vLLM greedy output matches, TRT-LLM repeats to max tokens
bugSomething isn't workingSomething isn't workingMultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16792 In NVIDIA/TensorRT-LLM;[Bug] DSpark speculative decoding: accept length collapses to ~1 at generation batch size > 1 in disaggregated serving
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16767 In NVIDIA/TensorRT-LLM;gRPC GenerateRequest.max_tokens is required — should be optional with an engine-side default (parity with LLM API)
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#16549 In NVIDIA/TensorRT-LLM;[New Model]: thinkingmachines/Inkling
new modelRequest to add a new modelRequest to add a new modelStatus: Open.#16507 In NVIDIA/TensorRT-LLM;[Bug]: find_input_mm_embeds is order-dependent for mixed full-prefill and partial VLM requests
MultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsStatus: Open.#16460 In NVIDIA/TensorRT-LLM;[Bug]: Preserve per-item processed metadata when computing multimodal token lengths
MultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsStatus: Open.#16459 In NVIDIA/TensorRT-LLM;KV cache connector hangs with Eagle speculative decoding when rewind crosses block boundary
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16448 In NVIDIA/TensorRT-LLM;