| pooled: unsloth-zoo + vllm + llama.cpp | [DeepEP V2] Fill invalid recv_topk_idx with -1 (#46432) · vllm | 10.0 | 0.0 |
| vllm-project/vllm | [DeepEP V2] Fill invalid recv_topk_idx with -1 (#46432) · vllm | 10.0 | 0.0 |
| ggml-org/llama.cpp | model : support granite multilingual embeddings R2 (ibm-granite/granite-embedding-{97,311}m-multilingual-r2) (#22716) · conversion | 9.0 | 0.0 |
| ggml-org/llama.cpp | convert : minor fixes for numpy 2.x (#23571) · examples | 9.0 | 0.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | model : support granite multilingual embeddings R2 (ibm-granite/granite-embedding-{97,311}m-multilingual-r2) (#22716) · conversion | 9.0 | 0.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | convert : minor fixes for numpy 2.x (#23571) · examples | 9.0 | 0.0 |
| ggml-org/llama.cpp | cuda : fix KQ mask offset integer overflow in fattn MMA kernel (#23610) · ggml | 10.0 | 1.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | cuda : fix KQ mask offset integer overflow in fattn MMA kernel (#23610) · ggml | 10.0 | 1.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Model]Fix MiniMaxM2ForCausalLM perf regression (#45935) · tests | 8.2 | 0.0 |
| vllm-project/vllm | [Model]Fix MiniMaxM2ForCausalLM perf regression (#45935) · tests | 8.2 | 0.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Kernel] Add swap AB optimization to fused_moe_kernel (#36559) · vllm | 10.0 | 2.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in sampler kernels (#46560) · tests | 9.0 | 1.0 |
| vllm-project/vllm | [Kernel] Add swap AB optimization to fused_moe_kernel (#36559) · vllm | 10.0 | 2.0 |
| vllm-project/vllm | [Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in sampler kernels (#46560) · tests | 9.0 | 1.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(moe_wna16): access tp_size via moe_config for RoutedExperts compatibility (#45404) · vllm | 9.3 | 2.0 |
| vllm-project/vllm | fix(moe_wna16): access tp_size via moe_config for RoutedExperts compatibility (#45404) · vllm | 9.3 | 2.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(moe-fp8): inline weight/quant-state lookup; narrow ImportError scope · unsloth_zoo | 9.0 | 3.0 |
| unslothai/unsloth-zoo | fix(moe-fp8): inline weight/quant-state lookup; narrow ImportError scope · unsloth_zoo | 9.0 | 3.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | Auto-install fused lm_head + cross_entropy forward across transformers · tests | 9.8 | 4.0 |
| unslothai/unsloth-zoo | Auto-install fused lm_head + cross_entropy forward across transformers · tests | 9.8 | 4.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [AMD][OCP MX][CI] Fix tests to not dispatch on `UNFUSED_TRITON` backend on MI300, improve w_mxfp4_a_fp8 emulation support (#46142) · tests | 10.0 | 4.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | Humming support for 2/3/5/6/7-bit pack-quantized weight-only inference (#46389) · vllm | 10.0 | 4.5 |
| vllm-project/vllm | [AMD][OCP MX][CI] Fix tests to not dispatch on `UNFUSED_TRITON` backend on MI300, improve w_mxfp4_a_fp8 emulation support (#46142) · tests | 10.0 | 4.5 |
| vllm-project/vllm | Humming support for 2/3/5/6/7-bit pack-quantized weight-only inference (#46389) · vllm | 10.0 | 4.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | add fixes for moe · unsloth_zoo | 10.0 | 5.2 |
| unslothai/unsloth-zoo | add fixes for moe · unsloth_zoo | 10.0 | 5.2 |
| pooled: unsloth-zoo + vllm + llama.cpp | Fix gptoss 4bit (#524) · unsloth_zoo | 6.5 | 2.0 |
| unslothai/unsloth-zoo | Fix gptoss 4bit (#524) · unsloth_zoo | 6.5 | 2.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) (#45544) · vllm | 5.8 | 1.5 |
| vllm-project/vllm | [Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) (#45544) · vllm | 5.8 | 1.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | Fix dense vLLM state dict parity · tests | 6.5 | 2.3 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Perf][1/N] Expand Triton kernel warmup coverage, DSv4 (#46634) · vllm | 9.5 | 5.3 |
| unslothai/unsloth-zoo | Fix dense vLLM state dict parity · tests | 6.5 | 2.3 |
| vllm-project/vllm | [Perf][1/N] Expand Triton kernel warmup coverage, DSv4 (#46634) · vllm | 9.5 | 5.3 |
| ggml-org/llama.cpp | sycl : support MUL_MAT and OUT_PROD with Q1_0 (#24721) · ggml | 9.7 | 5.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | sycl : support MUL_MAT and OUT_PROD with Q1_0 (#24721) · ggml | 9.7 | 5.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | Capture outputs fixes for transformers v5 (#713) · tests | 9.0 | 5.0 |
| unslothai/unsloth-zoo | Capture outputs fixes for transformers v5 (#713) · tests | 9.0 | 5.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | perf(moe-fp8): batched 3D dequant + FP8Experts dispatcher + Trainer guard · unsloth_zoo | 9.5 | 5.9 |
| unslothai/unsloth-zoo | perf(moe-fp8): batched 3D dequant + FP8Experts dispatcher + Trainer guard · unsloth_zoo | 9.5 | 5.9 |
| ggml-org/llama.cpp | ggml-webgpu: support non-square subgroup matrix configs for Intel GPUs (#21669) · ggml | 8.5 | 5.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(moe): loud-fail on silent fallbacks in MoE merge + FP8 forward paths · unsloth_zoo | 10.0 | 6.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | Fix review findings for PR #3: Gemma4 LoRA/BnB patches, GDN extraction, finalize_huggingface_model · unsloth_zoo | 9.5 | 6.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | ggml-webgpu: support non-square subgroup matrix configs for Intel GPUs (#21669) · ggml | 8.5 | 5.0 |
| unslothai/unsloth-zoo | fix(moe): loud-fail on silent fallbacks in MoE merge + FP8 forward paths · unsloth_zoo | 10.0 | 6.5 |
| unslothai/unsloth-zoo | Fix review findings for PR #3: Gemma4 LoRA/BnB patches, GDN extraction, finalize_huggingface_model · unsloth_zoo | 9.5 | 6.0 |
| ggml-org/llama.cpp | cuda: Q1_0 initial backend (#21629) · ggml | 9.7 | 6.3 |
| pooled: unsloth-zoo + vllm + llama.cpp | cuda: Q1_0 initial backend (#21629) · ggml | 9.7 | 6.3 |
| pooled: unsloth-zoo + vllm + llama.cpp | [XPU][MoE] Add WNA16 oracle backend for GPTQ sym-int4 (xpu_fused_moe) (#41426) · vllm | 9.3 | 5.9 |
| vllm-project/vllm | [XPU][MoE] Add WNA16 oracle backend for GPTQ sym-int4 (xpu_fused_moe) (#41426) · vllm | 9.3 | 5.9 |
| ggml-org/llama.cpp | [SYCL] Add BF16 support to GET_ROWS operation (#21391) · ggml | 10.0 | 7.3 |
| pooled: unsloth-zoo + vllm + llama.cpp | [SYCL] Add BF16 support to GET_ROWS operation (#21391) · ggml | 10.0 | 7.3 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix] Fix corrupt outputs in MoE FP8 LoRA responses and MoE base model responses when LoRAs are loaded (#42120) · tests | 3.5 | 0.8 |
| vllm-project/vllm | [Bugfix] Fix corrupt outputs in MoE FP8 LoRA responses and MoE base model responses when LoRAs are loaded (#42120) · tests | 3.5 | 0.8 |
| ggml-org/llama.cpp | ggml: `gguf_init_from_callback` and `gguf_init_from_buffer` (#22341) · ggml | 10.0 | 7.4 |
| pooled: unsloth-zoo + vllm + llama.cpp | ggml: `gguf_init_from_callback` and `gguf_init_from_buffer` (#22341) · ggml | 10.0 | 7.4 |
| ggml-org/llama.cpp | ggml: support concat for scalar types at cuda backend (#24011) · ggml | 10.0 | 7.5 |
| ggml-org/llama.cpp | [SYCL] support bf16 on bin_bcast OP and unary OPs (#24838) · ggml | 10.0 | 7.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | [DSv4 Perf] DSv4 flashinfer sparse index cache for metadata, 2%~4% TTFT improvement (#45863) · tests | 10.0 | 7.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | nixl_ep: Skip post-receive quantization for NVFP4 (#45606) · vllm | 9.0 | 6.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | ggml: support concat for scalar types at cuda backend (#24011) · ggml | 10.0 | 7.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | [SYCL] support bf16 on bin_bcast OP and unary OPs (#24838) · ggml | 10.0 | 7.5 |
| vllm-project/vllm | [DSv4 Perf] DSv4 flashinfer sparse index cache for metadata, 2%~4% TTFT improvement (#45863) · tests | 10.0 | 7.5 |
| vllm-project/vllm | nixl_ep: Skip post-receive quantization for NVFP4 (#45606) · vllm | 9.0 | 6.5 |
| ggml-org/llama.cpp | vulkan: Switch MUL_MAT_VEC to 4 K per iteration for F16/32 (#22887) · ggml | 10.0 | 8.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | vulkan: Switch MUL_MAT_VEC to 4 K per iteration for F16/32 (#22887) · ggml | 10.0 | 8.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | Harden fused-forward AST rewriter and adapter · unsloth_zoo | 7.5 | 9.2 |
| unslothai/unsloth-zoo | Harden fused-forward AST rewriter and adapter · unsloth_zoo | 7.5 | 9.2 |
| ggml-org/llama.cpp | sycl : fix failed ut cases of norm (#25044) · ggml | 10.0 | 8.5 |
| ggml-org/llama.cpp | ggml : fix ARM NEON nvfp4 dot product on non-dotprod targets (#21559) · ggml | 10.0 | 8.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(peft-param-wrapper): handle merge_and_unload for 4-bit MoE experts (B4) · unsloth_zoo | 10.0 | 8.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | sycl : fix failed ut cases of norm (#25044) · ggml | 10.0 | 8.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | ggml : fix ARM NEON nvfp4 dot product on non-dotprod targets (#21559) · ggml | 10.0 | 8.5 |
| unslothai/unsloth-zoo | fix(peft-param-wrapper): handle merge_and_unload for 4-bit MoE experts (B4) · unsloth_zoo | 10.0 | 8.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | feat(mlx): add save_method to save_pretrained_merged · unsloth_zoo | 10.0 | 8.6 |
| unslothai/unsloth-zoo | feat(mlx): add save_method to save_pretrained_merged · unsloth_zoo | 10.0 | 8.6 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(moe-quant): gate v5-only patches · tests | 10.0 | 8.7 |
| unslothai/unsloth-zoo | fix(moe-quant): gate v5-only patches · tests | 10.0 | 8.7 |
| ggml-org/llama.cpp | dflash: refactor draft model conversion (#25110) · conversion | 10.0 | 9.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(mlx): repair stub injection on Apple Silicon (3 sub-bugs) · unsloth_zoo | 10.0 | 9.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | Fp8 compressed (#358) · unsloth_zoo | 10.0 | 9.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix] Fix minimax_qk_norm_fusion (#44983) · vllm | 5.5 | 4.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | dflash: refactor draft model conversion (#25110) · conversion | 10.0 | 9.0 |
| unslothai/unsloth-zoo | fix(mlx): repair stub injection on Apple Silicon (3 sub-bugs) · unsloth_zoo | 10.0 | 9.0 |
| unslothai/unsloth-zoo | Fp8 compressed (#358) · unsloth_zoo | 10.0 | 9.0 |
| vllm-project/vllm | [Bugfix] Fix minimax_qk_norm_fusion (#44983) · vllm | 5.5 | 4.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix(moe-bnb): dequant Params4bit experts in transformers v5 grouped/batched MoE forward (B6) · unsloth_zoo | 10.0 | 9.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | Fix ModuleNotFoundError when loading gpt-oss models without triton_kernels (#4088) (#539) · unsloth_zoo | 10.0 | 9.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | [XPU] Fix Triton attn fp8/bf16 check failing (#45758) · vllm | 9.5 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Kernel][Bugfix] Fix INT8 per-token-head KV cache rounding in Triton reshape-and-cache (#45361) · tests | 9.5 | 10.0 |
| unslothai/unsloth-zoo | fix(moe-bnb): dequant Params4bit experts in transformers v5 grouped/batched MoE forward (B6) · unsloth_zoo | 10.0 | 9.5 |
| unslothai/unsloth-zoo | Fix ModuleNotFoundError when loading gpt-oss models without triton_kernels (#4088) (#539) · unsloth_zoo | 10.0 | 9.5 |
| vllm-project/vllm | [XPU] Fix Triton attn fp8/bf16 check failing (#45758) · vllm | 9.5 | 10.0 |
| vllm-project/vllm | [Kernel][Bugfix] Fix INT8 per-token-head KV cache rounding in Triton reshape-and-cache (#45361) · tests | 9.5 | 10.0 |
| ggml-org/llama.cpp | [SYCL] Add Q8_0 reorder optimization (~3x tg speedup on Intel Arc) (#21527) · ggml | 10.0 | 10.0 |
| ggml-org/llama.cpp | CUDA: Various fixes to `cpy.cu` (#25000) · ggml | 10.0 | 10.0 |
| ggml-org/llama.cpp | metal: Q1_0 backend (#21528) · ggml | 10.0 | 10.0 |
| ggml-org/llama.cpp | ggml: vectorize ggml_vec_dot_q4_1_q8_1 with WASM SIMD128 (#22209) · ggml | 10.0 | 10.0 |
| ggml-org/llama.cpp | fix: GLM-DSA crash in llama-tokenize when using vocab_only (#22102) · src | 10.0 | 10.0 |
| ggml-org/llama.cpp | metal : fix FA support logic (#21898) · ggml | 8.5 | 8.5 |
| pooled: unsloth-zoo + vllm + llama.cpp | Use torch.Tensor.reshape for non-contiguous tensor in ce loss function (#591) · unsloth_zoo | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix] Restrict FlashInfer cuDNN FP8 ViT attention gate to Blackwell (SM 100) (#45251) · vllm | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression (#46860) · tests | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Kernel] Enable TritonW4A16LinearKernel as CUDA fallback for non-Marlin-aligned W4A16 shapes (#43731) · vllm | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | Add weights padding for fp8 per-block online quantization (#44763) · vllm | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [Perf] Fix dsv3_router_gemm heuristic (#44217) · vllm | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | [SYCL] Add Q8_0 reorder optimization (~3x tg speedup on Intel Arc) (#21527) · ggml | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | CUDA: Various fixes to `cpy.cu` (#25000) · ggml | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | metal: Q1_0 backend (#21528) · ggml | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | ggml: vectorize ggml_vec_dot_q4_1_q8_1 with WASM SIMD128 (#22209) · ggml | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | fix: GLM-DSA crash in llama-tokenize when using vocab_only (#22102) · src | 10.0 | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | metal : fix FA support logic (#21898) · ggml | 8.5 | 8.5 |
| unslothai/unsloth-zoo | Use torch.Tensor.reshape for non-contiguous tensor in ce loss function (#591) · unsloth_zoo | 10.0 | 10.0 |
| vllm-project/vllm | [Bugfix] Restrict FlashInfer cuDNN FP8 ViT attention gate to Blackwell (SM 100) (#45251) · vllm | 10.0 | 10.0 |
| vllm-project/vllm | [Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression (#46860) · tests | 10.0 | 10.0 |
| vllm-project/vllm | [Kernel] Enable TritonW4A16LinearKernel as CUDA fallback for non-Marlin-aligned W4A16 shapes (#43731) · vllm | 10.0 | 10.0 |
| vllm-project/vllm | Add weights padding for fp8 per-block online quantization (#44763) · vllm | 10.0 | 10.0 |
| vllm-project/vllm | [Perf] Fix dsv3_router_gemm heuristic (#44217) · vllm | 10.0 | 10.0 |
| ggml-org/llama.cpp | model: mistral small 4 support (#20649) · convert_hf_to_gguf.py | - | 9.7 |
| ggml-org/llama.cpp | CUDA & CPU: support F32 kernel type for `CONV_TRANSPOSE_2D` (#17094) · ggml | - | 10.0 |
| ggml-org/llama.cpp | sycl: support reordered Q4_K/Q5_K/Q6_K MoE MUL_MAT_ID (#24452) · ggml | 9.7 | - |
| ggml-org/llama.cpp | sycl : fix the failed UT cases of conv_3d (#24900) · ggml | 10.0 | - |
| pooled: unsloth-zoo + vllm + llama.cpp | Fix LoRA scaling count mismatch on merge for Qwen2.5-VL exports (#2966) (#806) · tests | 8.3 | - |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix] Support non-power-of-2 top_k in legacy triton_kernels routing (#46406) · vllm | 10.0 | - |
| pooled: unsloth-zoo + vllm + llama.cpp | [Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 (#44692) · vllm | 10.0 | - |
| pooled: unsloth-zoo + vllm + llama.cpp | model: mistral small 4 support (#20649) · convert_hf_to_gguf.py | - | 9.7 |
| pooled: unsloth-zoo + vllm + llama.cpp | CUDA & CPU: support F32 kernel type for `CONV_TRANSPOSE_2D` (#17094) · ggml | - | 10.0 |
| pooled: unsloth-zoo + vllm + llama.cpp | sycl: support reordered Q4_K/Q5_K/Q6_K MoE MUL_MAT_ID (#24452) · ggml | 9.7 | - |
| pooled: unsloth-zoo + vllm + llama.cpp | sycl : fix the failed UT cases of conv_3d (#24900) · ggml | 10.0 | - |
| unslothai/unsloth-zoo | Fix LoRA scaling count mismatch on merge for Qwen2.5-VL exports (#2966) (#806) · tests | 8.3 | - |
| vllm-project/vllm | [Bugfix] Support non-power-of-2 top_k in legacy triton_kernels routing (#46406) · vllm | 10.0 | - |
| vllm-project/vllm | [Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 (#44692) · vllm | 10.0 | - |