| [2/N][Attention] Enable masked MHA for sparse MLA prefills (#48770) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: OpenAI Codex <noreply@openai.com> | 1 个月前 |
| Allow markdownlint to run locally (#36398) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> | 6 个月前 |
| [Sparse24] [Deprecation] Remove Sparse24 CT integration and kernels (#36799) Signed-off-by: Kyle Sayers <kylesayrs@gmail.com> | 5 个月前 |
| [Kernel] Fuse FP8 output quantization into merge_attn_states (#36518) Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com> | 5 个月前 |
| [Bugfix] Handle DeepseekV4ForCausalLM in benchmark_moe get_model_params (#52044) Co-authored-by: Do_it_now_! <23432123@users.noreply.github.com> | 1 个月前 |
| [Bench] benchmark_serving_multi_turn: make non-standard conversation_id payload opt-in (#43756) Signed-off-by: Change72 <cguo51@asu.edu> | 3 个月前 |
| [Chore]:Extract math and argparse utilities to separate modules (#27188) Signed-off-by: Yeshwanth Surya <yeshsurya@gmail.com> Signed-off-by: Yeshwanth N <yeshsurya@gmail.com> Signed-off-by: yeshsurya <yeshsurya@gmail.com> | 10 个月前 |
| [Kernel] ReplaySSM: cache SSM inputs for faster Mamba2 standard decode (#48018) Signed-off-by: Johnny-Liou <a897111@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk> | 1 个月前 |
| benchmarks: simplify test jsonschema (#14567) Signed-off-by: Russell Bryant <rbryant@redhat.com> | 1 年前 |
| [Docs] Update link to Benchmark CLI documentation (#33254) Signed-off-by: Eldar Kurtić <8884008+eldarkurtic@users.noreply.github.com> | 7 个月前 |
| [vLLM IR] Add IR op testing and benchmarking infrastructure (#40167) Signed-off-by: Yanan Cao <gmagogsfm@gmail.com> Co-authored-by: Theresa Shan <Theresa.Shan@amd.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> | 4 个月前 |
| Bump Transformers version to 5.10.4 (#41359) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> | 2 个月前 |
| [Chore] Update more locations to use attention_config.backend (#31153) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> | 8 个月前 |
| [Chore]:Extract math and argparse utilities to separate modules (#27188) Signed-off-by: Yeshwanth Surya <yeshsurya@gmail.com> Signed-off-by: Yeshwanth N <yeshsurya@gmail.com> Signed-off-by: yeshsurya <yeshsurya@gmail.com> | 10 个月前 |
| [Core] Add xxHash as a high-performance hash option for accelerating prefix caching (#29163) Signed-off-by: LuminolT <lumischen01@gmail.com> Signed-off-by: Lumis Chen <lumischen01@gmail.com> Co-authored-by: Russell Bryant <rbryant@redhat.com> | 9 个月前 |
| Hidden states extraction improvements (#43805) Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> | 3 个月前 |
| [CI/Build][Doc] Fully deprecate old bench scripts for serving / throughput / latency (#24411) Signed-off-by: Ye (Charlotte) Qi <yeq@meta.com> | 1 年前 |
| [Mypy] Better fixes for the mypy issues in vllm/config (#37902) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> | 5 个月前 |
| [Cleanup] Remove obsolete spec decoding compatibility logic (#32003) Signed-off-by: Nick Hill <nickhill123@gmail.com> | 8 个月前 |
| [Bug Fix] Allow pinned memory for WSL2 (#41496) Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com> | 3 个月前 |
| [Core] Add xxHash as a high-performance hash option for accelerating prefix caching (#29163) Signed-off-by: LuminolT <lumischen01@gmail.com> Signed-off-by: Lumis Chen <lumischen01@gmail.com> Co-authored-by: Russell Bryant <rbryant@redhat.com> | 9 个月前 |
| [Mypy] Better fixes for the mypy issues in vllm/config (#37902) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> | 5 个月前 |
| [Mypy] Better fixes for the mypy issues in vllm/config (#37902) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> | 5 个月前 |
| [CI/Build][Doc] Fully deprecate old bench scripts for serving / throughput / latency (#24411) Signed-off-by: Ye (Charlotte) Qi <yeq@meta.com> | 1 年前 |
| [Misc] Add common random prefix option to structured-output serving benchmark (#41632) Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> | 4 个月前 |
| [CI/Build][Doc] Fully deprecate old bench scripts for serving / throughput / latency (#24411) Signed-off-by: Ye (Charlotte) Qi <yeq@meta.com> | 1 年前 |
| Revert "[Platform] Replace torch.cuda.Event with torch.Event (#47140)" (#47668) Signed-off-by: Kunshang Ji <kunshang.ji@intel.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> | 2 个月前 |
| [Refactor] Remove dead or duplicate func utils or variables (#35318) Signed-off-by: yewentao256 <zhyanwentao@126.com> | 6 个月前 |
| [Core] Add kvcache watermark to reduce preemptions (#44594) Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 |
| [CI][BugFix] ShellCheck cleanup to remove baseline and preserve runtime behavior (#34514) Signed-off-by: junuxyz <216036880+junuxyz@users.noreply.github.com> | 6 个月前 |
| feat(benchmarks): Add Prefix Caching Benchmark to Serving Benchmark (#3277) | 2 年前 |