-
Notifications
You must be signed in to change notification settings - Fork 1
Run a reproducible NVIDIA vLLM serving benchmark #8
Copy link
Copy link
Open
Labels
area: evaluationBenchmarks, quality evaluation, and result artifactsBenchmarks, quality evaluation, and result artifactsarea: vllmvLLM integration workvLLM integration workhelp wantedExtra attention is neededExtra attention is neededlaunch blockerRequired before performance-focused official launchRequired before performance-focused official launchneeds GPURequires reproducible GPU validationRequires reproducible GPU validation
Description
Activity
Metadata
Metadata
Assignees
Labels
area: evaluationBenchmarks, quality evaluation, and result artifactsBenchmarks, quality evaluation, and result artifactsarea: vllmvLLM integration workvLLM integration workhelp wantedExtra attention is neededExtra attention is neededlaunch blockerRequired before performance-focused official launchRequired before performance-focused official launchneeds GPURequires reproducible GPU validationRequires reproducible GPU validation
Produce the serving evidence required for an official performance launch after the vLLM integration lands.
Acceptance criteria:
This issue depends on the vLLM integration design and implementation. Comment with available hardware before claiming it.