Skip to content

Run a reproducible NVIDIA vLLM serving benchmark #8

Description

@DaBestCode

Produce the serving evidence required for an official performance launch after the vLLM integration lands.

Acceptance criteria:

  • Compare dense vLLM, RandKV, and at least one scored-selector baseline at matched model, dtype, cache budget, request trace, and generation parameters.
  • Use a pinned container or lockfile and record GPU model/count, driver, CUDA, vLLM, PyTorch, model revision, and all command lines.
  • Include warm-up and repeated trials; synchronize timing where applicable.
  • Report request throughput, output-token throughput, TTFT, inter-token latency, peak GPU memory, and failure/OOM counts.
  • Publish every raw trial in machine-readable form plus the aggregation script.
  • Pair performance results with matched-budget quality results; do not optimize one by changing the other protocol.
  • State negative or inconclusive results without filtering them.

This issue depends on the vLLM integration design and implementation. Comment with available hardware before claiming it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: evaluationBenchmarks, quality evaluation, and result artifactsarea: vllmvLLM integration workhelp wantedExtra attention is neededlaunch blockerRequired before performance-focused official launchneeds GPURequires reproducible GPU validation

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions