A reproducible, neutral benchmark for choosing a long-context serving stack on NVIDIA GB10. It compares vLLM, SGLang, and direct TensorRT-LLM serving with the same prompts, client-side metric formulas, cache protocols, and concurrency matrix.
The checked-in cross-engine values are preliminary one-repetition evidence. They are useful for portfolio review and experiment planning, but are not a final universal ranking. Run the three-repetition matrix before making production decisions. The separate TensorRT-LLM prefill-budget study uses three 100-request repetitions and profiler-backed evidence.
- Engine-neutral long-context workload construction and exact token-count validation.
- Cold unique-prefix versus warm shared-prefix cache experiments.
- TTFT, TPOT, ITL, E2E latency, request/output throughput, cache evidence, and host telemetry.
- Sequential Docker orchestration with pinned model, tokenizer, dataset, and image evidence.
- Open-loop offered-load testing with bounded admission and explicit SLA pass/fail decisions.
- Decision-oriented reporting that separates observations from hypotheses and profiler-backed conclusions.
For the supplied warm shared-prefix comparison at C1/C2/C4, TensorRT-LLM has the lowest reported TTFT and E2E latency and the highest reported request throughput. The supplied cross-engine data uses one repetition per configuration.
The TensorRT-LLM cold-C4 anomaly has now been isolated separately. Across three 100-request repetitions, reducing the prefill budget from 8192 to 2048 improved P95 TTFT by 26.9% and P95 ITL by 42.1%, with a 2.1% output-throughput reduction. Nsight traces showed the longest attention kernel falling from 474.804 ms to 126.315 ms and the longest CUDA synchronization wait falling from 5.792 s to 1.485 s. See the prefill-budget case study.
See results analysis, limitations, and the compact summary artifacts for the evidence status and units.
| Dimension | Default control |
|---|---|
| Model | openai/gpt-oss-20b |
| Hardware target | One NVIDIA GB10 / DGX Spark system |
| Prompt / output | Exactly 120,000 input tokens / 512 generated tokens |
| Samples | 100 canonical records |
| Modes | cold, warm_shared |
| Concurrency | 1, 2, 4 |
| Repetitions | 3 for the final matrix |
| Memory target | 0.80 engine-specific fraction |
| Client | One neutral OpenAI-compatible streaming client |
Scheduler, batching, kernel, cache-block, and memory-allocation semantics differ across runtimes. The controls are matched by intent and documented as such; they are not claimed to be mechanically identical.
Install the project and developer tools:
python3 -m pip install -e ".[dev]"Run a GPU/Docker smoke test across all three engines:
./scripts/smoke_test.sh --overwrite --cooldown-seconds 5Run the full three-engine matrix:
./bench run --engines all --modes cold,warm_shared \
--concurrency 1,2,4 --repetitions 3Capture an Nsight Systems trace for the TensorRT-LLM cold-C4 investigation:
./bench run --config config/tensorrt-llm-profile-c4.yaml \
--profile-nsys --overwriteSee the profiling runbook for the baseline-first workflow, generated artifacts, and the questions the trace must answer.
Run the initial production-capacity discovery sweep:
./bench capacity --config config/tensorrt-llm-capacity-120k.yaml \
--rates 0.008,0.012,0.016,0.020,0.024 --requests 20 \
--results-dir results/capacity/tensorrt-llm-120k-discovery \
--skip-image-pullThis starts a fresh server for every cold load point, generates arrivals independently of request completion, applies bounded admission, and reports the highest rate satisfying every SLA. See the capacity runbook.
The smoke and full commands perform real inference and require Docker, NVIDIA Container Toolkit, model/data access, and sufficient disk. Ordinary CI runs tests, dry-run orchestration, package checks, and manifest verification only; it does not run GPU inference.
Generate accepted-run tables from the reporting pipeline:
./bench report
python3 scripts/generate_charts.py --input results/report/summary.csvThe report writes an autogenerated marker and results/report/provenance.json containing the benchmark commit, experiment-lock hash, accepted runs, and accepted repetitions. The supplied portfolio charts are stored under assets/charts/ with their compact source CSVs and preliminary-results labeling. The chart generator remains available for regenerating SVG decision views from accepted report summaries.
Do not hand-edit public numeric tables. The intended evidence flow is:
accepted run JSON + request timings
-> ./bench report
-> summary CSV + provenance
-> scripts/generate_charts.py
-> charts and portfolio report
- Methodology and fairness contract
- Architecture
- Results and analysis
- Profiling plan
- Prefill-budget case study
- SLA-constrained capacity runbook
- Context-length characterization runbook
- Context campaigns support immutable-contract resume and compatible multi-source
consolidation through
scripts/analyze_context_scaling.py. - Fairness and limitations
- Production recommendations
- Reproducibility and implementation details
- Result schema
- Design-to-implementation mapping
TensorRT-LLM uses the direct trtllm-serve backend with the same prepared prompts and neutral client. Triton integration is intentionally deferred and documented as planned work; it is not represented as a completed experiment.
GitHub Actions runs on Python 3.10, 3.11, and 3.12. It installs .[dev], runs pytest, validates the CLI and three-engine dry-run plan, and verifies PROJECT_MANIFEST.sha256. See .github/workflows/ci.yml.
The repository excludes model weights, Hugging Face caches, TensorRT engines, server logs, raw telemetry, and request JSONL from version control. Large evidence bundles belong in GitHub Releases with their lock file, image digests, environment manifest, and benchmark commit.
- vLLM integration
- SGLang integration
- TensorRT-LLM direct integration
- 120K cold and warm shared-prefix workloads
- Cache evidence and result provenance
- CI and packaging validation
- Three repetitions per matrix cell
- Nsight Systems capture integration
- Profiler-backed root cause for the TensorRT-LLM cold-C4 ITL anomaly
- TensorRT-LLM through Triton
- Context scaling: 8K, 32K, 64K, 120K (three-repetition validation)
- SLA-constrained throughput with validated production capacity evidence
Maintainer: Venkata Rami Reddy Kallu. See CITATION.cff for citation metadata. The project is licensed under MIT.