01 / BASELINE · DIRECT INFERENCE
One path in.
A prompt goes straight to one model server. The simplest useful inference setup.
CLIENT → vLLM → QWEN 8B → GPU 0
↗
A model, a prompt.
One serving path from your prompt to the GPU.
A SMALL EXPERIMENT
NO BENCHMARK SAMPLES YET| METRIC | DIRECT | SCALED |
|---|---|---|
| Requests / sec | — | — |
| P50 latency | — | — |
| P95 latency | — | — |
| Errors | — | — |