Qwen3.8-27B · RTX PRO 6000 Blackwell · 9 October 2026
The same model, the same prompts, the same GPU. Three ways of serving it.
Answer
vLLM FP8
50.6tok/s
1,000 tok/s total at 25 users
Ollama Q8_0
47.3tok/s
165 tok/s total at 25 users
vLLM BF16
28.8tok/s
605 tok/s total at 25 users
Single-request speed is close between the two 8-bit formats and far below on BF16. Under concurrent load the gap widens: vLLM FP8 delivers 6.1× Ollama's throughput at 25 users, because Ollama's slot count is fixed at load time and it stops improving past 8.
Prompt 1 of 3
[req <unique id>] Describe the purpose of installation qualification in two paragraphs.
| Backend | Prompt tokens | Output tokens | tok/s, 1 user | tok/s total, 25 users | First token |
|---|---|---|---|---|---|
| vLLM FP8 | 67 | 200 | 50.6 | 1,000.0 | 0.03 s |
| Ollama Q8_0 | 25 | ≤200 | 47.3 | 165.1 | 0.30 s |
| vLLM BF16 | 67 | 200 | 28.8 | 605.1 | 0.06 s |
Prompt 2 of 3
[req <unique id>] You are a pharmaceutical validation assistant. Summarise the key qualification steps implied by the data below in flowing prose. step_id|parameter|setpoint|measured|tolerance|result IQ-00000|chamber_temp_0|20.0|20.0|±0.5|PASS IQ-00001|chamber_temp_1|21.1|21.3|±0.5|PASS IQ-00002|chamber_temp_2|22.2|22.6|±0.5|PASS … 122 more rows of the same shape
| Backend | Prompt tokens | Output tokens | tok/s, 1 user | tok/s total, 25 users | First token |
|---|---|---|---|---|---|
| vLLM FP8 | 4,077 | 500 | 50.4 | 529.4 | 0.42 s |
| Ollama Q8_0 | 4,035 | ≤500 | 46.8 | 117.6 | 1.34 s |
| vLLM BF16 | 4,077 | 500 | 28.7 | 328.3 | 0.72 s |
Prompt 3 of 3
[req <unique id>] You are a pharmaceutical validation assistant. Summarise the key qualification steps implied by the data below in flowing prose. step_id|parameter|setpoint|measured|tolerance|result IQ-00000|chamber_temp_0|20.0|20.0|±0.5|PASS … 3,139 more rows, the size of a real IOPQ workbook
| Backend | Prompt tokens | Output tokens | tok/s, 1 user | tok/s total, 25 users | First token, 25 users |
|---|---|---|---|---|---|
| vLLM FP8 | 100,247 | 500 | 45.9 | 23.8 | 254 s |
| Ollama Q8_0 | 100,205 | ≤500 | 39.5 | 9.8 | 707 s |
| vLLM BF16 | 100,247 | 500 | 27.2 | 17.5 | 360 s |
At this prompt size reading the document dominates, so total throughput stops rising with more users and the wait before the first token grows instead. Reading rate: 5,298 tok/s on FP8, 3,890 on BF16, 2,369 on Ollama.
How it was run
Every figure comes from a real HTTP request to a running server that read the prompt and generated the tokens. Nothing is simulated, extrapolated or taken from a datasheet. The responses were streamed, so the time to the first token is separated from the generation rate that follows.
Each ran alone with the others stopped, since they cannot share the GPU. Sampling was pinned to temperature 0 with neutral top_p and top_k on all three, because the model ships a config that vLLM honours and Ollama does not. Each concurrent request carried a unique leading token, so no server could answer from a cached copy of a prompt it had already seen. A warmup request preceded every measurement and was discarded.
vLLM can be told to generate an exact token count, so its runs produced precisely 200 or 500. Ollama has no equivalent and treats the number as a ceiling, so a response may stop earlier. Generation rate is therefore computed per request as tokens divided by that request's own generation time, which makes it independent of how long each answer ran.
Not tested
Only three of the four possible combinations were measured. BF16 was tested on vLLM because that is the configuration under consideration; nobody proposed running BF16 through Ollama, so it was left out rather than spending four hours downloading a second 56 GB copy of the weights in a different file format.
It would serve as a control, showing that the BF16 penalty is the memory bus rather than something specific to vLLM. The existing numbers already point that way: BF16 is slower than FP8 at reading long prompts, which is a compute-bound task where a larger weight format has no excuse. If the question comes up, the run takes about an hour of machine time.