Qwen3.8-27B · RTX PRO 6000 Blackwell · 9 October 2026

Which Backend Is Faster

The same model, the same prompts, the same GPU. Three ways of serving it.

Answer

vLLM with FP8 weights, at every prompt size and every load level.

Fastest

vLLM FP8

50.6tok/s

1,000 tok/s total at 25 users

Ollama Q8_0

47.3tok/s

165 tok/s total at 25 users

vLLM BF16

28.8tok/s

605 tok/s total at 25 users

Single-request speed is close between the two 8-bit formats and far below on BF16. Under concurrent load the gap widens: vLLM FP8 delivers 6.1× Ollama's throughput at 25 users, because Ollama's slot count is fixed at load time and it stops improving past 8.

Prompt 1 of 3

Short

200 output tokens requested
[req <unique id>] Describe the purpose of installation qualification in two paragraphs.
One request at a time, then the same prompt from 25 clients at once.
BackendPrompt tokensOutput tokens tok/s, 1 usertok/s total, 25 usersFirst token
vLLM FP8 6720050.61,000.00.03 s
Ollama Q8_0 25≤20047.3165.10.30 s
vLLM BF16 6720028.8605.10.06 s

Prompt 2 of 3

Medium

~4,000 prompt tokens · 500 output tokens requested
[req <unique id>] You are a pharmaceutical validation assistant. Summarise the key
qualification steps implied by the data below in flowing prose.

step_id|parameter|setpoint|measured|tolerance|result
IQ-00000|chamber_temp_0|20.0|20.0|±0.5|PASS
IQ-00001|chamber_temp_1|21.1|21.3|±0.5|PASS
IQ-00002|chamber_temp_2|22.2|22.6|±0.5|PASS
… 122 more rows of the same shape
Pipe-delimited rows, the shape the platform's spreadsheet attachments take.
BackendPrompt tokensOutput tokens tok/s, 1 usertok/s total, 25 usersFirst token
vLLM FP8 4,07750050.4529.40.42 s
Ollama Q8_0 4,035≤50046.8117.61.34 s
vLLM BF16 4,07750028.7328.30.72 s

Prompt 3 of 3

Long document

~100,000 prompt tokens · 500 output tokens requested
[req <unique id>] You are a pharmaceutical validation assistant. Summarise the key
qualification steps implied by the data below in flowing prose.

step_id|parameter|setpoint|measured|tolerance|result
IQ-00000|chamber_temp_0|20.0|20.0|±0.5|PASS
… 3,139 more rows, the size of a real IOPQ workbook
The regime the platform actually runs in. Note the first-token column.
BackendPrompt tokensOutput tokens tok/s, 1 usertok/s total, 25 usersFirst token, 25 users
vLLM FP8 100,24750045.923.8254 s
Ollama Q8_0 100,205≤50039.59.8707 s
vLLM BF16 100,24750027.217.5360 s

At this prompt size reading the document dominates, so total throughput stops rising with more users and the wait before the first token grows instead. Reading rate: 5,298 tok/s on FP8, 3,890 on BF16, 2,369 on Ollama.

How it was run

Real prompts, over the network, to each live server

Were actual prompts fed to the models? Yes.

Every figure comes from a real HTTP request to a running server that read the prompt and generated the tokens. Nothing is simulated, extrapolated or taken from a datasheet. The responses were streamed, so the time to the first token is separated from the generation rate that follows.

Were the three backends measured the same way?

Each ran alone with the others stopped, since they cannot share the GPU. Sampling was pinned to temperature 0 with neutral top_p and top_k on all three, because the model ships a config that vLLM honours and Ollama does not. Each concurrent request carried a unique leading token, so no server could answer from a cached copy of a prompt it had already seen. A warmup request preceded every measurement and was discarded.

Why is Ollama's output column "≤"?

vLLM can be told to generate an exact token count, so its runs produced precisely 200 or 500. Ollama has no equivalent and treats the number as a ceiling, so a response may stop earlier. Generation rate is therefore computed per request as tokens divided by that request's own generation time, which makes it independent of how long each answer ran.

Not tested

One combination was not measured

BF16 on Ollama

Only three of the four possible combinations were measured. BF16 was tested on vLLM because that is the configuration under consideration; nobody proposed running BF16 through Ollama, so it was left out rather than spending four hours downloading a second 56 GB copy of the weights in a different file format.

It would serve as a control, showing that the BF16 penalty is the memory bus rather than something specific to vLLM. The existing numbers already point that way: BF16 is slower than FP8 at reading long prompts, which is a compute-bound task where a larger weight format has no excuse. If the question comes up, the run takes about an hour of machine time.