Environment
| item |
value |
| Host |
Dell XPS 15 7590, i7-9750H, NVIDIA GeForce GTX 1650 4 GB (Turing, sm_75), Ubuntu 26.04.1, kernel 7.0.0-31-generic |
| NVIDIA driver |
595.84 (open kernel module), CUDA 13.2 runtime reported by nvidia-smi |
| Docker |
Engine 29.8.0 (build 88096ef), Compose v5.5.1, containerd 2.3.4, nvidia-container-toolkit 1.20.0-1 |
| docker-model-plugin |
1.2.6-1ubuntu.26.04resolute (CLI v1.2.6) |
| Runner |
docker/model-runner:latest-cuda = sha256:a93d05a54b554fd2cc287206cef0eecdbf04cb8c51f14384c97db35a82634de1 (= tag v1.2.8-cuda), /version → v1.2.8, backend llama.cpp 72874f559 |
| Runner install |
docker model status on a fresh Docker CE host (auto-install), --gpu auto-detected as cuda; docker exec docker-model-runner nvidia-smi sees the GPU |
| Model |
ai/qwen3:4b-instruct-2507-q4_K_M = sha256:8cd2927dc3a33531aa57f2479e2dcfb16e3bf53e31de561a612fcbc3cf6a6893 (Qwen3-4B-Instruct-2507-Q4_K_M.gguf, 2.50 GB) |
Engine invocation (from ps) |
/app/llama-server -ngl 999 --metrics --model …/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --host inference-runner-0.sock --jinja (no --ctx-size; effective n_ctx 4096) |
| VRAM |
3096 MiB used by llama-server with the model loaded, all layers offloaded |
Server config (from docker model logs) |
load_model: initializing, n_slots = 4, n_ctx_slot = 4096, kv_unified = 'true'; during the bench the log shows decode: Context size has been exceeded. off = 0/1, n_batch = 1, ret = 1 followed by update_slots: decode() failed for pairs of task ids |
Reproduction A — the bench
docker model pull ai/qwen3:4b-instruct-2507-q4_K_M
docker model bench ai/qwen3:4b-instruct-2507-q4_K_M --concurrency 1,2 --duration 30s --json
Observed:
- 9 of 18 requests fail with
request failed during benchmark: request failed with status 500: {"error":{"code":500,"message":"Context size has been exceeded.","type":"server_error"}} (2 at concurrency 1, 7 at concurrency 2).
- The 9 requests that succeed stop at 120–137 tokens (a normal 100-word summary), so the model was producing coherent output at least part of the time.
- Bench JSON afterwards:
TPS 6.03 at c=1 and 3.21 at c=2, because failed requests' wall time is counted while their tokens are not.
- After the bench, every request to the still-loaded model returns only
? (0x3F) up to max_tokens, finish_reason: "length", never EOS, for example:
curl -s http://localhost:12434/engines/llama.cpp/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model":"ai/qwen3:4b-instruct-2507-q4_K_M",
"messages":[{"role":"user","content":"Reply with the single word OK."}],
"max_tokens":8,"temperature":0.2}'
→ {"content":"????????","finish_reason":"length", ...}
The same happens with temperature: 0, with a system prompt, and with the model addressed by tag, by docker.io/ prefix, or by @sha256: digest. A second model on the same runner (ai/nomic-embed-text-v1.5:137M-F16, embeddings) keeps working. docker model status still reports the runner healthy.
Reproduction B — a forced-continuation streaming request (no bench)
On a freshly loaded, verified-healthy model:
curl -s -N http://localhost:12434/engines/llama.cpp/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model":"ai/qwen3:4b-instruct-2507-q4_K_M",
"messages":[{"role":"user","content":"Write a comprehensive 100 word summary on whales and their impact on society."}],
"max_tokens":8192,"temperature":0.2,"ignore_eos":true,"stream":true}'
Observed: 4071 content deltas; deltas 1–6 are Wh|ales| play| a| vital| role, deltas 7–4071 are all ?. Stream ends with finish_reason: "length" (context full at 4096). Every subsequent request returns only ?, exactly as after the bench.
What does NOT reproduce it (each on a healthy model, each followed by an "OK" probe that passed)
- Non-streaming requests of any kind, including the same
ignore_eos + max_tokens 8192 request without stream (4071 coherent tokens, 0 ?, context filled to 4096, model healthy afterwards).
- Streaming with
max_tokens 8 or 160, and streaming with max_tokens 8192 but without ignore_eos (model stops at ~125 tokens).
- Client-cancelled streams (connection dropped after 4 s, ~150 deltas), single and two concurrent.
- Two concurrent bounded requests (121/124 tokens, 36 tok/s each).
So: ordinary streaming, cancellation, concurrency and a full context are harmless; the bench and forced-continuation streaming poison the loaded model. The exact request shape docker model bench sends was not reproduced from the API; the bench itself is the reliable reproducer.
Recovery
docker model unload --all, or the runner's idle unload (observed 251 s and 281 s after the last request). Reload after that takes < 2 s (GGUF in page cache) and the model answers normally again ("OK", "4", coherent 120–130-token summaries at ~40 tok/s).
Health probe used
{"messages":[{"role":"user","content":"Reply with the single word OK."}],"max_tokens":8,"temperature":0.2} → healthy iff content is OK and finish_reason is stop.
Environment
nvidia-smiubuntu.26.04resolute (CLIv1.2.6)docker/model-runner:latest-cuda=sha256:a93d05a54b554fd2cc287206cef0eecdbf04cb8c51f14384c97db35a82634de1(= tagv1.2.8-cuda),/version→v1.2.8, backend llama.cpp72874f559docker model statuson a fresh Docker CE host (auto-install),--gpuauto-detected as cuda;docker exec docker-model-runner nvidia-smisees the GPUai/qwen3:4b-instruct-2507-q4_K_M=sha256:8cd2927dc3a33531aa57f2479e2dcfb16e3bf53e31de561a612fcbc3cf6a6893(Qwen3-4B-Instruct-2507-Q4_K_M.gguf, 2.50 GB)ps)/app/llama-server -ngl 999 --metrics --model …/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --host inference-runner-0.sock --jinja(no--ctx-size; effective n_ctx 4096)docker model logs)load_model: initializing, n_slots = 4, n_ctx_slot = 4096, kv_unified = 'true'; during the bench the log showsdecode: Context size has been exceeded. off = 0/1, n_batch = 1, ret = 1followed byupdate_slots: decode() failedfor pairs of task idsReproduction A — the bench
Observed:
request failed during benchmark: request failed with status 500: {"error":{"code":500,"message":"Context size has been exceeded.","type":"server_error"}}(2 at concurrency 1, 7 at concurrency 2).TPS6.03 at c=1 and 3.21 at c=2, because failed requests' wall time is counted while their tokens are not.?(0x3F) up tomax_tokens,finish_reason: "length", never EOS, for example:The same happens with
temperature: 0, with a system prompt, and with the model addressed by tag, bydocker.io/prefix, or by@sha256:digest. A second model on the same runner (ai/nomic-embed-text-v1.5:137M-F16, embeddings) keeps working.docker model statusstill reports the runner healthy.Reproduction B — a forced-continuation streaming request (no bench)
On a freshly loaded, verified-healthy model:
Observed: 4071 content deltas; deltas 1–6 are
Wh|ales| play| a| vital| role, deltas 7–4071 are all?. Stream ends withfinish_reason: "length"(context full at 4096). Every subsequent request returns only?, exactly as after the bench.What does NOT reproduce it (each on a healthy model, each followed by an "OK" probe that passed)
ignore_eos+max_tokens 8192request withoutstream(4071 coherent tokens, 0?, context filled to 4096, model healthy afterwards).max_tokens8 or 160, and streaming withmax_tokens 8192but withoutignore_eos(model stops at ~125 tokens).So: ordinary streaming, cancellation, concurrency and a full context are harmless; the bench and forced-continuation streaming poison the loaded model. The exact request shape
docker model benchsends was not reproduced from the API; the bench itself is the reliable reproducer.Recovery
docker model unload --all, or the runner's idle unload (observed 251 s and 281 s after the last request). Reload after that takes < 2 s (GGUF in page cache) and the model answers normally again ("OK", "4", coherent 120–130-token summaries at ~40 tok/s).Health probe used
{"messages":[{"role":"user","content":"Reply with the single word OK."}],"max_tokens":8,"temperature":0.2}→ healthy iff content isOKandfinish_reasonisstop.