Skip to content

DeepSeek V4 Flash Vision on two DGX Spark systems

Runtime recipe · DeepSeek

This is the public, parameterized form of the accepted two-node NVIDIA DGX Spark deployment. Its image contains the pinned local-inference-lab/vLLM fork rather than an arbitrary upstream vLLM release. The recipe fixes the model revision, engine flags, environment, topology, and validated image. You supply the details that belong to your cluster or Docker hosts.

Validated imageghcr.io/randomvariable/vllm-b12x-multi@sha256:42d5d72bd79a371c3b5a756fe2a1814f3c9393dd9a73d65bcefa01eb8f28d726Configure
Model
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Model revision
6821d6ad3681a4b137b066b76094fa82ebd0a380
Topology
2 nodes · TP=2 · one DGX Spark per node
Context
1,000,000 tokens

The authored recipe preserves the accepted runtime command and tuning: B12X attention and linear backends, InstantTensor, a graph cap of 16 with capture sizes 1, 2, 3, 4, 8, 12, 16, GPU utilization 0.81, four sequences, and an 8,192-token batch budget. It does not present these constants as a general tuning matrix.

This is the LWS configuration I run. The Docker target uses the same engine definition, but I have not run it on two separate hosts.

The deployment benchmark guide explains how to run the benchmark against another deployment.

Estonia profile134,240 prompt tokensBenchmark v0.6.2Warm startsOne wave per concurrency
ConcurrencyStart time / TTFTPer-request decodeMean requestCompleted
173.72 sCold prefillWarm TTFT: 0.72 s47.93 tok/s83.88 s1 / 1
21.33 s39.61 tok/s107.78 s2 / 2
41.76 s26.56 tok/s99.15 s4 / 4
853.07 s26.17 tok/s168.75 s8 / 8
16251.78 s22.39 tok/s415.30 s16 / 16
30281.51 s22.91 tok/s404.02 s21 / 30
48253.56 s22.93 tok/s421.60 s17 / 48
64242.58 s22.69 tok/s353.98 s24 / 64

Cold prefill: the earlier concurrency-one run processed the 134,240-token prompt in 73.72 seconds at 1,821 tokens/s. Its three subsequent requests used the warm prefix cache and averaged 0.94 seconds TTFT.

Warm concurrency sweep: the identical 707,468-character prompt had already been processed before every result in the table. Every completed request passed. Per-request decode slowed as concurrency increased. The TTFT cliff began at concurrency 8. Concurrency 16 was the largest complete batch; at 30 and above, every incomplete request stalled before its first token.

The benchmark reports this decode value as aggregate_gen_tok_s, calculated as total generated tokens divided by the sum of each completed request’s generation duration. Concurrent intervals overlap, so this is a duration-weighted per-request rate, not total server throughput.

The later scout completed in 0.67 seconds because it hit the warm prefix cache. Its apparent 200,114 tokens/s is not physical prefill throughput. This is a one-wave warm-start concurrency characterization, not a stable statistical performance baseline.

Requests enter at rank zero and collectives cross the secondary network; the shape is the same for every recipe here and is described in traffic paths in a two-node TP=2 group.

Manifest generator

Choose a target and supply the details that belong to your hosts or cluster. Values stay in this page and are never sent to a service.

Loading recipe definition...

Deployment target

This is the two-node LeaderWorkerSet deployment I run.

Container image

Loading deployment parameters...

Required values gate all copy and download controls. No partial deployment is offered.

Rendered output

Operator steps

Current LWS setup

No manifests generated

Complete the required fields, then select Generate manifests.

  • For Kubernetes, complete the cluster prerequisite checklist and confirm the NVIDIA runtime, LWS controller, CNI, Multus, and RDMA allocator are ready.
  • For Docker, synchronize the model on both hosts before starting either serving container. Start rank zero first, then let the worker’s rendezvous check succeed before its headless server starts.
  • Stop a GB10 serving container gracefully with a 120-second timeout. Long kernel execution is not grounds for force-killing a GPU process.
  • A digest-qualified latest image is newer, not equivalent to the validated recipe image. Selecting it changes the validation status shown by the generator.