DeepSeek V4 Flash Vision on two DGX Spark systems
Runtime recipe · DeepSeek
This is the public, parameterized form of the accepted two-node NVIDIA DGX Spark deployment. Its image contains the pinned local-inference-lab/vLLM fork rather than an arbitrary upstream vLLM release. The recipe fixes the model revision, engine flags, environment, topology, and validated image. You supply the details that belong to your cluster or Docker hosts.
ghcr.io/randomvariable/vllm-b12x-multi@sha256:42d5d72bd79a371c3b5a756fe2a1814f3c9393dd9a73d65bcefa01eb8f28d726Configure- Model
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
- Model revision
- 6821d6ad3681a4b137b066b76094fa82ebd0a380
- Topology
- 2 nodes · TP=2 · one DGX Spark per node
- Context
- 1,000,000 tokens
What is fixed
Section titled “What is fixed”The authored recipe preserves the accepted runtime command and tuning: B12X attention and linear backends, InstantTensor, a graph cap of 16 with capture sizes 1, 2, 3, 4, 8, 12, 16, GPU utilization 0.81, four sequences, and an 8,192-token batch budget. It does not present these constants as a general tuning matrix.
This is the LWS configuration I run. The Docker target uses the same engine definition, but I have not run it on two separate hosts.
The deployment benchmark guide explains how to run the benchmark against another deployment.
Measured performance
Section titled “Measured performance”Cold prefill: the earlier concurrency-one run processed the 134,240-token prompt in 73.72 seconds at 1,821 tokens/s. Its three subsequent requests used the warm prefix cache and averaged 0.94 seconds TTFT.
Warm concurrency sweep: the identical 707,468-character prompt had already been processed before every result in the table. Every completed request passed. Per-request decode slowed as concurrency increased. The TTFT cliff began at concurrency 8. Concurrency 16 was the largest complete batch; at 30 and above, every incomplete request stalled before its first token.
The benchmark reports this decode value as aggregate_gen_tok_s, calculated as total generated tokens divided by the sum of each completed request’s generation duration. Concurrent intervals overlap, so this is a duration-weighted per-request rate, not total server throughput.
The later scout completed in 0.67 seconds because it hit the warm prefix cache. Its apparent 200,114 tokens/s is not physical prefill throughput. This is a one-wave warm-start concurrency characterization, not a stable statistical performance baseline.
Requests enter at rank zero and collectives cross the secondary network; the shape is the same for every recipe here and is described in traffic paths in a two-node TP=2 group.
Manifest generator
Choose a target and supply the details that belong to your hosts or cluster. Values stay in this page and are never sent to a service.
Loading recipe definition...
No manifests generated
Complete the required fields, then select Generate manifests.
Before applying output
Section titled “Before applying output”- For Kubernetes, complete the cluster prerequisite checklist and confirm the NVIDIA runtime, LWS controller, CNI, Multus, and RDMA allocator are ready.
- For Docker, synchronize the model on both hosts before starting either serving container. Start rank zero first, then let the worker’s rendezvous check succeed before its headless server starts.
- Stop a GB10 serving container gracefully with a 120-second timeout. Long kernel execution is not grounds for force-killing a GPU process.
- A digest-qualified latest image is newer, not equivalent to the validated recipe image. Selecting it changes the validation status shown by the generator.