← Blog

Running DeepSeek-V4-Flash across four free GPUs

A runbook for serving DeepSeek-V4-Flash on 4× L4 when the model is nearly twice your total VRAM — seven steps, with every FreeToken flag, startup log line and real error explained.

Title card: DeepSeek-V4-Flash across 4 free GPUs. A 167 GB MoE model streams experts from host RAM into four NVIDIA L4 cards with ~96 GB of VRAM between them, at 20.37 tok/s.

TL;DR

  • What: DeepSeek-V4-Flash (167 GB) served across 4× NVIDIA L4 (~96 GB VRAM total) on Kaggle's free tier — 20.37 tok/s, TTFT 6.11 s, 1331-token prompts prefilled in 16.3 s, model loaded in 8.8 min.
  • Why it fits: ~92% of the checkpoint is routed experts, and top_k 8 of 256 means a token touches ~3% of them. The experts stay in host RAM (~154 GB of 188) and stream over PCIe into a 4080-slot LRU cache in VRAM. Only 3.38 GiB of dense weights is resident per GPU.
  • Getting the GPUs: four L4s are not in Kaggle's accelerator dropdown. Attach a competition, set machine_shape: NvidiaL4, push from the CLI. kernel_type must be notebook or you silently get one P100 — and the box is then offline by construction, so everything has to be pre-staged.
  • Engine: FreeToken, because vLLM's expert offload is still an open RFC and nothing else supports DSV4's architecture. Released 0.1.2 refuses TP>1, so this needs PR #70 — plus a backport of PR #243, or long prompts hang forever on sm_89.
  • The four traps: --num-tokens silently overrides the memory solver instead of constraining it; a present-but-unimportable sgl_kernel kills all four workers on any prompt ≥ 32 tokens; flashinfer imports cleanly and then cannot JIT; and this is a reasoning model, so tokens arrive as reasoning_content, not content.
  • The ceiling: ~1.05 GB crosses PCIe per token per rank — about 85% of the time budget. More GPUs will not help; a better interconnect or a higher cache hit rate will.

Reproduce it: three Kaggle notebooks and kaggle kernels push -p .

What this post covers

Sooner or later everyone running big models hits the same wall: the checkpoint is larger than the GPU. The usual answer is to add GPUs, and that part everybody knows. What is less obvious is what to do when adding every GPU you can still leaves you short — four cards, 96 GB of VRAM, and a 167 GB model.

This is a runbook for exactly that case. It walks from an empty notebook to a served model, explains every flag in the launch command, shows the startup log lines that tell you whether your configuration is sane, and ends with the six errors you will actually hit and what each one means. It assumes you can read a stack trace and have used a serving engine before. It does not assume you know anything about mixture-of-experts offload.

Every result here was measured on Kaggle's free tier. Where a figure is borrowed or approximate the text says so. The serving notebook is at kaggle.com/code/hisiterbk/dsv4-tp4 — fork it and you get the launcher, the health-poll loop and the benchmark cells as they actually ran.

Three places things run, and the post labels every block with which.

  • Your machine — the kaggle CLI and the three kernel-metadata.json files. Nothing heavy happens here; you are only pushing folders.
  • Two build notebooks on Kaggle (ft-wheels, ft-pr70) — CPU, internet on. These produce the offline wheel bundle and the patched engine wheel.
  • One serve notebook on Kaggle (dsv4-tp4) — 4× L4, internet off. Everything from Step 4 onward is a cell in this notebook, including the commands written with a $ prompt.

Install pip install kaggle and authenticate with kaggle auth login before anything else.

Getting four L4s for free

What didn’t work
  • Picking the L4 from the accelerator dropdown. It is not in the list, and no amount of clicking puts it there.
  • kernel_type: "script". The push succeeds, the kernel runs, and you silently get one P100 — machine_shape is ignored entirely.
  • competition_sources with enable_internet: true. Hard 400 at push time. You cannot have the accelerator and a network.

This is the part most people do not know exists, so it gets its own section.

Kaggle's free accelerator dropdown offers a P100, two T4s, or a TPU. Four L4s are not in that list — the shape exists, but you cannot pick it in the UI. You unlock it by attaching a competition that permits it and then requesting the shape by name in notebook metadata, pushed from the CLI.

Four conditions, all of which have to hold at once:

File on your machine — dsv4-tp4/kernel-metadata.json

{
  "id": "<owner>/dsv4-tp4",
  "code_file": "tp4.ipynb",
  "kernel_type": "notebook",      <- 1. "script" silently downgrades to P100
  "enable_gpu": "true",
  "machine_shape": "NvidiaL4",    <- 2. the shape, requested by name
  "competition_sources": ["arc-prize-2026-arc-agi-2"],  <- 3. what unlocks it
  "enable_internet": "false",     <- 4. forced by the line above
  "kernel_sources": ["<owner>/ft-wheels", "<owner>/ft-pr70"],
  "model_sources": ["dangkhoa2016/deepseek-ai-deepseek-v4-flash-0731/transformers/default/1"]
}

Your machine — shell

$ kaggle kernels push -p .
Kernel version 9 successfully pushed.

Why each condition matters

  • kernel_type must be notebook. A script kernel ignores machine_shape entirely and hands you one P100. No warning and no error — the push succeeds and you only discover it from nvidia-smi. This is the single easiest way to waste a run.
  • Join the competition first, then attach it. Open its page, accept the rules, and only then does competition_sources work — a push referencing a competition you have not joined is rejected. Joining alone changes nothing either: it is the competition_sources entry on the kernel that flips the allocation. You need both.
  • enable_internet must be false. This is the real price. competition_sources together with enable_internet: true is rejected with a hard 400 at push time, so the box is offline by construction and every wheel, model and patch has to arrive pre-staged through kernel_sources or model_sources. Step 1 covers how.
  • Push from the CLI. The web editor will not offer you a shape that is not in its dropdown.

What lands

ResourceFree tier, L4 shape
GPUs4 × L4, 23,034 MiB each — ~96 GB (90 GiB)
Architecturesm_89 (Ada), driver 580.159.04
Host RAM188 GB (cgroup limit 173 GiB)
CPU46 vCPU
Scratch disk~1 TB on /kaggle/tmp
InterconnectPCIe, no NVLink
Model storageNFSv3 mount, read-only
Session limit12 hours

The 188 GB of host RAM is the underrated half of that table. It is what makes expert offload viable at all — it holds the ~154 GB of expert banks that will never fit in VRAM. A four-GPU box with 32 GB of system memory would not serve this model no matter how the VRAM was arranged.

Two honest caveats. We verified this with arc-prize-2026-arc-agi-2 and did not test every competition — not all of them permit the L4 shape, so check before you build around one. And the unlock changes which accelerator you may request, not how much of it you get: Kaggle's weekly GPU quota and the 12-hour session cap still apply.

The model

One detail from that table matters more than it looks: these GPUs are not connected by NVLink. Every cross-GPU collective and every expert fetch crosses PCIe. That shapes which parallelism strategy is worth using and, as the arithmetic below shows, sets the ceiling on decode speed.

The model is DeepSeek-V4-Flash-0731: 43 layers, 256 routed experts per layer, top_k 8. Dense weights are FP8; the routed experts are FP4, packed into I8 tensors with F8_E8M0 scale tensors alongside. On disk it is 48 safetensors shards, about 156 GiB. It advertises a 1,048,576-token context.

Will it fit? The ten-minute check

What didn’t work
  • The obvious arithmetic. 167 GB of weights against 96 GB of VRAM says no — and it is right, if every byte has to be resident.
  • Our first estimate of the non-expert weights. We guessed around 30 GB and concluded four GPUs were needed just to hold them. The real figure is 13.5 GiB in total, 3.38 GiB per rank — comfortable on a single card. That wrong number shaped the whole plan.

The arithmetic everybody does first is the one that says no. 167 GB of weights, 96 GB of VRAM across four cards. Tensor parallelism splits a model but does not shrink it — every byte still has to be resident on some GPU. On that math you are 71 GB short and there is no flag that fixes it.

The arithmetic that says yes starts from sparsity. With top_k 8 out of 256 experts, any given token touches about 3% of the expert weights. The other 97% sit idle in VRAM doing nothing but occupying space.

So do not put them in VRAM. FreeToken's offload MoE backend keeps the expert banks in host RAM and maintains a fixed-size LRU slot cache on the GPU, streaming across PCIe the experts each token actually selects. The fit check becomes:

What lives wherePer GPUAcross 4 GPUs
Dense weights (TP-sharded)3.38 GiB13.5 GiB
MoE slot cache (tunable)~14 GiB~56 GiB
KV pool (replicated)1.08 GiB1.08 GiB each
CUDA graphs + headroom~3.5 GiB—
Expert bankshost RAM~154 GB of 188 GB

The question changes from does the model fit to how much of it can you keep hot. That is a far better question, because it has a dial on it. Your real constraint is now host RAM: you need enough to hold the expert banks, and 188 GB comfortably holds ~154 GB.

The ten-minute version: add up your dense (non-expert) weights and divide by your GPU count. If that plus a few GiB of KV and graph headroom fits on one card, you can serve the model regardless of how large the experts are — provided host RAM holds the experts. For us: 13.5 GiB ÷ 4 = 3.4 GiB per card against 22.5 GiB usable. Comfortable.

The engine, and why not vLLM

What didn’t work

Four engines, before the one that worked.

  • Every general-purpose engine we would reach for first. vLLM has no expert-aware offload yet, SGLang none at all, llama.cpp wants GGUF, KTransformers wants Intel AMX.
  • Released FreeToken 0.1.2. Even the engine built for this refuses: TP > 1 is not supported for this expert format.

We served this with FreeToken, which is not the obvious choice. The obvious choice is vLLM. Here is why it was not.

What the expert cache actually does

FreeToken's offload backend keeps the expert banks — about 154 GB — in pinned host RAM and reserves a fixed number of expert slots in VRAM — not layers, not experts, slots. Each decode step, the router picks its top_k 8 experts per layer, the cache remaps those expert ids onto slots, and only the misses are fetched across PCIe:

FreeToken source — reference, not something you run

cache.ensure_experts(self.layer_id, topk_ids)  # in-place expert-id -> slot
cache.copy_missing()                           # only the misses cross PCIe
return self._expert_gemm(cache, hidden_states, topk_weights, topk_ids,
                         views=cache.bank_views(), ...)
Router — replicated, so every rank picks the same experts top_k 8 of 256, once per layer, 43 layers = 344 lookups per token Host RAM expert banks ~154 GB 11,008 experts, pinned TP-sharded per rank VRAM slot cache 4080 slots, LRU ~37% resident 8–18% decode hit rate FP4 expert GEMM on the GPU PCIe misses only hits: no transfer Per token, per rank: 344 lookups × ~3.5 MB per expert slice = ~1.2 GB if every one missed At a 13% hit rate that is ~1.05 GB across PCIe, for every single token
One decode step. The expert banks never move; only the misses do.

Eviction is LRU. The cache is global, not per-layer: 4080 slots shared across all 43 layers, against 11,008 layer-expert pairs in total (43 × 256), so about 37% of the model's experts can be resident at once — roughly 95 per layer on average, though the LRU is free to spend them unevenly. The num_experts minimum the engine enforces is a global floor of one layer's worth, not a per-layer guarantee. Two things follow from that design, and both matter more than the slot count:

  • Prefill and decode take different paths. Above a crossover of T × top_k ≥ num_experts, prefill switches from on-demand slot fetches to streaming whole expert layers, because past that point a chunk touches most of the layer anyway. Step 7 covers what happens when that second path is broken.
  • Residency is not hit rate. We did not measure our own hit rate — this figure is borrowed. FreeToken's issue tracker reports LRU realising roughly 8–18% decode hit rate on DSV4's fine-grained MoE — 256 small experts per layer scatter far worse than the 8–16 large experts of a Mixtral-style model. Holding 37% of the experts does not mean hitting 37% of the time.

If it holds here, that — not the GPU count — is what sets our 20 tok/s ceiling, because every miss is a PCIe round trip on a box with no NVLink. Measuring it on your own workload is the first thing worth doing.

The arithmetic that explains 20 tok/s

That last figure is worth carrying through, because it decides whether this approach suits your hardware. The chain is short, and every input is either measured or stated:

StepWorkingResult
One expert, all ranks154 GB ÷ (43 × 256)14.0 MB
One expert, this rank÷ 4 (TP-sharded)3.5 MB
Lookups per token43 layers × top_k 8344
If every one missed344 × 3.5 MB1.20 GB
At a 13% hit rate× 0.871.05 GB
Over PCIe Gen4 ×16÷ ~25 GB/s effective42 ms
Predicted1000 ÷ 42~24 tok/s
Measured226 tokens, timed client-side20.37 tok/s

Across the whole 8–18% hit-rate range the prediction lands between 22.6 and 25.3 tok/s, against 20.37 measured — so expert traffic accounts for roughly 85% of the per-token time budget. That is the claim "the limiter is PCIe, not compute" made checkable instead of asserted.

It also tells you how to shop. Doubling the GPU count buys nothing here; halving the bytes per miss, or raising the hit rate, buys almost everything. An NVLink box, or one with enough VRAM to hold a larger slice of the experts, moves this number. A faster GPU does not.

Assumptions, since two of these are not measured. The 8–18% hit rate is borrowed from FreeToken's issue tracker, not observed on our run. The ~25 GB/s is a reasonable effective figure for PCIe Gen4 ×16; we did not confirm the link width, generation, or whether the four cards contend for one root complex. Treat this as an order-of-magnitude argument that happens to agree with the measurement, not as a second measurement.

Why not vLLM, llama.cpp or KTransformers

EngineExpert offload todayWhy not here
vLLM --cpu-offload-gb is generic layer-wise offload, not expert-aware. Routing-driven expert cache with LFRU eviction is RFC #38256, still open PR 1 of 3 is unmerged and requires --enforce-eager, so no CUDA graphs; PRs 2–3 unimplemented. No DSV4 architecture support
SGLang No native expert offload; tracked as a KTransformers-integration request Same architecture gap
llama.cpp Real and shipping — --n-cpu-moe puts expert tensors in system RAM GGUF only, so the FP4 checkpoint would need converting; tensor parallelism across 4 cards is not its strength
KTransformers Purpose-built for this, with AMX/AVX-512 CPU expert kernels Its advantage is CPU-side compute on Intel AMX hardware; ours is a rented Kaggle box, and multi-GPU TP is not its main path

The decisive reason is narrower than any of those rows, though: architecture support. DeepSeek-V4-Flash is not a Mixtral clone. It needs MLA latent-KV attention, a Lightning Indexer, sparse attention, a sliding-window radix KV cache on 128-token window pages, and FP4 experts packed as I8 with F8_E8M0 scales. FreeToken ships a DSV4-specific attention backend (dsv4_sparse), KV pool and cost model. A generic engine would need all of that written first, and expert offload on top.

This is a selection rationale, not a benchmark. We did not run vLLM, llama.cpp or KTransformers on this checkpoint — in most cases we could not have, which is the point. Treat the table as "why we started here", and expect the vLLM row to age quickly once #38256 lands.

Step 1: Getting the model onto the machine

What didn’t work
  • Resolving the wheel closure on macOS. pip evaluated the Linux markers against the laptop and produced a complete, wrong wheel set with no error.
  • Forcing it with --platform + --only-binary=:all:. Aborted on one tag mismatch — most CUDA wheels are manylinux_2_17, nvidia-cuda-cupti is manylinux_2_25.
  • Docker under linux/amd64 emulation. Ten minutes, zero bytes downloaded.
  • Shipping the wheels as a Kaggle dataset. Kaggle strips + from filenames, so 0.1.2+gfb7f732de arrived unparseable and pip refused it.

The GPUs are sorted. Now the 156 GiB of weights and every wheel the engine needs have to reach a box with no network.

Check your disk and your network

Before anything else, confirm what you are reading the weights across. A 156 GiB checkpoint over a slow mount sets your load time floor no matter how fast the GPUs are.

Kaggle notebook cell — dsv4-tp4

$ nvidia-smi --query-gpu=index,name,memory.total,compute_cap --format=csv
index, name, memory.total [MiB], compute_cap
0, NVIDIA L4, 23034 MiB, 8.9
1, NVIDIA L4, 23034 MiB, 8.9
2, NVIDIA L4, 23034 MiB, 8.9
3, NVIDIA L4, 23034 MiB, 8.9

$ du -sh "$CKPT"
156G

$ mount | grep kaggle/input
192.168.7.2:/... on /kaggle/input type nfs (ro,hard,nolock,rsize=524288,timeo=600)

We measured that mount at roughly 130 MB/s on a single stream and 830 MB/s across eight parallel streams. That gap matters later: it is the difference between a serial and a parallel expert load.

What you actually get

Read the safetensors headers rather than trusting the model card. For this checkpoint each shard carries 768 expert tensors — 256 experts × three matrices (w1, w2, w3) — and the experts account for about 3.22 GB per layer:

PartFormatSize
Whole checkpoint on disk48 safetensors shards156 GiB (167 GB)
Routed expert banks, loaded to host RAMFP4 in I8 + F8_E8M0 scales~154 GB
Dense / attention / shared, resident on GPUFP83.38 GiB × 4 ranks
Those two, added up154 + 14.5~168 GB — the checkpoint

Those are measured, not derived — du for the first, the serve log for the other two. The split is the whole story. The two loaded parts account for the whole checkpoint within rounding, which is the check worth doing on any figure like this: about 92% of this model is routed experts, and routed experts are exactly the part that does not need to be on a GPU.

Staging an offline install

An air-gapped target changes one thing: you resolve dependencies on one machine and install them on another. Everything below follows from that split, and none of it is Kaggle-specific — the same applies to an offline cluster node or a locked-down CI runner.

The pattern is two environments: one has a network and does nothing but resolve and download (pip download -d ./wheels), the other has the GPUs and installs from that directory with the index switched off (pip install --no-index --find-links ./wheels).

On Kaggle that is a CPU notebook with internet writing into /kaggle/working, attached to the GPU notebook through kernel_sources. Elsewhere it is a build host and an artifact store. The mechanism differs; the rule does not.

Resolve on the platform you will install on

This is the part that catches people, and it is worth stating as a rule: pip download resolves against the machine running it, not the machine you are targeting. Python version, OS, architecture and libc version all feed into which wheel gets selected, and environment markers such as platform_system == "Linux" or sys_platform == "linux" are evaluated against the host. Resolve on a mismatched machine and you get a complete, self-consistent, entirely wrong set of wheels — and no error, because nothing failed.

pip does offer --platform, --python-version, --implementation and --abi to override this, but they come with a catch: they require --only-binary=:all:, and you must name a platform tag that every package in the closure actually publishes. We tried it, and the run aborted on a single mismatch — most CUDA wheels were manylinux_2_17, but nvidia-cuda-cupti ships manylinux_2_25. One wrong tag fails the whole batch.

The reliable answer is to resolve somewhere that matches the target: a container of the same base image, a CI runner on the same OS, or in our case simply a second notebook on the same platform as the GPU box. That took one run and produced 105 wheels, 3.49 GB.

If you are on macOS — or Windows, or an ARM laptop targeting x86 — do not resolve locally, and do not reach for emulation as the fix. Running the download under --platform linux/amd64 in Docker on Apple silicon fetched nothing in ten minutes before we gave up. Use a native machine that matches the target.

Then two CUDA-specific traps

  • Pin the compiler toolchain to one version. nvcc, crt and nvvm come from separate wheels and pip will happily mix them. Mix 13.0 and 13.4 and the runtime kernel build dies with ptxas: Unsupported .version 9.4; current version is 9.0. We aligned those three at 13.4.59. nvidia-cuda-cccl is headers only and publishes no matching 13.4 build, so it stays at 13.0.85 — it contributes no PTX and does not participate in the version skew.
  • Create the unversioned sonames. NVIDIA wheels ship only libcudart.so.13. Anything that compiles at runtime links -lcudart and the linker wants a bare libcudart.so, so symlink them in place before the first build. The same applies to lib vs lib64 directory names.

Kaggle strips + from dataset filenames. A PEP 440 local version like 0.1.2+gfb7f732de becomes the unparseable 0.1.2gfb7f732de and pip refuses it. Notebook output attached via kernel_sources preserves the character; a dataset does not. Use the former.

The two build notebooks

Everything above is the principle. Here is the actual shape of the two notebooks the serve notebook attaches, because without them nothing reproduces. Build them in this order — each must finish successfully before the next can attach its output.

Notebook A — resolve the wheels (ft-wheels)

CPU, internet on, no competition attached. That last part is not optional: the moment you add competition_sources you lose the network, which is the whole reason this is a separate notebook.

File on your machine — ft-wheels/kernel-metadata.json

# ft-wheels/kernel-metadata.json
{
  "id": "<owner>/ft-wheels",
  "code_file": "wheels.ipynb",
  "kernel_type": "notebook",
  "enable_gpu": "false",
  "enable_internet": "true",   <- possible only with NO competition_sources
  "competition_sources": []
}

The notebook fetches the engine's nightly manifest, verifies both wheels against the checksums it publishes, then resolves the whole dependency closure beside them:

Kaggle notebook cell — ft-wheels

import os, json, hashlib, subprocess
W = "/kaggle/working/wheels"; os.makedirs(W, exist_ok=True)
ENG = json.loads(subprocess.run(
    "curl -sL https://github.com/FlashML-org/FreeToken/releases"
    "/download/nightly/engine-linux_x86_64.json",
    shell=True, capture_output=True, text=True).stdout)

staged = {}
for key in ("runtime", "kernel_cache"):
    a = ENG[key]; dst = os.path.join(W, a["name"])
    subprocess.run(f'curl -sL -o "{dst}" "{a["url"]}"', shell=True, check=True)
    assert hashlib.sha256(open(dst,"rb").read()).hexdigest() == a["sha256"]
    staged[key] = dst

subprocess.run(["pip", "download", "-d", W, staged["runtime"], staged["kernel_cache"],
                # pin the CUDA toolchain to ONE version -- see the trap above
                "nvidia-cuda-nvcc==13.4.59", "nvidia-cuda-crt==13.4.59",
                "nvidia-nvvm==13.4.59",      "nvidia-cuda-cccl==13.0.85"])

Output: 105 wheels, 3.49 GB in /kaggle/working/wheels. The prebuilt kernel cache is the reason this works at all on a box with no CUDA toolkit — it ships cu130 kernels already compiled for sm_89 among its 88 .so files, so nothing has to be built for the GPU at serve time.

Notebook B — build the patched engine (ft-pr70)

PR #70 is not released, so there is no wheel to download. You build one. Same metadata as above — internet on, no competition — with code_file pointing at this notebook instead.

Clone at the exact commit rather than the branch head, so your build is reproducible:

Kaggle notebook cell — ft-pr70

import os, re, glob, subprocess
SRC, OUT = "/kaggle/working/src", "/kaggle/working/pr70"
SHA = "c066b38e1c9981c8c8adb0efb328049fe87aa8c0"   # PR #70
os.makedirs(OUT, exist_ok=True)
subprocess.run(["git","clone","-q","https://github.com/FlashML-org/FreeToken",SRC], check=True)
subprocess.run(["git","fetch","-q","origin","pull/70/head"], cwd=SRC, check=True)
subprocess.run(["git","checkout","-q",SHA], cwd=SRC, check=True)

Two things make the build succeed that are not obvious:

Kaggle notebook cell — ft-pr70

# 1. The build hardcodes -L<wheel>/lib and -L<wheel>/lib64 and links -lcudart,
#    but NVIDIA wheels ship only versioned sonames. Create the bare names in place.
#    You need this again in Step 3, so make it a function.
def link_sonames(DP):
    for so in glob.glob(DP + "/nvidia/**/lib*/lib*.so.*", recursive=True):
        base = re.sub(r"\.so\.\d+.*$", ".so", so)
        if not os.path.exists(base):
            try: os.symlink(so, base)
            except OSError: pass
    for d in glob.glob(DP + "/nvidia/*/lib"):              # and lib -> lib64
        l64 = os.path.join(os.path.dirname(d), "lib64")
        if not os.path.exists(l64): os.symlink(d, l64)

link_sonames("/usr/local/lib/python3.12/dist-packages")

# 2. Call the PEP 517 backend directly. `pip wheel` swallows the compiler error
#    and reports only "failed building wheel", which tells you nothing.
os.chdir(SRC)   # build_wheel builds the CWD; !cd in a cell does not change it
from setuptools import build_meta as bm
whl = bm.build_wheel(OUT)      # arg is the OUTPUT dir -> /kaggle/working/pr70/*.whl

That second point saved more time than anything else in this post. Running the backend by hand turns an opaque pip failure into the actual ld or nvcc message, which is how the missing sonames were found in the first place.

Output — ft-pr70

created 31 unversioned .so symlinks
BUILT: freetoken-0.1.2-cp312-cp312-linux_x86_64.whl

Push them, then wire them up

Your machine — shell, in this order

$ kaggle kernels push -p ft-wheels    # wait for it to finish
$ kaggle kernels push -p ft-pr70      # wait for it to finish
$ kaggle kernels push -p dsv4-tp4     # attaches both as kernel_sources

A kernel's output is only attachable once the run has completed, so the two pushes above are sequential, not parallel. And if you intend anyone else to fork the serve notebook, all three have to be public — a public notebook whose kernel_sources are private is not reproducible by anybody but you.

Step 2: Pick your split, then check it divides

What didn’t work
  • Expecting four GPUs to mean four times the context. The KV pool is replicated per rank, so TP buys none. We planned around a 256K-token pool and the solver settled on 32K.
  • Stock tensor parallelism. Released FreeToken gates TP>1 for every quantized expert format, so --tensor-parallel-size 4 fails before a layer is built.

There are three common ways to cut a model across GPUs, and for a MoE model with offloaded experts the choice is more constrained than usual.

SplitHow it cutsWhy we did or didn't
TensorEvery layer sliced across all GPUsUsed. Splits both dense weights and expert banks, so no rank stores a full copy
PipelineWhole layers assigned to each GPUNot supported for this runtime; would also idle 3 of 4 cards at batch size 1
ExpertExperts distributed across GPUsMoot once experts live in host RAM

Now the part that catches people. Released FreeToken 0.1.2 will not do tensor parallelism on this model at all:

FreeToken source — reference

# layers/quantization/moe/base.py:153
if not tp_ok and cfg.tp_size > 1:
    return "TP > 1 is not supported for this expert format"

Every quantized expert path — fp8_block, nvfp4, mxfp4 — passes tp_ok=False. Multi-GPU serving of this checkpoint requires PR #70, which adds a genuinely tensor-parallel DSV4 runtime: rank-sliced weights, rank-local FP4 expert banks, and collectives placed where rank-local partials become complete values.

What is sharded, and what is not

PR #70 publishes an ownership table, and you should read it before tuning anything, because one row changes what four GPUs buy you:

ComponentLayoutCompletion
Token embeddingVocabulary rowsMasked local lookup, then all-reduce
LM headVocabulary rowsLocal logits, then rank-ordered all-gather
MLA wq_bColumn-parallel over query headsDeterministic head block per rank
MLA wo_bRow-parallelOne all-reduce completes attention
Routed FP4 expertsSplit on moe_inter_dimAdded to shared partial, one MoE all-reduce
Latent KV + paged KV poolsReplicatedEvery local head reads the one latent KV
Router, compressors, indexerReplicatedEvery rank must choose the same experts

Tensor parallelism does not buy you context here. The KV pool is replicated, so every rank holds the whole thing and four GPUs give you exactly the context of one. What TP buys is not storing the expert banks four times over. If you came looking for a longer context window, this is not the lever.

Check it divides

TP=4 is not always legal. PR #70 validates up front and fails with a targeted config error rather than producing wrong math, and your TP size must divide all of:

  • attention head count and MLA o_groups — partial groups are rejected outright
  • the vocabulary
  • the quantized MoE intermediate dimension — FP8 128×128 blocks and FP4 32-element scale blocks cannot be cut at arbitrary boundaries

For this checkpoint, 1, 2 and 4 all divide cleanly. If your TP size does not, the run stops before any layer is built, which is the behaviour you want.

Step 3: Install it on the offline box

What didn’t work
  • Replacing LD_LIBRARY_PATH instead of appending to it. The TP rank workers inherit this environment, and the driver's libcuda.so.1 lives on the system path. Clobbering it gave us RuntimeError: cannot use CUDA device 3: only 0 device(s) visible.
  • Leaving the NVIDIA wheels' libcuda.so stub in place. It shadows the real driver library. Delete it.
  • Assuming the install is done when pip exits 0. Two of the packages it installs cannot run on this box at all — see below.

The box has no network, four mounts, and nothing installed. This step is the whole bridge between them, and it has five parts that all have to happen before ft serve will start.

Find the mounts

Do not hardcode these paths. One kernel_sources entry mounts at /kaggle/input/<kernel>/, two or more at /kaggle/input/notebooks/<owner>/<kernel>/, so glob for whichever you got:

Kaggle notebook cell — dsv4-tp4

import glob, os, re
FL   = next(iter(glob.glob("/kaggle/input/**/ft-wheels/wheels", recursive=True)), None)
PR   = glob.glob("/kaggle/input/**/ft-pr70/pr70/freetoken-*.whl", recursive=True)
PR_WHEEL = PR[0] if PR else None
CKPT = os.path.dirname(sorted(
    glob.glob("/kaggle/input/models/**/config.json", recursive=True), key=len)[0])
assert FL and PR_WHEEL, "build-notebook output not mounted"
print(FL, PR_WHEEL, CKPT, sep="\n")

Install from the local wheel directory

Kaggle's image is Python 3.12, so everything staged in Notebook A is cp312. --no-index is what guarantees you are installing what you staged and not silently reaching for a network that is not there:

Kaggle notebook cell — dsv4-tp4 (shell magic, runs on the GPU box)

!pip install --no-index --find-links {FL} {PR_WHEEL} \
      nvidia-cuda-nvcc==13.4.59 nvidia-cuda-crt==13.4.59 \
      nvidia-nvvm==13.4.59      nvidia-cuda-cccl==13.0.85

Remove the native packages that cannot run here

This is the step that is easy to skip and costs a whole run. Both packages install cleanly and neither works on this hardware, and because FreeToken probes for them with find_spec, their mere presence routes the model down a code path that dies mid-forward. Make "broken" mean "absent":

Kaggle notebook cell — dsv4-tp4

import importlib, importlib.util, subprocess, sys
FORCE_DROP = {"flashinfer"}          # imports fine; its JIT cannot compile here
for mod, dist in (("sgl_kernel","sglang-kernel"), ("flashinfer","flashinfer-python")):
    try:
        if mod in FORCE_DROP: raise RuntimeError("runtime JIT unusable on this box")
        importlib.import_module(mod)
    except BaseException:
        subprocess.run([sys.executable,"-m","pip","uninstall","-y",dist])
        importlib.invalidate_caches()
        assert importlib.util.find_spec(mod) is None   # or freetoken still takes the fused path

Put the CUDA toolchain on PATH

TP>1 compiles one kernel at runtime (pynccl.cu), so nvcc has to be findable even though nothing else is built here:

Kaggle notebook cell — dsv4-tp4

DP   = "/usr/local/lib/python3.12/dist-packages"
NVCC = sorted(glob.glob(DP + "/nvidia/**/bin/nvcc", recursive=True))[0]
os.environ["PATH"] = os.path.dirname(NVCC) + ":" + os.environ["PATH"]
os.environ["CUDA_HOME"] = os.path.dirname(os.path.dirname(NVCC))

link_sonames(DP)   # same helper as Notebook B -- the linker wants bare .so names

libdirs = sorted({os.path.dirname(x) for x in
                  glob.glob(DP + "/nvidia/**/lib*/lib*.so", recursive=True)})
os.environ["LIBRARY_PATH"] = ":".join(libdirs) + ":" + os.environ.get("LIBRARY_PATH","")
# APPEND -- never clobber. Rank workers inherit this, and libcuda.so.1 is a system lib.
os.environ["LD_LIBRARY_PATH"] = ":".join(
    [os.environ.get("LD_LIBRARY_PATH",""), "/usr/lib/x86_64-linux-gnu", *libdirs]).strip(":")

for stub in glob.glob(DP + "/nvidia/**/libcuda.so", recursive=True):
    if os.path.islink(stub): os.remove(stub)    # do not shadow the real driver

import torch; assert torch.cuda.device_count() == 4

Apply the sm_89 patch

If you are on an Ada card (L4, RTX 40-series) with torch < 2.12, do this or long prompts hang forever — Step 7 explains why. PR #70's tree predates the fix, so PR #243 touches exactly one file, and it is self-contained — so the cheapest route is to check that file out from the PR inside Notebook B (where the tree already exists), ship it alongside the wheel, and copy it over the installed copy after pip finishes. Do not take main's version of the file: it has drifted past #243 and drops imports PR #70 still uses.

Kaggle notebook cell — dsv4-tp4

# In Notebook B, where the PR #70 tree already exists, produce the patched file:
FILE = "python/freetoken/kernel/triton/fp8_pertensor_linear.py"
subprocess.run(["git","fetch","-q","origin","pull/243/head:pr243"], cwd=SRC, check=True)
# PR #243 touches only this one file, so take its post-merge version of it and
# re-apply just that file's diff onto the PR #70 checkout:
subprocess.run(["git","checkout","pr243","--",FILE], cwd=SRC, check=True)
# copy SRC/FILE into the ft-pr70 notebook output alongside the wheel

# Then in the serve notebook, after pip install:
import py_compile, shutil
DST = DP + "/freetoken/kernel/triton/fp8_pertensor_linear.py"
assert "rowwise_scaled_mm_ok" not in open(DST).read(), "already patched"
shutil.copy(PATCHED_FILE_FROM_MOUNT, DST)
py_compile.compile(DST, doraise=True)
from freetoken.kernel.triton.fp8_pertensor_linear import rowwise_scaled_mm_ok
print("row-wise enabled:", rowwise_scaled_mm_ok())   # must print False on sm_89 + torch<2.12

Patched, rowwise_scaled_mm_ok() returns False on sm_89 with torch < 2.12 and the fused projection runs one tensor-wise GEMM per part instead. On torch ≥ 2.12 it is a no-op, so it is safe to apply unconditionally.

Step 4: The command, and every flag in it

What didn’t work
  • --num-tokens, at four different values. Each one bypassed the fit solver rather than constraining it, so the whole ladder measured nothing.
  • --moe-cache-size 128. Below num_experts (256); rejected instantly.
  • --moe-cache-size 256. Passes that check, then fails a second one — prefill overlap needs 2 × num_experts.
  • --moe-backend hybrid. Chosen off a benchmark measured at TP=1 on the released engine, which told us nothing about TP=4 on a patched branch.

Here is the launch command that works, in full:

Kaggle notebook cell — dsv4-tp4 (runs on the GPU box, not your laptop)

!ft serve \
    --model-path {CKPT} \                # resolved in Step 3; never hardcode it
    --gpu 0,1,2,3 \
    --tensor-parallel-size 4 \
    --moe-backend offload \
    --expert-load serial \
    --memory-ratio 0.90 \
    --decode-log-interval 1

Seven flags. Note what is not there — no context length, no cache size, no page count. That omission is deliberate and is covered below.

Every flag, explained

FlagWhat it doesWhy this value
--model-pathLocal folder or HF repo idThe NFS mount; also accepts --model
--gpu 0,1,2,3One entry per TP rank; entry i is rank i. Accepts indices or GPU UUIDsMust have exactly as many entries as the TP size or it errors at parse
--tensor-parallel-size 4TP degree; alias --tp-sizeAll four cards. Must divide heads, vocab, o_groups, MoE inner dim
--moe-backend offloadauto / fused / offload / cpu / hybridfused keeps experts resident and will not fit. offload streams them from host RAM
--expert-load serialauto / serial / parallel host-RAM read strategyparallel is faster but needs a whole-shard buffer on top of the banks. serial is the low-memory read
--memory-ratio 0.90Fraction of free VRAM the engine may use for weights + MoE cache + KV combinedThe remaining 10% is CUDA-graph and activation headroom
--decode-log-interval 1Print a scheduler line every N decode forwards1 while tuning; raise it in production

Three more you will reach for, and one you should know exists:

FlagWhat it does
--moe-cache-size NExplicit GPU expert-slot count. Minimum is num_experts (256 here), and if prefill overlap is on the real minimum is 2 × num_experts = 512
--moe-cache-autoSolves slot count and KV pages from one budget. The default when you pass an offload backend with no sizing flag
--kv-reserve-tokens NKV floor reserved before auto fills experts. Only has effect under --moe-cache-auto
--expert-prefetch 1Shards queued ahead by the parallel loader — inert under --expert-load serial. Drop to 1 if you switch to parallel loading and host RAM is tight
--port 1919Where the OpenAI-compatible API listens. Default 1919; --host to bind elsewhere
--cuda-graph-max-bs NLargest batch size captured as a CUDA graph. Lower it if capture OOMs; 1 is the minimal-graph setting

The flag not to pass

This one cost us three runs, so it gets its own section: do not pass --num-tokens or --num-pages.

It reads like a constraint. It behaves like an override:

FreeToken source — reference

# kvcache/dsv4_paged_pool.py:343
num_pages = config.num_page_override
if num_pages is None:
    sizes = dsv4_solve_num_pages(available_memory, ...)   # fit-checked
else:
    floor = _dsv4_window_floor_pages(config, P)
    if num_pages < floor: raise ...                      # ONLY a lower bound
    sizes = _dsv4_pool_sizes(config, num_pages + 1)      # never consults memory

Pass it and the solver is skipped entirely. The engine validates a floor, then allocates the pool you asked for whether or not it fits, and you get an out-of-memory error deep inside buffer allocation. Leave it off and the solver sizes the pool against actual free memory, every time.

It gets worse in combination. Selecting --moe-backend offload with no cache-sizing flag silently switches on moe_cache_auto:

FreeToken source — paraphrased, see server/args.py

# server/args.py:700, condensed
if is_offload_moe_backend(kwargs["moe_backend"]) and _no_cache_flag:
    kwargs["moe_cache_auto"] = True        # you did not ask for this

Auto plans expert slots and KV pages against a single shared budget. But the engine only writes back auto's KV half if you have not set an override — so with --num-tokens present, the experts keep their greedy share and the KV pool allocates its pinned size against bytes that are already spent. Two sane behaviours composing into a double-spend.

Step 5: How to read the startup log

What didn’t work
  • Treating HTTP 200 from /health as ready. That endpoint is always 200 — the JSON status field is the gate, moving loading → ok → error. Polling for 200 gave us a "ready in 10s" on a 167 GB model, and one result we had to retract.
  • terminate() on the server process. It leaves the TP rank workers alive, and the next health check gets answered by the previous run. Kill the process group.

Launch it and wait for the right signal

ft serve never exits, so do not run it with a blocking call and a timeout — start it and poll. It listens on port 1919 by default (--port to change it), and the readiness signal is not the HTTP status:

Kaggle notebook cell — dsv4-tp4

import subprocess, requests, time, os, signal, json
URL = "http://127.0.0.1:1919"

def health():
    try: return requests.get(URL + "/health", timeout=5).json()
    except Exception: return None

log = open("/kaggle/working/serve.log", "w")   # /kaggle/working is saved as output
srv = subprocess.Popen(cmd, stdout=log, stderr=subprocess.STDOUT,
                       start_new_session=True)   # own process group, so we can kill it

t0 = time.time()
while time.time() - t0 < 7200:
    h = health()
    if h:
        if h.get("status") == "ok": break        # the gate -- NOT response.ok
        if h.get("status") == "error": raise RuntimeError(h.get("message"))
        pr = h.get("progress") or {}                # done_bytes / total_bytes while loading
    if srv.poll() is not None:                      # it died; the log has why
        raise RuntimeError(open("/kaggle/working/serve.log").read()[-4000:])
    time.sleep(5)
print("ready in", round(time.time() - t0), "s")

def hard_stop():
    # terminate() leaves the TP rank workers alive and they keep the port
    try: os.killpg(os.getpgid(srv.pid), signal.SIGTERM)
    except Exception: pass
    subprocess.run("pkill -9 -f 'ft serve' || true", shell=True)

Loading takes about nine minutes, so it pays to know which lines matter. These seven lines tell you everything about whether your configuration is sensible.

Output — serve log on the GPU box

[1] Resolved config: moe_backend='offload', attention_backend='dsv4_sparse',
                 cache_type='swa_radix', page_size=128
[2] torch intra-op threads: 48 -> 12 per rank (4 ranks share this host)
[3] Free memory before loading model: 20.70 GiB
[4] Weights: 3.32 GiB declared by parameters, 3.38 GiB measured on device
[5] --moe-cache-auto resolved moe_cache_size=4080 num_pages=500
[6] Allocating 32768 tokens for DSV4 KV cache (52 window pages), total = 1.08 GiB
[7] Start capturing CUDA graphs with sizes: [1, 2, 4]
    Free GPU memory after capturing CUDA graphs: 2.71 GiB
    STATUS OK in 525s
  1. Resolved config. Confirms the backend you asked for survived auto-resolution. dsv4_sparse attention and swa_radix cache were chosen for you; if you see fused here on a model this size, stop, because it will not fit.
  2. Thread split. Four ranks share one host, so each gets a quarter of the cores. If this line is missing, your ranks are each spawning a full-machine thread pool and fighting each other.
  3. Free memory before loading. Your real budget, measured after the CUDA context exists. 20.70 GiB of a nominal 22.5.
  4. Declared vs measured. The gap between what the parameters claim and what the load actually cost is resident-but-unaccounted memory — staging buffers, per-layer scratch. 60 MiB here, which is healthy. A large gap comes straight out of your cache budget.
  5. What auto decided. 4080 expert slots out of 11,008 total (43 layers × 256), and 500 KV pages. This is the line to watch when tuning.
  6. Actual KV pool. 32,768 tokens for 1.08 GiB. Note this is nowhere near the model's advertised 1M context — the solver sized it to the memory that was left.
  7. Graph capture. Batch sizes 1, 2 and 4, costing about 0.8 GiB. If capture fails here, lower --cuda-graph-max-bs.

Then check the cards agree with each other:

Kaggle notebook cell — dsv4-tp4

$ nvidia-smi --query-gpu=index,memory.used,memory.total --format=csv,noheader
0, 19789 MiB, 23034 MiB
1, 19789 MiB, 23034 MiB
2, 19789 MiB, 23034 MiB
3, 19789 MiB, 23034 MiB

Four identical figures is the signature of real sharding. If one rank is much larger than the others, you are replicating something you meant to split.

Step 6: Benchmark it

What didn’t work
  • The engine’s own throughput counters. On short replies they reported 0.13, then 5.42, then 1392.89 tok/s for one request. We nearly filed the first as a performance bug.
  • Reading delta.content from the stream. This is a reasoning model; tokens arrive as delta.reasoning_content. We logged thirty minutes of "hang" on a server that was working perfectly.

The client

The served model id is the basename of --model-path, so here it is literally "1". Check it rather than guessing: GET /v1/models. And read both content channels — this is a reasoning model, so on many prompts every token arrives as reasoning_content and a client watching only content sees an empty stream forever:

Kaggle notebook cell — dsv4-tp4

MID = requests.get(URL + "/v1/models", timeout=30).json()["data"][0]["id"]

t0 = time.time(); first = None; n = 0; out = []
with requests.post(URL + "/v1/chat/completions", stream=True, timeout=900, json={
        "model": MID, "stream": True, "max_tokens": 256, "temperature": 0.0,
        "messages": [{"role": "user", "content": "Explain in 150 words why MoE "
                      "expert offload lets a 167GB model serve from 96GB of VRAM."}]}) as r:
    for line in r.iter_lines():
        if not line or not line.startswith(b"data: "): continue
        if line == b"data: [DONE]": break
        d = (json.loads(line[6:]).get("choices") or [{}])[0].get("delta", {})
        tok = d.get("content") or d.get("reasoning_content")   # BOTH channels
        if tok:
            if first is None: first = time.time() - t0
            out.append(tok); n += 1

dt = time.time() - t0
print(f"TTFT {first:.2f}s   decode {(n-1)/(dt-first):.2f} tok/s   {n} tokens")

Time it client-side like this, over a couple of hundred tokens. The engine's own counters are not a substitute, for the reason below.

The memory side

19,789 MiB used of 23,034 on every card, with about 2.7 GiB free after graph capture. The expert cache holds 4080 of 11,008 experts, roughly 37%. That leftover VRAM is the tuning headroom: more slots means a higher cache hit rate and fewer PCIe fetches per token.

The speed side

MeasureResultConditions
Load time525 sSerial expert read from NFS
Prefill (request A)1331 tokens / 16.3 sNon-streamed, 31 tokens out; end-to-end
TTFT (request B)6.11 sDifferent request — a ~25-token prompt, not the 1331-token one above
Decode (request B)20.37 tok/s226 tokens, greedy, batch 1, timed client-side
Sampled decodeOKtemperature 0.6, 63 tokens

About 20 tokens per second on hardware that can hold barely half the model. The limiter is PCIe expert traffic, not compute, which is exactly what you would expect from the design.

One number we threw away, and why

Our first decode measurements came from the scheduler's own log lines during three-token replies:

Output — serve log, and why you should not trust it

Decode batch, #running-req: 1, #token: 128, gen throughput (token/s): 0.13
Decode batch, #running-req: 0, #token: 0,   gen throughput (token/s): 5.42
Decode batch, #running-req: 0, #token: 0,   gen throughput (token/s): 1392.89

0.13, then 5.42, then 1392.89 tok/s for the same request. None of those are the model's speed. The first includes prefill in its window, the last two are empty batches dividing by an interval where nothing was generated. A throughput counter needs a sustained run to mean anything.

We nearly reported 0.13 tok/s as a catastrophic result and went hunting for a performance bug that did not exist. The honest number came from timing 226 streamed tokens end to end, client-side, which is the only measurement here we would defend.

What these numbers are not. Single run, one prompt shape, batch size 1, no concurrency testing. With 37% of experts cached and 2.7 GiB per card unused, 20.37 tok/s is a floor rather than a tuned result.

Step 7: Errors you will actually hit

"TP > 1 is not supported for this expert format"

KernelSelectionError: triton: TP > 1 is not supported for this expert format

You are on released FreeToken rather than PR #70. Every quantized expert path gates TP>1. There is no flag for it; you need the patched runtime.

"moe_cache_size=128 is too small"

ValueError: moe_cache_size=128 is too small: need at least num_experts=256 slots

The offload cache needs at least one slot per expert per layer. And if prefill overlap is enabled — it is, by default — the true minimum is double that, which surfaces as a separate assertion about borrowing two full expert-layer buffers. Either pass at least 2 × num_experts, or use --moe-cache-auto and let it decide.

OutOfMemoryError in _alloc_buffers

OutOfMemoryError: Tried to allocate 206.00 MiB. GPU 0: 22.03 GiB total, 163.06 MiB free

Almost always --num-tokens, per Step 4. You pinned a KV pool larger than the memory left after the expert cache took its share. Remove the flag and let the solver size it.

A silent hang with no log line at all

(no error, no "Prefill batch" line, GPU pinned at 100% utilisation)

On sm_89 with torch < 2.12, row-wise torch._scaled_mm launches its CUTLASS stream-K kernel off the current stream. Issued from a side stream at M ≥ 256 it stalls the GPU outright, and DSV4's FP8 fused q/k/v projection takes exactly that path. This is pytorch/pytorch#177651, fixed in torch 2.12.

FreeToken kept its 2.11 pin and shipped an in-tree workaround in PR #243. PR #70 branched before that landed, so a build from it reintroduces the hang — apply #243's diff to PR #70's copy of fp8_pertensor_linear.py, not to main's, which has diverged past it.

L4 is sm_89 and the default wheel is torch 2.11.0+cu130. If you serve FP8 models on Ada cards, check this pair before you debug anything else.

All four workers die, but only on prompts over 32 tokens

ImportError: sgl_kernel/sm100/common_ops.abi3.so: undefined symbol: _ZNR5torch7Library4_def...

A 32-token cliff looks like a batching bug. It is a packaging bug. FreeToken decides whether fused CUDA kernels are available like this:

FreeToken source — reference

# kernel/backend.py
def _importable(name) -> bool:
    try:
        return importlib.util.find_spec(name) is not None   # no import side effects
    except Exception:
        return False

find_spec proves a package exists; it never imports it. Our sgl_kernel existed and raised on import — it was loading its sm100 (Blackwell) variant on an Ada card, and that .so carried an undefined torch symbol from a different torch build. FreeToken read it as available, the MoE took the fused path, and the real import blew up inside the forward pass.

The 32 comes from the routing crossover:

FreeToken source — reference

# models/deepseek_v4/moe.py :: _prefill_routed
if (hidden_states.shape[0] * self.top_k >= self.num_experts
        or cache.is_unpinned_layer(self.layer_id)):
    return super()._prefill_routed(...)   # whole-layer streaming -> fused path

With 256 experts and top_k 8 that is 8T ≥ 256, so T ≥ 32. Below the crossover, prefill uses the on-demand slot path and never touches the broken package. Five- and seven-token smoke tests pass and prove nothing. Fix: uninstall it, so the pure-Triton fallback engages as designed.

An 1800-second stall during generation

tvm_ffi/include/tvm/ffi/string.h(373): error: namespace "std" has no member "memcmp"

Having learned the lesson above, we wrote a probe that imported each optional native package and removed any that failed. flashinfer imported fine, so it stayed — then JIT-compiled a kernel through nvcc at generation time, and its tvm_ffi headers do not compile against CUDA 13.4.

"Imports cleanly" was still the wrong health check. For a library that compiles at runtime, whether it works is not knowable at import time. Remove it; Triton covers the same ops.

One that is not an engine error at all. If a streamed request returns nothing while the server looks healthy, check which field you are reading. DSV4-Flash is a reasoning model and FreeToken auto-selects reasoning_parser='deepseekv32', so tokens arrive as delta.reasoning_content, not delta.content. We logged thirty minutes of "hang" that was entirely our own client.

Wrapping up

Four things to carry out of this, all of them checks you can run in a minute.

  1. Count your dense weights, not your total weights. For a MoE model with offload, the fit question is whether the non-expert weights divided by your GPU count fit on one card, and whether host RAM holds the experts. 167 GB against 96 GB of VRAM sounds impossible and is not.
  2. Let the solver size the KV pool. A flag that names a resource is not always a constraint on it — --num-tokens overrides the fit check rather than informing it. If a knob and a solver both exist, find out which one wins before you tune.
  3. Check your architecture against your reference's. sm_86 and sm_89 are one generation apart and differ on exactly the kernel path that hung for us. "It works on their four GPUs" is not portable evidence.
  4. A broken optional package is worse than a missing one. find_spec proves a directory exists, nothing more. A present-but-unimportable native extension suppresses the fallback written for its absence, and a JIT library that imports may still fail to compile. Probe by importing, and for JIT libraries, by executing.

The answer for this hardware is plain --tensor-parallel-size 4 with --moe-backend offload and no cache flags at all — let auto solve it. And one last caution about the interconnect: with no NVLink, every expert miss is a PCIe round trip. That single fact is why the expert cache hit rate, not the GPU count, is the number worth tuning next.

All of it reproduces from three notebooks: one that resolves the offline wheel bundle, one that builds the patched engine wheel, and the serve notebook, which attaches the other two as kernel_sources. Fork it, or clone it and push with kaggle kernels push -p .

Credits and references

Measured September 2026 on Kaggle free tier — 4× NVIDIA L4 23,034 MiB, sm_89, driver 580.159.04, 188 GB host RAM, 46 vCPU, PCIe (no NVLink). Engine: FreeToken 0.1.2 at PR #70 (c066b38) with PR #243 backported.