AI & Agents · Systems

swift-qwen3.8-rtx3090

A reproducible pipeline that rebuilds a 27B finetune into a fast single-RTX-3090 variant. It derives the draft vocabulary and int4 GPTQ calibration from the model's own outputs, which lifts draftable-token coverage from 96.7% to 99.8% and decode from 94.0 to 98.4 tokens per second, level with syv's fast build of base Qwen.

Active · 2026

Python · vLLM · GPTQ · RTX 3090

swift-qwen3.8-rtx3090

My coding agent runs against a local model, not an API. The machine is a single RTX 3090, and the stack that makes 27B usable on it is syv’s HyperQwen. It pins and patches , keeps the at , and requantises the and the draft module so the whole thing fits in with a 114k .

Stock Qwen3.8-27B reasons too much for my use. On two open-ended code tasks it spent the full 8,192- and never wrote any code. Swift, a community tuned for shorter reasoning, finished the same tasks in 2,500 to 7,100 thinking tokens and produced working code. No -speed trick closes a gap like that, so I switched.

The syv stack will serve Swift, but as a third-party , and those get a lower tier of support than base Qwen. Base Qwen has a “fast variant” with an and draft module, plus a counted over the model’s own outputs. That variant is worth about 15 percent. Third-party checkpoints get int8 heads and base Qwen’s draft vocabulary. Swift ran fine that way. I wanted to know how much of the fast variant’s gain I could get back by rebuilding it for Swift.

Body, heads and drafter

A checkpoint in this stack has three parts, and the fast variant leaves the biggest one alone.

Diagram of the model. Token ids go through the 248,320-row embedding table (int8, unchanged) and the 4-bit AWQ body (unchanged), and the last hidden state goes to the lm_head (int4 GPTQ, recalibrated), which scores every token. The MTP layer (int4 GPTQ, recalibrated) takes that hidden state plus the embedding of the previous drafted token, and runs once per drafted token. Its draft head, 25,879 lm_head rows for Swift's draft vocabulary, proposes candidate tokens. A dashed line takes the candidate block back to the target model, which checks it in one forward pass through the embedding table, body and lm_head and keeps the candidates it agrees with up to the first miss. A box groups the draft vocabulary, draft head and MTP weights as the speculative artefacts.
The parts of the fast variant. The body stays as it is. The draft vocabulary is counted from Swift's outputs, and the lm_head and MTP weights are requantised against Swift's own hidden states.

The body is the stack of transformer layers in the middle, and it holds most of the 27B parameters. Each layer reads a of 5,120 numbers per token and writes an updated one. Swift’s body is a 4-bit export. I never change it.

The heads sit at either end of the body, with one row per token in the 248k vocabulary. Strictly, only the lm_head is a head, but the syv stack quantises the two tables together and calls them the heads, so I do too. The turn a token into its first 5,120 numbers. The lm_head turns the last hidden state into a score for every token. At 248k by 5,120, each table is about 1.3 billion parameters, and the AWQ export leaves both unquantised. Swift’s int8 build quantises them to int8.

Only one of them is worth shrinking further. vLLM reads the whole lm_head on every decode step, so the fast variant takes it to int4. A token only reads its own row of the embeddings, so they stay at int8.

The drafter is the MTP module. It is one extra transformer layer plus an input projection, eight in all. The projection mixes two inputs, the body’s last hidden state and the embedding of the previous drafted token. The drafter runs once per guess, feeding each drafted token back in to guess the next. To score those guesses it uses a draft head, a slice of the lm_head with rows only for the tokens in the draft vocabulary. For Swift that is 25,879 rows.

The draft vocabulary, the draft head and the MTP weights exist only for , so I call them the speculative artefacts. None of them are newly trained. The vocabulary is a count over Swift’s outputs, and the MTP weights are Swift’s own, requantised against its hidden states. The full model then checks the whole block of drafted tokens in one forward pass through the embeddings, body and lm_head, and keeps them up to the first one it disagrees with. A bad artefact makes the model slower, never wrong.

Speculative artefacts depend on the model

The MTP drafter scores a 40,960-row slice of lm_head instead of the full 248k vocabulary. It can never draft a token outside that slice. Every miss is a guaranteed rejection, and a rejection ends the speculation chain, so misses cost more than their share of tokens.

The shipped slice was counted over base Qwen’s outputs. Against what Swift actually generates, it covered 96.7 percent of tokens. That sounds fine until you look at where the other 3.3 percent sits. It is code and syntax, which is most of what my agent emits.

GPTQ calibration has the same problem. The you quantise against comes from the hidden states that feed the layer. States captured from base Qwen are the wrong calibration for a finetune that writes differently.

So I redid the fast-variant recipe with every input derived from Swift. The syv repo already had the recipe as scripts under drafter/, and I ran them unmodified. The work was feeding them the right corpus and handling a checkpoint layout they did not expect. Swift’s AWQ export splits lm_head, embeddings and MTP across separate , and upstream’s quantisation scripts assumed base Qwen’s single-shard layout.

Corpus, outputs, vocabulary

Upstream built its draft vocabulary from a Danish-heavy chat mix. That suits the people who wrote it and says little about a Rust coding agent. I wrote a generator that emits 5,500 prompts weighted by my workload. 35 percent is Rust implementation work, 16 percent TypeScript, 12 percent debugging from symptoms, 10 percent code edits and 8 percent tool-use transcripts. The rest is architecture questions, reasoning puzzles and general chat. About 60 percent of prompts run with thinking on, because that is how I use the model day to day. The generator is seeded, so the same config always produces the same corpus.

Running 3,072 of those prompts through Swift took about 1.7 GPU-hours and produced 4.2 million output tokens. I stopped there because coverage had flattened. A vocabulary built from the first quarter of the data already covered 99.17 percent of held-out outputs, and the full run only reached 99.81.

Line chart of held-out draft-vocabulary coverage against the number of counted output tokens. The Swift-derived list starts at 99.17 percent and flattens at 99.81 percent, well above the base Qwen list's flat 96.7 percent.
Held-out coverage as the counted corpus grows. Generation stopped once more tokens stopped moving the curve.

The vocabulary is a frequency count over 90 percent of those outputs, with special tokens forced in and everything else ranked by count. The other 10 percent, 417k tokens, is the held-out set. Against it, coverage went from 96.7 to 99.81 percent, and to 99.86 on code prompts.

The overlap surprised me more. Only 18,701 of base Qwen’s 40,960 ids appear in Swift’s list at all. Swift’s corpus only ever emitted 25,879 distinct tokens, so the right list is shorter than the slot count. Any check that demands exactly 40,960 rows is wrong for this model, which mattered when I came to write the verifier.

GPTQ heads

With outputs in hand, the rest is upstream’s recipe. Capture hidden states for every token, dump the MTP Hessians, then GPTQ the lm_head and the eight MTP linears to int4 against them. I judged the lm_head by from the head on held-out states. int4 scored 0.00707. GPTQ on Swift’s own Hessian scored 0.00234, lower than the 0.0029 that base Qwen’s fast variant shipped at on its own states.

Bar chart of lm_head KL divergence from bf16. Round-to-nearest int4 scores 0.00707, GPTQ int4 calibrated on Swift scores 0.00234, and the base Qwen fast variant's reference is 0.0029.
lm_head KL divergence from the bf16 head on held-out states. Lower is better.

The bug that earned a verifier

The fast directory shares body shards with the int8 directory through , so a 15.2 GB variant costs almost nothing extra on disk. The catch is that vLLM reads every key of every shard it opens, not only the the index assigns to that shard.

My first build hardlinked model-nonquant.safetensors, which still carried the int8 MTP tensors from the int8 build. The new int4 tensors lived in model_extra_tensors.safetensors. vLLM loaded both, and the boot died on a shape assertion that named neither file. Every existence check passed, because every mapped tensor existed in the right shard exactly once. Nothing looked at the other shards. Had the shapes matched, whichever copy loaded last would have won, and I would have served the wrong weights without an error.

Upstream’s verify.sh couldn’t help. It hard-required an int8 lm_head, so it rejected every int4 fast variant, base Qwen’s included. Until I had something better, my fast directory booted with VERIFY=0.

So I wrote a structural verifier before making the fast variant my default boot target. It reads headers directly and needs no torch. It checks that packed and scale shapes match the bit widths the config declares, that the draft head has one row per id in the draft vocabulary, and that every index-mapped tensor lives in exactly one shard, the one the index names. That last check catches the collision before anything boots. The repo ships a stdlib-only selftest that builds toy model directories, corrupts each one in a specific way, and asserts the verifier rejects all of them. The stale-tensor collision is one of the fixtures.

Results

All five configurations ran the same night with four repetitions each, and the table shows medians. Decode is tokens per second. is the fraction of drafted tokens the target kept, read from vLLM’s own speculative-decoding counters. Tok/step is tokens produced per .

Paired bar chart of decode throughput and speculative acceptance for Swift int8, Swift-fast and Qwen fast under MTP long context, and Swift-fast and Qwen fast under DFlash2.
Decode throughput and speculative acceptance across the five configurations, one RTX 3090 at 250 W.
variantdecodeacceptancetok/stepVRAM MiB
Swift int8, MTP long94.00.6302.8922085
Swift-fast, MTP long98.40.6602.9822225
Qwen fast, MTP long98.20.6342.9022191
Swift-fast, DFlash2156.40.3853.7023081
Qwen fast, DFlash2167.40.3873.7123007

MTP with the long context is my daily profile. There, Swift-fast decodes 4.7 percent faster than Swift’s int8 build and matches base Qwen’s fast variant, 98.4 against 98.2 tokens per second, with higher acceptance. A separate quote benchmark asks the model to reproduce 8,000 tokens of its own context verbatim. Swift-fast drafts it at 0.998 acceptance and 3.99 tokens per step against a ceiling of 4, up from 0.962 on the int8 build.

The profile is faster on paper, and I still don’t use it as the default. It caps context at 46k tokens, and my agent’s wall time depends on how many reasoning tokens a task takes and whether it finishes, more than on raw decode. In my agent-latency runs, Swift-fast and Qwen-fast finished the same share of tasks, and long agent sessions need the 114k context more than the extra speed.

The new vocabulary helped DFlash2 anyway. On an earlier night, Swift’s DFlash2 acceptance trailed Qwen’s, 0.367 against 0.387. It now sits at 0.385, which is noise. The remaining gap of 156 against 167 tokens per second comes from the AWQ body reading about 0.24 GiB more weight bytes per step than Qwen’s body. No drafter fixes that. Requantising the body would, but it means a 55 GB source download and more GPU hours to win back 7 percent on a profile I don’t run daily. I skipped it.

Quality held. On a nine-task battery with mechanical checks, the fast variant passed 8 of 9, the same set as the int8 build. The failure is a hard combinatorics problem that both models get right on some samples and wrong on others. I saw no regressions in how reasoning terminates.

Splitting the work

Two of the fixes weren’t specific to Swift, so I sent them upstream. PR 122 makes verify.sh accept the int4 lm_head layout and checks the packed shape against the declared bit width. PR 123 adds the duplicate-tensor scan. syv merged both.

The target-specific pipeline stayed in my repo, swift-qwen3.8-rtx3090. A workload mix is an opinion about one person’s usage, and HyperQwen is deliberately model-agnostic. The repo runs upstream’s scripts rather than forking them, and adds the corpus generator, vocabulary builder, head quantisation and verifier around them.

Publishing the weights

The finished model is on the as liamwh/Swift-Qwen3.8-27B-W4A16-syv-fast. It serves directly on the syv stack with no preparation step. The model card carries both licences the Swift Open License requires of derivatives, and traces the provenance from Alibaba through UkisAI and TheUnderscore to this pipeline.

The upload was much smaller than the model. The AWQ body shards are byte-for-byte identical to TheUnderscore’s checkpoint, so deduplication only transferred what I changed: the int4 heads, the MTP tensors, the extra tensors and the model card. About 2 GB of upload published a 15.2 GB model.

The pipeline repo still contains no weights. If you want to run Swift-fast, download the published model. If you want to rebuild it, or build the same thing for another Qwen3.8-family finetune, use the pipeline with a checkpoint you’re licensed to use. Edit the workload mix first. The draft vocabulary is only as good as the corpus behind it.

Open chat

Interested in working together? Reach out.

Strategy, architecture, and implementation — from workflow to production.

© 2026 Liam Woodleigh. All rights reserved.