Skip to content

Latest commit

 

History

History
155 lines (116 loc) · 8.84 KB

File metadata and controls

155 lines (116 loc) · 8.84 KB

Which model? Sizes, versions and what fits

Strata runs one model, Qwen3.8-Flash-Next, in several sizes (the same model, compressed more or less) and several versions (the original, a coding version, a fine-tune). The installer recommends one for your PC; this page explains the choice. Back to the README.

On this page: Pick by RAM · Speed · The sizes · Will it fit? · The versions · Adding another model

Pick by RAM

Your RAM Take Why
32 GB Coder it fits 32 GB; best at code, weaker at everything else (why). With a 24 GB card, Q2_0 and IQ2_XS run too (low-RAM mode) and are the better pick for general use
48 GB IQ2_XS (or Q2_0, the fastest) the larger sizes do not fit
64 GB IQ2_XS (recommended), or IQ3_XXS / IQ3_S every size fits (IQ3_S with little else open)
96 GB or more IQ3_S, or Unsloth's UD-IQ4_XS (~4-bit) room for the largest sizes

Not sure? Take IQ2_XS. The Coder is the one that fits a 32 GB PC, but it keeps only 256 of the 512 experts, chosen on code data, so it is weaker outside code and in languages other than English. For general use, or whenever your RAM allows, take a full model (Q2_0, IQ2_XS or IQ3_S).

How fast is each size

Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM:

Size Writes answers (short chat) Writes answers (128K context) Reads your prompt
Q2_0 94 tokens/s 76 tokens/s 2,650 tokens/s
IQ2_XS 79 tokens/s 63 tokens/s 2,090 tokens/s
IQ3_XXS 62 tokens/s 49 tokens/s 1,750 tokens/s
IQ3_S 53 tokens/s 46 tokens/s 1,620 tokens/s
Coder (IQ1_M) 55 tokens/s 43 tokens/s 2,180 tokens/s

Measured on an AMD RX 9070 XT (16 GB), a Ryzen 9 3900X and 47 GB of RAM (Linux, setup's own install; Q2_0 is above setup's RAM estimate for 47 GB and was installed with --model Q2_0 --yes):

Size Writes answers (short chat) Writes answers (128K context) Reads your prompt
Q2_0 60 tokens/s 48 tokens/s 1,160 tokens/s
IQ2_XS 52 tokens/s 36 tokens/s 1,110 tokens/s
Coder (IQ1_M) 44 tokens/s 33 tokens/s 1,420 tokens/s
  • Writes answers = how fast the reply appears (tokens per second; a token is about ¾ of a word).
  • Reads your prompt = how fast it takes in what you send (long documents, code, chat history), measured on a 32K-token prompt; a 4K prompt reads at 910-1,580 tokens/s. A 32K prompt takes about 15 seconds with Q2_0.

A card with more VRAM is faster, because more of the model fits on the GPU: an RTX 3090 (24 GB) should do roughly 100-140 tokens per second. All measurements, long-context numbers and estimates for other cards are in the details. The AMD measurements per card (RX 9070 XT, Radeon AI PRO R9700, RX 7800 XT, RX 9060 XT, RX 6900 XT) are in AMD_HIP.md.

Every PC is different: START-HERE.bat --calibrate measures a few engine settings on yours and keeps the fastest (about 5-10 minutes; on the PC above it made the Coder 7% faster; NVIDIA cards for now). Measured Strata on your own PC? See Community benchmark results for a report template and how to share your results in a pull request.

The sizes

Model RAM+VRAM Requirements Speed Quality
Q2_0 37.6 GB fastest good
IQ2_XS 39.2 GB fast better (recommended)
IQ3_XXS 47.0 GB slower great
IQ3_S 54.8 GB slowest best: matches the full model on the published tests (original model only)

The download is 66-76 GB for the three smaller sizes (details); the first start also fetches the MTP draft layer (~6 GB, +1 GB with images).

Will it fit?

Shard 1 is the part of the model that gets loaded when it starts: its experts go into your RAM, the rest onto your graphics card (the second shard, a 29 GB lookup table, stays on the SSD). So it fits when your RAM is at least the experts + about 10 GB for Windows and your other programs - the experts are most of shard 1: 34 GB for Q2_0, 35.5 for IQ2_XS, 43 for IQ3_XXS, 50 for IQ3_S, 23 for the Coder (the dense weights in shard 1 go to the graphics card). With 64 GB of RAM every size fits (IQ3_S with little else open); with 48 GB, Q2_0 and IQ2_XS. A bigger graphics card makes it faster, but it doesn't lower the RAM needed

  • except in the low-RAM mode below.

A big graphics card and little RAM

On a PC whose RAM cannot hold the experts beside the system, setup maps them from the model's files instead and keeps in RAM only what the graphics card does not hold (the low-RAM mode, chosen by setup). For example, a 32 GB PC with a 24 GB GPU runs Q2_0, IQ2_XS and the Coder this way, and a 32 GB PC with a 12-16 GB GPU the Coder. With a small card most experts then come from the SSD and it is much slower (setup says so). START-HERE.bat --setup --low-ram on|off overrides the choice.

The versions

Qwen3.8-Flash-Next

The original, in all four sizes. With images, and with the experimental speed projection as an option.

Coder

Coder - ISTA-DASLab's coding version: it keeps 256 of each layer's 512 experts, chosen on code (with agentic and image) data, and drops the rest (91% of the full model's SWE-bench Verified score, 99% of LiveCodeBench, by its authors). One size (IQ1_M: its experts stored like IQ3_S): shard 1 is 29.6 GB, so it fits a PC with 32 GB of RAM, runs 262K context on 64 GB, and reads long prompts the fastest of all. With half the experts gone it is weaker outside code and in languages other than English, Chinese and other CJK text included (#438: Chinese answers came out wrong or looping where English was fine). For general use, or whenever your RAM allows, take a size that keeps every expert: Q2_0, IQ2_XS or IQ3_S. More: details.

START-HERE.bat --setup --family coder

Swift 1.5

Swift 1.5 - a fine-tune by UkisAI that thinks much shorter before answering, so you get the answer sooner, with about the same quality. Same speed per token, and about the same RAM as the same size of the original (no IQ3_S). Its own license applies (see its page). More: details.

START-HERE.bat --setup --family swift --model IQ2_XS

Unsloth UD-IQ4_XS

Unsloth's UD-IQ4_XS (~4-bit) is the fourth version in setup's menu (--family unsloth, its first size), a regular choice from 0.1.39: a 94 GB download with 59.5 GB of experts (IQ3_S and IQ4_NL), between IQ3_S and UD-Q4_K_XL in quality. Strata keeps your RAM minus 24 GB of the experts in RAM and reads the rest from the SSD while it answers: on a 64 GB PC part of them come from the SSD (slower; an NVMe SSD helps), from ~80 GB of RAM all of them stay in RAM. It needs 48 GB of RAM or more and engine 0.1.38 or newer; images are an option, as with the other models. Details: UD-IQ4_XS.

START-HERE.bat --setup --family unsloth --model UD-IQ4_XS

Unsloth UD-Q4_K_XL (experimental)

Unsloth's 4-bit UD-Q4_K_XL (experimental) is the second size of the same family (--family unsloth): the closest to the full model, but a 111 GB download whose 77 GB of experts do not fit in RAM. Strata keeps your RAM minus 24 GB of them in RAM and reads the rest from the SSD while it answers: 7-8.5 tokens/s on a 64 GB PC with a 12 GB GPU, several times slower than the sizes above, and long prompts are slow. It needs 48 GB of RAM or more, an NVMe SSD and one NVIDIA GPU (no images yet). Details and measurements: UD-Q4_K_XL.

START-HERE.bat --setup --family unsloth --model UD-Q4_K_XL

OrcaRouter Uncensored IQ3_XXS

For OrcaRouter's Flash-Next Uncensored IQ3_XXS, see the manual compatibility setup. It needs an explicit packing conversion and is not an installer menu option.

Adding or switching models

You can add another model any time with SETUP.bat (the same as START-HERE.bat --setup; on Linux ./setup.sh --setup). Files another model shares are not downloaded again (the Coder uses the original's shard 2 and vision encoder). With more than one model installed, START-HERE.bat asks which one to start; run-<model>.bat (Linux: run-<model>.sh) starts one directly.