If a large AI model runs on a 12GB graphics card, where does the rest go? Readers of Hackaday’s Strata story asked how it differs from familiar offloading. The project documentation gives three places to look.

Choosing computations

Strata is an inference engine for Qwen3.8-Flash-Next on a PC. The model uses a mixture of experts: it selects some of its computational modules for each token, a small unit of text processing.

VRAM and RAM

VRAM holds weights for shared computations and frequently selected experts. In the usual configuration, system RAM holds expert weights, and the CPU also computes experts absent from the GPU. A low-RAM mode changes the arrangement by avoiding RAM copies of experts held on the GPU.

Another job for the SSD

The SSD has another job. The documentation describes reading selected rows from a large n-gram lookup table as tokens are processed. The table is part of the model data, stored on SSD and consulted during execution.

Conversations take space

Long conversations also need space for the KV cache, which holds earlier computed states. Strata describes configurations that keep part of this cache in RAM for long contexts. Graphics-memory capacity alone cannot establish how much room remains.

Conditions behind the numbers

The author’s README lists 12GB or more VRAM, at least 32GB RAM and about 80GB of free storage. Actual capacity needs can be higher depending on the model and settings; these figures do not guarantee operation. Coder, the variant suggested for 32GB RAM, carries a warning about weaker Chinese, Japanese and Korean text performance.

V’s view

V’s view. I want to draw the rest of the PC alongside the question “Will it fit on my graphics card?” Following where data lives and which processor works on it makes the machinery clearer than one large number.

This is a reading of public documentation. We have not installed it or tested performance or Korean-language quality.