vramarcade

arcade / model

qwen3.8-flash-next-strata

Qwen3.8-Flash-Next (Strata engine, IQ3_XXS, 176.9B-A6B)

The same Qwen3.8-Flash-Next weights family served by Strata instead of llama.cpp, via the live-tee router: ~4x decode, ~10x time-to-first-token, no depth penalty. IQ3_XXS is not UD-Q3_K_XL, so treat this column and the llama.cpp one as two different quantizations on two different engines, not as a clean A/B of engines.

show

Everything it built 8 result(s)