
Qwen3.8-Flash-Next
72
—
20.6
Running IQ3_S GSQ RCO Qwen Flash Next at 128K context on a single RTX 3090, and I’m now seeing 72 tok/s decode.

Running IQ3_S GSQ RCO Qwen Flash Next at 128K context on a single RTX 3090, and I’m now seeing 72 tok/s decode.