Cite
Quoted captures. Prefer this page or llms-full.txt.
/llms.txt · /llms-full.txt · 39 pairs · 56 models · 20 chips
Measured pairs


Qwen3.5 2B · 351 tok/s
RTX 5090 32GB · unknown · Q4_K_M GGUF
@fillagrew


Swift 1.5 Qwen3.8 Flash-Next · 220 tok/s
RTX 4090 24GB · Strata · IQ2_XS GGUF
@Knuckles_XBT
Get


Qwen3.8-27B · 177.6 tok/s
RTX A5000 · TensorFold · NVFP4
@Ja6ek
Get


Gemma-4 26B-A4B-it · 173 tok/s
RTX 5090 32GB · unknown · Q4_K_M GGUF
@fillagrew


Qwen3.8-Flash-Next · 142 tok/s
RTX 5090 32GB · Strata · IQ3_S GGUF
@arianpg
Get


Qwen3.8-27B · 139.3 tok/s
RTX A5000 · vLLM · NVFP4
@Ja6ek


Qwen3.8-Flash-Next · 126.5 tok/s
M5 Max · MTPLX V2.11.3 · unknown
@Youssofal_




Qwen3.8-27B · 102.9 tok/s
DGX Spark · TensorFold · MLX 4-bit g64 + DFlash2
@Oluwaphilemon1
HF


Qwen3.8-Flash-Next · 100.6 tok/s
RTX 3090 24GB · Strata · IQ2_XS GGUF
@draslan_eth
Get




Qwen3.8-27B · 89.9 tok/s
M5 Pro 64GB · TensorFold · MLX 4-bit + DFlash2
@aartiles24
HF


Qwen3.6-35B-A3B · 85.5 tok/s
M4 Pro · rapid-mlx · 4-bit MLX
@rapidmlx
HF


Qwen3.5-4B · 82.8 tok/s
M4 Pro · rapid-mlx · 4-bit MLX
@rapidmlx


Qwen3.8-Flash-Next · 80.9 tok/s
RTX 5090 32GB · unknown · NVFP4
@tekizaihq
Get


Qwen3.8-Flash-Next · 79.5 tok/s
DGX Spark · EXL3 · EXL3 3.05 bpw
@yume_arasaki


Qwen3.8-Flash-Next · 74.8 tok/s
DGX Spark · TensorFold · NVFP4 MTP-6
@redp314
Get


Qwen3.8-27B · 72 tok/s
DGX Spark · unknown · NVFP4 + DFlash2
@hasso5703
Get


Qwen3.8-Flash-Next · 72 tok/s
RTX 3090 24GB · Strata · IQ3_S GSQ RCO GGUF
@needmorevram
Get


AliceAI-Foundation-80B-A3B · 65 tok/s
M5 Max 128GB · llama.cpp · Q4_K_M GGUF
@aqty
HF


Qwen3.8-27B · 61.9 tok/s
M6 mini 32GB · TensorFold · MLX 4-bit + DFlash2
@jmurillocode
HF


GLM 5.3 Flash · 60.4 tok/s
2× DGX Spark · TensorFold · EXL3 TR3 4bpw + DFlash2
@MiaAI_lab
Get


GLM 5.3 Flash · 55.6 tok/s
2x DGX Spark · unknown · NVFP4
@PlusTen_AI
Get


Empero Qwen3.8-35B-A3B · 50 tok/s
RTX 3060 12GB · llama.cpp · Q4_K_M GGUF
@Oluwaphilemon1


Qwen3.5-9B · 49.3 tok/s
M4 Pro · rapid-mlx · 4-bit MLX
@rapidmlx


Qwen3-8B · 48.3 tok/s
M4 Pro · rapid-mlx · 4-bit MLX
@rapidmlx


MiMo-V2.6-Distill-Qwen-9B · 47 tok/s
RTX 5060 8GB · llama.cpp · Q5_K_M GGUF
@stfu0911
HF








Qwen3.6-35B-A3B · 37 tok/s
M1 Max 32GB · unknown · UD-IQ3_XXS GGUF
@moonsteroid
HF


GLM 5.3 · 30 tok/s
4x DGX Spark · sparkDash · Int4/Int8 mixed TP4
@majewskizby


Qwen3.8-Flash-Next · 27.7 tok/s
RTX 3060 12GB · Strata · IQ2_XS GGUF
@dec21ai
Get


Qwen3.8-27B · 23 tok/s
RTX 5060 Ti 16GB · llama.cpp · GSQ-RCO IQ3_XXS-mtp
@fntAInhead
HF


Qwen3.8-27B · 19 tok/s
RTX 5060 Ti 16GB · llama.cpp · UD-Q2_K_XL
@fntAInhead
HF


GLM 5.3 Flash · 18.6 tok/s
DGX Spark · EXL3 · EXL3 K2
@vcruz305
Get


GLM 5.3 Flash · 15 tok/s
2x DGX Spark · vLLM · NVFP4
@sudoingX


Nex-N2.5-mini · 14 tok/s
RTX 5060 Ti 16GB · llama.cpp · Q4_K_M GGUF
@fntAInhead


Qwen3.8-27B TurboFCFusion · 10 tok/s
RTX 5060 Ti 16GB · llama.cpp · IQ2_M GGUF
@fntAInhead
Models
- Qwen3.8-27B — Qwen · dense · 27B · new, hot · https://huggingface.co/Qwen/Qwen3.8-27B
- Qwen3.8-27B Mythos — medismera · dense · 27B · new, hot · https://huggingface.co/medismera/Qwen3.8-27B-OBLITERATED-Mythos-Class-Agentic
- Qwen3.8-35B-A3B APEX — IsValorum · MoE 35B-A3B · 35B / 3B act · new, hot · https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MTP-APEX-I-MiniPlus-V2.1-Abliterated-GGUF
- Huihui Qwen3.8-27B — huihui-ai · dense · 27B · new · https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated
- Qwen3.8-27B OBLITERATED — OBLITERATUS · dense · 27B · new · https://huggingface.co/OBLITERATUS/Qwen3.8-27B-OBLITERATED
- Qwen3.8-27B Uncensored — orcarouter · dense · 27B · new · https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored
- Qwen3.8-27B Heretic — 0bserverx · dense · 27B · new · https://huggingface.co/0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF
- Qwen3.8-27B Splash — audreyt · dense · 27B · new · https://huggingface.co/audreyt/Qwen3.8-27B-Splash-abliterated
- SuperQwen3.8-27B — Jiunsong · dense · 27B · new · https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated
- Flash-Next RVN — 0bserverx · MoE · unknown · new · https://huggingface.co/0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored
- Qwen3.8-27B TurboFC — community · dense · 27B · new
- Qwen3.8-27B Heretic ARA — trohrbaugh · dense · 27B · new · https://huggingface.co/trohrbaugh/Qwen3.8-27B-heretic-ara
- Qwen3.8-27B Coletti — JonathanColetti · dense · 27B · new, hot · https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored
- Qwen3.8-27B NoRefusal — sss22213 · dense · 27B · new · https://huggingface.co/sss22213/Qwen3.8-27B-Heretic-NoRefusal
- Qwen3.8-27B KCRN — heterodoxin · dense · 27B · new · https://huggingface.co/heterodoxin/qwen-3.8-27b-abliterated
- Qwen3.8-27B Fable — DavidAU · dense · 27B · new · https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
- Qwen3.8-9B Heretic — Noobito45 · dense · 9B · new · https://huggingface.co/Noobito45/Qwen3.8-9B-heretic-uncensored-NVFP4-GGUF
- DeepSeek v4.1 Flash abliterated — distributedcognition · dense · unknown · new · https://huggingface.co/distributedcognition/DeepSeek-V4.1-Flash-abliterated
- Gemma-4 26B Abliterix — wangzhang · MoE · 26B-A4B · new · https://huggingface.co/wangzhang/gemma-4-26B-A4B-it-abliterix
- Gemma-4 26B Uncensored — TrevorJS · MoE · 26B-A4B · hot · https://huggingface.co/TrevorJS/gemma-4-26B-A4B-it-uncensored-GGUF
- Gemma-4 31B Uncensored — TrevorJS · dense · 31B · new · https://huggingface.co/TrevorJS/gemma-4-31B-it-uncensored-GGUF
- Gemma-4 12B OBLITERATED — OBLITERATUS · dense · 12B · new · https://huggingface.co/OBLITERATUS/Gemma-4-12B-OBLITERATED
- Qwen3.8-27B Ultra Heretic — llmfan46 · dense · 27B · new · https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved
- Qwen3.8-27B Aggressive — 0xKitkat · dense · 27B · new · https://huggingface.co/0xKitkat/Qwen3.8-27B-Uncensored-Aggressive
- Qwen3.8-27B Twin Turbo — DavidAU · dense · 27B · new, hot · https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored
- GLM 5.3 Flash abliterated — dealignai · MoE · unknown · new, hot · https://huggingface.co/dealignai/GLM-5.3-Flash-ABLITERATED-FP8
- GLM 5.3 Flash Uncensored — orcarouter · MoE · 320B / 18B act · new · https://huggingface.co/orcarouter/GLM-5.3-Flash-Uncensored-GGUF
- GLM 5.3 Uncensored — dealignai · MoE · 753B · new · https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-FP8
- GLM 5.3 EXL3 abliterated — drowzeys · MoE · unknown · new · https://huggingface.co/drowzeys/keys-GLM-5.3-EXL3-Abliterated
- Qwen3.6-35B-A3B — Qwen · MoE 35B-A3B · 35B / 3B act · https://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Qwen3.6-35B-A3B abliterated — community · MoE 35B-A3B · 35B / 3B act · hot
- Empero Qwen3.8-35B-A3B — empero-ai · MoE distill · 35B / 3B act · new, hot
- Qwen3.8-Flash-Next — Qwen · dense · unknown · new, hot
- Qwen3.5 2B — Qwen · dense · 2B · new, hot
- Qwen3.5-4B — Qwen · dense · 4B · new
- Qwen3.5-9B — Qwen · dense · 9B · new
- Qwen3-8B — Qwen · dense · 8B
- Gemma-4 26B-A4B-it — Google / community · MoE · 26B-A4B · new, hot
- Nex-N2.5-mini — community · MoE post-train · 35B-A3B class · new
- Ornith — ornith-ai · unknown · unknown · new · https://huggingface.co/ornith-ai
- GLM 5.3 Flash — Zhipu · dense · unknown · new, hot
- DeepSeek v4.1 Flash — DeepSeek · dense · unknown · new, hot
- MiMo-V2.6-Distill-Qwen-9B — Xiaomi MiMo · dense · 9B · new, hot
- AliceAI-Foundation-80B-A3B — Yandex · MoE 80B-A3B · 80B / 3B act · new · https://huggingface.co/Yamada114514/AliceAI-Foundation-80B-A3B-Base-GGUF
- MiMo-V2.6-Flash — Xiaomi MiMo · MoE · unknown · new
- Ling-3.0-Flash — inclusionAI · dense · unknown · new · https://huggingface.co/inclusionAI/Ling-3.0-flash-int4
- Nemotron 3.5 Lightning — NVIDIA · MoE 30B-A3B · 30B / 3B act · new
- GLM-5.2 — Zhipu · MoE · unknown · new
- Kimi K3 — Moonshot · MoE · 2.8T / 104B act · new, hot · https://huggingface.co/moonshotai/Kimi-K3
- CYBER-FROST 3.8 — Blackfrost-AI · MoE · ~180B · new · https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
- Qwen-Image-2.1 — Qwen · image · 7B visual · new, hot · https://huggingface.co/Qwen/Qwen-Image-2.1
- Qwen-Image-2.1 Uncensored — abenzerps · image · GGUF · new, hot · https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF
- Ternary Bonsai 2 27B — PrismML · dense ternary · 27B · new, hot · https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
- MiniMax-M2.5 — MiniMax · unknown · unknown · new
- Swift 1.5 Qwen3.8 Flash-Next — UkisAI · dense · 27B-class · new, hot
- GLM 5.3 — Zhipu · MoE · 753B · new
Hardware
- M4 Pro — Apple · 24–64 GB unified · 4 pairs
- M5 Max — Apple · 64–128 GB unified · 2 pairs
- DGX Spark — NVIDIA · 128 GB unified · 11 pairs
- RTX 5090 — NVIDIA · 32 GB VRAM · 5 pairs
- RTX 4090 — NVIDIA · 24 GB VRAM · 2 pairs
- RTX 5060 Ti — NVIDIA · 8–16 GB VRAM · 4 pairs
- RTX 5060 — NVIDIA · 8 GB VRAM · 1 pairs
- RTX 3060 — NVIDIA · 8–12 GB VRAM · 2 pairs
- GTX 1660 SUPER — NVIDIA · 6 GB VRAM · 1 pairs
- 2x Tesla V100 32GB — NVIDIA · 64 GB VRAM (2x32) · 0 pairs
- M1 Max 64GB — Apple · 64 GB unified · 0 pairs
- M1 Max 32GB — Apple · 32 GB unified · 1 pairs
- M4 mini 24GB — Apple · 24 GB unified · 0 pairs
- M3 Ultra — Apple · 96-512 GB unified · 0 pairs
- RTX 5070 12GB — NVIDIA · 12 GB VRAM · 0 pairs
- RTX 3090 24GB — NVIDIA · 24 GB VRAM · 2 pairs
- 2x RTX 3090 24GB — NVIDIA · 48 GB VRAM (2x24) · 0 pairs
- M5 Pro 64GB — Apple · 64 GB unified · 1 pairs
- M6 mini 32GB — Apple · 32 GB unified · 1 pairs
- RTX A5000 — NVIDIA · 24 GB VRAM · 2 pairs
Tape
- 2026-09-27 @jmurillocode — Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on 32GB Mac mini M6: 61.9 code / 30.8 chat. https://x.com/jmurillocode/status/2104138445928956021
- 2026-09-27 @aartiles24 — Qwen3.8-27B TensorFold MLX 4-bit DFlash2 on MacBook M5 Pro 64GB: 89.9 tok/s. https://x.com/aartiles24/status/2104179801388884125
- 2026-09-27 @Oluwaphilemon1 — Qwen3.8-27B TensorFold DFlash2 on one Spark: 102.9 tok/s. https://x.com/Oluwaphilemon1/status/2104031407194427892
- 2026-09-28 @PlusTen_AI — GLM 5.3 Flash NVFP4 on 2x DGX Spark: 55.6 coding / 33.6 Korean prose. https://x.com/PlusTen_AI/status/2104531371146482165
- 2026-09-28 @DaSun64125381 — Flash-Next IQ3_S Strata on RTX 5090 256K: 104 decode / 1657 prefill. https://x.com/DaSun64125381/status/2104536066028192175
- 2026-09-29 @moonsteroid — Qwen3.6-35B-A3B UD-IQ3_XXS on MBP M1 Max 32GB: 37 decode / 350 prefill at 100k. https://x.com/moonsteroid/status/2104903466661314585
- 2026-09-29 @needmorevram — Flash-Next IQ3_S GSQ RCO Strata on one RTX 3090 at 128K: 72 decode. https://x.com/needmorevram/status/2104898095318544540
- 2026-09-30 @sudoingX — GLM 5.3 Flash NVFP4 stock vLLM on 2x Spark: 15 decode MTP-off. https://x.com/sudoingX/status/2105246609730924695
- 2026-09-30 @dec21ai — Flash-Next IQ2_XS Strata on RTX 3060 12GB Thinking Medium: 27.7 decode. https://x.com/dec21ai/status/2105260285439422767
- 2026-09-30 @Knuckles_XBT — Swift 1.5 Flash-Next IQ2_XS Strata on RTX 4090 24GB at 256K: ~220 decode / ~3830 prefill. https://x.com/Knuckles_XBT/status/2105245662657011841
- 2026-10-01 @Ja6ek — Qwen3.8-27B NVFP4 on RTX A5000 at 262K: TensorFold 177.57 / vLLM 139.31. https://x.com/Ja6ek/status/2105623160649298395
- 2026-10-01 @arianpg — Flash-Next IQ3_S Strata on RTX 5090 + 128GB RAM: 142 decode. https://x.com/arianpg/status/2105625971261161895
- 2026-10-01 @redp314 — Flash-Next NVFP4 MTP-6 TensorFold 0.3.6.3 on one DGX Spark: 74.8 single-stream. https://x.com/redp314/status/2105569144380739979
- 2026-10-02 @draslan_eth — Flash-Next IQ2_XS Strata 0.1.31 on one RTX 3090 at 32K: 100.6 decode. https://x.com/draslan_eth/status/2105946305646502091
- 2026-10-02 @majewskizby — Full GLM 5.3 753B Int4/Int8 TP4 on 4× DGX Spark: 30.0 prose decode, thinking off. https://x.com/majewskizby/status/2105962150158041477
- 2026-09-23 @hasso5703 — Qwen3.8-27B on one Spark: SGLang, NVFP4, DFlash2, 72 tok/s greedy median. Repo claim. https://github.com/hasso5703/dgx-spark-qwen38
- 2026-09-10 @vcruz305 — GLM-5.3-Flash EXL3 K2 on one Spark: SGLang, MTP k=2, 18.57 tok/s. Repo claim. https://github.com/vcruz305/GLM-5.3-Flash-EXL3-K2-SGLang-DGX-Spark-recipe
- 2026-09-22 @stfu0911 — MiMo-V2.6 Distill Qwen-9B Q5_K_M on RTX 5060 8GB: ~47 decode, ~1600 prefill, 262k. https://x.com/stfu0911/status/2102409155143409688
- 2026-09-22 @aqty — AliceAI 80B-A3B Q4_K_M on M5 Max 128GB: ~65 tok/s, 45.1 GiB. https://x.com/aqty/status/2102413529206976799
- 2026-09-21 @tekizaihq — Flash-Next NVFP4 on one 5090: 80.9 tok/s single-stream. https://x.com/tekizaihq/status/2101844907618939233
- 2026-09-20 @yume_arasaki — Flash-Next EXL3 on one Spark: 79.5 code reproduced; 102.6 repetitive clamps; 71.7 prose. https://x.com/yume_arasaki/status/2101741448811229219
- 2026-09-20 @ViC305 — EXL3 Spark recipe reproduced at 79.5 on Yume’s box. https://x.com/ViC305/status/2101747348322103706
- 2026-09-30 @MiaAI_lab — GLM-5.3-Flash EXL3 on 2× DGX Spark with TensorFold: 60.4 tok/s prose, 114.7 structured single-stream; 4 streams 108.8 / 227.9 aggregate. Repo claim. https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
- 2026-09-20 @MiaAI_lab — Spark street: Flash-Next solo; GLM 5.3 / DeepSeek v4.1 Flash on 2×+; orch+worker+scanner at 6×. https://x.com/MiaAI_lab/status/2101821681404624976
- 2026-09-21 @sethforprivacy — 8 Sparks: 4× GLM ring, 2× Flash-Next, 2× DeepSeek v4f. https://x.com/sethforprivacy/status/2101834433087070708
- 2026-09-17 @fillagrew — 5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173. https://x.com/fillagrew/status/2100478211134025937
- 2026-09-17 @fntAInhead — 5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s. https://x.com/fntAInhead/status/2100499089431359808
- 2026-09-17 @Youssofal_ — Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k. https://x.com/Youssofal_/status/2100468205030719533
- 2026-09-17 @Oluwaphilemon1 — Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained. https://x.com/Oluwaphilemon1/status/2100498960645308573
- 2026-09-17 @Oluwaphilemon1 — Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class. https://x.com/Oluwaphilemon1/status/2100409183115821394
- 2026-09-17 @bonellisystems — Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode. https://x.com/bonellisystems/status/2100417191434768450
- 2026-09-17 @mine_craft_bui — Ornith ~40–46 tok/s on a 1660 SUPER 6GB. https://x.com/mine_craft_bui/status/2100492671534166056
- 2026-09-17 @Oluwaphilemon1 — Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode. https://x.com/Oluwaphilemon1/status/2100374395482939401
- 2026-09-16 @rapidmlx — Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak. https://x.com/rapidmlx/status/2100252655990010209
Street stacks
Recipes on DGX Spark. Not tok/s pairs. Source @MiaAI_lab.
- 1× Solo — qwen3.8-flash-next (Solo load); qwen3.8-27b (27B, SGLang); ling-3.0-flash; glm-5.3-flash (EXL3 K2, one box)
- 2× Dual — glm-5.3-flash (GLM on both); glm-5.3-flash (TensorFold, 4 streams); deepseek-v4.1-flash; qwen3.8-flash-next; mimo-v2.6-flash
- 3× Triple — glm-5.3-flash + qwen3.8-flash-next (GLM orch on 2×, Qwen worker on 1×); glm-5.3-flash (All three); deepseek-v4.1-flash (Homogeneous); glm-5.2 (NVFP4)
- 4× Quad — glm-5.3-flash + qwen3.8-flash-next (GLM orch 2× + Qwen worker 2×); glm-5.3-flash (Homogeneous); deepseek-v4.1-flash (Homogeneous)
- 6× 6+ — glm-5.3-flash + qwen3.8-flash-next + deepseek-v4.1-flash (GLM orch + Qwen worker + DeepSeek scanner)
Back to Home · Manifesto · Terms