Pushing Large MoE Models Beyond VRAM Limits on Consumer Hardware.
How far can FreeToken push large Mixture-of-Experts (MoE) models beyond nominal VRAM limits on consumer hardware?
gpt-oss-120b: ~54.4 tok/s decode, ~92.9 GB peak RAM.
FreeToken uses host RAM + expert offload to make very large MoE inference practical on 24GB NVIDIA GPU.
Worked: FreeToken installed, RX 7900 XTX detected as gfx1100. Triton kernels compiled. Large-model MoE/offload initialization progressed.
Failed: Qwen3.6-35B: HIP device fault during GPU warm-up. gpt-oss-120b: First large-MoE execution failed in Windows/gfx1100 HIP kernel path.
Control experiments prove hardware capability:
Blocker is specifically FreeToken execution path on Windows + RDNA3/gfx1100, NOT the hardware.
Hybrid MoE backend:
Counterexample to "12 GB VRAM is too small for serious local AI." 35B MoE at ~20 tok/s using only ~7 GB VRAM.
FreeToken automatically selects hybrid backend when VRAM is insufficient:
Inactive experts stored in system RAM.
Active experts cached on GPU.
When PCIe bandwidth is bottleneck, CPU executes experts directly.
Only 3.3B parameters activated per forward pass (Qwen3.6-35B-A3B).
Test quantization levels: Q4_K_M, Q5_K_M, Q8_0.
Report FreeToken + gfx1100 issue to upstream maintainers.
Benchmark FreeToken vs vLLM on same hardware.
Multiple MoE models simultaneously with shared expert cache.
Integrate with FOL-Lab pipeline for production testing.
/root/freetoken-benchmarks//root/freetoken-configs/FreeToken democratizes large MoE models by enabling consumer hardware to run 35B-120B parameter models through intelligent hybrid CPU/GPU execution. The 12GB VRAM breakthrough on Morefine M900 proves that accessible hardware can handle serious AI workloads.