Milestone 01 — 2026.09.01-04 // Multi-Hardware
NVIDIA OK AMD Partial 12GB VRAM OK

FreeToken testing

Pushing Large MoE Models Beyond VRAM Limits on Consumer Hardware.

David's original post my tests: №1-1 №1-2 №2 №3 →next: FOL-Lab revolut.hot

// Research Question During My Sept 2026 Week Sprint:

How far can FreeToken push large Mixture-of-Experts (MoE) models beyond nominal VRAM limits on consumer hardware?

// Test Environment

[1] Olares One OK

  • OS: Linux / Olares OS
  • GPU: NVIDIA RTX 5090 Laptop
  • VRAM: 24 GB
  • RAM: 96 GB

[2] Minisforum M1-Lite PARTIAL

  • OS: Windows 11
  • GPU: AMD RX 7900 XTX
  • VRAM: 24 GB
  • RAM: 96 GB

[3] Morefine M900 OK

  • OS: Windows 11
  • GPU: RTX 3060 12GB eGPU
  • VRAM: 12 GB
  • RAM: 48 GB

// Key Results

Olares One [1] Breakthrough

135.5
tok/s decode
Qwen3.6-35B
139.7
tok/s prefill
~98s
cold load

gpt-oss-120b: ~54.4 tok/s decode, ~92.9 GB peak RAM.

FreeToken uses host RAM + expert offload to make very large MoE inference practical on 24GB NVIDIA GPU.

Minisforum [2] AMD/Windows Blocker

Worked: FreeToken installed, RX 7900 XTX detected as gfx1100. Triton kernels compiled. Large-model MoE/offload initialization progressed.

Failed: Qwen3.6-35B: HIP device fault during GPU warm-up. gpt-oss-120b: First large-MoE execution failed in Windows/gfx1100 HIP kernel path.

Control experiments prove hardware capability:

  • LM Studio / llama.cpp / Vulkan: Both models run coherently
  • Ollama / ROCm: Qwen3.6-35B runs at 100% GPU

Blocker is specifically FreeToken execution path on Windows + RDNA3/gfx1100, NOT the hardware.

Morefine M900 [3] 12GB VRAM Breakthrough

20.53
tok/s
Qwen3.6-35B
7.07
GB VRAM used
85.2%
CPU execution

Hybrid MoE backend:

  • PCIe expert gather: ~3.3 GB/s
  • CPU expert bandwidth: ~19.9 GB/s
  • Split: 14.8% PCIe fetch / 85.2% CPU execution

Counterexample to "12 GB VRAM is too small for serious local AI." 35B MoE at ~20 tok/s using only ~7 GB VRAM.

// Architecture

FreeToken automatically selects hybrid backend when VRAM is insufficient:

Expert Offload

Inactive experts stored in system RAM.

GPU Caching

Active experts cached on GPU.

CPU Execution

When PCIe bandwidth is bottleneck, CPU executes experts directly.

MoE Routing

Only 3.3B parameters activated per forward pass (Qwen3.6-35B-A3B).

// Learnings

  • 12GB VRAM + CPU offload is viable for 35B MoE models (specific configuration)
  • FreeToken excels on NVIDIA hardware with proper expert offload
  • Consumer hardware can run models previously thought to require datacenter GPUs
  • Windows + AMD RDNA3 has specific FreeToken compatibility issues (not hardware limitation)

// Next Steps

01 // Optimize Morefine [3]

Test quantization levels: Q4_K_M, Q5_K_M, Q8_0.

02 // AMD Debug

Report FreeToken + gfx1100 issue to upstream maintainers.

03 // vLLM Comparison

Benchmark FreeToken vs vLLM on same hardware.

04 // Multi-Model Serving

Multiple MoE models simultaneously with shared expert cache.

05 // Real-World Workloads

Integrate with FOL-Lab pipeline for production testing.

// Artifacts

// Bottom Line

FreeToken democratizes large MoE models by enabling consumer hardware to run 35B-120B parameter models through intelligent hybrid CPU/GPU execution. The 12GB VRAM breakthrough on Morefine M900 proves that accessible hardware can handle serious AI workloads.