The machine in the basement.

A MINISFORUM AI X1 Pro-370 under a desk in Metro Atlanta. One language model and one image model share a single integrated GPU. The header dot is this machine, right now.

MINISFORUM AI X1 Pro-370 standing vertically on its stand

Checking the box.

The numbers

Measured on this box. Not vendor claims.

Language

  • 33 tok/sgeneration, Qwen3-30B-A3B Q4
  • ~3B activeof 30B. Mixture of experts. That is why it fits.
  • 64GB DDR5both channels full. More RAM would buy nothing.
  • ~4 at oncethen latency falls over. One machine, no autoscale.
  • 50s cold, 0.3s warmprefix cache on a ~7,700-token prompt.
  • Radeon 890Mno discrete GPU. Bandwidth is the ceiling.

Images

  • ~22stypical 1024² render, SDXL Lightning
  • 8 stepseuler, CFG 2.0. ComfyUI. About 1.86s/step.
  • ~7GB checkpoint~20s to load cold. Unloaded after 2 minutes idle.
  • One at a time10/day per IP. A render owns the GPU.

Why a 30B fits in 64GB

Generating a token means reading the active weights. On this hardware the bus is the limit, about 60–70GB/s in practice. A dense 32B reads ~19GB per token and crawls. This mixture-of-experts model activates about 3B, reads ~1.5–2GB, and measures 33 tok/s.

A smaller dense 8B would be slower. Model size is the wrong metric here. Active parameters are the right one.

What it will do

  • Answer at conversational speed, when warm. About 33 tokens a second from a 30B-class model on an iGPU. A primed prefix makes the wait before the first word ~0.3s, not ~50s.
  • Draw a picture. SDXL Lightning, about 22 seconds, eight steps you can watch. One at a time, same GPU.
  • Tell the truth about itself. The header dot is live: primed, priming, drawing, or off. No fake “online.”
  • Get out of the way when it cannot. Other sites that call it fail open to a cloud model. Nobody waits on a cold box.

What it will not do

  • Beat a free cloud chat. Faster, stronger models exist with no queue. Capability is not the pitch.
  • Long-context interactive work. A ~16,000-token prompt takes minutes. Prefill is compute-bound; MoE does not help it.
  • Two jobs at once, well. A render and an answer share one GPU. The machine does one thing at a time.
  • Scale. Five people chatting is unusable. Ten is down.

How it was built

The full runbook is still one page. It will be cut into modules.

  • Home AI box setup

    MINISFORUM AI X1 Pro-370 on Windows: measured tok/s, Tailscale, a Cloudflare tunnel, and what the box will not do.