The machine in the basement.

connecting…

A MINISFORUM AI X1 Pro-370 in a basement on a desk in North Metro Atlanta. One language model and one image model share a single integrated GPU. The line above is this machine, right now.

It is not a demo rig. It answers rabinforest.com (opens in a new tab) — a working portfolio site — every time somebody asks it a question.

MINISFORUM AI X1 Pro-370 standing vertically on its stand
MINISFORUM AI X1 Pro-370 · 64GB DDR5
rabinforest.com, showing its status strip reporting this box as warm with a prefix cachedOpen rabinforest.com (opens in a new tab)

This box runs rabinforest.com

A portfolio you can talk to, and the assistant behind it is the machine above — not a hosted API with a local-looking label. Look at the strip across the top of that screenshot: it is this box, naming the model it has loaded and whether the prompt is still cached.

When the box is cold, busy with a render, or mid-swap, the site hands off to Gemini and says so on that same strip. That is the arrangement, stated rather than hidden — a basement machine cannot promise an SLA.

Try it

Switch the model

Checking whether the box can switch…

Two engines, one GPU

They take turns. A render owns the GPU while it runs, which is why a question asked mid-render can come back from the cloud instead.

Language

gpt-oss-20b

Generation
25 tok/s
Active per token
~3.6B of 21B
First word
under 1s warm, ~27s cold
Primed with
~9,200 tokens
Quantisation
MXFP4

Swappable — not the one loaded

Images

SDXL Lightning

Typical render
~22s, 1024²
Steps
8 steps, ~1.86s/step
Sampler
euler, CFG 2.0
Checkpoint
~7GB
Throughput
one at a time, 10/day per IP

Fixed — there is only ever one, so nothing to report live.

A neon ramen shop at night, reflected in wet pavement
ramen shop
A lighthouse on rocks in a storm, waves breaking
lighthouse
A cat asleep in a sunlit bookshop window
bookshop cat
Orange koi swimming in a stone pond
koi pond

Rendered on this box — SDXL Lightning, 1024², 8 steps, about ~22s each. Generate your own (opens in a new tab).

Wondering how the two language models compare? Both are measured, side by side.

Where this box does the work

Four live pages on rabinforest.com. Each one is answered by the machine described above — the same GPU, the same model, the same queue.

More

  • How it works

    Why models this big run at all on an integrated GPU, and what they cost.

  • Operating notes

    What the first month teaches you: the failures that recur, and why.

  • Home AI box setup

    MINISFORUM AI X1 Pro-370 on Windows: measured tok/s, Tailscale, a Cloudflare tunnel, and what the box will not do.