Language
gpt-oss-20b
- Generation
- 25 tok/s
- Active per token
- ~3.6B of 21B
- First word
- under 1s warm, ~27s cold
- Primed with
- ~9,200 tokens
- Quantisation
- MXFP4
Swappable — not the one loaded
connecting…
A MINISFORUM AI X1 Pro-370 in a basement on a desk in North Metro Atlanta. One language model and one image model share a single integrated GPU. The line above is this machine, right now.
It is not a demo rig. It answers rabinforest.com (opens in a new tab) — a working portfolio site — every time somebody asks it a question.

Open rabinforest.com (opens in a new tab)A portfolio you can talk to, and the assistant behind it is the machine above — not a hosted API with a local-looking label. Look at the strip across the top of that screenshot: it is this box, naming the model it has loaded and whether the prompt is still cached.
When the box is cold, busy with a render, or mid-swap, the site hands off to Gemini and says so on that same strip. That is the arrangement, stated rather than hidden — a basement machine cannot promise an SLA.
Try it
Checking whether the box can switch…
They take turns. A render owns the GPU while it runs, which is why a question asked mid-render can come back from the cloud instead.
Language
Swappable — not the one loaded
Images
Fixed — there is only ever one, so nothing to report live.




Rendered on this box — SDXL Lightning, 1024², 8 steps, about ~22s each. Generate your own (opens in a new tab).
Wondering how the two language models compare? Both are measured, side by side.
Four live pages on rabinforest.com. Each one is answered by the machine described above — the same GPU, the same model, the same queue.
The front door of rabinforest.com, and this box is its default engine. Ask it something and the answer comes from the basement — unless the box is cold or busy, in which case it quietly hands off to the cloud and says so.
One question put to this box and to hosted models at the same time, side by side. The speed difference is the whole point, and it is not flattering in every direction.
Paste a URL and a claim; the box reads the page and judges whether it holds up. It will say "can't tell" when a source is silent, which is the harder and more honest answer.
SDXL Lightning on the same integrated GPU. About ~22s a picture, and the language model goes quiet while it draws.
Why models this big run at all on an integrated GPU, and what they cost.
What the first month teaches you: the failures that recur, and why.
MINISFORUM AI X1 Pro-370 on Windows: measured tok/s, Tailscale, a Cloudflare tunnel, and what the box will not do.