Skip to main content
PartForge
DEMO DATALive connectors are not configured. No prices or benchmark records on this screen are live.
AI Desktop / Separate workspace

Build local AI systems without mixing them into the normal PC page.

Deterministic model fit, multi-GPU topology, VRAM, RAM, PCIe, power, and cooling checks. All values shown here are demo fixtures.

AI build fingerprintDual-GPU inference server

2× NVIDIA RTX 6000 Ada · 256 GB ECC RAM

Aggregate VRAM96 GB

vLLM sharding enabled

Visible local models6

0 hidden by the no-fit gate

Peak power1420 W

240 V circuit review required

AI-only component plan

Server build checklist

Demo catalog
Server partSelected AI-capable optionRoleDemo price
Platform boardSupermicro M12SWA-TFWRX80 · 7× PCIe x16$1,149 demo
CPU socketsAMD Threadripper PRO 5975WX32 cores · 128 PCIe lanes$2,299 demo
ECC memory256 GB DDR4 ECC RDIMM8-channel · 8×32 GB$824 demo
Model storage2× 4 TB Gen4 NVMeLibrary + scratch$650 demo
AI acceleratorsNVIDIA RTX 6000 Ada 48 GBSelectable topology below$7,500 each demo
Chassis & airflow4U GPU server chassisFront-to-back high pressure$899 demo
Power & PDU2× 1600 W redundant PSU240 V circuit review required$1,450 demo
AI benchmarks

Fit and throughput, with uncertainty

Workload: Serving / agents

Estimated tokens/sec by visible model

Qwen3 8B74–110
Gemma 3 12B50–74
Qwen3 30B-A3B99–147
gpt-oss-20b99–147
Qwen3 32B19–28

VRAM residency and headroom

Qwen3 8B11.2 GB
Gemma 3 12B15.8 GB
Qwen3 30B-A3B24.2 GB
gpt-oss-20b17.6 GB
Qwen3 32B38.4 GB
Concurrent users58

Modeled demo range at batch-friendly load

Power efficiency77 tok/s/kW

Directional only; no measured power log

RAM headroom136 GB

After the largest visible model estimate

Models that fit

Runnable local model table

6 visible · no cloud-only models
ModelArchitectureRequired memoryFitEstimated throughputBackend
Qwen3 8BChat · coding · RAG · Apache-2.0Dense8B total / 8B active11.2 GB VRAM16 GB RAMLOCAL FIT72/100 modeled confidence74–110 tok/sDemo estimatevLLMQ5 · 32K
Gemma 3 12BChat · vision · GemmaDense12B total / 12B active15.8 GB VRAM20 GB RAMLOCAL FIT72/100 modeled confidence50–74 tok/sDemo estimatevLLMQ5 · 32K
Qwen3 30B-A3BAgents · long context · Apache-2.0MoE30B total / 3B active24.2 GB VRAM36 GB RAMLOCAL FIT72/100 modeled confidence99–147 tok/sDemo estimatevLLMQ5 · 32K
gpt-oss-20bReasoning · tools · Apache-2.0MoE20B total / 3.6B active17.6 GB VRAM27 GB RAMLOCAL FIT72/100 modeled confidence99–147 tok/sDemo estimatevLLMQ5 · 32K
Qwen3 32BCoding · reasoning · Apache-2.0Dense32B total / 32B active38.4 GB VRAM38 GB RAMLOCAL FIT72/100 modeled confidence19–28 tok/sDemo estimatevLLMQ5 · 32K
gpt-oss-120bReasoning · agents · Apache-2.0MoE120B total / 5.1B active87.2 GB VRAM120 GB RAMLOCAL FIT SHARDED72/100 modeled confidence99–147 tok/sDemo estimatevLLMQ5 · 32K