Why you cannot simply drop a gaming card into a server
This is the first thing you need to be able to explain to a customer, because the question "why not a 4090 or 5090, they're cheaper" comes up every single time. There are six differences, and every one of them is about operations rather than benchmark scores.
Passive cooling
NVIDIA lists the thermal solution of the server cards as passive and describes the 6000 SE as designed for multi-GPU server deployments requiring passive cooling. There are no fans on the board; the chassis moves the air. Outside a server with engineered airflow the card will not run.
ECC memory
The 6000 SE specification states 96 GB of GDDR7 with error-correcting code. Under a 24/7 load — a multi-day inference run or a render job — a single bit flip without ECC means a corrupted result with no error message of any kind.
Different output sets
The 6000 SE lists four DisplayPort 2.1 outputs, so it can also work as a graphics card in the data centre. The 4500 SE has no display output row in its specification at all: it is a purely headless accelerator.
Partitioning: MIG and vGPU
The card can be carved into isolated instances with their own memory, cache and cores, or shared with remote users through vGPU. Up to four instances on the 6000 SE, up to two at 16 GB on the 4500 SE.
System certification
RTX PRO Servers hold NVIDIA-Certified status and are validated with partner tools as part of the AI Factory Validated Design. That is a guarantee of compatibility with the AI Enterprise stack and NVIDIA networking.
Configurable power
The 6000 SE's power draw is published as a 400–600 W range. The card is tuned to the chassis thermal budget rather than the other way round. Consumer cards offer nothing equivalent.
The argument to make with a customer is not performance, it is predictability and manageability. A gaming card is cheaper to buy and more expensive to own: it cannot be shared between users, it has no ECC, it cannot be power-capped to the chassis, and it is not part of a certified configuration backed by the server vendor and NVIDIA.
Four generational changes people pay for
Blackwell is the generation that succeeded Ada Lovelace. For selling server cards, exactly four of its properties matter; everything else is detail for engineers.
| Technology | What NVIDIA says | What the customer gets |
|---|---|---|
| 5th-gen Tensor Cores and FP4 | Up to 3× the performance of the previous generation, adding FP4 precision. The 6000 SE carries the second-generation Transformer Engine | Model weights occupy a quarter of the memory they would in FP16. This, and not the CUDA core count, is where the multi-fold gain in LLM inference comes from |
| GDDR7 memory | Significantly increased bandwidth and capacity | Made 96 GB on a single card possible — a capacity previously available only on expensive SXM platforms |
| 4th-gen RT Cores | Up to 2× the performance of the previous generation; RTX Mega Geometry enables up to 100× more ray-traced triangles | Rendering, digital twins, simulation, Omniverse. Precisely what purely computational accelerators lack |
| NVENC 9 and NVDEC 6 | 4:2:2 support in H.264 and HEVC, improved AV1 quality, double the H.264 decode throughput | Video analytics, transcoding, streaming, media production — segments where the card sells without any AI conversation at all |
The key to understanding the whole range: RTX PRO cards can do AI compute and graphics. That versatility is their core competitive property and the main lever in a sales conversation.
The two cards, side by side
Both are built on Blackwell, both are passive, both are PCIe Gen5 x16, both support MIG. The similarity ends there: these are cards for different thermal budgets and different jobs.
| Specification | RTX PRO 6000 Blackwell SE | RTX PRO 4500 Blackwell SE |
|---|---|---|
| CUDA cores | 24,064 | 10,496 |
| RT Cores (4th gen) | 188 | 82 |
| Memory | 96 GB GDDR7 with ECC | 32 GB GDDR7 |
| Memory bandwidth | 1,597 GB/s | 800 GB/s |
| Memory interface | 512-bit | 256-bit |
| FP4 Tensor Core | 4 PFLOPS | 1.6 PFLOPS |
| FP8 Tensor Core | 2 PFLOPS | 811 TFLOPS |
| FP16 / BF16 Tensor Core | 1 PFLOPS | 406 TFLOPS |
| TF32 Tensor Core | 234 TFLOPS | 203 TFLOPS |
| Single precision FP32 | 120 TFLOPS | 51 TFLOPS |
| Peak RT Core performance | 355 TFLOPS | 154 TFLOPS |
| Power consumption | 400–600 W (configurable) | 165 W |
| Form factor | Air: dual-slot, 4.4" × 10.5" · Liquid: single-slot, FHXL | Single-slot, 4.4" × 10.5" |
| Thermal solution | Passive | Passive |
| Interface | PCIe Gen 5 x16 | PCI Express 5.0 x16 |
| MIG | Up to 4 isolated instances | Up to 2 instances at 16 GB |
| Video engines | 4 NVENC / 4 NVDEC | 3 NVENC / 3 NVDEC |
| Display outputs | 4 × DisplayPort 2.1 | Not listed in the specification |
| Confidential compute | Not listed in the specification | Capable |
| Power connector | Not listed in the specification | 1 × PCIe CEM5 16-pin |
| Predecessor | NVIDIA L40S | NVIDIA L4 |
The "not listed" rows are deliberate: NVIDIA publishes slightly different parameter sets for the two cards, and filling the gaps by analogy is exactly how wrong figures find their way into reviews.
Careful: three different products with nearly identical names
This is the most common error in specifications and quotations. "RTX PRO 6000 Blackwell" is a family of three cards, with the server-class 4500 sitting separately alongside.
| Specification | Server Edition | Workstation Edition | Max-Q Workstation |
|---|---|---|---|
| Memory | 96 GB GDDR7 ECC | 96 GB GDDR7 ECC | 96 GB GDDR7 ECC |
| Power consumption | 400–600 W | 600 W | 300 W |
| Thermal | Passive | Double flow-through | Active |
| Dimensions | 4.4" × 10.5", dual slot | 5.4" × 12", dual slot | 4.4" × 10.5", dual slot |
| Display outputs | 4 × DisplayPort 2.1 | 4 × DisplayPort 2.1 | 4 × DisplayPort 2.1 |
| Bus | PCIe Gen 5 x16 | PCIe Gen 5 x16 | PCIe Gen 5 x16 |
| Purpose per NVIDIA | Multi-GPU servers requiring passive cooling: inference, fine-tuning, distributed rendering, HPC, virtual workstations | Maximum performance in a single-GPU workstation | Dense workstation configurations, up to four GPUs |
Max-Q is a practical compromise for a customer who needs density but has no server rack: half the heat at the same memory capacity, up to four cards in a workstation. Note also NVIDIA's wording for the Server Edition — fine-tuning is named explicitly, so it is within the stated use case, unlike training from scratch.
The memory-per-watt ratio of the two server cards is close; what differs radically is density per slot: 165 W in one slot against 600 W in two. Choosing between them is really choosing between "one large model" and "many parallel streams". On compute the gap is two- to two-and-a-half-fold; on power draw it is nearly fourfold.
Position within the NVIDIA data centre range
The most valuable skill in selling these cards is being able to say when they are the wrong answer. The data centre range splits into two branches: universal RTX PRO PCIe cards, and compute accelerators built for training.
| Product | Role as described by NVIDIA | Graphics and RT | Interconnect |
|---|---|---|---|
| RTX PRO 4500 SE | Inference on small and medium models, data processing and data science, video analytics, virtualisation. For data centre, edge and cloud | Yes | PCI Express 5.0 x16 |
| RTX PRO 6000 SE | Inference, fine-tuning, distributed rendering, HPC, virtual workstations. The universal data centre GPU | Yes | PCIe Gen 5 x16 |
| H series and B series | Separate NVIDIA product lines for scenarios that require GPU-to-GPU connectivity over NVLink | Limited | NVLink / NVSwitch |
The specifications of both cards list PCI Express only as the interconnect; there is no NVLink row. The cards communicate over PCIe and their memory does not pool. Eight cards at 96 GB is not "768 GB for one model" — it is eight accelerators that are independent in memory. The large per-card capacity compensates for this in parallel inference, but a single model that does not fit in 96 GB needs either tensor parallelism with PCIe overhead, or an NVLink platform.
Arguments for the generational move
- Against the L40S: NVIDIA states that RTX PRO servers run simulation and synthetic data generation up to 4× faster than L40S systems, and describes the 6000 SE as delivering a significant leap over the L40S in multimodal agentic and generative applications.
- Against the L4: for the 4500 SE NVIDIA claims over 5× the performance of the previous-generation L4 in inference on small and medium models, when using NVIDIA's optimised frameworks and NIM microservices.
- Against CPU-only systems: for AI video understanding NVIDIA claims up to 100× the performance of CPU-only systems, with a reduction in footprint and energy consumption of more than 95%. For vector databases, up to 50×. This is the strongest argument available when talking to a customer running a fleet of CPU servers.
An important caveat on those last two numbers: the 50× and 100× comparisons are between an eight-card 4500 SE server and a CPU system (for vector databases, an AMD 9654 with 192 vCPUs, 33 million vectors in Milvus 2.5.25, HNSW against cuVS CAGRA; for video, a dual-socket AMD 9654 running Qwen3-VL-8B). That is a legitimate platform comparison, but it is not card against card, and a competent engineer on the customer's side will ask.
The features that actually sell the card
Specifications sell to an engineer. A commercial buyer is sold on the items below — they are what turns hardware into a service that can be resold or shared across departments.
| Technology | Mechanism | Value to the customer |
|---|---|---|
| MIG | Hardware partitioning of the GPU into isolated instances, each with its own high-bandwidth memory, cache and compute cores plus guaranteed quality of service. Up to four on the 6000 SE, up to two at 16 GB on the 4500 SE | Guaranteed quality of service per tenant. One card serves several independent workloads with no interference — the basis for billing and multi-tenancy |
| vGPU | Paired with NVIDIA RTX Virtual Workstation and Virtual PC software, the card is virtualised for remote users. MIG-backed time-slicing lets AI and graphics workloads run simultaneously. vGPU 20 adds the AI Virtual Workstation Toolkit and fixed-share scheduling | Engineers and designers work on thin clients while CAD licences and data never leave the perimeter. Very often this use case, not AI, is what pays for the purchase |
| ECC | Error correction on graphics memory, stated for the 96 GB of GDDR7 on the 6000 SE | A precondition for 24/7 production workloads and a standard line item in infrastructure tenders |
| Confidential computing | Listed as capable in the 4500 SE specification table | An argument in regulated industries: protecting data and models during processing |
| NVIDIA AI Enterprise | A suite of tools, libraries and frameworks including NIM and NeMo microservices | A separate line on the invoice and, at the same time, the answer to "who is going to support this" |
| NVIDIA-Certified | RTX PRO Servers hold the status and are validated with partner tools under the AI Factory Validated Design | Guaranteed compatibility with NVIDIA networking, BlueField-3 and the AI Enterprise stack. Removes integration risk from customer and integrator alike |
| NVIDIA Run:ai | Workload and GPU orchestration that raises utilisation and overall throughput | Filling idle cards is a direct argument for payback |
Present MIG in money, not gigabytes. One RTX PRO 6000 SE split into four instances is four AI developer workstations, or four independent inference services, on a single physical card. And the inverse rule: you cannot plan four users onto a 4500 SE — it supports two instances at most. This detail is regularly lost when quotations are drafted.
The server as a system, not a box with cards in it
This is where "knowing the hardware" ends and "understanding server technology" begins. The cards are the smaller part of the engineering problem. The rest is chassis, network, power and cooling. NVIDIA documents the reference configurations; the specific machines are documented by their makers.
NVIDIA Four reference configurations
NVIDIA defines the RTX PRO Server as a universal data centre platform and gives four examples built on modular MGX reference designs. These are the benchmark against which any vendor proposal should be judged.
| Configuration | GPUs | Networking | DPU |
|---|---|---|---|
| MGX 6U | 8 × RTX PRO 6000 SE, liquid-cooled | ConnectX-8 SuperNICs with PCIe Gen 6 switch | BlueField-3 |
| MGX 4U | 8 × RTX PRO 6000 SE | ConnectX-8 SuperNICs with PCIe Gen 6 switch | BlueField-3 |
| MGX 2U | 2 × RTX PRO 6000 SE | ConnectX-7 or BlueField-3 SuperNIC | BlueField-3 |
| MGX 2U | 8 × RTX PRO 4500 SE | ConnectX-7 SuperNICs | BlueField-3 |
The last row is easy to overlook and shouldn't be: eight of the smaller cards fit into two rack units. That is 256 GB of memory and 12.8 PFLOPS of FP4 in 2U, drawing 1.3 kW on the GPUs against 4.8 kW for eight of the larger cards. For video analytics and small-model inference it is the densest official configuration.
NVIDIA What the switch inside the SuperNIC buys you
NVIDIA describes the major architectural change of recent generations as follows: the ConnectX-8 SuperNIC PCIe Switch is a board integrating multiple SuperNICs, each combining ultra-fast networking with a 48-lane PCIe Gen 6 switch. It eliminates discrete PCIe switches, doubles inter-GPU bandwidth, simplifies board design and enables up to 800 Gb/s. Alongside it sits the BlueField-3 DPU, which per NVIDIA offloads networking, data access and security from the CPU and provides zero-trust isolation. Spectrum-X completes the networking subsystem: through RoCE adaptive routing and congestion control it accelerates storage performance by nearly 50%.
Eight-GPU PCIe servers come in two topologies, and the difference is worth being able to explain. In a switched design the cards connect to dedicated PCIe switch silicon which then uplinks to the CPUs: any card reaches any other at full width without leaving the switch, and there are more lanes available than the CPUs themselves provide. In a direct-attach design each card plugs into a CPU root complex: cheaper and one hop closer to host memory, but the cards divide between sockets and traffic between the halves crosses the inter-socket link. The AGS-6220V2 is direct-attach; the AGS-4UMGX-R1 is switched. That this is not pedantry is visible from both sides of the industry: NVIDIA built the ConnectX-8 SuperNIC precisely to remove discrete switches while keeping their benefit, and AWS separately notes the very low peer-to-peer latency between GPUs sharing a PCIe switch.
Vendors Three classes of machine
2U, 2–4 GPUs
Conventional rack servers from Cisco, Dell, HPE, Lenovo and Supermicro — all of them listed among NVIDIA's leading system partners. They go into the customer's existing rack. The entry point into accelerated computing for a company that already runs its own data centre.
4U MGX, 8 GPUs
A purpose-built platform on the NVIDIA MGX reference design with ConnectX-8 and BlueField-3. A cluster building block rather than "a server with graphics cards". Examples: INNO3D AGS-4UMGX-R1, Gigabyte XL44-SX2-AAS1, ZOTAC ZRS-MGX-R1, Supermicro SYS-522GA-NRT.
6U, 8 GPUs
Platforms such as the INNO3D AGS-6220V2. The same eight cards in half again the volume: airflow headroom, serviceability and price instead of density and high-speed networking. A self-sufficient working machine, not a cluster node.
Vendor One platform in detail: INNO3D AGS-6220V2
It is instructive to set a machine of a different philosophy next to the MGX reference. This is a 6U chassis for eight RTX PRO 6000 Blackwell Server Edition cards, and nearly every decision in it runs opposite to the MGX logic. Figures are from the INNO3D datasheet for SKU AGS-6220V2P2-B00.
| Subsystem | Implementation | What it means in practice |
|---|---|---|
| Supply type | Barebone. Chassis, motherboard, four PSUs, a 25 Gb/s OCP module, VROC key, rails and heatsinks pre-installed. No CPUs, memory or drives | A platform, not a finished server. CPUs, DDR5, NVMe and the eight cards are separate quotation lines with separate lead times |
| Form factor | 6U, 900 × 438 × 264 mm | Eight cards in 6U against eight in 4U for the reference. The 900 mm depth means checking rack depth including cable management |
| GPU slots | 8 × PCIe 5.0 x16, FHFL, double width. Direct CPU attach, no PCIe switch | Every card gets a genuine x16. But 8 × x16 is 128 lanes across two sockets: the cards split 4+4 and traffic between the halves crosses the inter-socket link. For tensor parallelism that is a bottleneck |
| CPUs | 2 × Intel Xeon Scalable 4th and 5th generation, LGA 4677, C741 chipset. 350 W TDP limit per socket | The ceiling rules out some higher-end SKUs — verify the chosen processor before ordering |
| Memory | 32 DDR5 RDIMM / RDIMM-3DS slots, 8 channels per CPU. RDIMM up to 96 GB, RDIMM-3DS up to 256 GB per module. 5,600 MT/s at 1DPC and 4,400 at 2DPC on 5th-gen Xeon | A ceiling near 8 TB. Filling every slot drops the speed from 5,600 to 4,400 |
| Networking | One OCP slot with a 25 Gb/s module included, plus a 1GbE BMC management port | The decisive difference from the reference. No ConnectX-8, no BlueField-3, no 400 Gb/s. A standalone machine, not a cluster element |
| Storage | 12 × 3.5/2.5" bays (8 × SATA-3 + 4 × NVMe), one internal M.2 NVMe, Intel VROC RAID 0/1/10/5 | Eight of the twelve bays are SATA only, leaving four NVMe bays for a working dataset |
| Power | 4 × 2,700 W CRPS, 3+1 redundant, 80 PLUS Platinum | 8,100 W usable. The arithmetic: 8 × 600 W of GPU plus 2 × 350 W of CPU plus memory, drives and fans is around 6.5 kW. The feed must be 200–240 V with C19/C20 connectors |
| Cooling | 16 hot-swap fans: 8 internal and 8 at the rear | The main advantage of 6U: plenty of air, fans replaceable without downtime |
| Management | Dedicated IPMI port, VGA, 2 × USB 3.2 Gen1. A TPM header with SPI interface; the TPM 2.0 module itself is optional | Where a tender requires TPM, the module is ordered separately |
| Operating systems | Windows Server 2022/2019, RHEL 9.2–8.6, SLES 15 SP4, Ubuntu 22.04 LTS, VMware ESXi 8.0U1 / 7.0U3i, Citrix Hypervisor 8.2 LTSR CU1 | ESXi and Citrix support signals the platform is also intended for VDI |
The title page is issued for SKU AGS-6220V2P2-B00 (item 0K219-B001) while the specification page header reads AGS-6220V2P2-B10. Likely a typo or an adjacent bundle revision, but the part number must be confirmed in writing when ordering: it determines what is actually pre-installed. There is also a minor memory inconsistency — one line states 4,800/4,400 MHz while the line below gives 5,600 MT/s at 1DPC for 5th-generation Xeon.
Vendor Same maker, different philosophy: INNO3D AGS-4UMGX-R1
INNO3D also builds a second platform for the same cards — a 4U on the NVIDIA MGX reference design, with ConnectX-8 SuperNICs and BlueField-3. Comparing two machines from one vendor is more instructive than any cross-brand comparison: you see exactly what the extra money buys, with no allowance for differences in build quality or service.
| Subsystem | AGS-4UMGX-R1 · MGX | AGS-6220V2 · 6U |
|---|---|---|
| Form factor | NVIDIA MGX 4U, 858 × 440 × 175 mm | 6U, 900 × 438 × 264 mm |
| GPU topology | 8 × PCIe Gen 5 x16 for eight RTX PRO 6000 SE, plus a dedicated Gen 5 x16 slot for BlueField-3 | 8 × PCIe 5.0 x16, direct CPU attach with no switch |
| Networking | ConnectX-8 SuperNIC: eight 400 Gb/s QSFP ports, plus 2 × 10GbE and a 1GbE BMC port | A 25 Gb/s OCP module plus a 1GbE BMC port |
| DPU | BlueField-3 in a dedicated slot | None |
| CPUs | 2 × Intel Xeon 6 (Granite Rapids / Sierra Forest), 6500 and 6700 series, LGA 4710 | 2 × Intel Xeon Scalable 4th and 5th generation, LGA 4677, 350 W cap |
| Memory | 32 DDR5 RDIMM / MRDIMM slots, 8 channels per CPU. 6,400 MT/s at 1DPC, 5,200 at 2DPC. Modules of 32, 64, 96 and 128 GB | 32 DDR5 RDIMM / RDIMM-3DS slots. 5,600 MT/s at 1DPC, 4,400 at 2DPC |
| Storage | 8 hot-swap 2.5" bays taking up to eight 25 mm E1.S drives; 2 × M.2 NVMe Gen 5 x4 internally | 12 bays (8 × SATA-3 + 4 × NVMe), 1 × internal M.2 NVMe |
| Power | 4 × 3,200 W CRPS, 3+1, 80 PLUS Titanium | 4 × 2,700 W CRPS, 3+1, 80 PLUS Platinum |
| Cooling | 10 hot-swap 80 mm fans split by zone: five for the GPU zone, five for the HPM zone | 16 fans: 8 internal and 8 at the rear |
| TPM | On-board TPM 2.0 with a Microchip CEC1736 hardware root of trust | A TPM header only; the module is optional |
| Operating envelope | 0–35 °C, 8–90% humidity | 10–35 °C, 8–80% humidity |
| Chassis | Built by Chenbro | Not stated |
| Purpose per vendor | NVIDIA Omniverse Enterprise, large-scale digital twins, HPC, LLM inference and fine-tuning | Transfer AI training, fine-tuning, LLM training and inference, AI cloud, HPC |
The difference between the machines is not the card count but the connectivity and the class of everything around it. The MGX version gives eight 400 Gb/s ports against one at 25, a dedicated BlueField-3, CPUs a generation newer, memory a step faster, power supplies of a higher efficiency class and greater capacity, and an on-board TPM rather than an optional one. And it is smaller: 4U against 6U, and 858 mm deep against 900. The "six rack units for the airflow" argument looks weaker in this light: the MGX platform has fewer fans, ten against sixteen, but splits them into GPU and HPM zones — at half the volume and with more power on tap, that points to a better-engineered cooling scheme rather than a more generous one. The 6U machine remains a sensible choice where a cheaper standalone node is wanted, but the case for it now rests on price rather than engineering.
| Criterion | MGX 4U reference NV | INNO3D 6U Vendor | Mainstream 2U Vendor |
|---|---|---|---|
| GPUs | 8 | 8 | 2–4 |
| Rack density | High | Low | Medium |
| Inter-node networking | ConnectX-8, up to 800 Gb/s, BlueField-3 | OCP 25 Gb/s | Configuration dependent |
| PCIe topology | Switched Gen 6 fabric inside the SuperNIC | Direct attach, cards split 4+4 across sockets | Direct attach |
| Supply type | Usually a finished system | Barebone | Finished system with vendor support |
| Purpose | Cluster building block | Self-contained node: inference, rendering, VDI | First step into accelerated computing |
| Serviceability | Dense layout | Modular, volume to spare, hot-swap fans | Standard rack service |
The line for a quotation: an MGX-reference configuration is bought when the customer plans to grow into a cluster. A 6U platform is bought when they need a lot of compute in one machine, and the platform price difference goes into cards instead. If a second server and a distributed workload appear within two years, the absence of high-speed networking becomes a dead end.
| Vendor | Model | Format | GPUs | Distinguishing feature per vendor |
|---|---|---|---|---|
| INNO3D | AGS-4UMGX-R1 | 4U MGX | 8 | ConnectX-8 SuperNIC with eight 400 Gb/s ports, BlueField-3 in a dedicated slot, Xeon 6, 4 × 3,200 W Titanium, on-board TPM 2.0 |
| INNO3D | AGS-6220V2 | 6U | 8 | Barebone, direct attach without a switch, 25 Gb/s OCP, 4 × 2,700 W Platinum, 16 fans |
| Lenovo | ThinkSystem SR650a V4 | 2U | 4 | The card can be power-capped to 450 W in order to fit four into the chassis |
| Supermicro | SYS-522GA-NRT, SYS-422GL-NR | 4U / 5U | 8 | Among the first to adopt the MGX PCIe Switch board with ConnectX-8 |
| Gigabyte | XL44-SX2-AAS1 | 4U MGX | 8 | Dual Xeon 6700/6500, 32 DDR5 slots, 3+1 redundant 3,200 W Titanium supplies, BlueField-3 and 4 × ConnectX-8 |
| ZOTAC | ZRS-MGX-R1 | 4U MGX | 8 | 6th-gen Xeon, 32 DIMM slots, eight 400 Gb/s ports on ConnectX-8, 3+1 redundant 3,200 W supplies |
| ASUS | ESC8000A-E13 | 4U | 8 | Built on the AMD EPYC platform |
It is the chassis thermal budget, not the card specification, that determines the final configuration. Lenovo provides the illustration: the card can be power-capped to 450 W precisely in order to fit four into the chassis. This is exactly what NVIDIA made the 400–600 W range configurable for. Before promising a customer eight cards, calculate rack power delivery and the room's ability to remove heat.
NVIDIA provides its own tools: the Qualified System Catalog, a catalog of systems qualified with a specific GPU and filterable by card model; the NVIDIA-Certified Systems programme; and the AI Factory Validated Design as a configuration benchmark. One observation that matters for procurement: the system partner list on the RTX PRO Server page names Cisco, Dell, HPE, Lenovo and Supermicro as leading partners, plus Advantech, ASRock Rack, ASUS, Compal, Eviden, Foxconn, GIGABYTE, Inventec, MiTAC, MSI, Pegatron, QCT, Wistron and Wiwynn. INNO3D is not on that list — even though the company builds a platform in the NVIDIA MGX form factor with ConnectX-8 SuperNICs and BlueField-3, matching NVIDIA's reference configuration in composition. Absence from the partner list and the existence of an MGX product are compatible facts: the list need not be exhaustive. The conclusion is unchanged, though: confirmation must come from NVIDIA's catalog rather than from a product page. Check the Qualified System Catalog filtered by the relevant card and ask the distributor for documentary evidence of status.
AGS-6220V2: what to lead with, and what not to argue about
This layer is written for the conversation with the customer. Every argument comes with the boundary beyond which it stops working: the engineer on the customer's side will find those boundaries anyway, and naming them first turns the meeting from an audit into a joint sizing exercise.
Eight load-bearing arguments
The cheapest way into eight cards
$121,700 against $158,210 for an MGX platform — a saving of $36,510, close to a quarter of the budget. Both platforms are barebone, so the prices compare directly. That money goes into CPUs, memory and drives rather than into a chassis.
Boundary: the saving evaporates if the customer needs 400 Gb/s networking — external adapters and a switch will consume the difference.
Barebone as freedom
CPUs, memory and NVMe come through the customer's own contracts at their own prices. With the current DDR5 shortage that is real leverage: they can fit memory from existing stock.
Boundary: this works against turnkey servers but not against MGX platforms — those ship barebone too. Do not present it as a difference from the AGS-4UMGX-R1.
Two Xeon generations to choose from
The LGA 4677 socket takes both 4th and 5th generation Xeon Scalable. The customer can use an existing fleet or buy more affordable parts. MGX platforms on LGA 4710 offer no such option.
Boundary: the 350 W per-socket cap rules out some higher-end SKUs.
Twelve bays instead of eight
8 × SATA plus 4 × NVMe. For media archives, warm data and datasets that do not need NVMe read speeds, SATA is several times cheaper per terabyte.
Boundary: only four bays remain for a fast working dataset.
Serviceability and spare volume
16 hot-swap fans, modular layout, easy internal access. Lower skill requirements for the field engineer and less downtime when replacing parts.
Boundary: in colocation billed per rack unit, 6U costs half again as much as 4U.
Direct CPU attach
Every card gets a genuine x16, one hop closer to host memory, a simpler firmware stack and fewer points of failure. For eight independent inference streams a switch adds nothing.
Boundary: the cards split 4+4 across sockets and traffic between the halves crosses the inter-socket link.
Nearly a quarter of headroom
8,100 W usable in a 3+1 arrangement against a calculated draw of around 6.5 kW. Room to grow into higher-TDP CPUs, more drives and future cards.
Boundary: it needs a 200–240 V feed with C19/C20 connectors — in older racks that is a separate conversation.
Payback in six months
Against AWS on-demand at round-the-clock load the machine pays for itself in 6.2 months, and it beats renting from 21.7% utilisation upward. This is the strongest numeric argument available — see Layer 08.
Boundary: below roughly a fifth of the time, the honest recommendation is cloud.
Workloads you can sell it into with confidence
| Workload | Why it fits | What to cite |
|---|---|---|
| LLM inference up to roughly 70B parameters | 96 GB and 1,597 GB/s per card, native FP4 | AWS and Microsoft independently place this threshold as the single-card boundary |
| Many independent models and multi-tenancy | MIG, up to four isolated instances per card | Up to 32 isolated instances per system with guaranteed quality of service — the basis for billing |
| Agentic services and RAG | Memory capacity for long context and KV cache | Microsoft names RAG on models under 70B as a target scenario for this card |
| Fine-tuning existing models | Sufficient memory plus FP8/BF16 | NVIDIA names fine-tuning explicitly among Server Edition use cases — this is not a stretch |
| VDI and virtual workstations | vGPU over MIG, time-slicing between AI and graphics | The datasheet lists VMware ESXi 8.0U1 and Citrix Hypervisor 8.2 — the platform was built with this in mind |
| Rendering, Omniverse, OpenUSD, digital twins | 188 fourth-generation RT Cores, 355 TFLOPS peak | Exactly what purely computational accelerators lack: a competitor on the H series cannot close this scenario |
| Video analytics, transcoding, streaming | Four NVENC and four NVDEC engines per card | 32 encode and 32 decode engines per system, with 4:2:2 support in H.264 and HEVC |
| Vector search, analytics, data science | Acceleration through cuVS and the CUDA-X stack | NVIDIA claims up to 50× against a CPU system on index building |
| Scientific computing and FP32 simulation | 120 TFLOPS of FP32 per card | One node covers what used to require a small cluster |
| On-premise AI in regulated industries | Data and models never leave the customer's perimeter | The cloud comparison in Layer 08 shows this is also cheaper under sustained load |
Workloads not to sell this machine into
The list below is more useful than the one above. A system sold for the wrong workload comes back along with your reputation, while a limitation named in time almost always converts into a different line from your own portfolio.
| Workload | Why it does not fit | What to offer instead |
|---|---|---|
| Training large models from scratch | The cards have no NVLink and memory does not pool. Training sits outside NVIDIA's stated use cases for both RTX PRO cards | NVLink platforms on the H or B series, or rented capacity for the duration of training |
| Distributed training or inference across nodes | Networking is a single 25 Gb/s OCP module — an order of magnitude short for inter-node traffic | INNO3D AGS-4UMGX-R1: eight 400 Gb/s ConnectX-8 ports plus BlueField-3 |
| Tensor parallelism across all eight cards under a strict SLA | 128 lanes across two sockets means a 4+4 split; traffic between halves crosses the inter-socket link | AGS-4UMGX-R1 with the switched PCIe Gen 6 fabric inside the SuperNIC |
| Storage access via GPUDirect Storage and RDMA over fabric | No DPU: nothing to offload networking and data access onto | AGS-4UMGX-R1 with BlueField-3 in a dedicated slot |
| Dense colocation billed per rack unit | 6U for eight cards against 4U for MGX — 50% more space rental for the same compute | AGS-4UMGX-R1, or an MGX 2U configuration on eight RTX PRO 4500 SE |
| A tender with mandatory TPM 2.0 | The board carries only an SPI header; the module itself is optional | Add the module, or offer AGS-4UMGX-R1, where TPM 2.0 is on-board with a Microchip CEC1736 root of trust |
| Pipelines with large datasets on fast drives | Only four of twelve bays are NVMe; there is a single internal M.2 | AGS-4UMGX-R1 with eight E1.S bays, or external storage |
| A confidential computing requirement | NVIDIA lists support in the RTX PRO 4500 SE specification; there is no such row for the 6000 SE | An MGX 2U configuration on the RTX PRO 4500 SE, where it is officially stated |
| Heavy CPU-side data preprocessing | 4th and 5th generation Xeon with a 350 W cap, memory at 5,600 MT/s at 1DPC | AGS-4UMGX-R1 on Xeon 6 with memory at 6,400 MT/s |
| Edge deployment with an unstable climate | Operating range 10–35 °C, humidity to 80% non-condensing | Purpose-built edge platforms, or improvements to the room |
| Racks shallower than 950 mm | A 900 mm chassis plus cable management does not physically fit | AGS-4UMGX-R1 at 858 mm deep |
Of the eleven unsuitable scenarios, seven are covered by another machine in your own portfolio. The right response to "we need distributed workloads" or "our colocation bills per rack unit" is not to defend the 6220V2 but to move to the AGS-4UMGX-R1. The deal gets larger in the process: the MGX platform costs 30% more. Sell the portfolio, not the single line.
Objections you will hear, and honest answers
| Objection | Honest answer |
|---|---|
| "Why not gaming cards, they're cheaper" | Passive cooling matched to chassis airflow, ECC, MIG and vGPU, power configurable from 400 to 600 W to fit the thermal budget, and system certification. A gaming card is cheaper to buy and more expensive to own: it cannot be shared between users and is not part of a supported configuration |
| "A competitor fits the same eight cards in 4U and you need 6U" | Concede it directly. Our advantages here are price, serviceability and twice the storage bays. If density and networking matter more to the customer, offer the AGS-4UMGX-R1 — that is also our machine |
| "INNO3D is not on NVIDIA's system partner list" | Concede it. The list on a product page is not exhaustive, and we do have a platform in the MGX form factor with ConnectX-8 and BlueField-3. Offer a check in the NVIDIA Qualified System Catalog and request written confirmation of status through the distributor — that settles the question on paper rather than verbally |
| "This is barebone, we need a finished server" | List the missing lines up front: two CPUs, memory, drives. Offer integration and give lead times. Note as well that during a memory shortage, buying memory themselves often wins on both price and delivery |
| "Too expensive" | Move from price to cost of ownership. Against AWS on-demand the payback is 6.2 months at round-the-clock load, and the advantage starts at 21.7% utilisation. Ask how many hours a day the cards will actually be busy — that question closes the deal more reliably than a discount |
| "We'll use the cloud" | Agree for the pilot and for peaks — it is honest and it builds trust. Then show the threshold: above roughly a quarter utilisation, buying is cheaper. The optimal pattern is to buy for the baseline and burst into the cloud |
| "We need to train our own models" | Clarify: from scratch or fine-tuning. NVIDIA names fine-tuning explicitly among Server Edition use cases, and there the machine belongs. Training from scratch is not our scenario, and it is better said immediately |
Each takes a minute to check and costs credibility across the whole proposal. Do not say "768 GB for one model" — memory does not pool; these are eight cards independent in memory. Do not mention NVLink — both card specifications list PCI Express only. Do not promise NVIDIA-Certified status before checking the specific configuration in the catalog. And do not transfer the 50× and 100× figures to this machine: they were measured on a server with eight RTX PRO 4500 SE against a CPU system, not on the 6000 SE and not against another GPU.
Sizing a configuration
The commonest sizing error is counting teraflops. For inference, memory capacity comes first: if the model weights do not fit entirely into GPU memory, the system spills into system RAM and throughput collapses.
| Customer workload | Card | Platform | What to verify |
|---|---|---|---|
| LLM inference up to roughly 70B parameters | 6000 SE | 2U with two cards | Both AWS and Microsoft place this threshold as the single-card boundary. Confirm context length and KV cache size |
| Inference on larger models | 6000 SE | 4U or 6U with eight cards | Tensor parallelism; remember the PCIe penalty from the absence of NVLink |
| Many small models | 4500 SE | MGX 2U with eight cards | Size by density per slot and parallel stream count, not aggregate TFLOPS |
| Video analytics, transcoding | 4500 SE | MGX 2U, edge chassis | Three NVENC and three NVDEC engines per card is the governing figure, not CUDA cores |
| Data processing, vector search | 4500 SE | MGX 2U with eight cards | NVIDIA claims up to 50× against CPU on vector index building |
| Virtual workstations | Either | 2U or 4U | vGPU licences and MIG instance count: four on the larger card, two on the smaller |
| Rendering, Omniverse, simulation | 6000 SE | MGX 4U or 6U | RT Cores are mandatory; peak RT on the larger card is more than double |
| Fine-tuning existing models | 6000 SE | 4U or 6U | Fine-tuning is named explicitly by NVIDIA among Server Edition use cases |
| A pilot with no capital expenditure | 6000 SE | AWS G7e, Azure NCv6, Google Cloud G4 | See Layer 08 |
| Training large models from scratch | Outside the stated use cases for either card. The interconnect listed is PCI Express only, with no NVLink row | ||
Questions to answer before writing the specification
- Which models are being run, at what precision? That sets the required VRAM; the roughly 70B mark at FP8 is the single-card boundary for the larger card.
- What context length and how many concurrent sessions? That sets the KV cache, which frequently drives the choice.
- Training or inference only? If fine-tuning occupies a small share of the time, size the hardware for inference and rent capacity for the heavy work.
- Are graphics and ray tracing needed? If so, there is effectively no alternative to RTX PRO at this budget.
- How many independent consumers? Four MIG instances on the 6000 SE against two at 16 GB on the 4500 SE.
- How many kilowatts are delivered to the rack and how is the heat removed? This constraint trims the configuration more often than anything else.
- Air or liquid? On the 6000 SE these are different form factors: dual-slot against single-slot FHXL.
- Is a second server planned? Then choose the network now: 25GbE is adequate for a standalone machine and a dead end for distributed workloads.
- Finished system or barebone? This determines the number of quotation lines and the lead times.
Economics and market
This has to start with a caveat: NVIDIA does not publish prices for these products. Any price figures in reviews are third-party estimates, and they are not in this reference. Prices are requested from a partner or through NVIDIA Marketplace. Comparable economics can, however, be built from cloud provider documentation.
| Platform | Configuration | Notable features |
|---|---|---|
| AWS EC2 G7e | Up to 8 cards totalling 768 GB of GPU memory, Intel Emerald Rapids processors, up to 192 vCPUs, up to 2,048 GiB of system memory, up to 15.2 TB of local NVMe and up to 1,600 Gb/s of networking via EFA | Up to 2.3× the inference performance of G6e. GPUDirect P2P between cards over PCIe, GPUDirect RDMA over EFA, GPUDirect Storage with FSx for Lustre. Per AWS, a single card handles a 70-billion-parameter model at FP8 |
| Azure NC RTX PRO 6000 BSE v6 | Host on Intel Granite Rapids with all-core turbo up to 4.2 GHz; fractional and full configurations based on MIG | Cards are exposed via SR-IOV, which limits capture of some GPU telemetry. AKS supported. Target workloads: digital twins and Omniverse, inference and RAG under 70B, rendering, VDI on RTX Virtual Workstation, FP32 scientific visualisation |
| Google Cloud G4 | Fractional VMs with vGPU profiles at 12, 24, 48 and 96 GB | Per NVIDIA's technical blog, the profiles target uses from streaming to high-fidelity 3D rendering and robotics sensor simulation |
Three consumption models the customer chooses between
Buying the server
Justified under constant 24/7 load, data residency requirements and a predictable operating horizon. Gives control over data and a fixed cost.
Renting a dedicated server
The middle path: no capital expenditure, but a fixed configuration. MIG profiles let the provider sell fractions of a card and the customer take exactly the memory they need.
Public cloud
The card is available from three hyperscalers. Good for peaks, testing and a pilot before purchase; under sustained load it is usually more expensive than owned hardware.
What belongs in the TCO
Card cost, AI Enterprise and vGPU licences, electricity and cooling (eight larger cards draw up to 4.8 kW on GPUs alone against 1.3 kW for eight smaller), rack space, vendor support, cost of downtime. Density per slot often wins the TCO case for the 4500 SE where the 6000 SE wins on peak performance.
Vendor Provider The numbers worked through
The model below is built on a real platform price and verifiable tariffs. Every assumption is stated in the table — change them and you get your own answer rather than someone else's.
| Item | AGS-6220V2 · 6U | AGS-4UMGX-R1 · MGX |
|---|---|---|
| CPUs | None | None, listed as a configuration option |
| Memory | None | None, listed as a configuration option |
| Drives | None | None; internal and front-panel both listed as options |
| Power supplies | 4 pre-installed | 4 pre-installed |
| Fans | 16 | 10 pre-installed |
| CPU heatsinks and carriers | 2 heatsinks with mounts | 2 heatsinks and 4 carriers |
| Rack rails | Included | Included, L-shaped |
| Network module | 25 Gb/s OCP pre-installed | ConnectX-8 appears in the specification; not itemised in the contents — confirm |
| DPU | None | BlueField-3 appears in the specification; not itemised in the contents — confirm |
| RAID key | Intel VROC pre-installed | Not listed |
| Power cords | 4 × C19–C20 | Not listed |
The point that settles a frequent question: both platforms are barebone. Memory and drives are bought separately either way, so comparing $121,700 with $158,210 is legitimate — both figures rest on the same basis. The small items favour the 6U: its network module, VROC key and power cords are already in the box.
| Parameter | Value | Provenance |
|---|---|---|
| 6U platform with 8 × RTX PRO 6000 SE, barebone | $121,700 | Quoted price. Competing 6U models taken at the same figure |
| MGX 4U platform with 8 cards, barebone | $158,210 | Taken as +30% over the 6U. Also without CPUs, memory or drives — see Table 16 |
| Completing the 6U: 2 × 4th- or 5th-gen Xeon, 1 TB DDR5, NVMe | ≈ $20,000 | Estimate. Confirm with the supplier |
| Completing the MGX 4U: 2 × Xeon 6, 1 TB DDR5 or MRDIMM, E1.S | ≈ $25,000 | Estimate. Dearer for three reasons: CPUs for LGA 4710, memory at 6,400 MT/s, and data-centre-class E1.S drives |
| If the customer needs bulk cold storage | the gap widens | The 6U has eight SATA bays — the cheapest terabytes available. The MGX takes E1.S only. At tens of terabytes the completion gap easily doubles |
| System power draw | 6.0 kW | 8 × 600 W GPU + 2 × 350 W CPU + 0.5 kW for memory, drives and fans |
| PSU efficiency | 94% / 96% | 80 PLUS Platinum on the 6U, Titanium on the MGX 4U — per datasheets |
| Facility PUE | 1.4 | Assumption. A modern data centre can do better |
| Electricity | $0.20 / kWh | Assumption. Sensitivity shown below |
| AWS g7e.48xlarge, 8 cards, on-demand | $33.14 / hr | us-east-1. Consistent across four independent AWS price trackers; verify in the AWS calculator |
| AWS g7e.48xlarge, spot | from $12.78 / hr | Lowest price; varies by availability zone and interruption risk |
| Horizon | 3 years | Assumption |
| Excluded | NVIDIA AI Enterprise and vGPU licences, extended vendor support, rack space or colocation, switches and optics outside the server, staff, traffic. AWS bills egress separately | |
| Item | 6U (INNO3D and peers) | MGX 4U | AWS on-demand | AWS spot |
|---|---|---|---|---|
| Capital expenditure | $141,700 | $183,210 | — | — |
| Electricity per year | $15,656 | $15,330 | included | included |
| Three-year total at 24/7 | $188,669 | $229,200 | $871,032 | $335,803 |
| Cost per GPU-hour at 100% utilisation | $0.90 | $1.09 | $4.14 | $1.60 |
| Cost per GPU-hour at 50% utilisation | $1.79 | $2.18 | $4.14 | $1.60 |
| Payback against on-demand at 24/7 | 6.2 months | 8.0 months | — | — |
The governing asymmetry: owned hardware is almost entirely fixed cost, cloud is almost entirely variable. Cloud is therefore not more expensive in general — it is more expensive at high utilisation. The 50% row makes this visible: the owned machine's hourly cost doubles while the cloud rate does not move.
| Scenario | 6U | MGX 4U |
|---|---|---|
| Utilisation threshold against on-demand | 21.7% | 26.3% |
| Utilisation threshold against spot | 56.2% | 68.3% |
| In hours over three years, against on-demand | 5,692 of 26,280 | 6,915 of 26,280 |
| At electricity of $0.10 / kWh | threshold 19.0% | — |
| At electricity of $0.30 / kWh | threshold 24.4% | — |
The last two rows show that the electricity tariff matters far less than utilisation: tripling the rate moves the threshold by only five percentage points. The first question to ask a customer is not what they pay per kilowatt-hour but how many hours a day the cards will actually be busy.
First. If the cards are busy more than a quarter of the time, buying beats on-demand rental, and under round-the-clock load the machine pays for itself in roughly six months. Against spot pricing the threshold is much higher — around 56% — but spot offers no availability guarantee and is unsuitable for production. Second. Cloud wins in three situations: a pilot before purchase, peak load on top of an owned fleet, and projects shorter than a year. The sensible pattern is to buy for the baseline and burst into the cloud. Third. The difference in threshold between the 6U and the MGX 4U is under five percentage points. If the customer clears the utilisation bar anyway, the choice between platforms is decided not by money but by whether they need the networking to grow into a cluster.
Useful arithmetic for the conversation, but it has to be done carefully. Eight cards at NVIDIA's marketplace price come to $106,000. That implies roughly $15,700 for the 6U chassis itself and roughly $52,210 for the MGX 4U. The premium is $36,510.
What is in it. An MGX 4U chassis, a motherboard for Xeon 6 on LGA 4710 with memory to 6,400 MT/s, Titanium-class 3,200 W supplies instead of Platinum 2,700 W, eight 400 Gb/s ConnectX-8 ports and a dedicated BlueField-3. That works out at about $4,560 per 400 Gb/s port — before the customer buys a switch.
What is not in it — an important correction. The Xeon 6 processors themselves are not covered by the premium: neither delivery includes CPUs. The premium buys the board that accepts Xeon 6; the processors are purchased separately and will cost more on the MGX than 4th or 5th generation Xeon on the 6220V2. That adds to the gap rather than sitting inside it.
What must be confirmed in writing. Neither ConnectX-8 nor BlueField-3 is itemised in the AGS-4UMGX-R1 contents: both are named only in the specification rows, as "Configured with". It reads as "included in the configured system", but with $36,510 at stake it needs to come from the supplier on paper. The direct question: are the ConnectX-8 SuperNIC board and the BlueField-3 module included in the price of the AGS-4UMGX-R1, or billed separately? If separately, the arithmetic above collapses and the comparison has to be rebuilt.
What the premium definitely does not repay. Energy efficiency: moving from Platinum to Titanium saves $326 a year, or about $980 over three years — 2.7% of the premium. Titanium should be sold as thermal headroom and reliability, not as a saving on the electricity bill.
According to NVIDIA, the Llama Nemotron Super model in NVFP4 on a single RTX PRO 6000 delivers up to 3× better price-performance than FP8 on an H100. That inverts the usual logic: the cheaper universal card turns out to be the better buy for inference — provided the model fits in memory. The market corroborates it indirectly: AWS independently measured up to 2.3× the inference performance on G7e instances relative to the previous L40S-based generation.
Where to go next
NVIDIA Source of all card and technology data
- RTX PRO 6000 Blackwell Server Edition · RTX PRO 4500 Blackwell Server Edition — specifications and performance claims.
- RTX PRO 6000 Blackwell Series — the three-edition comparison table, source for Table 3.
- RTX PRO Server — four reference configurations, networking technologies, system partner list.
- Datasheets: 6000 SE · 4500 SE · RTX PRO Server · Data Center GPU Line Card.
- MGX · Multi-Instance GPU · GPU virtualisation · ConnectX-8 SuperNIC · BlueField-3 · Spectrum-X.
- Technical blog on the 4500 SE and vGPU 20 · Press release on RTX PRO Servers.
- Verification tools: Qualified System Catalog · NVIDIA-Certified Systems · AI Factory Validated Design · NVIDIA Marketplace.
- MLPerf results — independently verifiable benchmarks.
Vendors Source of all server data
- INNO3D AGS-4UMGX-R1 and its PDF datasheet — the MGX 4U platform.
- INNO3D AGS-6220V2 and its PDF datasheet — the 6U platform.
- Lenovo Press: 6000 SE in ThinkSystem and 4500 SE — compatibility matrices and slot power caps. One such document replaces ten reviews.
- Gigabyte: RTX PRO solutions · ZOTAC ZRS-MGX-R1 · Supermicro: system portfolio.
Providers Source of cloud data
- AWS EC2 G7e and the AWS blog announcement.
- Microsoft Learn: NC RTX PRO 6000 BSE v6 series and the series overview.
Further reading not a source of figures
The material below is useful for a first pass at the subject, but any figures in it should be checked against the primary sources above — discrepancies in exactly these publications are what prompted this reference to be rebuilt.
- ENServeTheHome — PCIe GPU in Server Guide, the reference treatment of these platforms; their YouTube channel is the best video material on the subject.
- ENigor'sLAB on the 4500 SE — explains why the card exists: rack density, watts and TCO rather than TFLOPS.
- RUCytrix — RTX PRO 6000 Server Edition, the most server-oriented of the Russian-language treatments.
- RUServerMall — 1, 2, 4 or 8 GPUs in a server, on chassis selection logic.
- UKAlfa-Server — RTX PRO 4500 SE review with a recommended server configuration.
Glossary
| Term | Expansion | Meaning |
|---|---|---|
| MIG | Multi-Instance GPU | Hardware partitioning of a card into isolated GPUs with guaranteed quality of service |
| vGPU | Virtual GPU | GPU virtualisation for VDI; vWS and vPC licences apply |
| MGX | NVIDIA MGX | Modular reference design for server platforms |
| SuperNIC | ConnectX-8 SuperNIC | Network adapter with a 48-lane PCIe Gen 6 switch |
| DPU | Data Processing Unit | BlueField-3: offloads networking, data access and security from the CPU |
| GPUDirect P2P | Peer to Peer | Direct GPU-to-GPU transfer over PCIe without going through system memory |
| SR-IOV | Single Root I/O Virtualization | A method of exposing a device to virtual machines; used in Azure NCv6 |
| FP4 / NVFP4 | 4-bit format | Numeric format that packs weights four times denser than FP16 |
| NVENC / NVDEC | Encoder / Decoder | Hardware video encode and decode engines |
| KV cache | Key-Value cache | Memory holding conversational context; often the real driver of VRAM consumption |
| FHFL / FHXL | Full Height Full Length / Extra Length | A full-size expansion card and its extended variant for the liquid version |
| Barebone | — | A platform supplied without CPUs, memory or drives |
| Headless | — | Operating with no display attached |