Field reference · Revision of 17 August 2026 · Angle: selling INNO3D systems

NVIDIA server GPUs:
RTX PRO 6000 and RTX PRO 4500 Blackwell Server Edition

A working reference for an INNO3D GPU server salesperson. From first principles to platform architecture, the case for and against the AGS-6220V2, and a cost-of-ownership model. Arguments come with their boundaries: the engineer on the customer side will find them anyway, and it is better to name them first.

ClassUniversal PCIe GPUs for the enterprise data centre
What they succeedL40S → PRO 6000 · L4 → PRO 4500
Key sectionLayer 06: arguments for and against, objection handling

NVIDIA NVIDIA on its own products  ·  Vendor server makers on theirs  ·  Provider cloud providers on theirs

Layer 00 · FundamentalsHow a server GPU differs from any other

Why you cannot simply drop a gaming card into a server

This is the first thing you need to be able to explain to a customer, because the question "why not a 4090 or 5090, they're cheaper" comes up every single time. There are six differences, and every one of them is about operations rather than benchmark scores.

Difference 01 · NVIDIA

Passive cooling

NVIDIA lists the thermal solution of the server cards as passive and describes the 6000 SE as designed for multi-GPU server deployments requiring passive cooling. There are no fans on the board; the chassis moves the air. Outside a server with engineered airflow the card will not run.

Difference 02 · NVIDIA

ECC memory

The 6000 SE specification states 96 GB of GDDR7 with error-correcting code. Under a 24/7 load — a multi-day inference run or a render job — a single bit flip without ECC means a corrupted result with no error message of any kind.

Difference 03 · NVIDIA

Different output sets

The 6000 SE lists four DisplayPort 2.1 outputs, so it can also work as a graphics card in the data centre. The 4500 SE has no display output row in its specification at all: it is a purely headless accelerator.

Difference 04 · NVIDIA

Partitioning: MIG and vGPU

The card can be carved into isolated instances with their own memory, cache and cores, or shared with remote users through vGPU. Up to four instances on the 6000 SE, up to two at 16 GB on the 4500 SE.

Difference 05 · NVIDIA

System certification

RTX PRO Servers hold NVIDIA-Certified status and are validated with partner tools as part of the AI Factory Validated Design. That is a guarantee of compatibility with the AI Enterprise stack and NVIDIA networking.

Difference 06 · NVIDIA

Configurable power

The 6000 SE's power draw is published as a 400–600 W range. The card is tuned to the chassis thermal budget rather than the other way round. Consumer cards offer nothing equivalent.

The commercial takeaway

The argument to make with a customer is not performance, it is predictability and manageability. A gaming card is cheaper to buy and more expensive to own: it cannot be shared between users, it has no ECC, it cannot be power-capped to the chassis, and it is not part of a certified configuration backed by the server vendor and NVIDIA.

Layer 01 · Architecture NVIDIAWhat in Blackwell actually delivers the gain

Four generational changes people pay for

Blackwell is the generation that succeeded Ada Lovelace. For selling server cards, exactly four of its properties matter; everything else is detail for engineers.

Table 1 · Blackwell changes and what they mean commercially
TechnologyWhat NVIDIA saysWhat the customer gets
5th-gen Tensor Cores and FP4Up to 3× the performance of the previous generation, adding FP4 precision. The 6000 SE carries the second-generation Transformer EngineModel weights occupy a quarter of the memory they would in FP16. This, and not the CUDA core count, is where the multi-fold gain in LLM inference comes from
GDDR7 memorySignificantly increased bandwidth and capacityMade 96 GB on a single card possible — a capacity previously available only on expensive SXM platforms
4th-gen RT CoresUp to 2× the performance of the previous generation; RTX Mega Geometry enables up to 100× more ray-traced trianglesRendering, digital twins, simulation, Omniverse. Precisely what purely computational accelerators lack
NVENC 9 and NVDEC 64:2:2 support in H.264 and HEVC, improved AV1 quality, double the H.264 decode throughputVideo analytics, transcoding, streaming, media production — segments where the card sells without any AI conversation at all

The key to understanding the whole range: RTX PRO cards can do AI compute and graphics. That versatility is their core competitive property and the main lever in a sales conversation.

Layer 02 · Product NVIDIASpecifications from NVIDIA's product pages

The two cards, side by side

Both are built on Blackwell, both are passive, both are PCIe Gen5 x16, both support MIG. The similarity ends there: these are cards for different thermal budgets and different jobs.

Table 2 · RTX PRO 6000 SE against RTX PRO 4500 SE
SpecificationRTX PRO 6000 Blackwell SERTX PRO 4500 Blackwell SE
CUDA cores24,06410,496
RT Cores (4th gen)18882
Memory96 GB GDDR7 with ECC32 GB GDDR7
Memory bandwidth1,597 GB/s800 GB/s
Memory interface512-bit256-bit
FP4 Tensor Core4 PFLOPS1.6 PFLOPS
FP8 Tensor Core2 PFLOPS811 TFLOPS
FP16 / BF16 Tensor Core1 PFLOPS406 TFLOPS
TF32 Tensor Core234 TFLOPS203 TFLOPS
Single precision FP32120 TFLOPS51 TFLOPS
Peak RT Core performance355 TFLOPS154 TFLOPS
Power consumption400–600 W (configurable)165 W
Form factorAir: dual-slot, 4.4" × 10.5" · Liquid: single-slot, FHXLSingle-slot, 4.4" × 10.5"
Thermal solutionPassivePassive
InterfacePCIe Gen 5 x16PCI Express 5.0 x16
MIGUp to 4 isolated instancesUp to 2 instances at 16 GB
Video engines4 NVENC / 4 NVDEC3 NVENC / 3 NVDEC
Display outputs4 × DisplayPort 2.1Not listed in the specification
Confidential computeNot listed in the specificationCapable
Power connectorNot listed in the specification1 × PCIe CEM5 16-pin
PredecessorNVIDIA L40SNVIDIA L4

The "not listed" rows are deliberate: NVIDIA publishes slightly different parameter sets for the two cards, and filling the gaps by analogy is exactly how wrong figures find their way into reviews.

Careful: three different products with nearly identical names

This is the most common error in specifications and quotations. "RTX PRO 6000 Blackwell" is a family of three cards, with the server-class 4500 sitting separately alongside.

Table 3 · RTX PRO 6000 Blackwell editions
SpecificationServer EditionWorkstation EditionMax-Q Workstation
Memory96 GB GDDR7 ECC96 GB GDDR7 ECC96 GB GDDR7 ECC
Power consumption400–600 W600 W300 W
ThermalPassiveDouble flow-throughActive
Dimensions4.4" × 10.5", dual slot5.4" × 12", dual slot4.4" × 10.5", dual slot
Display outputs4 × DisplayPort 2.14 × DisplayPort 2.14 × DisplayPort 2.1
BusPCIe Gen 5 x16PCIe Gen 5 x16PCIe Gen 5 x16
Purpose per NVIDIAMulti-GPU servers requiring passive cooling: inference, fine-tuning, distributed rendering, HPC, virtual workstationsMaximum performance in a single-GPU workstationDense workstation configurations, up to four GPUs

Max-Q is a practical compromise for a customer who needs density but has no server rack: half the heat at the same memory capacity, up to four cards in a workstation. Note also NVIDIA's wording for the Server Edition — fine-tuning is named explicitly, so it is within the stated use case, unlike training from scratch.

An applied rule of thumb

The memory-per-watt ratio of the two server cards is close; what differs radically is density per slot: 165 W in one slot against 600 W in two. Choosing between them is really choosing between "one large model" and "many parallel streams". On compute the gap is two- to two-and-a-half-fold; on power draw it is nearly fourfold.

Layer 03 · Positioning NVIDIAWhere the limits of applicability run

Position within the NVIDIA data centre range

The most valuable skill in selling these cards is being able to say when they are the wrong answer. The data centre range splits into two branches: universal RTX PRO PCIe cards, and compute accelerators built for training.

Table 4 · Roles within the range
ProductRole as described by NVIDIAGraphics and RTInterconnect
RTX PRO 4500 SEInference on small and medium models, data processing and data science, video analytics, virtualisation. For data centre, edge and cloudYesPCI Express 5.0 x16
RTX PRO 6000 SEInference, fine-tuning, distributed rendering, HPC, virtual workstations. The universal data centre GPUYesPCIe Gen 5 x16
H series and B seriesSeparate NVIDIA product lines for scenarios that require GPU-to-GPU connectivity over NVLinkLimitedNVLink / NVSwitch
The limitation to disclose early

The specifications of both cards list PCI Express only as the interconnect; there is no NVLink row. The cards communicate over PCIe and their memory does not pool. Eight cards at 96 GB is not "768 GB for one model" — it is eight accelerators that are independent in memory. The large per-card capacity compensates for this in parallel inference, but a single model that does not fit in 96 GB needs either tensor parallelism with PCIe overhead, or an NVLink platform.

Arguments for the generational move

  • Against the L40S: NVIDIA states that RTX PRO servers run simulation and synthetic data generation up to 4× faster than L40S systems, and describes the 6000 SE as delivering a significant leap over the L40S in multimodal agentic and generative applications.
  • Against the L4: for the 4500 SE NVIDIA claims over 5× the performance of the previous-generation L4 in inference on small and medium models, when using NVIDIA's optimised frameworks and NIM microservices.
  • Against CPU-only systems: for AI video understanding NVIDIA claims up to 100× the performance of CPU-only systems, with a reduction in footprint and energy consumption of more than 95%. For vector databases, up to 50×. This is the strongest argument available when talking to a customer running a fleet of CPU servers.

An important caveat on those last two numbers: the 50× and 100× comparisons are between an eight-card 4500 SE server and a CPU system (for vector databases, an AMD 9654 with 192 vCPUs, 33 million vectors in Milvus 2.5.25, HNSW against cuVS CAGRA; for video, a dual-socket AMD 9654 running Qwen3-VL-8B). That is a legitimate platform comparison, but it is not card against card, and a competent engineer on the customer's side will ask.

Layer 04 · Stack NVIDIAThe capabilities a business pays for

The features that actually sell the card

Specifications sell to an engineer. A commercial buyer is sold on the items below — they are what turns hardware into a service that can be resold or shared across departments.

Table 5 · Enterprise-grade capabilities
TechnologyMechanismValue to the customer
MIGHardware partitioning of the GPU into isolated instances, each with its own high-bandwidth memory, cache and compute cores plus guaranteed quality of service. Up to four on the 6000 SE, up to two at 16 GB on the 4500 SEGuaranteed quality of service per tenant. One card serves several independent workloads with no interference — the basis for billing and multi-tenancy
vGPUPaired with NVIDIA RTX Virtual Workstation and Virtual PC software, the card is virtualised for remote users. MIG-backed time-slicing lets AI and graphics workloads run simultaneously. vGPU 20 adds the AI Virtual Workstation Toolkit and fixed-share schedulingEngineers and designers work on thin clients while CAD licences and data never leave the perimeter. Very often this use case, not AI, is what pays for the purchase
ECCError correction on graphics memory, stated for the 96 GB of GDDR7 on the 6000 SEA precondition for 24/7 production workloads and a standard line item in infrastructure tenders
Confidential computingListed as capable in the 4500 SE specification tableAn argument in regulated industries: protecting data and models during processing
NVIDIA AI EnterpriseA suite of tools, libraries and frameworks including NIM and NeMo microservicesA separate line on the invoice and, at the same time, the answer to "who is going to support this"
NVIDIA-CertifiedRTX PRO Servers hold the status and are validated with partner tools under the AI Factory Validated DesignGuaranteed compatibility with NVIDIA networking, BlueField-3 and the AI Enterprise stack. Removes integration risk from customer and integrator alike
NVIDIA Run:aiWorkload and GPU orchestration that raises utilisation and overall throughputFilling idle cards is a direct argument for payback
An applied technique

Present MIG in money, not gigabytes. One RTX PRO 6000 SE split into four instances is four AI developer workstations, or four independent inference services, on a single physical card. And the inverse rule: you cannot plan four users onto a 4500 SE — it supports two instances at most. This detail is regularly lost when quotations are drafted.

Layer 05 · Platform NVIDIA VendorsFrom the card to the rack

The server as a system, not a box with cards in it

This is where "knowing the hardware" ends and "understanding server technology" begins. The cards are the smaller part of the engineering problem. The rest is chassis, network, power and cooling. NVIDIA documents the reference configurations; the specific machines are documented by their makers.

NVIDIA Four reference configurations

NVIDIA defines the RTX PRO Server as a universal data centre platform and gives four examples built on modular MGX reference designs. These are the benchmark against which any vendor proposal should be judged.

Table 6 · RTX PRO Server reference configurations
ConfigurationGPUsNetworkingDPU
MGX 6U8 × RTX PRO 6000 SE, liquid-cooledConnectX-8 SuperNICs with PCIe Gen 6 switchBlueField-3
MGX 4U8 × RTX PRO 6000 SEConnectX-8 SuperNICs with PCIe Gen 6 switchBlueField-3
MGX 2U2 × RTX PRO 6000 SEConnectX-7 or BlueField-3 SuperNICBlueField-3
MGX 2U8 × RTX PRO 4500 SEConnectX-7 SuperNICsBlueField-3

The last row is easy to overlook and shouldn't be: eight of the smaller cards fit into two rack units. That is 256 GB of memory and 12.8 PFLOPS of FP4 in 2U, drawing 1.3 kW on the GPUs against 4.8 kW for eight of the larger cards. For video analytics and small-model inference it is the densest official configuration.

NVIDIA What the switch inside the SuperNIC buys you

NVIDIA describes the major architectural change of recent generations as follows: the ConnectX-8 SuperNIC PCIe Switch is a board integrating multiple SuperNICs, each combining ultra-fast networking with a 48-lane PCIe Gen 6 switch. It eliminates discrete PCIe switches, doubles inter-GPU bandwidth, simplifies board design and enables up to 800 Gb/s. Alongside it sits the BlueField-3 DPU, which per NVIDIA offloads networking, data access and security from the CPU and provides zero-trust isolation. Spectrum-X completes the networking subsystem: through RoCE adaptive routing and congestion control it accelerates storage performance by nearly 50%.

Switched against direct-attach topology: the difference

Eight-GPU PCIe servers come in two topologies, and the difference is worth being able to explain. In a switched design the cards connect to dedicated PCIe switch silicon which then uplinks to the CPUs: any card reaches any other at full width without leaving the switch, and there are more lanes available than the CPUs themselves provide. In a direct-attach design each card plugs into a CPU root complex: cheaper and one hop closer to host memory, but the cards divide between sockets and traffic between the halves crosses the inter-socket link. The AGS-6220V2 is direct-attach; the AGS-4UMGX-R1 is switched. That this is not pedantry is visible from both sides of the industry: NVIDIA built the ConnectX-8 SuperNIC precisely to remove discrete switches while keeping their benefit, and AWS separately notes the very low peer-to-peer latency between GPUs sharing a PCIe switch.

Vendors Three classes of machine

Mainstream

2U, 2–4 GPUs

Conventional rack servers from Cisco, Dell, HPE, Lenovo and Supermicro — all of them listed among NVIDIA's leading system partners. They go into the customer's existing rack. The entry point into accelerated computing for a company that already runs its own data centre.

AI factory

4U MGX, 8 GPUs

A purpose-built platform on the NVIDIA MGX reference design with ConnectX-8 and BlueField-3. A cluster building block rather than "a server with graphics cards". Examples: INNO3D AGS-4UMGX-R1, Gigabyte XL44-SX2-AAS1, ZOTAC ZRS-MGX-R1, Supermicro SYS-522GA-NRT.

Standalone node

6U, 8 GPUs

Platforms such as the INNO3D AGS-6220V2. The same eight cards in half again the volume: airflow headroom, serviceability and price instead of density and high-speed networking. A self-sufficient working machine, not a cluster node.

Vendor One platform in detail: INNO3D AGS-6220V2

It is instructive to set a machine of a different philosophy next to the MGX reference. This is a 6U chassis for eight RTX PRO 6000 Blackwell Server Edition cards, and nearly every decision in it runs opposite to the MGX logic. Figures are from the INNO3D datasheet for SKU AGS-6220V2P2-B00.

Table 7 · INNO3D AGS-6220V2 per the manufacturer's datasheet
SubsystemImplementationWhat it means in practice
Supply typeBarebone. Chassis, motherboard, four PSUs, a 25 Gb/s OCP module, VROC key, rails and heatsinks pre-installed. No CPUs, memory or drivesA platform, not a finished server. CPUs, DDR5, NVMe and the eight cards are separate quotation lines with separate lead times
Form factor6U, 900 × 438 × 264 mmEight cards in 6U against eight in 4U for the reference. The 900 mm depth means checking rack depth including cable management
GPU slots8 × PCIe 5.0 x16, FHFL, double width. Direct CPU attach, no PCIe switchEvery card gets a genuine x16. But 8 × x16 is 128 lanes across two sockets: the cards split 4+4 and traffic between the halves crosses the inter-socket link. For tensor parallelism that is a bottleneck
CPUs2 × Intel Xeon Scalable 4th and 5th generation, LGA 4677, C741 chipset. 350 W TDP limit per socketThe ceiling rules out some higher-end SKUs — verify the chosen processor before ordering
Memory32 DDR5 RDIMM / RDIMM-3DS slots, 8 channels per CPU. RDIMM up to 96 GB, RDIMM-3DS up to 256 GB per module. 5,600 MT/s at 1DPC and 4,400 at 2DPC on 5th-gen XeonA ceiling near 8 TB. Filling every slot drops the speed from 5,600 to 4,400
NetworkingOne OCP slot with a 25 Gb/s module included, plus a 1GbE BMC management portThe decisive difference from the reference. No ConnectX-8, no BlueField-3, no 400 Gb/s. A standalone machine, not a cluster element
Storage12 × 3.5/2.5" bays (8 × SATA-3 + 4 × NVMe), one internal M.2 NVMe, Intel VROC RAID 0/1/10/5Eight of the twelve bays are SATA only, leaving four NVMe bays for a working dataset
Power4 × 2,700 W CRPS, 3+1 redundant, 80 PLUS Platinum8,100 W usable. The arithmetic: 8 × 600 W of GPU plus 2 × 350 W of CPU plus memory, drives and fans is around 6.5 kW. The feed must be 200–240 V with C19/C20 connectors
Cooling16 hot-swap fans: 8 internal and 8 at the rearThe main advantage of 6U: plenty of air, fans replaceable without downtime
ManagementDedicated IPMI port, VGA, 2 × USB 3.2 Gen1. A TPM header with SPI interface; the TPM 2.0 module itself is optionalWhere a tender requires TPM, the module is ordered separately
Operating systemsWindows Server 2022/2019, RHEL 9.2–8.6, SLES 15 SP4, Ubuntu 22.04 LTS, VMware ESXi 8.0U1 / 7.0U3i, Citrix Hypervisor 8.2 LTSR CU1ESXi and Citrix support signals the platform is also intended for VDI
The one divergence inside the datasheet itself

The title page is issued for SKU AGS-6220V2P2-B00 (item 0K219-B001) while the specification page header reads AGS-6220V2P2-B10. Likely a typo or an adjacent bundle revision, but the part number must be confirmed in writing when ordering: it determines what is actually pre-installed. There is also a minor memory inconsistency — one line states 4,800/4,400 MHz while the line below gives 5,600 MT/s at 1DPC for 5th-generation Xeon.

Vendor Same maker, different philosophy: INNO3D AGS-4UMGX-R1

INNO3D also builds a second platform for the same cards — a 4U on the NVIDIA MGX reference design, with ConnectX-8 SuperNICs and BlueField-3. Comparing two machines from one vendor is more instructive than any cross-brand comparison: you see exactly what the extra money buys, with no allowance for differences in build quality or service.

Table 8 · Two INNO3D platforms for the same cards
SubsystemAGS-4UMGX-R1 · MGXAGS-6220V2 · 6U
Form factorNVIDIA MGX 4U, 858 × 440 × 175 mm6U, 900 × 438 × 264 mm
GPU topology8 × PCIe Gen 5 x16 for eight RTX PRO 6000 SE, plus a dedicated Gen 5 x16 slot for BlueField-38 × PCIe 5.0 x16, direct CPU attach with no switch
NetworkingConnectX-8 SuperNIC: eight 400 Gb/s QSFP ports, plus 2 × 10GbE and a 1GbE BMC portA 25 Gb/s OCP module plus a 1GbE BMC port
DPUBlueField-3 in a dedicated slotNone
CPUs2 × Intel Xeon 6 (Granite Rapids / Sierra Forest), 6500 and 6700 series, LGA 47102 × Intel Xeon Scalable 4th and 5th generation, LGA 4677, 350 W cap
Memory32 DDR5 RDIMM / MRDIMM slots, 8 channels per CPU. 6,400 MT/s at 1DPC, 5,200 at 2DPC. Modules of 32, 64, 96 and 128 GB32 DDR5 RDIMM / RDIMM-3DS slots. 5,600 MT/s at 1DPC, 4,400 at 2DPC
Storage8 hot-swap 2.5" bays taking up to eight 25 mm E1.S drives; 2 × M.2 NVMe Gen 5 x4 internally12 bays (8 × SATA-3 + 4 × NVMe), 1 × internal M.2 NVMe
Power4 × 3,200 W CRPS, 3+1, 80 PLUS Titanium4 × 2,700 W CRPS, 3+1, 80 PLUS Platinum
Cooling10 hot-swap 80 mm fans split by zone: five for the GPU zone, five for the HPM zone16 fans: 8 internal and 8 at the rear
TPMOn-board TPM 2.0 with a Microchip CEC1736 hardware root of trustA TPM header only; the module is optional
Operating envelope0–35 °C, 8–90% humidity10–35 °C, 8–80% humidity
ChassisBuilt by ChenbroNot stated
Purpose per vendorNVIDIA Omniverse Enterprise, large-scale digital twins, HPC, LLM inference and fine-tuningTransfer AI training, fine-tuning, LLM training and inference, AI cloud, HPC
What a same-vendor comparison reveals

The difference between the machines is not the card count but the connectivity and the class of everything around it. The MGX version gives eight 400 Gb/s ports against one at 25, a dedicated BlueField-3, CPUs a generation newer, memory a step faster, power supplies of a higher efficiency class and greater capacity, and an on-board TPM rather than an optional one. And it is smaller: 4U against 6U, and 858 mm deep against 900. The "six rack units for the airflow" argument looks weaker in this light: the MGX platform has fewer fans, ten against sixteen, but splits them into GPU and HPM zones — at half the volume and with more power on tap, that points to a better-engineered cooling scheme rather than a more generous one. The 6U machine remains a sensible choice where a cheaper standalone node is wanted, but the case for it now rests on price rather than engineering.

Table 9 · Three philosophies applied to the same eight cards
CriterionMGX 4U reference NVINNO3D 6U VendorMainstream 2U Vendor
GPUs882–4
Rack densityHighLowMedium
Inter-node networkingConnectX-8, up to 800 Gb/s, BlueField-3OCP 25 Gb/sConfiguration dependent
PCIe topologySwitched Gen 6 fabric inside the SuperNICDirect attach, cards split 4+4 across socketsDirect attach
Supply typeUsually a finished systemBareboneFinished system with vendor support
PurposeCluster building blockSelf-contained node: inference, rendering, VDIFirst step into accelerated computing
ServiceabilityDense layoutModular, volume to spare, hot-swap fansStandard rack service

The line for a quotation: an MGX-reference configuration is bought when the customer plans to grow into a cluster. A 6U platform is bought when they need a lot of compute in one machine, and the platform price difference goes into cards instead. If a second server and a distributed workload appear within two years, the absence of high-speed networking becomes a dead end.

Table 10 · Platforms worth studying and comparing
VendorModelFormatGPUsDistinguishing feature per vendor
INNO3DAGS-4UMGX-R14U MGX8ConnectX-8 SuperNIC with eight 400 Gb/s ports, BlueField-3 in a dedicated slot, Xeon 6, 4 × 3,200 W Titanium, on-board TPM 2.0
INNO3DAGS-6220V26U8Barebone, direct attach without a switch, 25 Gb/s OCP, 4 × 2,700 W Platinum, 16 fans
LenovoThinkSystem SR650a V42U4The card can be power-capped to 450 W in order to fit four into the chassis
SupermicroSYS-522GA-NRT, SYS-422GL-NR4U / 5U8Among the first to adopt the MGX PCIe Switch board with ConnectX-8
GigabyteXL44-SX2-AAS14U MGX8Dual Xeon 6700/6500, 32 DDR5 slots, 3+1 redundant 3,200 W Titanium supplies, BlueField-3 and 4 × ConnectX-8
ZOTACZRS-MGX-R14U MGX86th-gen Xeon, 32 DIMM slots, eight 400 Gb/s ports on ConnectX-8, 3+1 redundant 3,200 W supplies
ASUSESC8000A-E134U8Built on the AMD EPYC platform
The practical detail that decides the deal

It is the chassis thermal budget, not the card specification, that determines the final configuration. Lenovo provides the illustration: the card can be power-capped to 450 W precisely in order to fit four into the chassis. This is exactly what NVIDIA made the 400–600 W range configurable for. Before promising a customer eight cards, calculate rack power delivery and the room's ability to remove heat.

How to confirm compatibility on paper

NVIDIA provides its own tools: the Qualified System Catalog, a catalog of systems qualified with a specific GPU and filterable by card model; the NVIDIA-Certified Systems programme; and the AI Factory Validated Design as a configuration benchmark. One observation that matters for procurement: the system partner list on the RTX PRO Server page names Cisco, Dell, HPE, Lenovo and Supermicro as leading partners, plus Advantech, ASRock Rack, ASUS, Compal, Eviden, Foxconn, GIGABYTE, Inventec, MiTAC, MSI, Pegatron, QCT, Wistron and Wiwynn. INNO3D is not on that list — even though the company builds a platform in the NVIDIA MGX form factor with ConnectX-8 SuperNICs and BlueField-3, matching NVIDIA's reference configuration in composition. Absence from the partner list and the existence of an MGX product are compatible facts: the list need not be exhaustive. The conclusion is unchanged, though: confirmation must come from NVIDIA's catalog rather than from a product page. Check the Qualified System Catalog filtered by the relevant card and ask the distributor for documentary evidence of status.

Layer 06 · Selling VendorThe case for and against the AGS-6220V2

AGS-6220V2: what to lead with, and what not to argue about

This layer is written for the conversation with the customer. Every argument comes with the boundary beyond which it stops working: the engineer on the customer's side will find those boundaries anyway, and naming them first turns the meeting from an audit into a joint sizing exercise.

Eight load-bearing arguments

Argument 01 · Price

The cheapest way into eight cards

$121,700 against $158,210 for an MGX platform — a saving of $36,510, close to a quarter of the budget. Both platforms are barebone, so the prices compare directly. That money goes into CPUs, memory and drives rather than into a chassis.
Boundary: the saving evaporates if the customer needs 400 Gb/s networking — external adapters and a switch will consume the difference.

Argument 02 · Procurement

Barebone as freedom

CPUs, memory and NVMe come through the customer's own contracts at their own prices. With the current DDR5 shortage that is real leverage: they can fit memory from existing stock.
Boundary: this works against turnkey servers but not against MGX platforms — those ship barebone too. Do not present it as a difference from the AGS-4UMGX-R1.

Argument 03 · CPUs

Two Xeon generations to choose from

The LGA 4677 socket takes both 4th and 5th generation Xeon Scalable. The customer can use an existing fleet or buy more affordable parts. MGX platforms on LGA 4710 offer no such option.
Boundary: the 350 W per-socket cap rules out some higher-end SKUs.

Argument 04 · Storage

Twelve bays instead of eight

8 × SATA plus 4 × NVMe. For media archives, warm data and datasets that do not need NVMe read speeds, SATA is several times cheaper per terabyte.
Boundary: only four bays remain for a fast working dataset.

Argument 05 · Operations

Serviceability and spare volume

16 hot-swap fans, modular layout, easy internal access. Lower skill requirements for the field engineer and less downtime when replacing parts.
Boundary: in colocation billed per rack unit, 6U costs half again as much as 4U.

Argument 06 · Topology

Direct CPU attach

Every card gets a genuine x16, one hop closer to host memory, a simpler firmware stack and fewer points of failure. For eight independent inference streams a switch adds nothing.
Boundary: the cards split 4+4 across sockets and traffic between the halves crosses the inter-socket link.

Argument 07 · Power

Nearly a quarter of headroom

8,100 W usable in a 3+1 arrangement against a calculated draw of around 6.5 kW. Room to grow into higher-TDP CPUs, more drives and future cards.
Boundary: it needs a 200–240 V feed with C19/C20 connectors — in older racks that is a separate conversation.

Argument 08 · Money

Payback in six months

Against AWS on-demand at round-the-clock load the machine pays for itself in 6.2 months, and it beats renting from 21.7% utilisation upward. This is the strongest numeric argument available — see Layer 08.
Boundary: below roughly a fifth of the time, the honest recommendation is cloud.

Workloads you can sell it into with confidence

Table 11 · Suitable scenarios and how to back them
WorkloadWhy it fitsWhat to cite
LLM inference up to roughly 70B parameters96 GB and 1,597 GB/s per card, native FP4AWS and Microsoft independently place this threshold as the single-card boundary
Many independent models and multi-tenancyMIG, up to four isolated instances per cardUp to 32 isolated instances per system with guaranteed quality of service — the basis for billing
Agentic services and RAGMemory capacity for long context and KV cacheMicrosoft names RAG on models under 70B as a target scenario for this card
Fine-tuning existing modelsSufficient memory plus FP8/BF16NVIDIA names fine-tuning explicitly among Server Edition use cases — this is not a stretch
VDI and virtual workstationsvGPU over MIG, time-slicing between AI and graphicsThe datasheet lists VMware ESXi 8.0U1 and Citrix Hypervisor 8.2 — the platform was built with this in mind
Rendering, Omniverse, OpenUSD, digital twins188 fourth-generation RT Cores, 355 TFLOPS peakExactly what purely computational accelerators lack: a competitor on the H series cannot close this scenario
Video analytics, transcoding, streamingFour NVENC and four NVDEC engines per card32 encode and 32 decode engines per system, with 4:2:2 support in H.264 and HEVC
Vector search, analytics, data scienceAcceleration through cuVS and the CUDA-X stackNVIDIA claims up to 50× against a CPU system on index building
Scientific computing and FP32 simulation120 TFLOPS of FP32 per cardOne node covers what used to require a small cluster
On-premise AI in regulated industriesData and models never leave the customer's perimeterThe cloud comparison in Layer 08 shows this is also cheaper under sustained load

Workloads not to sell this machine into

The list below is more useful than the one above. A system sold for the wrong workload comes back along with your reputation, while a limitation named in time almost always converts into a different line from your own portfolio.

Table 12 · Unsuitable scenarios, reasons, and where to redirect
WorkloadWhy it does not fitWhat to offer instead
Training large models from scratchThe cards have no NVLink and memory does not pool. Training sits outside NVIDIA's stated use cases for both RTX PRO cardsNVLink platforms on the H or B series, or rented capacity for the duration of training
Distributed training or inference across nodesNetworking is a single 25 Gb/s OCP module — an order of magnitude short for inter-node trafficINNO3D AGS-4UMGX-R1: eight 400 Gb/s ConnectX-8 ports plus BlueField-3
Tensor parallelism across all eight cards under a strict SLA128 lanes across two sockets means a 4+4 split; traffic between halves crosses the inter-socket linkAGS-4UMGX-R1 with the switched PCIe Gen 6 fabric inside the SuperNIC
Storage access via GPUDirect Storage and RDMA over fabricNo DPU: nothing to offload networking and data access ontoAGS-4UMGX-R1 with BlueField-3 in a dedicated slot
Dense colocation billed per rack unit6U for eight cards against 4U for MGX — 50% more space rental for the same computeAGS-4UMGX-R1, or an MGX 2U configuration on eight RTX PRO 4500 SE
A tender with mandatory TPM 2.0The board carries only an SPI header; the module itself is optionalAdd the module, or offer AGS-4UMGX-R1, where TPM 2.0 is on-board with a Microchip CEC1736 root of trust
Pipelines with large datasets on fast drivesOnly four of twelve bays are NVMe; there is a single internal M.2AGS-4UMGX-R1 with eight E1.S bays, or external storage
A confidential computing requirementNVIDIA lists support in the RTX PRO 4500 SE specification; there is no such row for the 6000 SEAn MGX 2U configuration on the RTX PRO 4500 SE, where it is officially stated
Heavy CPU-side data preprocessing4th and 5th generation Xeon with a 350 W cap, memory at 5,600 MT/s at 1DPCAGS-4UMGX-R1 on Xeon 6 with memory at 6,400 MT/s
Edge deployment with an unstable climateOperating range 10–35 °C, humidity to 80% non-condensingPurpose-built edge platforms, or improvements to the room
Racks shallower than 950 mmA 900 mm chassis plus cable management does not physically fitAGS-4UMGX-R1 at 858 mm deep
The main consequence of that table

Of the eleven unsuitable scenarios, seven are covered by another machine in your own portfolio. The right response to "we need distributed workloads" or "our colocation bills per rack unit" is not to defend the 6220V2 but to move to the AGS-4UMGX-R1. The deal gets larger in the process: the MGX platform costs 30% more. Sell the portfolio, not the single line.

Objections you will hear, and honest answers

Table 13 · Handling objections
ObjectionHonest answer
"Why not gaming cards, they're cheaper"Passive cooling matched to chassis airflow, ECC, MIG and vGPU, power configurable from 400 to 600 W to fit the thermal budget, and system certification. A gaming card is cheaper to buy and more expensive to own: it cannot be shared between users and is not part of a supported configuration
"A competitor fits the same eight cards in 4U and you need 6U"Concede it directly. Our advantages here are price, serviceability and twice the storage bays. If density and networking matter more to the customer, offer the AGS-4UMGX-R1 — that is also our machine
"INNO3D is not on NVIDIA's system partner list"Concede it. The list on a product page is not exhaustive, and we do have a platform in the MGX form factor with ConnectX-8 and BlueField-3. Offer a check in the NVIDIA Qualified System Catalog and request written confirmation of status through the distributor — that settles the question on paper rather than verbally
"This is barebone, we need a finished server"List the missing lines up front: two CPUs, memory, drives. Offer integration and give lead times. Note as well that during a memory shortage, buying memory themselves often wins on both price and delivery
"Too expensive"Move from price to cost of ownership. Against AWS on-demand the payback is 6.2 months at round-the-clock load, and the advantage starts at 21.7% utilisation. Ask how many hours a day the cards will actually be busy — that question closes the deal more reliably than a discount
"We'll use the cloud"Agree for the pilot and for peaks — it is honest and it builds trust. Then show the threshold: above roughly a quarter utilisation, buying is cheaper. The optimal pattern is to buy for the baseline and burst into the cloud
"We need to train our own models"Clarify: from scratch or fine-tuning. NVIDIA names fine-tuning explicitly among Server Edition use cases, and there the machine belongs. Training from scratch is not our scenario, and it is better said immediately
Four claims to avoid

Each takes a minute to check and costs credibility across the whole proposal. Do not say "768 GB for one model" — memory does not pool; these are eight cards independent in memory. Do not mention NVLink — both card specifications list PCI Express only. Do not promise NVIDIA-Certified status before checking the specific configuration in the catalog. And do not transfer the 50× and 100× figures to this machine: they were measured on a server with eight RTX PRO 4500 SE against a CPU system, not on the 6000 SE and not against another GPU.

Layer 07 · PracticeFrom the customer's problem to a specification

Sizing a configuration

The commonest sizing error is counting teraflops. For inference, memory capacity comes first: if the model weights do not fit entirely into GPU memory, the system spills into system RAM and throughput collapses.

Table 14 · Workload → configuration
Customer workloadCardPlatformWhat to verify
LLM inference up to roughly 70B parameters6000 SE2U with two cardsBoth AWS and Microsoft place this threshold as the single-card boundary. Confirm context length and KV cache size
Inference on larger models6000 SE4U or 6U with eight cardsTensor parallelism; remember the PCIe penalty from the absence of NVLink
Many small models4500 SEMGX 2U with eight cardsSize by density per slot and parallel stream count, not aggregate TFLOPS
Video analytics, transcoding4500 SEMGX 2U, edge chassisThree NVENC and three NVDEC engines per card is the governing figure, not CUDA cores
Data processing, vector search4500 SEMGX 2U with eight cardsNVIDIA claims up to 50× against CPU on vector index building
Virtual workstationsEither2U or 4UvGPU licences and MIG instance count: four on the larger card, two on the smaller
Rendering, Omniverse, simulation6000 SEMGX 4U or 6URT Cores are mandatory; peak RT on the larger card is more than double
Fine-tuning existing models6000 SE4U or 6UFine-tuning is named explicitly by NVIDIA among Server Edition use cases
A pilot with no capital expenditure6000 SEAWS G7e, Azure NCv6, Google Cloud G4See Layer 08
Training large models from scratchOutside the stated use cases for either card. The interconnect listed is PCI Express only, with no NVLink row

Questions to answer before writing the specification

  1. Which models are being run, at what precision? That sets the required VRAM; the roughly 70B mark at FP8 is the single-card boundary for the larger card.
  2. What context length and how many concurrent sessions? That sets the KV cache, which frequently drives the choice.
  3. Training or inference only? If fine-tuning occupies a small share of the time, size the hardware for inference and rent capacity for the heavy work.
  4. Are graphics and ray tracing needed? If so, there is effectively no alternative to RTX PRO at this budget.
  5. How many independent consumers? Four MIG instances on the 6000 SE against two at 16 GB on the 4500 SE.
  6. How many kilowatts are delivered to the rack and how is the heat removed? This constraint trims the configuration more often than anything else.
  7. Air or liquid? On the 6000 SE these are different form factors: dual-slot against single-slot FHXL.
  8. Is a second server planned? Then choose the network now: 25GbE is adequate for a standalone machine and a dead end for distributed workloads.
  9. Finished system or barebone? This determines the number of quotation lines and the lead times.
Layer 08 · Money ProvidersConsumption models and the state of the market

Economics and market

This has to start with a caveat: NVIDIA does not publish prices for these products. Any price figures in reviews are third-party estimates, and they are not in this reference. Prices are requested from a partner or through NVIDIA Marketplace. Comparable economics can, however, be built from cloud provider documentation.

Table 15 · Cloud offerings with the RTX PRO 6000 SE
PlatformConfigurationNotable features
AWS EC2 G7eUp to 8 cards totalling 768 GB of GPU memory, Intel Emerald Rapids processors, up to 192 vCPUs, up to 2,048 GiB of system memory, up to 15.2 TB of local NVMe and up to 1,600 Gb/s of networking via EFAUp to 2.3× the inference performance of G6e. GPUDirect P2P between cards over PCIe, GPUDirect RDMA over EFA, GPUDirect Storage with FSx for Lustre. Per AWS, a single card handles a 70-billion-parameter model at FP8
Azure NC RTX PRO 6000 BSE v6Host on Intel Granite Rapids with all-core turbo up to 4.2 GHz; fractional and full configurations based on MIGCards are exposed via SR-IOV, which limits capture of some GPU telemetry. AKS supported. Target workloads: digital twins and Omniverse, inference and RAG under 70B, rendering, VDI on RTX Virtual Workstation, FP32 scientific visualisation
Google Cloud G4Fractional VMs with vGPU profiles at 12, 24, 48 and 96 GBPer NVIDIA's technical blog, the profiles target uses from streaming to high-fidelity 3D rendering and robotics sensor simulation

Three consumption models the customer chooses between

Model 01

Buying the server

Justified under constant 24/7 load, data residency requirements and a predictable operating horizon. Gives control over data and a fixed cost.

Model 02

Renting a dedicated server

The middle path: no capital expenditure, but a fixed configuration. MIG profiles let the provider sell fractions of a card and the customer take exactly the memory they need.

Model 03

Public cloud

The card is available from three hyperscalers. Good for peaks, testing and a pilot before purchase; under sustained load it is usually more expensive than owned hardware.

Argument

What belongs in the TCO

Card cost, AI Enterprise and vGPU licences, electricity and cooling (eight larger cards draw up to 4.8 kW on GPUs alone against 1.3 kW for eight smaller), rack space, vendor support, cost of downtime. Density per slot often wins the TCO case for the 4500 SE where the 6000 SE wins on peak performance.

Vendor Provider The numbers worked through

The model below is built on a real platform price and verifiable tariffs. Every assumption is stated in the table — change them and you get your own answer rather than someone else's.

Table 16 · What each platform actually ships with
ItemAGS-6220V2 · 6UAGS-4UMGX-R1 · MGX
CPUsNoneNone, listed as a configuration option
MemoryNoneNone, listed as a configuration option
DrivesNoneNone; internal and front-panel both listed as options
Power supplies4 pre-installed4 pre-installed
Fans1610 pre-installed
CPU heatsinks and carriers2 heatsinks with mounts2 heatsinks and 4 carriers
Rack railsIncludedIncluded, L-shaped
Network module25 Gb/s OCP pre-installedConnectX-8 appears in the specification; not itemised in the contents — confirm
DPUNoneBlueField-3 appears in the specification; not itemised in the contents — confirm
RAID keyIntel VROC pre-installedNot listed
Power cords4 × C19–C20Not listed

The point that settles a frequent question: both platforms are barebone. Memory and drives are bought separately either way, so comparing $121,700 with $158,210 is legitimate — both figures rest on the same basis. The small items favour the 6U: its network module, VROC key and power cords are already in the box.

Table 17 · Inputs and assumptions
ParameterValueProvenance
6U platform with 8 × RTX PRO 6000 SE, barebone$121,700Quoted price. Competing 6U models taken at the same figure
MGX 4U platform with 8 cards, barebone$158,210Taken as +30% over the 6U. Also without CPUs, memory or drives — see Table 16
Completing the 6U: 2 × 4th- or 5th-gen Xeon, 1 TB DDR5, NVMe≈ $20,000Estimate. Confirm with the supplier
Completing the MGX 4U: 2 × Xeon 6, 1 TB DDR5 or MRDIMM, E1.S≈ $25,000Estimate. Dearer for three reasons: CPUs for LGA 4710, memory at 6,400 MT/s, and data-centre-class E1.S drives
If the customer needs bulk cold storagethe gap widensThe 6U has eight SATA bays — the cheapest terabytes available. The MGX takes E1.S only. At tens of terabytes the completion gap easily doubles
System power draw6.0 kW8 × 600 W GPU + 2 × 350 W CPU + 0.5 kW for memory, drives and fans
PSU efficiency94% / 96%80 PLUS Platinum on the 6U, Titanium on the MGX 4U — per datasheets
Facility PUE1.4Assumption. A modern data centre can do better
Electricity$0.20 / kWhAssumption. Sensitivity shown below
AWS g7e.48xlarge, 8 cards, on-demand$33.14 / hrus-east-1. Consistent across four independent AWS price trackers; verify in the AWS calculator
AWS g7e.48xlarge, spotfrom $12.78 / hrLowest price; varies by availability zone and interruption risk
Horizon3 yearsAssumption
ExcludedNVIDIA AI Enterprise and vGPU licences, extended vendor support, rack space or colocation, switches and optics outside the server, staff, traffic. AWS bills egress separately
Table 18 · Three-year cost of ownership
Item6U (INNO3D and peers)MGX 4UAWS on-demandAWS spot
Capital expenditure$141,700$183,210
Electricity per year$15,656$15,330includedincluded
Three-year total at 24/7$188,669$229,200$871,032$335,803
Cost per GPU-hour at 100% utilisation$0.90$1.09$4.14$1.60
Cost per GPU-hour at 50% utilisation$1.79$2.18$4.14$1.60
Payback against on-demand at 24/76.2 months8.0 months

The governing asymmetry: owned hardware is almost entirely fixed cost, cloud is almost entirely variable. Cloud is therefore not more expensive in general — it is more expensive at high utilisation. The 50% row makes this visible: the owned machine's hourly cost doubles while the cloud rate does not move.

Table 19 · Thresholds above which buying beats renting
Scenario6UMGX 4U
Utilisation threshold against on-demand21.7%26.3%
Utilisation threshold against spot56.2%68.3%
In hours over three years, against on-demand5,692 of 26,2806,915 of 26,280
At electricity of $0.10 / kWhthreshold 19.0%
At electricity of $0.30 / kWhthreshold 24.4%

The last two rows show that the electricity tariff matters far less than utilisation: tripling the rate moves the threshold by only five percentage points. The first question to ask a customer is not what they pay per kilowatt-hour but how many hours a day the cards will actually be busy.

Three conclusions you can take into a negotiation

First. If the cards are busy more than a quarter of the time, buying beats on-demand rental, and under round-the-clock load the machine pays for itself in roughly six months. Against spot pricing the threshold is much higher — around 56% — but spot offers no availability guarantee and is unsuitable for production. Second. Cloud wins in three situations: a pilot before purchase, peak load on top of an owned fleet, and projects shorter than a year. The sensible pattern is to buy for the baseline and burst into the cloud. Third. The difference in threshold between the 6U and the MGX 4U is under five percentage points. If the customer clears the utilisation bar anyway, the choice between platforms is decided not by money but by whether they need the networking to grow into a cluster.

The 30% premium: what is in it and what is not

Useful arithmetic for the conversation, but it has to be done carefully. Eight cards at NVIDIA's marketplace price come to $106,000. That implies roughly $15,700 for the 6U chassis itself and roughly $52,210 for the MGX 4U. The premium is $36,510.

What is in it. An MGX 4U chassis, a motherboard for Xeon 6 on LGA 4710 with memory to 6,400 MT/s, Titanium-class 3,200 W supplies instead of Platinum 2,700 W, eight 400 Gb/s ConnectX-8 ports and a dedicated BlueField-3. That works out at about $4,560 per 400 Gb/s port — before the customer buys a switch.

What is not in it — an important correction. The Xeon 6 processors themselves are not covered by the premium: neither delivery includes CPUs. The premium buys the board that accepts Xeon 6; the processors are purchased separately and will cost more on the MGX than 4th or 5th generation Xeon on the 6220V2. That adds to the gap rather than sitting inside it.

What must be confirmed in writing. Neither ConnectX-8 nor BlueField-3 is itemised in the AGS-4UMGX-R1 contents: both are named only in the specification rows, as "Configured with". It reads as "included in the configured system", but with $36,510 at stake it needs to come from the supplier on paper. The direct question: are the ConnectX-8 SuperNIC board and the BlueField-3 module included in the price of the AGS-4UMGX-R1, or billed separately? If separately, the arithmetic above collapses and the comparison has to be rebuilt.

What the premium definitely does not repay. Energy efficiency: moving from Platinum to Titanium saves $326 a year, or about $980 over three years — 2.7% of the premium. Titanium should be sold as thermal headroom and reliability, not as a saving on the electricity bill.

The strongest commercial argument of the generation

According to NVIDIA, the Llama Nemotron Super model in NVFP4 on a single RTX PRO 6000 delivers up to 3× better price-performance than FP8 on an H100. That inverts the usual logic: the cheaper universal card turns out to be the better buy for inference — provided the model fits in memory. The market corroborates it indirectly: AWS independently measured up to 2.3× the inference performance on G7e instances relative to the previous L40S-based generation.

AppendixSources by tier, and glossary

Where to go next

NVIDIA Source of all card and technology data

Vendors Source of all server data

Providers Source of cloud data

Further reading not a source of figures

The material below is useful for a first pass at the subject, but any figures in it should be checked against the primary sources above — discrepancies in exactly these publications are what prompted this reference to be rebuilt.

Glossary

Table 20 · Terms
TermExpansionMeaning
MIGMulti-Instance GPUHardware partitioning of a card into isolated GPUs with guaranteed quality of service
vGPUVirtual GPUGPU virtualisation for VDI; vWS and vPC licences apply
MGXNVIDIA MGXModular reference design for server platforms
SuperNICConnectX-8 SuperNICNetwork adapter with a 48-lane PCIe Gen 6 switch
DPUData Processing UnitBlueField-3: offloads networking, data access and security from the CPU
GPUDirect P2PPeer to PeerDirect GPU-to-GPU transfer over PCIe without going through system memory
SR-IOVSingle Root I/O VirtualizationA method of exposing a device to virtual machines; used in Azure NCv6
FP4 / NVFP44-bit formatNumeric format that packs weights four times denser than FP16
NVENC / NVDECEncoder / DecoderHardware video encode and decode engines
KV cacheKey-Value cacheMemory holding conversational context; often the real driver of VRAM consumption
FHFL / FHXLFull Height Full Length / Extra LengthA full-size expansion card and its extended variant for the liquid version
BareboneA platform supplied without CPUs, memory or drives
HeadlessOperating with no display attached