Key takeaways
- The Tesla T4 (Turing, 16GB GDDR6, 70 W, single slot, no external power) fits almost any server with a free PCIe slot and is the standard card for inference, video transcoding and virtual desktop.
- The Tesla V100 (Volta, 16GB or 32GB HBM2, 250 W PCIe or 300 W SXM2) has the memory bandwidth and tensor cores for training, HPC and heavier inference, but needs a GPU-ready chassis with airflow and auxiliary power.
- Both are passively cooled datacenter cards: neither has a fan or a display output, so they belong in servers, not workstations.
- Choose by the model that has to fit in memory and by the chassis you already have; the price gap between the two is a fraction of a current-generation card.
Why used datacenter GPUs are back on shopping lists
Current-generation accelerators are allocated to the largest buyers and priced accordingly; at the same time, a growing share of AI work is inference on models that already exist, and inference on a mid-sized model does not need an 80GB card. The Tesla T4 and Tesla V100 came out of hyperscale and enterprise fleets in volume, so they are in stock, they have mature driver support, and they run the same CUDA software stack. That makes them the practical route to on-premises inference, transcoding and virtual desktop capacity for a small team, a lab or a regional host. Our Tesla collection lists both, with the T4 and V100 on their own pages.
Tesla T4: the card that fits anywhere
The T4 is a Turing-generation card with 16GB of GDDR6, a 70 W power limit, a single-slot low-profile board and no external power connector. Those four facts are the reason it is everywhere: it draws its power from the slot, it fits in 1U and 2U servers that were never designed for accelerators, and it can be installed two, four or eight to a chassis without a power-supply upgrade. It carries tensor cores for INT8 and FP16 inference and hardware video encode and decode blocks, which is why it is the default choice for inference serving, video transcoding farms and virtual desktop (vGPU) hosts.
What it is not: a training card. 16GB of GDDR6 and a 70 W envelope put a ceiling on batch sizes and throughput; models that fit in 16GB run well, larger ones do not fit.
Tesla V100: HBM2 bandwidth and tensor cores for real compute
The V100 is the Volta-generation datacenter accelerator with 16GB or 32GB of HBM2, first-generation tensor cores and much higher memory bandwidth than a GDDR6 card. It ships as a 250 W dual-slot PCIe card and as a 300 W SXM2 module for NVLink systems such as the DGX-1 and the HPE and Dell SXM2 chassis. The PCIe version needs an 8-pin auxiliary power feed and a chassis that moves enough air through a passive heatsink; the SXM2 version only fits a board designed for it.
It is the choice for training and fine-tuning models that fit in 16GB or 32GB, for HPC codes that use FP64, and for inference on larger models where memory and bandwidth, not power, are the constraint. Our V100 listings state the memory size and form factor in the product name; check the form factor before ordering.
Choosing between them
- Does the model fit? If the model and its working set fit in 16GB, the T4 does the job at 70 W. If you need 32GB, or the workload is training, the V100 32GB is the answer.
- What chassis do you have? A server without auxiliary GPU power or a GPU-rated airflow path takes the T4 only. GPU-ready 2U and 4U chassis take either.
- How many cards per node? Four T4s draw 280 W; two V100 PCIe cards draw 500 W. Power supply class decides.
- Virtual desktop or transcoding? T4. Those workloads want the encode and decode blocks and low power more than bandwidth.
What to check on a used card
Confirm the exact part number, the memory size (16GB or 32GB for the V100) and the form factor (PCIe or SXM2). Ask for the condition per listing, and plan for the driver branch the card needs: both are supported by current NVIDIA datacenter driver branches and run the standard CUDA and container tooling, but Volta support is on notice for retirement in a future branch, so check the support matrix of the driver you intend to run before committing to V100. For a node with several cards, or a larger fleet, request a volume quote with the chassis model and the number of cards, and ask for part numbers, condition and a price-valid date on the quote.
Related: Quadro RTX workstation cards for actively cooled towers, and the graphics card category for everything else.
