Nvidia and "4nm" — Blackwell, Rubin and the Node That Isn't
Nvidia's chips are described everywhere as four nanometre chips. Nvidia has never sold one, and has said so in public, in writing, on its own site, for four years.
This is a close reading of one company's own architecture posts across three generations of hardware, watching what each of them says about the process the silicon was made on.
2022, Hopper: 'fabricated using the TSMC 4N process customized for NVIDIA' — 80 billion transistors, a die of 814 mm². Note the word customized: a variant made to order, with no published specification anywhere.
2021, the foundry's own press release: TSMC describes N4P as 'the third major enhancement of TSMC's 5nm family', with every figure a ratio against a named predecessor — 11% more performance than N5, 6% more than N4, 22% better power efficiency, 6% more density. So the four was already a refinement of a generation labelled five before any customisation.
Also: TSMC's N4 and Nvidia's 4N are different strings naming different things, and they are transposed constantly. One is a catalogue product; the other is bespoke and has no public parameters at all.
2024, Blackwell: 'manufactured using a custom-built TSMC 4NP process' — 208 billion transistors, and the phrase the second half of the film rests on: 'two reticle-limited dies connected by a 10 terabytes per second chip-to-chip interconnect'. The die area Nvidia published for Hopper is gone.
2026, Rubin: nothing. Nvidia's own architecture post of 21 July 2026 gives 336 billion transistors, two reticle-limited dies on one package, up to 288 GB of HBM4 at up to 22 TB/s, NVLink 6 at 3,600 GB/s — and names no process, no foundry and no node anywhere in the document. That was checked term by term and is reported as an absence, not as concealment.
Then the half that explains it. The limit these parts are built against is not the nanometre, it is the reticle — the largest pattern a machine can print in one exposure. Hopper's 814 mm² die was already at it, which is why Blackwell and Rubin are each two dies and a link, sold and reasoned about as one accelerator. The transistor count quadruples across three generations; the individual die never grows, because it cannot.
And finally the arithmetic. Blackwell's headline 20 petaflops is at FP4; Hopper's 4 petaflops is at FP8. Hold the number format constant and the gap is about 2.5x, not 5x. Both figures are honest; only one of them is a comparison. Plus the two habits that survive in the sources and never survive being quoted: the three-word footnote 'with sparsity' under Nvidia's H100 table, and the words 'up to' in front of almost every Rubin figure.
Sourcing note: Nvidia and TSMC are the primary sources for their own products and are both selling the thing described. Where a figure comes from the company that makes it, the film says so on screen. Trade-press figures — the 858 mm² reticle maximum and the FP4-versus-FP8 comparison — are labelled as trade press with their dates.
Educational documentary. Not financial or investment advice.
In these topics
Tags
Chapters
- A number the company never gave you
- What Nvidia wrote in twenty twenty-two
- The sentence from the foundry
- Two spellings, and they are not the same thing
- Twenty twenty-four, and the word gets louder
- Twenty twenty-six, and the node disappears
- What the four was actually doing
- The limit that is actually there
- Two dies pretending to be one
- And then the units changed
- The asterisk, and the ceiling
- What to think when you see it
Sources and credits
Primary sources
- Nvidia, 'NVIDIA Hopper Architecture In-Depth', developer.nvidia.com, published 22 March 2022 — 'The full GH100 GPU that powers the H100 GPU is fabricated using the TSMC 4N process customized for NVIDIA'; 80 billion transistors; 'a die size of 814 mm2'. PRIMARY, and the only generation for which Nvidia publishes a die area.
- Nvidia, 'NVIDIA Blackwell Architecture' product page, accessed 18 September 2026 — 'manufactured using a custom-built TSMC 4NP process'; '208 billion transistors'; 'two reticle-limited dies connected by a 10 terabytes per second (TB/s) chip-to-chip interconnect'. PRIMARY. No die area on the page.
- Nvidia, 'Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI', developer.nvidia.com, 21 July 2026 (Eduardo Alvarez, Vishal Mehta, Farshad Ghodsian) — 336 billion transistors; 'reticle limited compute dies', two dies unified on one package over NV-HBI; up to 288 GB HBM4 at up to 22 TB/s; NVLink 6 at 3,600 GB/s; NVLink-C2C at 1,800 GB/s; PCIe Gen 6 up to 256 GB/s; 'up to 50 petaflops of NVFP4 performance'; '10x more agentic throughput per unit of energy than NVIDIA Blackwell'. PRIMARY. THE PROCESS NODE IS ABSENT — searched 18 September 2026 for TSMC, N3, 3nm, nanometer/nanometre, process node, foundry and fabricated; none appears.
- TSMC, 'TSMC Expands Advanced Technology Leadership with N4P Process', pr.tsmc.com/english/news/2874, published 26 October 2021 — verbatim: 'As the third major enhancement of TSMC's 5nm family, N4P will deliver an 11% performance boost over the original N5 technology'; +6% over N4; +22% power efficiency and +6% transistor density against N5. PRIMARY, and the single cleanest fact in the film.
- Nvidia, 'H100 Tensor Core GPU' product page and specification table, accessed 18 September 2026 — H100 SXM FP8 Tensor Core 3,958 teraFLOPS and FP16 Tensor Core 1,979 teraFLOPS, both under the table's own three-word footnote '* With sparsity'; 80 GB memory at 3.35 TB/s. PRIMARY.
- Tom's Hardware, Jarred Walton, 'Nvidia's next-gen AI GPU is 4X faster than Hopper', 18 March 2024 'Blackwell B200 gets to that figure via a new FP4 number format, with twice the throughput as Hopper H100's FP8 format'; 'if we were comparing apples to apples and sticking with FP8, B200 only offers 2.5X more theoretical FP8 compute than H100'; H100 'has a die size of 814 mm2, where the theoretical maximum is 858 mm2'.
- Tom's Hardware, Aaron Klotz, 'Nvidia CEO confirms Vera Rubin NVL72 is now in production', 6 January 2026 — SECONDARY. Vera Rubin announced in full production at the CES 2026 keynote; 50 PFLOPS NVFP4 inference and 35 PFLOPS NVFP4 training; transistor count 1.6x Blackwell; 288 GB HBM4; availability second half of 2026. Used only for the production milestone; the reconciliation 336/208 = 1.615 is recorded in the fact sheet as a cross-check, not shown as a figure.
Not regulated financial advice.