Forwards or Backwards

Topics › Artificial Intelligence (AI)

IBM NorthPole — Inference on a 12nm Chip

Published · 20 min

Watch on YouTube

IBM's NorthPole is built on a twelve nanometre GlobalFoundries process. In 2023, that was not a leading-edge node and was not close to one. On the inference workloads it was designed for, it outran parts built on far smaller nodes.

It did that by removing something rather than adding it.

Running a trained neural network is not a hard arithmetic problem. It is a waiting problem: the weights live in memory chips outside the processor, fetching them is far slower than the arithmetic, and most of the engineering in a modern accelerator exists to hide that wait rather than remove it. NorthPole runs inference without off-chip memory. 256 cores, each with its own memory beside it, 224 MB on the chip in total, 13 TB/s of on-chip memory bandwidth per chip. There is no fetch in the inner loop, so there is no wait to conceal. Dharmendra Modha's summary: architecture trumps Moore's law.

The film covers the chip in full - 22 billion transistors, 800 mm2, 12 nm, 256 cores, 2,048 threads, and the mixed 8-, 4- and 2-bit precision that is not a trick but the thing that makes staying on-chip possible at all. Then the results: a 2024 single 2U server of 16 cards running a three-billion-parameter variant of Granite-8b-code-base-4k at 4-bit, reaching 28,356 tokens per second; and the 2025 scaled system of 288 cards across 18 servers, 115 peta-ops at INT4, 3.7 PB/s, 30 kW, 730 kg, in 0.67 square metres.

And then the part most coverage of this chip skips. The two famous multipliers - 46.9x and 72.7x - are IBM's measurements of IBM's chip against GPUs the source does not name, and no independent benchmark of NorthPole was found for this film. The newest, largest and most formally presented result, a November 2025 preprint with 31 authors, carries no GPU comparison at all. The film states all of that on screen.

It also refuses one graphic it could easily have made: 2024's sub-millisecond latency next to 2025's 2.8 ms. Those are not comparable - 2.8 ms is per user with 28 users served at once on 288 cards - and putting them side by side would have been the easiest and most dishonest frame in the film.

What survives all of it is the finding that does not need a benchmark: on this workload, where the memory is mattered more than how small the transistors were.

The November 2025 arXiv paper is a preprint and is labelled as one on screen the first time it is used. Nothing in this film claims NorthPole has displaced the GPU.

Educational documentary. Not financial or investment advice.

In these topics

Tags

Chapters

  1. The chip that went the wrong way
  2. The problem is not arithmetic
  3. What NorthPole took out
  4. The chip itself
  5. Precision, and why it is not cheating
  6. Twenty twenty-four: one server
  7. The two numbers this film will not repeat carelessly
  8. Twenty twenty-five: the rack
  9. Two point eight is not slower than one
  10. What is missing from the evidence
  11. The constraint nobody has removed
  12. What this actually demonstrates
  13. Forwards or backwards

More from Forwards or Backwards on YouTube

Sources and credits

Primary sources

  • Modha et al., 'Neural inference at the frontier of energy, space, and time', Science, published 19 October 2023, doi 10.1126/science.adh1174 - the chip: 22 billion transistors, 800 mm2, 12 nm GlobalFoundries, 256 cores, 224 MB on-chip (192 MB core array + 32 MB frame buffer), 2,048 threads, 2,048/4,096/8,192 ops per core per cycle at 8/4/2-bit, over 4,096 wires crossing each core in each direction.
  • '11.4 IBM NorthPole: An Architecture for Neural Network Inference with a 12nm Chip', ISSCC 2024, IEEE - the architecture paper.
  • IBM Research blog, 'NorthPole LLM inference results', published 26 September 2024 - 16 cards in one 2U server; a three-billion-parameter variant of Granite-8b-code-base-4k at 4-bit weights and activations; latency below 1 ms per token; 28,356 tokens/s system throughput; 46.9x and 72.7x against UNNAMED GPUs. Quotes Dharmendra Modha: 'What is essential here is qualitative orders of magnitude in improvement.'
  • arXiv 2511.15950, 'A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference', 31 authors, IBM, submitted 20 November 2025 288 cards across 18 2U servers; 115 peta-ops at 4-bit integer; 3.7 PB/s; 30 kW; 730 kg; 0.67 m2 42U rack; 3 instances of Granite-3.3-8b-instruct at 2,048 context serving 28 simultaneous users at 2.8 ms per-user inter-token latency; or 18 instances of a 3B model, or one 70B instance. This paper is also the primary source that settled the die area, the per-chip bandwidth unit and the identity of the 2024 model.

Note: Every performance figure in this film traces to IBM or to IBM's own authors.

Note: The comparison GPUs behind the 46.9x and 72.7x figures are not named in the source, and the film does not imply any named modern GPU.

Not regulated financial advice.