- EC - Stories That Build the Future
- Posts
- Two Gaming Cards in a Bedroom: How AlexNet Started Modern AI
Two Gaming Cards in a Bedroom: How AlexNet Started Modern AI
In 2012, two students with $1,000 of gaming hardware beat the best-funded labs in the world. It only worked because of a bet NVIDIA had made six years earlier, for a market that barely existed.

In 2012, the best-funded computer vision labs in the world were racing to solve the same problem: how do you teach a computer to recognize what it is looking at? The approach at the time was painstaking. Researchers hand-built algorithms to detect edges, corners, and textures. Every year, the field gathered at the ImageNet Large Scale Visual Recognition Challenge to see whose hand-engineered system performed best. Progress was incremental.
Then two graduate students from the University of Toronto entered, and won by a margin nobody thought possible. Their hardware was two consumer gaming cards. The machine ran in a bedroom.
This is the story of that win, and of the bet that made it possible. A bet a chip company had made six years earlier when almost nobody understood it.

A NVIDIA GeForce GTX 580 gaming GPU, the same model AlexNet was trained on. The network ran on two of these cards, each costing about $500 at launch.
The Outsiders
Alex Krizhevsky and Ilya Sutskever were graduate students under Geoffrey Hinton, one of the founders of the field that would come to be called deep learning. Hinton had spent decades arguing that neural networks, which learn features on their own rather than being hand-programmed, were the right approach. For most of that time, the AI field disagreed.
Sutskever convinced Krizhevsky, who had an unusual talent for wringing performance out of GPUs, to train a neural network for the ImageNet competition. ImageNet was a dataset of 1.2 million labeled photographs assembled by Fei-Fei Li (now a Stanford professor) starting in 2006, crowdsourced through Amazon Mechanical Turk. Without it, there was nothing comparably large to train AlexNet on. Krizhevsky had already written CUDA code for a smaller neural network, called cuda-convnet, trained on a dataset called CIFAR-10. He extended it to support multiple GPUs and scaled it up to the much larger ImageNet dataset.
The training ran on a computer with two NVIDIA GTX 580 cards in a small setup at Krizhevsky's parents' house, often described as his bedroom. Each card cost about $500. Training took roughly six days.

Geoffrey Hinton with graduate students Alex Krizhevsky and Ilya Sutskever. Their network, AlexNet, started the modern AI era.
The Win
AlexNet won the 2012 ImageNet competition with a top-5 error rate of 15.3%. The second-place entry scored 26.2%. A margin of nearly 11 percentage points in a field that usually moved by fractions of a percent.
When Krizhevsky presented the results in Florence that October, the room was standing-room only. The computer vision community understood immediately what had happened. The hand-built approach was over. Deep neural networks trained on GPUs were the future.
Within months, Hinton, Krizhevsky, and Sutskever formed a company called DNNresearch. Google acquired it for $44 million. Sutskever later co-founded OpenAI. Hinton would go on to win the Nobel Prize in Physics in 2024. AlexNet was written in CUDA and C++.
The code itself is public. In 2025, the Computer History Museum released the original 2012 AlexNet source code, the actual file Krizhevsky wrote and ran on those two GPUs in his bedroom. If you have never read machine learning code before, it is worth a look: you can see the layer definitions, the training parameters, and the GPU memory split that the paper describes. It is shorter than most people expect. The repository is at github.com/computerhistory/AlexNet-Source-Code.
The architecture split its layers across two GTX 580 cards, an engineering constraint that shaped the design.

AlexNet architecture from the original 2012 NeurIPS paper by Krizhevsky, Sutskever, and Hinton. The two rows show the network split across two GPUs.
Training a neural network is, at the hardware level, a massive matrix multiplication problem. The network processes millions of images, adjusting billions of numerical weights across thousands of training iterations. CPUs handle a few complex tasks at a time, sequentially. GPUs handle thousands of simple tasks simultaneously, in parallel, which is exactly what matrix math at scale requires. A GTX 580 had 512 CUDA cores running in parallel. A comparable CPU had 4 to 8. For this kind of work, that is not a marginal difference.
Here is the part most retellings skip. Those two gaming cards could only train a neural network because of a decision NVIDIA had made six years earlier.
In 2000, GPUs were graphics chips. Their massive parallelism was locked behind low-level graphics APIs that only specialists could use. A Stanford PhD student named Ian Buck, working under Pat Hanrahan, built a programming language called Brook that let ordinary programmers write parallel code on a GPU. NVIDIA hired Buck in 2004. By 2005, he was building CUDA inside the company with a small team.
In November 2006, NVIDIA shipped the G80. It had 681 million transistors and was manufactured by TSMC. Crucially, it unified its 128 shader processors into a single programmable array that could run general-purpose code, not just graphics. It was the first gaming GPU ever architected for this. CUDA, the software layer, launched publicly in 2007.
The bet was expensive. The features that made the G80 good at general-purpose computing added cost and complexity without helping its core gaming business. As Jensen Huang told graduates at NTU in 2023:
"The added cost of CUDA was very high. Nvidia's profits took a huge hit. For many years, our market cap hovered just above $1 billion. We suffered many years of poor performance. Our shareholders were skeptical of CUDA, and preferred we focus on improving profitability. But we persevered."
AMD watched and made the opposite call, treating general-purpose GPU computing as a niche. For five years, CUDA lived in scientific labs. There was no obvious market. Then AlexNet happened. As Jensen put it in the same speech: "The famous AlexNet trained on GeForce GTX 580 started the big bang of AI."

The NVIDIA G80 engineering sample, launched November 2006. The first GPU built for general-purpose computing.
The Moat, and the Live Debate
After AlexNet, every AI lab on the planet realized GPUs trained neural networks dramatically faster than CPUs. They all bought NVIDIA. Every framework that followed, Theano, Caffe, TensorFlow, PyTorch, JAX, was built primarily on CUDA. By the time AMD recognized what was happening, the entire field had standardized on a software stack that only ran on NVIDIA hardware.
That is the moat. Not the chips. The ecosystem. Eighteen years of researchers, graduate students, libraries, documentation, and textbooks, all built on CUDA. AMD's competing framework, ROCm, is technically credible and still struggles for adoption, because the ecosystem is not there.
In 2012, NVIDIA was worth less than $10 billion. Today it is worth around $5 trillion. The question now is whether that moat holds. Two weeks ago, the journalist Dwarkesh Patel spent 103 minutes challenging Jensen Huang on exactly this: are TPUs going to commoditize the hardware, will AI-generated code make CUDA's developer lock-in irrelevant, why is Anthropic making massive bets on Google TPUs and AWS Trainium alongside NVIDIA. Michael Burry, the Big Short investor, shared the interview and said every NVIDIA bull should watch it.
Jensen Huang argued that the install base, the developer tooling, and the 18 years of libraries are the real product: "If you want to build on any computer first," he said, "building on CUDA first is incredibly smart."
The cracks Dwarkesh pointed to are real. The moat is also real. Both things are true at once.

The NVIDIA H100, a modern descendant of the G80 architecture, 18 years on.
What NVIDIA Saw
In 2006, the question NVIDIA answered was not "how do we build a better graphics card?" It was: what if GPUs become the dominant compute platform of the next era, and we make it as easy as possible to program them? Almost nobody believed that in 2006. Pat Hanrahan did. Bill Dally did. Ian Buck did. Jensen Huang did. They did not know about AI specifically. What they bet on was that someday, someone would find a workload GPUs were uniquely good for, and the software would already be there when that day came.
That day was September 30, 2012, in a bedroom in Toronto, on two gaming cards that cost a thousand dollars. The hardware did not win the AI war. An 18-year software bet did.
The question now is whether the next eighteen years break the moat, or extend it.
This is Part 4 of 5 in The Semiconductor Series at Engineering Community. Next week: The Traitorous Eight, how eight engineers walked out of one bad boss's office in 1957 and built Silicon Valley.
If this was worth your time, share it with someone who would appreciate the engineering.
Image credits:
Hinton/Krizhevsky/Sutskever trio photo: University of Toronto, via Computer History Museum (computerhistory.org)
AlexNet architecture diagram: Daniel Voigt Godoy, via Wikimedia Commons (CC BY 4.0)
Got a story worth telling?
Engineers, scientists, and builders - if you've been part of something that changed the game, we want to hear about it. If it's a fit, we'll share it with 300,000+ readers who live for this stuff.