Draw something!
I trained a neural network to recognize 10 kinds of doodles. Try one of them, or draw something completely different and see what it thinks.
- Cat
- Dog
- Wine Bottle
- Skateboard
- Snake
- Dumbbell
- Shark
- Guitar
- Car
- Palm Tree
Start drawing and the model's answer shows up here.
The model predicts:
01How It Works
When you draw above, five things happen in a split second, all inside your browser. Here's one real cat drawing from the dataset going through the same steps, with the model's actual output for it.
-
01
You draw
Each stroke is saved as a path of points, not as a screenshot.
-
02
It becomes a 28 ร 28 image
Scaled, centered and redrawn as 784 grayscale pixels, the same format as every drawing the model trained on.
-
03
The network looks for patterns
Early layers react to simple strokes like edges and curves. Later layers combine those into bigger shapes like ears, fins or wheels.
Above: 8 of the 32 tiny 3 ร 3 filters the first layer actually learned.
-
- Cat97.4%
- Dog2.5%
- Wine Bottle<0.1%
- Skateboard<0.1%
- Snake<0.1%
- Dumbbell0.1%
- Shark<0.1%
- Guitar<0.1%
- Car<0.1%
- Palm Tree<0.1%
04
It scores all 10 things it knows
Every class gets a score, and the scores add up to 100%. It can only choose from these 10, so a drawing of a house still lands on one of them.
-
The model predicts: Cat 97.4%
05
The highest score wins
The top score is shown as the prediction, with the next three underneath it.
02How I Trained It
A neural network starts out knowing nothing. It learns from labeled examples: drawings that come with the right answer attached.
All of mine came from Google's Quick, Draw! dataset, a public collection of doodles people drew while playing Google's online drawing game. I picked 10 categories and downloaded every drawing Google has for each one, between 120,451 (guitar) and 182,764 (car) per category. Every drawing is a 28 ร 28 grayscale image: white lines on a black background.
From each category I randomly picked 24,000 drawings and split them into three groups that never overlap:
- Training set 20,000 ร 10 = 200,000
- The drawings the network actually learns from.
- Cross-validation set 2,000 ร 10 = 20,000
- Drawings it never learns from. I checked accuracy on these while building the model to compare versions and decide what to change next.
- Test set 2,000 ร 10 = 20,000
- Locked away until every decision was made, then used once. That makes the final score an honest measure of how it does on drawings it has never seen.
My first experiments used a smaller training set of 6,000 drawings per category (60,000 total). The same cross-validation and test drawings were used throughout.
The training loop
The network is made of 121,930 adjustable numbers called weights. Training means finding values for them that turn a drawing into the right answer, by repeating one small loop over and over:
- 1Show examples128 drawings at a time
- 2It guesses10 scores each
- 3Measure the errorloss: cross-entropy
- 4Adjust the weightsoptimizer: Adam
repeat, 1,563 times per pass through the training set
The error measurement is called the loss: a single number that's high when the network is confidently wrong and low when it's confidently right. After each batch, every weight is nudged slightly in the direction that lowers the loss. One full pass through all 200,000 training drawings is called an epoch, and each model trained for up to 20 of them.
03Improving the Model
I didn't start with the final model. I built five versions, changing one thing at a time and scoring each one on the cross-validation drawings.
-
Model 1Basic neural network
A simple fully connected (dense) network: every pixel feeds into every neuron, with no built-in idea of which pixels sit next to each other. Trained on 60,000 drawings.
After 20 epochs it scored 95.94% on the drawings it trained on but only 81.99% on new ones. That gap is called overfitting: the model had become very good at the examples it had already seen by memorizing them, instead of learning patterns that carry over to unfamiliar drawings.
83.31%
-
Model 2Added regularization
Regularization (L2, ฮป = 0.001) adds a penalty for large weights, which discourages the network from leaning hard on a few specific pixels and pushes it toward simpler, more general patterns.
The gap between training and validation accuracy shrank from about 14 points to about 6 (89.59% vs 83.64% after 20 epochs), but accuracy on new drawings only inched up.
83.83%
-
Model 3More training data
Same network, but 20,000 training drawings per category instead of 6,000 (200,000 total).
With more variety, memorizing gets harder and general patterns pay off more. Training and validation accuracy finished less than a point apart (87.32% vs 86.41%). A follow-up test of different regularization strengths showed the extra data was doing most of the work: with no regularization at all, the same network reached 86.79%.
86.48%
-
Model 4Convolutional network
A convolutional neural network (CNN) slides small 3 ร 3 filters across the image, so it learns local visual patterns like edges, curves and corners, and can spot them anywhere in the drawing. A dense network has to learn the same stroke separately for every position.
The biggest improvement of any change: 5.5 points higher than the best dense network. But after epoch 5, validation accuracy started slipping (down to 91.21% by epoch 15) while training accuracy kept climbing to 96.90%. Overfitting again.
92.30%
-
Final modelEarly stopping
Same CNN, with early stopping: training stops on its own once performance on the validation drawings stops improving, instead of carrying on until the model starts memorizing.
Validation loss stopped improving after epoch 7, so training halted at epoch 9 and rolled back to the epoch 7 version. That's the model running on this page.
92.49%
Results
92.40%
accuracy on the test set: 18,481 of 20,000 drawings the model never saw during training or tuning.
conv 32 (3ร3) โ pool โ conv 64 (3ร3) โ pool โ dense 64 โ 10 outputs ยท 121,930 weights
overall 92.4%
- Cats and dogs were the biggest source of confusion, mostly with each other: 229 test cats were called dogs and 179 dogs were called cats.
- Palm trees were the easiest at 96.7%, with skateboards close behind at 96.2%.
04From the Trained Model to This Website
I trained the model in Python with TensorFlow and Keras, which saved it as doodle_classifier.keras. Browsers can't run that file, so it had to be converted.
- doodle_classifier.kerasPython / Keras, trained in the notebook
- model.json + weights .binTensorFlow.js format: the layer layout plus 488 KB of weights
- TensorFlow.js in your browserruns the model on your device
The converted files hold the same 121,930 weights as the original, and I checked that they match it exactly. When this page loads, TensorFlow.js (the JavaScript version of TensorFlow) downloads the model once and runs it on your device's CPU. There's no server doing the predicting, so your drawing never leaves your browser.
Your drawing, step by step
The model has only ever seen Quick, Draw! images. Models do best on inputs that look like what they trained on, so simply shrinking the canvas to 28 ร 28 would be a problem: your lines would be a different thickness, size and position than anything the model has seen. Instead, the site redraws your strokes the same way Google made the dataset images:
-
Record your strokes
Each stroke is stored as a list of points as you draw.
-
Simplify and scale
The drawing is moved to the corner and scaled, keeping its proportions, so its longest side spans 0 to 255. Then extra points are trimmed, matching the dataset's "simplified" stroke format.
Ramer-Douglas-Peucker, ฮต = 2
-
Redraw it like the dataset
It's centered in a 256-unit square with padding, then redrawn in white on black at the same line thickness Google used.
line 16 ยท padding 16, in 256-unit space
-
Shrink to 28 ร 28
Each of the 784 pixels becomes the average brightness of the area it covers, rounded to 0 to 255 like the dataset, then divided by 255 so every value is between 0 and 1, the same normalization used in training.
-
Run the model
TensorFlow.js feeds the 28 ร 28 image to the CNN, which returns 10 raw scores (called logits).
-
Turn scores into an answer
A softmax function converts the raw scores into percentages that add up to 100%. They're sorted, and the highest one is shown as the prediction with the next three below it. This happens automatically about a tenth of a second after you lift your pen.
99.1%
To check this matches, I took 2,000 dataset drawings (the first 200 in each category), where Google publishes both the simplified pen strokes and the official 28 ร 28 image. I ran the strokes through this site's pipeline and compared: pixels differed by less than half a percent on average, and the model gave the same top answer for both versions 99.1% of the time.
Curious? Add ?debug to this page's URL to see the exact 28 ร 28 image the model receives from your drawing.