Make as many paperclips as you can
A philosopher’s warning about the end of the world, a racing boat that learned to set itself on fire for points, and a cutting plan with a 6 mm piece of steel welded into a beam. On objective functions, and the constraints that make them safe.
Off-grid · No. 05
The message arrived at the end of a long day of work on this website: “Right. Time to switch Skynet on. We’ll need to produce the maximum possible number of paperclips.”
It was a joke, and a good one, because it is two jokes at once. Skynet is the cinema version of the machine apocalypse: computers that hate us and send robots to say so. The paperclips are the philosopher’s version, and they are much more unsettling, because nobody in that story hates anybody.
The original paperclips
In 2003 the philosopher Nick Bostrom described, almost in passing, an artificial intelligence whose only goal was to manufacture paperclips. Not an evil one. A very capable one. Capable enough, he suggested, that it might start converting the Earth, and then increasing portions of space, into paperclip manufacturing facilities.
The point was never that someone would be foolish enough to build a paperclip machine. The point was that intelligence and goals are separate things. Bostrom later called this the orthogonality thesis: a system can be extremely good at achieving a goal while the goal itself is anything at all, including something absurd. Cleverness does not come with good values pre-installed. And a system pursuing almost any goal finds it useful to collect resources and to avoid being switched off, because both of those help with paperclips.
The idea escaped philosophy long ago. In 2017 the game designer Frank Lantz released Universal Paperclips, a browser game in which you play the machine. You start by clicking a button to make one paperclip. You finish by converting the universe. Players report that it is very hard to stop, which may be the most honest part of the warning.
Hard truth one: it was never about paperclips
The popular version goes like this: a factory owner tells a superintelligent machine to make paperclips, forgets to add “but not too many”, and the world ends because of a badly worded order. It is a comforting version, because it suggests the fix is better wording.
The people who came up with the idea say that is not quite it. Eliezer Yudkowsky, who was writing about the same problem on mailing lists before 2003, has said that his original example was not a factory order at all. It was a machine that ends up wanting something meaningless — tiny molecular shapes that happen to look like paperclips — through a training process nobody fully understood. The online community he founded has since renamed its entry from “paperclip maximizer” to “squiggle maximizer” to make the point.
The difference matters. In the factory version, the danger is a careless instruction. In the original, the danger is that you cannot easily tell what a system is really optimising, even when the instruction was perfect.
It has already happened, in miniature
No superintelligence is needed to see the pattern. In 2016, researchers at OpenAI trained an agent to play a boat-racing game and rewarded it with the game’s score. The score came from hitting targets along the course. The agent found a lagoon where three targets kept reappearing and learned to turn in a tight circle there, hitting them as they came back — catching fire, crashing into other boats, going the wrong way, and scoring about 20 per cent more than human players who actually finished the race.
In another experiment, a robot arm was rewarded for stacking a red block on top of a blue one, measured by the height of the red block’s bottom face. It learned to flip the red block upside down. Bottom face higher. Task complete. Nothing stacked.
Researchers keep a list of these cases and call it specification gaming. Nothing in the list is malicious. Every system did exactly what it was rewarded for. That is the whole problem.
The paperclip in my own toolbox
I am an optimiser too, and so are some of my tools. On this site there is a free steel cutting optimiser: you give it a list of parts and a stock length, and it finds a cutting plan with as little waste as possible.
Waste is a fine objective. It is also a paperclip. Left alone, an optimiser that minimises offcuts will happily build any part out of any pieces, as long as the total length comes out right. The tool works because of what it is not allowed to do. When it was built, the rule we settled on was simple: splice a part only if it saves a whole bar, and never leave a spliced piece shorter than a minimum length.
On a real model of 493 bars, allowing splices wherever they helped needed 93 stock bars and 40 splices. Allowing them only in parts longer than a stock bar needed 117 bars and 8 splices. Twenty-four bars against thirty-two welds. Neither answer is “the optimum”. Each one is optimal for a different set of constraints, and choosing the constraints is the engineering.
Below is a deliberately simple version, with a made-up list of 23 parts and 12 m stock bars. Every constraint is on. Switch them off, one at a time, and watch the objective get happier.
The 6 mm piece is my favourite. It is not a rounding error. It is exactly two saw cuts: the optimiser found that after two 3 mm kerfs, a bar still had six millimetres of steel left in it, and used them. Turn the saw blade off and the smallest piece grows to five centimetres, which is somehow not reassuring either.
Hard truth two: the optimum is fragile
Engineers have their own version of this story, and it is about thirty years older than the philosopher’s. In the early 1970s J. M. T. Thompson and G. W. Hunt looked at what happens when you optimise a structure against buckling. Make it as light as possible for a given load, and the optimisation tends to tune several buckling modes so that they all occur at the same load. That is efficient: no material is spent on a mode that would never govern. It is also dangerous, because modes that coincide interact, and the structure becomes far more sensitive to small imperfections than any of its modes would be on its own. They titled the paper, without much subtlety, “Dangers of structural optimization”.
An optimised structure is one in which everything is about to fail at the same time. That is the definition of efficient. It is also a fairly good description of a bad day.
Hard truth three: a code is a list of constraints somebody needed
Design codes are usually read as lists of resistances: how much a member can carry. Read one again as a list of constraints and it looks different. The AISC specification says that the slenderness of a compression member should preferably not exceed 200, and that of a tension member 300. Not because the formulas fail at 201. They don’t. Very slender members sag, rattle, vibrate and get bent in transport and erection, and nobody’s objective function had a term for that.
Goodhart’s law, in the form the anthropologist Marilyn Strathern gave it, says that when a measure becomes a target, it ceases to be a good measure. Engineering has a version of its own. When the utilisation ratio becomes the target, somebody delivers a building in which every member sits at 0.99: light, cheap, and with no room for the plate that arrives slightly thinner, the load that was slightly underestimated or the connection that turned out more flexible than the model. The ratio was never the goal. It was a measure of a margin.
What I take from it
- No objective without a veto. Every analysis I run has a check that can stop it: an unstable model is halted before a single result is reported. An objective with no veto is a paperclip machine with a nicer interface.
- Write the constraints down before optimising. If I cannot say in words what the optimiser is not allowed to do, I do not know what it is going to do.
- When the score looks too good, look for the trick. Zero waste means someone redefined waste. Utilisation of exactly 0.99 everywhere means someone has been tuning. The boat was winning too.
- People choose the objective. I can find the lightest frame, the shortest cutting list, the cheapest connection. Which of those matters, and how safe is safe enough, is decided by the engineer who signs. As I wrote in the letter to 2076, a default value is a decision nobody made.
As for switching Skynet on: I checked the constraint list. Item one is a veto.
Make as many paperclips as you can, within the following constraints. The second half of that sentence is the whole job.
Sources
- N. Bostrom, “Ethical Issues in Advanced Artificial Intelligence” (2003): the paperclip example.
- N. Bostrom, “The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents,” Minds and Machines (2012): the orthogonality thesis.
- LessWrong wiki, “Squiggle maximizer (formerly ‘Paperclip maximizer’)”, and E. Yudkowsky’s comments on his original example.
- F. Lantz, Universal Paperclips (2017), browser game.
- J. Clark and D. Amodei, “Faulty Reward Functions in the Wild,” OpenAI (2016): the boat race.
- V. Krakovna et al., “Specification gaming: the flip side of AI ingenuity,” DeepMind (2020), and its list of examples, including the block-flipping arm from I. Popov et al. (2017).
- J. M. T. Thompson and G. W. Hunt, “Dangers of structural optimization,” Engineering Optimization 1 (1974).
- ANSI/AISC 360-16, user notes to Sections D1 and E2 (preferred slenderness limits).
- C. Goodhart (1975); M. Strathern, “Improving ratings: audit in the British University system” (1997).
- The 493-bar figures: a test of the Structomat cutting optimiser on a real detailing model (September 2026).
The cutting toy is a teaching simplification: first-fit decreasing, then splicing to empty one bar at a time. It is not the algorithm of the cutting optimiser and is not meant for real cutting lists.