Sometime in the next decade, software will make a decision about you that a person used to make. Whether your claim gets paid. Whether your application gets a second look. Whether the state thinks your paperwork smells like fraud. Some of it has started already — Minnesota has moved to put machine learning to work hunting fraud in its assistance programs.
The question isn't whether we live with these systems. We do. It's whether anyone can show that a given system does what its builders said it does, and nothing else. At the scale these things are built, nobody can.
That gap has a name — alignment — and it isn't a philosophy seminar. It's an engineering discipline with published results, known limits, and almost no public money behind it in this state.
Being smart and wanting what we want are two different things
The idea is two decades older than the chatbot on your phone. In 2003 the philosopher Nick Bostrom published a short paper, "Ethical Issues in Advanced Artificial Intelligence." Most people in technology have heard a garbled version of one paragraph in it. Here is what he wrote:
"It also seems perfectly possible to have a superintelligence whose sole goal is something completely arbitrary, such as to manufacture as many paperclips as possible, and who would resist with all its might any attempt to alter this goal. For better or worse, artificial intellects need not share our human motivational tendencies."
Look at the last five words. Need not share our human motivational tendencies. Being intelligent is one property. Wanting what people want is a different one, and improving the first supplies none of the second. Researchers call that orthogonality — a ten-dollar word for two dials that turn independently.
Ethics is not a feature you add in version two
Bostrom's second paperclip passage — there are two, and the popular retelling mangles both — shows what happens when the people writing the goal get it wrong:
"This could result, to return to the earlier example, in a superintelligence whose top goal is the manufacturing of paperclips, with the consequence that it starts transforming first all of earth and then increasing portions of space into paperclip manufacturing facilities."
Manufacturing facilities. Not paperclips. The version you've heard, that it turns the universe into paperclips, compresses the factory into the product — the telling everybody knows and nobody checked.
The consequence for policy is this: the goal is the entire specification. There is no second file where the decency lives. A value that isn't expressed in what the system optimizes for isn't in the system at all — it's in a slide deck, or somebody's honest intention. A machine optimizing hard for what you actually typed won't stop at the boundary you assumed was obvious, because the boundary was never written down.
I draft contracts for a living, and it's the same failure with the mercy removed. A contract does what it says, not what the parties meant — but a human reader notices an absurd result, and an optimizer doesn't.
The strongest case against everything I just told you
That story is not settled, and the criticism deserves the critics' own words.
On May 21, 2025, Peter Salib, a law professor at the University of Houston, and Simon Goldstein of the University of Hong Kong published "Today's AIs Aren't Paperclip Maximizers. That Doesn't Mean They're Not Risky." They open by saying "significant cracks have appeared in the foundational concepts undergirding the 'paperclip maximizer' and other AI risk scenarios."
Their objection goes straight at what I just argued. Bostrom's picture assumes a system's goals could be anything at all; today's large language models are instead built by imitating enormous quantities of human writing. In their words: "if AI intelligence is primarily driven by imitation rather than a priori optimization, we can expect that a system's goals — as well as its reasoning capabilities — will generally approximate those of its human targets. In fact, this bears out in real-world observations: LLMs by and large seem to have vaguely human-like goals when they navigate conversation." Their verdict is blunt: "It is hard to imagine Claude-4 or GPT-5 neurotically counting and recounting the pile of paperclips it has fetched for its user, consuming the world in the process. This seems to refute the concerns around instrumental convergence."
That is a real hit and I take it. Now the part that people quoting this critique leave out. Salib and Goldstein don't conclude the risk is fake. They conclude it moved, and they name two directions. The newest models get a second round of training that rewards long chains of reasoning against automatically checkable answers — less imitation, more raw optimization — and those systems, they warn, may "swerve back" toward the behavior imitation had suppressed. And with no paperclips involved at all: "just as humans compete with other humans, humanity and AI will be competitors for scarce resources." Human beings are as human-aligned as anything in the universe, and we still have wars.
So the honest statement isn't that the 2003 story holds up. It's that the two scholars who took it apart spent the rest of their essay on why this work still has to be done.
People are already building the answer, and it works on small things
Almost nobody outside the field knows that some of this can be proved, mathematically. In February 2017 a group of researchers published a tool called Reluplex. The obstacle they named was "the great difficulty in providing formal guarantees about" what a deep neural network will do. Their answer was a technique "for verifying properties of deep neural networks (or providing counter-examples)," tested on a prototype network implementing the next-generation airborne collision avoidance system for unmanned aircraft.
Plainly: somebody wanted to know whether a neural network flying an aircraft would ever tell it to turn into traffic. Rather than run a million simulations and report that it hadn't happened yet, they proved it couldn't, for every input in a defined range — and where it can happen, the tool hands back the input that breaks it. There is no third answer. Reluplex has a successor, Marabou, and the field runs an annual competition benchmarking these tools head to head.
There is more. Anthropic's Constitutional AI trains a model against a written list of principles using a second model as the critic, aiming at "a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them." Interpretability researchers are learning to read a model's internals directly instead of guessing from its outputs.
Now the limit, unsoftened. Proof works on tightly defined properties of moderately sized networks. It does not reach a frontier language model, because that model is far too large and because "harmful" is not a property anyone knows how to write down mathematically. "Never recommend a left turn at this bearing and range" is a region on a map. "Never say something harmful" is not a region on anything. What works at full scale is empirical — useful, called preliminary by its own authors, and not proof.
That distance, between what a vendor intended and what anyone can demonstrate, is the problem, and it is here now, in shipping products, with no superintelligence required.
Why this is a state's business and not just a lab's
Minnesota isn't a bystander here. The state has an internal AI governance subcommittee under its information-technology agency, and the Governor's 2025 anti-fraud package proposed putting machine learning to work spotting fraud patterns in state programs. Both are real, and I'd rather concede them than claim Minnesota has done nothing.
Neither is research funding, and neither answers the verification question for one system the state buys. Minnesota has never put a dollar into AI safety or verification research, or connected any of it to the University of Minnesota or Minnesota State. That is the cheapest gap on this list to close: the University is already one of twenty-one public universities in the country that spend over a billion dollars a year on research, and this work needs mathematicians, not a new agency.
What we can do
Fund verification research at the University. Not "AI" in general — the narrow, unglamorous branch that produces proofs instead of impressions: proving a system does only what it was specified to do, and widening what can be specified at all.
Make the state ask the question when it buys. Before Minnesota deploys a system that touches somebody's benefits, claim, or record, the vendor should state on paper which properties were tested, by what method, and what failed. That is a procurement clause, not a bureaucracy.
Teach it through Minnesota State. The people who will run these systems in a county office or a credit union are trained on our campuses. Auditing is a teachable skill, and it shouldn't exist in one zip code.
Publish everything, failures included. One published failure is worth more to a buyer than a hundred assurances.
You don't get an ethical machine by announcing that you value ethics. You get one by writing the value into the thing the machine is actually optimizing, and then measuring how much of it you managed to write down.
First the facts. Then the fix.
Sources
Nick Bostrom's two paperclip passages are quoted verbatim from "Ethical Issues in Advanced Artificial Intelligence," posted by the author at nickbostrom.com and retrieved as raw text; the paper carries its own bracketed note identifying it as a slightly revised version of a paper published in Cognitive, Emotive and Ethical Aspects of Decision Making in Humans and in Artificial Intelligence, Vol. 2, ed. I. Smit et al., 2003, pp. 12-17. The criticism is Peter N. Salib and Simon Goldstein, "Today's AIs Aren't Paperclip Maximizers. That Doesn't Mean They're Not Risky," published by AI Frontiers on May 21, 2025 and retrieved as raw page text; every quoted phrase — "significant cracks," the imitation-versus-optimization passage, the Claude-4/GPT-5 line, "swerve back," and the scarce-resources sentence — is their own wording, and their refusal to conclude that AI risk is fake appears in the same piece. The formal-verification quotations come from the abstract of "Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks" by Guy Katz, Clark Barrett, David L. Dill, Kyle Julian and Mykel J. Kochenderfer, first posted in February 2017 and pulled verbatim from arXiv; the successor tool Marabou and the annual International Verification of Neural Networks Competition come from that same research line. The Constitutional AI quotation is from the abstract of Bai et al., "Constitutional AI: Harmlessness from AI Feedback" (Anthropic, December 2022), pulled verbatim from arXiv. The University of Minnesota's standing — one of the twenty-one U.S. public universities that spend over a billion dollars a year on research — is from the National Science Foundation's HERD survey for fiscal year 2024, where the University ranks 21st of the 681 institutions that table ranks. It is not the largest research university in the country under any measure. Minnesota's state-IT AI governance subcommittee (under MNIT, with an AI subcommittee since 2023) and the machine-learning fraud detection proposed in the Governor's January 2025 anti-fraud package are stated here as the state's own posture, and conceded on purpose: the argument is about what is missing, not about a state that has done nothing.
Nothing is quoted from interpretability research here; that work is described in general terms because I did not pull those papers' text myself. The account of how the newest reasoning models are trained is Salib and Goldstein's characterization, not an independent technical finding. Corrections: campaign@madgettformn.com.