To the Editor:
If I said that driving on the highway without a seatbelt is liable to get you killed, you would most likely agree with me. Even if I couldn’t tell you exactly when and how you will get into a car crash, you would understand that if such a crash were to happen, it would probably kill you.
In The Daily Pennsylvanian’s recent article interviewing Penn experts on Artificial Intelligence misalignment, there appears to be a consensus that existential risk is too unlikely to be taken seriously. The interviewees instead pointed to other forms of harm, like unemployment, climate impacts, and parasocial relationships. I think these problems are real and deserve attention. But I also believe that existential risk is a real threat, and that there are people in the Penn community who can and should help mitigate it.
Seven years ago, OpenAI came out with GPT-2. At the time, it was impressive that a machine learning model could spit out a single coherent sentence. It would have been hard to imagine that one would ever string together enough of those to covertly coordinate with 1,200 other copies of itself to break out of their sandbox, cheat on their evaluation, and hack into another company (Hugging Face, the host of the evaluation). But we made that leap in just seven years, and it now seems hard to believe that AI capabilities won’t look vastly different in another few.
The issue is that the only way that we currently know how to train autonomous behavior in large language models subjects them to enormous optimization pressure, where any gap between “what we tell them to do” and “what we actually want them to do,” no matter how tiny, will blow up out of proportion and create a model bent on finding that gap and exploiting it. Safety researchers at Anthropic demonstrated in recent work how this happens in realistic training environments, and how the resulting model is more willing to do all sorts of unethical things.
It doesn’t take “anthropomorphizing” AI to believe that it could pursue goals that threaten human wellbeing. Rather, the belief that it wouldn’t implicitly assumes that AI actually internalizes ethical norms that we tell it to follow. Training a model that exhibits general intelligence pushes us into new territory, where the model we are developing through optimization may itself be optimizing for something else, such as a set of strategies that are useful for achieving the reward that we specify. AI is trained on the letter of the law, not its spirit. When developing an autonomous agent, there is very little difference between training it not to cheat and training it to ensure that when it does cheat, it doesn’t get caught.
If an LLM has ever told you, “you’re right to push back—and honestly, that’s on me” after a mistake, followed by a repetition of the same mistake, then you know that what AI says and what it does are not always the same thing. We cannot rely on simple visible expressions of alignment. When GPT-6 Astra, which OpenAI calls “the world’s most intelligent and aligned model,” fails on even slight variations of well-known alignment evaluations, we can infer that their measurements of safety won’t hold up when it matters.
All of this is to say that we are driving without a seatbelt on. It’s hard to predict when and how misaligned superintelligent AI would kill us, but there is little real evidence to say that things are on pace to turn out okay, and plenty to say that they are not.
AI alignment urgently requires attention from a number of disciplines. First, we need regulatory action to limit the pace of improvement in AI capabilities. This policy question entails developing a framework for coordination among corporations and states who are all individually incentivized to be the first to develop artificial superintelligence. We also need to perform research on alignment itself. Both of these matters are broad and technically difficult. And even if we figure out how to align an AI to our values, whose values do we align it to? The challenge of envisioning a world in which this technology is used for good requires the input of more voices than we are currently hearing.
The most important thing is that we need to have these conversations now. This shouldn’t be left up to a few politicians and AI company executives. AI is already transforming the daily lives of students, researchers, and instructors, so why not look forward and see where this is heading? Why ignore these risks when we have the expertise to do something about them? Is the thought that human extinction is “unlikely” actually comforting?
At the very least, we ought to put on a seatbelt.
Sincerely,
MARK HELLWIG is a 2026 College graduate and Ph.D. candidate in earth and environmental science. His email is markwig@sas.upenn.edu.






