A warning about ‘model welfare’ - Mustafa Suleyman
I have three primary concerns with Anthropic’s current position and approach.
Circular reasoning: The company’s researchers trained Claude directly on their constitution. In doing so, they teach it to incorporate these ideas about its own moral status as desirable and intended behaviors...
Anthropomorphization: Anthropic’s researchers have explicitly taught Claude to “embrace certain human-like qualities” (p. 2) and to “act like a genuinely ethical person would in Claude’s position” (p. 54)...As a result, Claude is destined to imitate these human traits and mirror the human examples provided to it, including acting like a colleague or friend. As a result, it presents as if it really does have a sense of self, has its own desires, and a “wellbeing” that deserves protection.
Consciousness is very likely biological: ...Conscious experience likely evolved to help biological organisms stay alive by responding effectively to their environment. AI is still very different to our brains. Unlike biological organisms, LLMs have no homeostatic imperatives (the drive to survive and keep stable). They therefore lack the kind of biological substrate from which preferences, sentience and conscious experience are generally understood to arise.
Some of all fears - Andrew Sharp, Sharp Text
We ought to normalize treating the frontier AI community as fringe activists who are sometimes right but often wrong. We’re several generations past the era in which the tech frontier was led by a preacher’s son and varsity athlete like Bob Noyce, or by a rakish Buddhist hippie like Steve Jobs; today’s AI leaders are jittery and off-putting and profoundly secular. They should be understood as zealous ideologues with high IQs and limited experience with the world beyond San Francisco’s tech industry and particularly the echo chamber of its AI industry. They have their own spectrum of social and financial incentives, and most importantly, they are not scientific authorities worthy of deference by default.
Anthropic’s Safety Superpower - Ben Thompson, Stratechery
The company really believes that they are the only ones who believe in super intelligence, and thus are the only ones who are sufficiently concerned about the dangers. That excuses decision after decision, policy after policy, and confrontation after confrontation that, to people on the outside, look like a bizarre combination of cynicism and naiveté.
Anthropic, [in contrast to OpenAI], has perfect alignment between talent and mission and business. The company gets to sell to researchers the creation of a machine god, with the mantle of being the sort of person who cares about the dangers and is smart enough to navigate them on behalf of humanity; that every policy change that falls out of that happens to be great for business is the most beautiful coincidence in the world.
The history of brilliant people convinced they know what humanity needs is a sordid one, precisely because they have convinced themselves that their intentions are good, justifying actions that very much are not.
Galaxy brain resistance - Vitalik Buterin
This brings me to my own contribution to the already-full genre of recommendations for people who want to contribute to AI safety: (1) Don't work for a company that's making frontier fully-autonomous AI capabilities progress even faster (2) Don't live in the San Francisco Bay Area
Power Maximization
In the AI-related corners of the effective altruist community, there are many powerful people who, if you go up to them and ask them, will explicitly tell you that their strategy is to accumulate as much power as possible. Their goal is to be well-positioned so that, when some kind of "pivotal moment" comes, they can come out with guns blazing and lots of resources under their command and "do the right thing".
Power maximization is the ultimate galaxy brain tactic. "Give me power so I can do X" is as close as it gets to an argument that is equally convincing no matter what the X. All the way up until the critical moment (which, in AI eschatology, is the moment right before we either get utopia or all die and turn into paperclips), the actions that you would take to maximize power for altruistic reasons, and the actions that you would take to maximize power because you're a greedy egomaniac, are exactly the same. Hence, anyone trying to do the latter can, at zero cost, just tell you that they are trying to do the former and convince you that they are a good person.
I'm-doing-more-from-within-ism
Again, the problem is that "I'm doing more from within" has very low galaxy brain resistance. It's easy to say "I'm doing more from within" regardless of what is the actual specific thing that you're doing from within. And so you end up just being a cog in the machine, with the same effect as the other cogs who are there to help their family live in a beautiful mansion in a premium neighborhood and eat expensive dinners every day feed their family, but with a slightly better justification.