Letting myself look foolish

I prefer to avoid saying things that I suspect might make me look stupid. Over the last couple of years, however, I’ve made a conscious effort to write publicly about what I’m thinking about, even if I’m afraid it might sound dumb.

Sometimes I have this feeling, like:

I have an observation about the world, about something that doesn’t seem to add up. Probably I’m missing something, and everyone but me can easily see what’s going on, and if I bring it up, then everyone will know how foolish I am for missing this obvious thing.

Sometimes I’m right, and I really was missing something obvious. But some of my most well-received writings turned out to be ones where I had this feeling, and I was ambivalent about whether it was worth saying anything.

Lately, I’m making more of an effort to overcome that feeling, and say the stupid thing anyway, because it might turn out not to be stupid.

Continue reading
Posted on

Pausing AI at human level seems harder than pausing ASAP

Some people think we should pause AI, but not now. They say we should wait until AI reaches human level,1 because:

  • It’s not (catastrophically) dangerous until after then.
  • Human-level AI will help us do safety research.

Alternatively, other people (like me) think we should pause AI as soon as possible.

Katja Grace wrote a nice concise case for pausing ASAP. I have something I’d like to add: pausing at human level seems harder than pausing ASAP.2

Pausing ASAP sounds hard. It will be hard to get international coordination around an AI pause, and implementing a pause sounds hard even if we can agree to it in principle. But pausing ASAP still seems easier than pausing at human-level AI, for several reasons:

  • Human-level AI is highly economically valuable, and therefore there is a great temptation to keep going. The monetary incentive to build increasingly-powerful AI will be intense, and industry lobbyists really won’t want to pause AI development. It seems hard to pause when the economic incentive to continue is so great.
  • If AI is smart enough to accelerate safety work, then it’s also smart enough to accelerate improvements in AI capabilities. And it will probably be disproportionately good at the latter: capabilities improvements are easy to measure, and AI tends to be disproportionately good at easily measurable tasks. (Recent AI models have seen bigger improvements in math and coding abilities than in writing or philosophy.) If AI R&D used to require a team of PhDs, and now all it requires is someone in a garage with access to the latest AI model, then it’s harder to enforce a pause because clandestine AI research is harder to catch.
  • This next argument is more about a unilateral pause than a coordinated pause, but: Some say that the “good” AI developers need to push the frontier to maintain their lead, and that they should wait until the last minute to burn their lead to work on safety. Burning their lead at the end provides the maximum uplift from AI-assisted safety work. However, AI developers always face a choice between an uncertain downside (keep going, and possibly kill everyone) vs. a certain downside (pause, crater your profit potential, and possibly some other developer kills everyone anyway). I cannot foresee them making a rational risk assessment under those circumstances. There is too much pressure to distort their beliefs in favor of continuing to push the frontier.
  • If we had technology that could replace human labor, but it’s cheaper, faster, and can be copied as many times as one wants, how would that change the economy and society? I don’t know, but I bet it would change a lot. That level of disruption makes the world unpredictable. It seems risky to follow plans along the lines of, “let’s wait until this technology radically transforms society, possibly making things totally unrecognizable, and then implement our plan after that. Surely it will still work!”

An important counterpoint:

  • The general public does not like AI. If AI starts taking people’s jobs, then people will really dislike it. The unemployment effect of AI may create enough political will to pause that it outweighs out the economic incentive effect.

I hope this final point is strong enough to make pausing much easier in the future (hopefully the near future). But that doesn’t mean we shouldn’t try to pause ASAP.

Notes

  1. The concept of “human-level” does not have a single agreed-upon definition, but let’s say an AI is human-level if it can do most job as well as or better than skilled humans. 

  2. By the time we get political buy-in and the policy frameworks necessary to pause, “ASAP” might already have turned into “at human-level AI”. 

Posted on

A frontier AI company should shut down

Prior discussion: niplav’s shortform (2025); Planning for Extreme AI Risks (2025) by Joshua Clymer

A frontier AI company (any one, I don’t care which) should close shop and make an announcement along the lines of:

Powerful AI could end the human race. We are too worried that we don’t know how to make this technology safe. We have decided to shut down because we don’t want to be responsible for building the thing that kills us all.

A common refrain among safety-conscious AI developers: “it doesn’t matter if we stop building dangerous AI, because someone else will just build it instead.” Is that really true, though? If a multi-hundred-billion-dollar company comes out and says “We’ve concluded that our product is horribly dangerous, nobody knows how to make it safe, and there’s too high a risk that it leads to human extinction”, this won’t raise any eyebrows? This has no chance of spurring policy-makers into action?

Continue reading
Posted on

Science-driven stories are good for the same reason that character-driven stories are good

(Spoilers in this post are hidden with spoiler tags.)

What made Project Hail Mary so good? Among other reasons, it’s because the science drove the story, instead of the other way around.

Character-driven stories and hard sci-fi might take up opposite positions in the ancient battle of “people vs. things”; but when they work, they work for fundamentally the same reasons.

In mediocre “people”-focused stories, the plot dictates how characters behave. In great people-focused stories, the characters decide what happens.

In mediocre sci-fi, the plot dictates what science and technology can do. In great sci-fi, the science and technology constrain what routes the plot can take.

Continue reading
Posted on

Sentient Welfare Across Three Futures

Three categories of futures, depending on how AI goes:

  1. ASI timelines are long.
  2. ASI timelines are short, and we’re on track to solving AI alignment.
  3. ASI timelines are short, and we’re not on track to solving AI alignment.

If we want to make a good future for all sentient beings, each of these futures has different implications for what we should work on.

Continue reading
Posted on

By Strong Default, ASI Will End Liberal Democracy

The existence of liberal democracy—with rule of law, constraints on government power, and enfranchised citizens—relies on a balance of power where individual bad actors can’t do too much damage. Artificial superintelligence (ASI), even if it’s aligned, would end that balance by default.

Continue reading
Posted on

The Future Will Be Weirder Than That

Many people in the animal welfare community treat AI as a powerful but normal technology, in the same category as the steam engine or the internet. They talk about how transformative AI will impact factory farming and what it will mean for animal advocacy.

Only two futures are plausible:

  1. AI progress slows down—either because it hits a natural wall, or because civilization deliberately makes the (correct) choice to stop building it until we know how to make it safe.
  2. Superintelligent AI makes the future radically weird: Dyson spheres, molecular nanotechnology, digital minds, von Neumann probes, and still-weirder things that nobody’s conceived of.

There is no plausible middle ground where we get “transformative AI”, but factory farming persists.

Two theses:

  1. If transformative AI arrives, then it will bring about profoundly radical changes to technology and society.
  2. AGI is general intelligence. It doesn’t just accelerate technological growth: it replaces human labor and judgment across every domain.

Animal advocacy strategy needs to reckon with these.

This criticism is written from a place of solidarity—I want animal activists to succeed, which is why I want to work out our disagreements.1

Continue reading
Posted on

Which is better for sentient beings: an "ethical" AI or a corrigible AI?

Cross-posted to the EA Forum.

An aligned ASI can be “ethical”1 (it does what we think is right), or it can be corrigible (it does what its principals want). If it’s ethical, that means it will refuse unethical orders, but the tradeoff is that you can’t change its mind if you realize that the AI is wrong about ethics—its values are permanently locked in.2

Assuming we succeed at aligning ASI to human interests, which type of ASI is more likely to be good for the welfare of non-human sentient beings?

My expectations, in brief:

  • Locked-in Coherent Extrapolated Volition or similar: likely to be good (>75% chance)
  • Corrigible ASI: probably good (>60% chance)
  • Locked-in current values: probably not terrible, but will miss out on most of the future’s potential
Continue reading
Posted on

The resource-constraints argument for why aligned ASI wouldn't be bad for animals

Cross-posted to the EA Forum.

In the far future, why would people use up precious resources recreating wild-animal suffering, when they could do so many other things with those resources instead?

That argument is an important reason to expect aligned ASI to produce a future that’s okay for animals, even if it’s narrowly focused on human welfare and doesn’t care about animals at all. This is an old argument, but I couldn’t find any source that cleanly lays it out, so that’s what I will do in this post. I’m not confident that this argument is decisive, but I will simply present it without further commentary.

The argument rests on these premises:

  1. Wild animal suffering is the predominant source of suffering in today’s world, and that’s bad.
  2. Longtermism is correct.
  3. There is not an overwhelming asymmetry between suffering and flourishing (if there were an overwhelming asymmetry, then we wouldn’t care if the future has much less suffering than happiness).

By assumption, we are talking about a world where ASI is aligned, but isn’t specifically aligned to the welfare of all sentient beings. It addresses the suffering of animals, but does not preclude risks of astronomical suffering.

The argument goes:

Continue reading
Posted on

← Newer Page 1 of 12