AI Safety Acculturation is Neglected
At the local AI safety co-working space, there are ~two kinds of regulars.
There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists.
Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals.
There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80k hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now.
There's a large culture gap between the rationalists and the professionals. Robust mutual understanding seems important if we want to work together on this project of not having AI blow up our civilization, and it is only by working together that there is a chance that we might succeed.
Unfortunately, there seems to be an underlying assumption that all you really need to do is stick the two groups in the same room for long enough and the acculturation will happen by default. This seems wrong to me; my experience so far is that there are many unspoken assumptions that are held by the two groups that sort of never come up in conversation, except in weird awkward eruptions when a fundamental assumption is violated. When that happens, the risk is that it pulls the two groups further apart, instead of closer together.
Some examples from recent events and discussions:
- A recent facilitated discussion where we were encouraged to process our emotions around AI safety, and some people said things like "I feel existential terror at the thought that maybe everyone I love and also our entire species is entirely wiped out in ten years' time" and other people said things like "I am feeling extremely inspired by the amount of driven and intelligent people in the room". I assume that both parties were kind of baffledly thinking "c'mon, dude, read the room" at the other party.
- A high-effort, small-group serious policy discussion of AI 2040 where most of the room simply could not take the idea of space property rights seriously because it seems like some sort of demented SF libertarian fever dream
- Me feeling very anxious when I contemplate sharing this piece in the co-working slack, because it's a moderate violation of professional norms that might result in the groups further feeling alienated from each other rather than the opposite
But also, like, timelines, so it seems useful to ask: what might deliberate acculturation look like?
Here are some ideas for how to more deliberately bridge fundamental culture gaps:
- Beyond speed friending, perhaps a regular AMA/interview type session, where you sit down a singular person in the network for an hour perhaps over lunch, and ask them questions such as: what made you care about AI safety? What's your work/research history like, what's one thing you've done that you're most proud of? What do you think is the most effective lever to change?
- Encouraging rationalists to hold basic sessions on concepts like "ask vs guess culture", "steelmanning", "double crux"
- Encouraging professionals to hold basic sessions on concepts like... hmm. This really just doesn't work at all. Rationality has a foundation of explicit, compressible concepts, while professional competence is mostly tacit and procedural, so there really isn't a straightforward analogue. The closest thing might be to get the professionals to share war stories instead, but I think sharing war stories outside of EITHER small groups or up on stage at some several hundred dollar conference is very Not Done in that culture?
- Maybe I should probably host a session that's just "here are all of the weird idiosyncratic things that rationalists are more likely to believe, ask me questions about them and I won't get mad".
- God, I really wish there was a professional counterpart to me who can host a session of "here are all of the weird things that professionals are more likely to believe, please ask questions"
- Establishing common knowledge that acculturation and mutual understanding is a task that needs to be worked on and will not passively happen by default, and people who have a foot in both worlds should consider if it would be useful for them to do something about this
Having written this list out, I am dissatisfied with it because I'm not actually that confident that doing all those things will result in the degree of trust and understanding that I think is necessary for us to do good work together.
What's the actual tension? Perhaps it's that it seems to me like the culturally rationalist are the gatekeepers of the funding resources for work on catastrophic risks, and this feels like a thing that is unsayable and somewhat anti-inductive (in that to explain it is to give an answer key to what funders want to hear, which is not ideal).
Recently a big tent animal welfare conference happened in town and I hosted a mixer for the EAs in attendance (around a third of the ~500-person conference attendees were EA). One of them, a long-established EA, gave me an interesting rundown of the wider animal welfare funding ecosystem. They said that EAs used to not be that accepted in the wider animal welfare movement, but as it became common knowledge that EAs control a large portion of the funding, the community began to embrace them more and adopt more EA frameworks and ways of thinking - but not without some amount of grumbling and resentment and bad feeling in at least some contingents. If you're not EA pilled you will simply resent the fact that the EAs are fundamentally not interested in funding the obviously morally important work of running your local donkey sanctuary, and there's literally nothing you can do about it. And if you are running a local donkey sanctuary you are sort of by definition not EA pilled.
Fortunately for AI safety, the smart policy person who wants to work on compute governance or export controls isn't proposing the AI-safety equivalent of a donkey sanctuary. And unlike the donkey sanctuary owner they have leverage in things rationalists can't buy at any price - things like institutional legibility, standing relationships, tacit knowledge about how to make key organisations do things. And perhaps the problem will also partially solve itself through funder pluralism; the rationalists are not literally the only people on the planet interested in funding alignment research and policy.
Still, I don't want to overstate the case for optimism.
Funder pluralism might result in parallel communities that don't collaborate in the same rooms at all. Animal welfare maybe only avoided this outcome because EA funding was a very significant portion of all funding, but it seems like what the rationalists are gatekeeping is one specific flavour of alignment funding.
Plenty of things professionals consider serious AI work (bias, misinformation, labor displacement, near-term harms) are donkey sanctuaries to at least some parts of the rationalist funding ecosystem, and perhaps it would be useful to make that common knowledge, but that legibility will come at the cost of resentment.
In animal welfare, EAs were newcomers who bought influence into an existing movement. In AI safety, rationalists are more likely to be ~founders, which tilts the balance even more in their favour, which brings with it a deeper (more resentment-building/polarizing) form of gatekeeping.
Bad near-term equilibria that might result (/maybe are already starting to appear) include:
- The professionals adopt rationalist frameworks because that's where some money is, and quietly grumble, and do not actually fully mean to do what they say they will do in the funding applications
- The rationalists are required to lossily launder their ideas through professional idiom, and where professional idiom fails them there is no way for them to convey their beliefs in an institutionally legible way
I'm squarely in the rationalist camp, which means that when I see a problem my default response is to write a post and publish it publicly. A skilled professional, on the other hand, would solve problems like these in a way that never produces a public artifact, and because I can't read about how they did it, I can't model what they'd do.
I'm cognizant of the fact that I'm holding only half the puzzle pieces here. If you identify more as a professional than as a rationalist, and reading this post made you feel some type of way, please come find me and give me the other half of the puzzle.