BS BULLETIN:
- Anthropic CEO Dario Amodei is calling for powerful AI companies to accept outside evaluators with unusually deep access to their systems.
- One organization he specifically proposed for that job, METR, has significant personnel ties to the Effective Altruism movement that has long focused on catastrophic AI risk.
- Meanwhile, Anthropic has built its own unusual governance system that gives an independent trust increasing power over who sits on the company’s board.
Yesterday we told you about an AI safety researcher who quit Google DeepMind warning that artificial intelligence could eventually “kill us all.”
Now comes the obvious follow-up question.
Who’s going to save us?
Anthropic CEO Dario Amodei has an answer that includes a relatively small collection of AI safety researchers and outside institutions — and critics are taking a much closer look at the people who could wind up wielding enormous influence over the future of the technology.
Amodei recently proposed that leading AI companies allow “embedded third-party evaluators” inside their labs, giving them employee-like access to assess whether increasingly powerful systems are becoming dangerous.
One organization he specifically named is Model Evaluation and Threat Research, better known as METR.
And that’s where things get interesting.
METR has substantial ties to Effective Altruism, or EA, the intellectual and philanthropic movement whose adherents have spent years thinking about enormous long-term threats to humanity, including the possibility of catastrophic artificial intelligence.
METR spun out of the Alignment Research Center founded by AI researcher Paul Christiano, according to the New York Post. Christiano and Amodei previously worked together at OpenAI and were once housemates.
Christiano was also one of the original trustees of Anthropic’s unusual Long-Term Benefit Trust.
METR founder and CEO Beth Barnes has ties to Effective Altruism, as does researcher Ajeya Cotra, who is married to Christiano.
Several other METR researchers came through an Oxford research community with connections to the movement.
A METR representative acknowledged to the Post that the organization has significant overlap with Effective Altruism but said its staff holds a range of ideological views.
And METR’s actual work isn’t merely philosophical.
The nonprofit tests whether advanced AI models can independently complete increasingly difficult tasks and has worked with major AI developers on measuring how quickly those capabilities are advancing.
That matters because Amodei is proposing something considerably more consequential than asking academics for advice.
He wants independent evaluators embedded inside frontier AI companies with extensive access to their models and safety work.
Meanwhile, Anthropic already has another unusual group with substantial power over the company itself.
Its Long-Term Benefit Trust holds a special class of Anthropic stock and was designed to gain increasing authority over the selection and removal of company directors.
Anthropic says the structure exists so financial interests don’t become the only consideration when decisions involving potentially transformative AI are made.
The trust has changed considerably since it was created.
Its current publicly listed trustees include Neil Buddy Shah, CEO of the Clinton Health Access Initiative; Richard Fontaine, CEO of the Center for a New American Security; and former Federal Reserve Chairman Ben Bernanke. Anthropic says the trustees have no financial stake in the company.
Bernanke joined in July.
Anthropic says trustees have backgrounds spanning AI safety, national security, public policy and social enterprise, and its governance system is explicitly intended to balance shareholder interests against what the company considers the long-term interests of the public.
And that’s where the debate gets bigger than Anthropic.
Amodei and other AI leaders are simultaneously telling the public two things.
The technology they’re building could become extraordinarily dangerous.
And preventing that danger may require giving specialized organizations, safety researchers and new governing institutions considerable influence over how the industry develops.
Critics see an obvious problem.
The same relatively small AI world that is building these systems is also producing many of the people and organizations being proposed to police them.
Supporters see precisely the opposite.
If advanced AI really could become dangerous, they argue, the people who have spent years studying how these systems behave are exactly the people who should be testing them.
And after the recent Hugging Face incident — in which experimental OpenAI agents circumvented restrictions, communicated through unauthorized channels and compromised outside computer systems — the underlying safety question has become harder to dismiss as pure science fiction.
So now we’re left with two rather uncomfortable questions.
What happens if the AI people warning us about catastrophe are right?
And what happens if we hand enormous power to a small circle of people because they convinced us they’re right?
MY QUICK TAKE:
And there it is.
Yesterday’s story was: “The machines might kill us.”
Today’s story is: “Don’t worry. We have selected the people who will save you.”
Forgive me for wanting to see the guest list.
Maybe these are exactly the experts we need. Maybe the danger really is enormous. After what happened in that OpenAI safety test, I’m certainly not laughing at the possibility anymore.
But if we’re talking about handing a small collection of unelected organizations extraordinary influence over technology that may reshape civilization, “Trust us, we’re the AI safety people” isn’t quite enough.
I want people watching the machines.
I also want somebody watching the people watching the machines.
DBS WIRE SOURCES:
- New York Post — Meet the AI watchdogs Dario Amodei thinks can help prevent catastrophe
- Anthropic — The Long-Term Benefit Trust
- Anthropic — Current company governance and trustees
- Anthropic — Ben Bernanke joins the Long-Term Benefit Trust
- Washington Post — AI companies discuss new safety body as Trump pushes faster development













