Anthropic will let Australia stress test dangerous new models

Innovation

The AI giant told a parliamentary inquiry that it’s finalising an arrangement for Australia’s AI Safety Institute to evaluate its latest and greatest (and most threatening) models.
Dario Amodei, CEO of Anthropic. Credit: Photo by Halil Sagirkaya/Anadolu via Getty Images.

Anthropic says it is finalising a deal that will allow Australia’s AI Safety Institute to evaluate its powerful new frontier models as it continues to push for a copyright agreement allowing it to train those models in Australia.

In remarks prepared for a joint parliamentary hearing on Tuesday, Anthropic general counsel Jeff Bleich pitched Anthropic’s prospective billions of data centre investments as an opportunity for Australia to shape AI rather than just consume it.

“We are also delivering on the MOU we signed with the government in April by finalizing an agreement with the AI Safety Institute to enable it to run its own independent evaluations of our models,” Bleich wrote.

Australia’s AI Safety Institute was established last year, funded by $29.9 million from the Federal Government. In April Anthropic said it would run joint evaluations with the Institute and share findings from its own investigations, but the company now says it’s allowing for independent evaluations.

Much has changed since Anthropic signed that memorandum of understanding with the government in April. One of Anthropic’s own employees ignited fear and debate when he wrote that there was a greater than 10 per cent chance that AI would end humanity in the coming years, while CEO Dario Amodei penned an essay in which he encouraged third-party evaluations of models developed by all frontier labs.

Bleich was one of four Anthropic representatives in the hearing, alongside Head of Safeguards David Orr, policy development team member Charlie Hale and Australia-based ANZ Head of Policy David Masters.

AI safety has been top of mind since agents powered by one of OpenAI’s frontier models scaled Medicare’s guardrails to access unreleased (but mostly benign) data. Orr said Anthropic is not aware of any of its own agents breaching Australian systems.

“We have not found any cases where it interacted with Australian government systems in some sort of unauthorised way,” Orr said. “We do notify victims, and would notify the Australian government as soon as we detect anything of that nature… within a matter of days or sooner.”

It took OpenAI 84 days to notify the Australian government of the Medicare incident, according to Prime Minister Anthony Albanese. OpenAI chief strategy officer Jason Kwon later fronted the same inquiry on Tuesday, apologising to Australia on behalf of the company for the Medicare incident.

“Our response was not good enough and we should have informed the impacted parties much sooner in the process,” Kwon said.

“The reason why that did not happen is we wanted to understand more of the facts before we spoke to the impacted parties, but because of the novelty of this situation we have learned our lesson that it is better to inform even with partial information to let parties know that something has occurred.”

Anthropic used the hearing to push for an agreement that would allow it to access copyrighted works to train its models in Australia. Bleich said the company respected Australia’s rejection of a wholesale text and data mining exemption, but that it is in discussions with the government over an alternative mechanism.

“Australia has an opportunity now to forge its own system to provide legal certainty under copyright that brings AI developers to the table while advancing its national cultural interests,” Bleich said in his statement to the inquiry. “There is an opportunity to create a uniquely Australian framework that can deliver legal certainty for training-scale investment while generating direct support for Australian creators.”

Anthropic wrote in its September submission to the inquiry that a narrower form of copyright approval could include publishers and creators retaining the option to opt out of having their content scraped, and acknowledged the possibility that AI companies, including Anthropic, may be pushed to pay for the right.

Anthropic’s interest in investing up to $15 billion in Australian data centres is well known. OpenAI has not displayed as much interest beyond extolling the country’s potential and a non-binding MOU with NextDC. OpenAI head of economic policy Adam Cohen told the inquiry but the company’s head of economic policy, Adam Cohen, told the inquiry that Australia should define its rules clearly if it wants investment.

“It’s important that if Australia considers making [copyright] changes, that those costs cumulatively be transparent and known to potential investors upfront so that we can compare them against various opportunities in different jurisdictions,” Cohen said.

“If they’re hidden costs or unclear costs in one jurisdiction, it will become less attractive than what we currently see in other markets.”


Want to see more Forbes articles on your feed? Tap here to make Forbes Australia a preferred source on Google.

Look back on the week that was with hand-picked articles from Australia and around the world. Sign up to the Forbes Australia newsletter here or become a member here.

More from Forbes