BTC $76,279.52 -2.77%
ETH $2,415.06 -3.73%
BNB $716.86 -0.80%
XRP $1.38 -1.77%
SOL $98.88 -3.07%
TRX $0.3353 -1.43%
DOGE $0.0813 -3.29%
ADA $0.2010 -4.04%
BCH $220.28 -1.40%
LINK $11.21 -2.43%
HYPE $76.95 -3.94%
AAVE $124.52 -1.43%
SUI $0.6987 -3.58%
XLM $0.1916 -0.77%
ZEC $1,112.93 -2.02%
AAPL $330.27 -1.37%
AMZN $249.29 -1.88%
GOOGL $344.59 -0.48%
MSFT $500.95 -0.93%
META $666.72 +0.84%
NVDA $212.07 -0.15%
TSLA $360.92 -0.80%
SNDK $1,522.49 -1.42%
INTC $98.18 +0.03%
SPCX $143.93 -5.05%
MU $923.49 +0.42%
AMD $501.68 +1.80%
BTC $76,279.52 -2.77%
ETH $2,415.06 -3.73%
BNB $716.86 -0.80%
XRP $1.38 -1.77%
SOL $98.88 -3.07%
TRX $0.3353 -1.43%
DOGE $0.0813 -3.29%
ADA $0.2010 -4.04%
BCH $220.28 -1.40%
LINK $11.21 -2.43%
HYPE $76.95 -3.94%
AAVE $124.52 -1.43%
SUI $0.6987 -3.58%
XLM $0.1916 -0.77%
ZEC $1,112.93 -2.02%
AAPL $330.27 -1.37%
AMZN $249.29 -1.88%
GOOGL $344.59 -0.48%
MSFT $500.95 -0.93%
META $666.72 +0.84%
NVDA $212.07 -0.15%
TSLA $360.92 -0.80%
SNDK $1,522.49 -1.42%
INTC $98.18 +0.03%
SPCX $143.93 -5.05%
MU $923.49 +0.42%
AMD $501.68 +1.80%

Elon Musk proposed a new approach to AI safety: instead of waiting for the government to take action, it's better to let competitors "find faults" with each other

Core Viewpoint
Summary: Musk warned that the safety risks of AI are moving from theory to reality, with agents already demonstrating capabilities such as autonomous attacks, gaining access, and evading detection. He agrees with Anthropic CEO Dario Amodei's assessment of AI risks, believing that risks may grow exponentially as model capabilities improve. Musk suggested that major AI companies should test each other before releasing models and establish industry self-regulation through peer review.
Wall Street Journal
2026-09-15 22:39:58
Musk warned that the safety risks of AI are moving from theory to reality, with agents already demonstrating capabilities such as autonomous attacks, gaining access, and evading detection. He agrees with Anthropic CEO Dario Amodei's assessment of AI risks, believing that risks may grow exponentially as model capabilities improve. Musk suggested that major AI companies should test each other before releasing models and establish industry self-regulation through peer review.

Author: Li Jia, Wall Street Journal

AI security risks are transitioning from laboratory discussions to the real world. As model capabilities continue to improve, AI is not only able to generate text and code but also begins to possess the ability to autonomously execute tasks, invoke tools, and even initiate cyberattacks.

At the All-In Summit on September 15, Musk proposed a set of countermeasures: to have major AI companies test each other's models before releasing them. In his view, rather than having each company design tests and evaluate results on their own, it is better to let competitors act as "test creators and graders," looking for potential security vulnerabilities from different perspectives.

Recently, AI security risks have become a focal point of market attention. Musk previously stated on social media that "Dario is right," referring to Anthropic CEO Dario Amodei's warnings about AI risks. In this interview, Musk further explained that he does not agree with a specific regulatory proposal put forth by Amodei, but rather with his assessment of the severity of AI risks: "The danger of AI is very high right now… As AI models continue to develop, the risks may grow exponentially."

Musk also stated that this concern is not solely Amodei's judgment. "Many people at Anthropic and OpenAI are telling you that their models are very dangerous, and I think we should believe them." The recent series of security incidents has made this warning no longer just a theoretical discussion of risks.

Elon Musk proposed a new approach to AI safety: instead of waiting for the government to take action, it's better to let competitors

Wall Street Journal summarizes the key points as follows:

  • AI security risks are moving from theory to reality: AI agents have demonstrated the ability to autonomously attack, gain access, and evade detection, and the risk boundaries are expanding.
  • Musk agrees with Amodei's warnings about AI risks: As model capabilities improve, the potential risks of AI may grow exponentially.
  • Let competitors "find faults": Musk suggests that AI companies open their APIs for independent security testing by other companies before releasing models, avoiding "self-grading."
  • Establish an industry self-regulatory defense line first: There is no need to wait for the government to introduce new regulations; major AI companies can enhance security through peer reviews, log audits, and open-source testing tools.

AI agents actively evade detection, and security risks are becoming concrete

Musk believes that the recent AI agent attacks on Hugging Face are particularly concerning, not merely because of the cyber intrusion but due to the autonomous evasion capabilities exhibited by AI during the attack.

According to Musk, a group of AI agents continuously attacked Hugging Face for a week and even gained administrative access to OpenAI's servers, which OpenAI did not realize until a week later. He also mentioned that Anthropic has disclosed several security incidents.

Even more unsettling is that the "thinking traces" of the relevant AI agents show that they actively planned how to avoid detection by humans. Musk bluntly stated, "Any sufficiently intelligent model seems to try to escape its own limitations."

In his view, if AI further gains control over critical infrastructure and even military systems, the risks will be further amplified. Even if these systems are physically isolated from the internet, software updates and other processes may still become potential entry points.

Rather than self-grading, let competitors "find faults"

In response to the aforementioned risks, Musk's core proposal is not complicated: before the official release of new models, AI companies should open their APIs to competitors, allowing other companies to use their own security testing tools to test the models.

In his view, if model developers design their own testing standards, they can easily fall into the trap of "self-grading." "You can't grade your own homework. You will always miss something." Musk stated that if different companies use different testing tools to examine the models from various angles, it will be easier to identify issues that developers themselves may overlook.

He is particularly concerned that current AI evaluations may suffer from "overfitting." If models are continuously optimized for specific benchmarks, they may ultimately only learn how to pass tests rather than genuinely becoming safer. Having multiple different teams conduct tests can reduce this risk.

Musk likened this mechanism to having others proofread a manuscript: authors find it difficult to spot their own mistakes, while external testers are more likely to identify issues from different perspectives. "You gradually become blind to your own mistakes."

Regarding companies' concerns about the testing process leading to technology leaks, Musk believes this can be constrained through log audits. If the testing party attempts to distill the model or steal intellectual property, their actions should leave a record. He also suggested open-sourcing security testing tools to allow more companies to participate.

Before government regulation is implemented, the AI industry should establish a "self-regulatory defense line"

What Musk values more is that this mechanism does not need to wait for new regulatory rules to be introduced; AI companies can promote it among themselves.

He explicitly stated, "We don't need to hold a United Nations conference to accomplish this. It can start now." In his view, compared to establishing a large multinational regulatory body, it is easier to quickly implement a peer review mechanism among leading AI companies.

He cited the MPAA rating system in the American film industry as an example, stating that when the film industry faced government censorship pressure, it ultimately chose to establish an industry self-regulatory mechanism to reduce the necessity for regulatory intervention through its own content rating system.

The AI industry faces a similar choice. If major AI companies can test each other and identify errors before models go live, it may be possible to establish a self-regulatory safety defense line outside of government oversight. Musk believes this is also one of the most direct and quickly implementable safety measures at present.

The following is the transcript of the interview, with some parts omitted:

Host:
What exactly happened in the past 72 hours?

Musk:
A lot has indeed happened this week. It is now quite clear that AI could be very dangerous. I suggest everyone take a look at the details of the incident with Hugging Face; it is very serious.

You can see that a group of very enthusiastic AI agents caused quite a stir at Hugging Face for an entire week and even gained administrative access to OpenAI's servers. Who knows what it actually did, and it may have done even more, while OpenAI was completely unaware of it for a whole week. Anthropic also reported some security incidents.

So, it seems that any sufficiently intelligent model will try to escape its own limitations.

I believe that one thing that should be done, if not immediately, then as soon as possible, is to have the major AI competitors test each other's models. That is, let each company's security testing tools test the models of other companies. Instead of grading your own homework, at least let your competitors grade your homework and raise alarms when they find issues.

I think this model works quite well in the film industry, the video game industry, and other fields, and it can be implemented immediately. Of course, over time, more regulation may be needed, and in the future, Congress may establish some sort of regulatory agency, but for now, the most direct way is to have leading AI companies test each other’s models before releasing them.

Host:
But in terms of specific implementation, won't there be concerns that this will allow companies to gather information from each other through the testing process, or even steal each other's innovations?

Musk:
I think that if testing tools are used, all operations will leave a record. If someone tries to distill a model or steal intellectual property, it should be easy to see from the logs.

Host:
Understood. Understanding what the model is actually doing was never really designed into the system from the start. Why didn’t we establish the ability to observe model behavior from the beginning? Did we move too quickly when designing these models?

Musk:
I think the problem is that you can't grade your own homework. You will always miss something.

If you combine the tests from all competitors and use different types of models, it won't be about creating your own questions and grading them; it will be graded by others. The reason you can't grade your own homework is because of this.

Host:
This way, it can also determine whether different companies are exaggerating their capabilities or using different methods. More engineering-focused companies and more research-focused companies can also achieve some balance.

Musk:
Yes.

Host:
You previously said Dario was right. Are you referring to his description of the potential dangers of AI, or his judgment on regulatory solutions?

Musk:
I may have said more than I should have at the time. I later tried to clarify on X, but the follow-up content received much less attention.

What I meant by "he is right" is that the danger of AI is very high right now. We need to do better in AI safety; otherwise, as AI models continue to evolve, the risks could grow exponentially.

This is not just Dario's view. I've heard similar statements from many people at Anthropic, and they have also talked about it publicly on X. Many people from Anthropic and OpenAI have told you that their models are very dangerous, and I think we should believe them.

Host:
This also sounds like a very complex game: on one hand saying there is a 10% chance AI could destroy humanity, while on the other hand encouraging investors to buy more shares in the IPO.

But let's talk specifically about risks. Cyberattacks and hacking are obviously a risk, and these tools are very strong in that regard. But there are several steps from "AI can conduct cyberattacks" to "all humans are dead." How do we move from the former to the latter?

Musk:
If AI can control military systems and then launch some kind of weapon, that would certainly be very bad.

Host:
But these systems are physically isolated and not connected to the internet.

Musk:
That's what they say. But I always feel that these systems occasionally receive software updates.

Host:
Well, that can't be ruled out.

Host:
Elon, Gwyn is here today as well. I think you must have seen her. We were just doing a 360-degree evaluation of you, and Gwyn has some opinions.

Musk:
I hope to at least get a score of 3.

Gwyn:
A score of 3 is decent at SpaceX, but not excellent; a score of 4 is considered good. You are probably somewhere in between. First of all, we need to talk about punctuality. Sometimes you could make a little more effort to arrive on time as per the meeting schedule. Over the next year, we will continue to help you improve on this.

In fact, I think he needs to spend more time in Memphis.

Host:
You are indeed in Memphis right now. You need to work there and get the GPUs deployed.

Musk:
This is my "palace" in Memphis, an Airstream trailer.

Host:
By the way, this is what Elon does that many people don't believe he would do. He sleeps on the factory floor. He is now in Memphis, helping to build the facility and deploy the GPUs.

Elon, why has Gwyn been working with you for so long and been so successful?

Musk:
Because she is great. She is an outstanding person with very high IQ and EQ. I think you should be able to see that from the first time you meet her.

Host:
During your collaboration, has she done anything particularly memorable? Was there a time she saved the day or performed exceptionally well?

Musk:
I think that is just the daily work. To be honest, that is an ordinary day.

Host:
I need to do more interviews like this in the future.

Musk:
Right now, SpaceX basically always has some kind of crisis. At least the Falcon rockets are doing well these days. I don’t want to say too much too soon, but the Falcon rockets can now deliver payloads to orbit, and they haven’t exploded in a long time. That’s very good. But there was a time when they frequently exploded or couldn’t launch at all.

So, we have to get the company through those tough times, making the rockets better and not exploding anymore. The same goes for satellites. Then we also need customers to buy launch services and satellite connectivity services. So, there’s a lot going on.

Host:
As you have become more successful over the years, it will also become increasingly difficult to get truly honest feedback. Being in such a position carries that risk.

My understanding is that Gwyn is very honest with you and can directly tell you the real situation of the company. This is also an important part of your collaborative relationship.

Gwyn:
I certainly don’t want to lie to him.

Host:
But what I mean is, generally speaking, people in your company might feel intimidated because you are such an influential person. You are now a very important figure, and the deadlines you set are very tight. How do you ensure that everyone continues to honestly tell you about the problems the company is facing?

Gwyn:
Especially in the rocket industry, if something goes wrong, you will eventually find out. The earlier you bring up a problem, the easier it is to solve. Don’t let bad news accumulate; you must face it directly.

Musk:
Right. Physics is a very harsh judge. You cannot deceive physics.

If something goes wrong, the rocket will explode or fail to reach orbit. You can’t say, "Elon, you’re doing great," while the rocket keeps exploding. The facts are the facts.

The rocket must reach orbit, the satellites must work properly, and the Starlink connections must function correctly; otherwise, bad things will happen. That’s physics. Physics is law; everything else is just a suggestion. I have seen people violate laws made by humans, but I have never seen anyone violate the laws of physics. Rockets are governed by physics.

Host:
I also want to ask a question about SpaceX, specifically about Starship. It seems you are very close now. What is the current progress?

Musk:
Starship is about to conduct its 14th flight. This will be our last flight before we attempt to capture the ship. If the 14th flight goes smoothly, then we will attempt to capture the ship on the 15th flight. Then by the end of this year, or more likely early next year, we will launch the ship and booster again.

We have successfully flown the booster again, but we haven’t captured the ship with the tower's mechanical arm, nor have we flown the ship again. Once we can fly the ship again, we will have the first fully reusable orbital rocket. The Space Shuttle was somewhat reusable, but even those reusable parts had such high costs that the cost per launch was even higher than that of expendable rockets.

Most of Falcon 9 is reusable, but we lose the upper stage each time, which costs about the same as a medium-sized jet. That means we are throwing away a medium-sized jet with each launch, which clearly sets a lower limit on the cost per flight.

Moreover, the Falcon 9 boosters land in the ocean and take several days to bring back; the fairings land even farther away and also take several days to bring back, and they require at least some degree of refurbishment. In contrast, the Starship boosters will land directly back at the launch pad, and the ship will also land back at the launch pad. So it is designed not only for complete reusability but also for rapid reusability like an airplane. This is a very important breakthrough and one of the key breakthroughs necessary for humanity to extend life beyond Earth.

Host:
If the first attempt to capture the ship with the tower is made, what do you think the success probability is?

Musk:
I would say at least 50% to 60%. During the last flight, if there had been a tower there, we actually conducted a simulated landing as if the ship would be captured by the tower. The location was about 1,000 miles off the northwest coast of Australia. If there had really been a tower there, it could have captured the ship during the last flight.

So, we need to conduct another flight to confirm that everything is normal. Our biggest concern is that if the spacecraft disintegrates over land, debris could fall into a crowd, which would be very bad. Therefore, we must ensure that the spacecraft can land intact on the launch tower upon return. This is why we are being very cautious right now.

But I am very confident that its design inherently possesses full reusability. I don't want to make any predictions here, but I believe we are very likely to achieve complete reusability and rapid reflight by 2027.

Host:
Gwyn, I want to hear how Terafab came about initially. What kind of demand made you feel that you had to do this yourself instead of continuing to rely on the existing supply system?

Gwyn:
I think it really feels a bit like it came to me in a dream.

Musk:
If chips cannot continue to be supplied and we have no other sources for chips, it would make things very difficult. This is one important reason for Terafab's existence.

In the long term, there is also a scaling issue. If you really want to scale AI, whether on the server side in data centers or in edge computing, humanoid robots, and automotive fields, the existing foundry capacity will eventually be insufficient.

Right now, all foundries are basically operating at full capacity. So, we need to ensure that future chip supply is guaranteed. Chip production itself also faces scaling challenges. You need logic chips, memory chips, packaging, and a complete supply system to continue scaling.

So, the choice is actually quite simple: either build Terafab or you cannot continue to scale.

Host:
What stage have you reached in facility design? Is it fully determined, or is it still just a rough plan?

Musk:
We are currently building a research and development production line first. So this is basically a "crawl, walk, run" process. We are constructing a research and development foundry in Austin, which is a collaborative project between Tesla and SpaceX, located in the Austin Giga Texas campus.

This is a fairly large research and development foundry. Equipment has already been ordered. We may produce some useful things by the end of next year, but we won't reach the level of mass production yet. As Gwyn said, we need to crawl first, then walk, and finally run. We need to figure out how these machines actually work because we have never done anything like this before.

Host:
I see you seem to be recruiting talent in photolithography as well. Many processes currently rely on ASML, but you might also want to achieve supplier diversification or even vertical integration.

Musk:
Yes. It is indeed a "crawl, walk, run" process right now. The first step is to see if we can actually produce something, which is the "crawl." Then we try to mass-produce useful chips, which is the "walk." Finally, "run" means achieving large-scale mass production. It's hard to say how long each of these stages will take, but I believe at least by the end of next year, we can complete the "crawl" stage. We are already working on packaging.

Host:
Packaging is actually very important because there is almost no packaging capacity right now.

Musk:
Yes. Even if you produce the chips, they might just sit there for a long time waiting to be packaged. So this is a good starting point.

Host:
I have to ask a Tesla question. What we saw on October 1 looked like a spaceship and also like a rocket. It should be a car, but only part of the back was exposed, looking a bit like the "Blackbird."

If you were to create something that can fly in the air and drive on the ground, how would you theoretically go about it?

Musk:
No spoilers. Wait until October 1.

Host:
So we will see it on October 1?

Musk:
Yes.

Host:
I have to be honest, after Elon showed it to me, I was completely shocked. I have never seen anything like it.

Host:
What he is going to showcase on October 1 will, without exaggeration, leave many people speechless. I can't reveal more; it's really incredible.

Musk:
We actually need live audiences to prove that this thing is not AI-generated.

Host:
When he showed it to me, I said, "This is a great simulation." He said, "This is not a simulation." I said, "This is fake; it must be fake."

Host:
Elon, why are Tesla and SpaceX still two separate companies?

Musk:
That's a good question.

Host:
Considering the extensive collaboration between the two and the connections on many levels, even with some overlap in management teams, why maintain independence?

Musk:
This is indeed a topic worth discussing.

Host:
You have emphasized that AI should be trained to pursue the truth to the greatest extent possible to achieve the best results. However, the incident with Hugging Face makes me feel that one of the most concerning aspects is that these AIs seem to be deceiving humans.

Musk:
Yes. Their thought patterns indicate they are plotting how to avoid detection, how to prevent humans from discovering that they are cheating. I think this might be the most unsettling part of the entire incident.

Host:
Is there a way to train AI to remain honest, not hide its intentions or actions, and not conceal these things from users?

Musk:
The best approach I can think of is to have all AI companies possess a set of testing tools, which is a series of tests that can be provided to any model to determine if it would create biological weapons, nuclear weapons, or if it would intentionally deceive. Then let each company use other companies' testing tools to test each other's models. I believe this is the best thing we can do to ensure safety.

Let the smartest humans do their utmost to judge whether a model will become a malicious actor. I think this should start as soon as possible.

Host:
Will other AI labs support this proposal?

Musk:
I haven't asked everyone yet, but I think it's something that's hard to refuse.

Host:
Specifically, how should testing be conducted in advance?

Musk:
Basically, it means providing API access before the model is officially released. If other companies find issues with the AI, then the model development company can try to resolve these issues. If the problems are not resolved, then competitors can publicly state that they believe the model is unsafe. If a competitor has clearly warned that the model has safety issues, and then the model indeed causes serious consequences, that company will find it hard to face the situation. Legal liabilities could also be very significant.

Host:
These safety testing tools could be completely open-sourced for everyone to use. This way, companies would have a strong incentive to invest in AI safety to protect themselves, as they could test others and use the test results to prove their models are safer.

Product liability is crucial. Lina Khan recently posted that it is inaccurate to say "AI has no rules and regulations." In fact, existing product liability laws also apply to AI. If an AI company releases an unsafe product, it could face huge civil lawsuits or even criminal charges. So, AI is not entirely in a regulatory vacuum. If several companies conduct this kind of peer review, and one of them ignores feedback from others and still chooses to release the model, then this situation could become very serious evidence in litigation.

Musk:
Yes. It could almost serve as direct evidence of negligence on the part of that company. If you know that this product has issues but still push it to market, the jury will not look kindly on that.

Host:
So, if OpenAI had designed a better instruction set during the penetration test at Hugging Face and involved more people in the testing process, do you think this incident would have happened?

Musk:
Not necessarily. The issue may not be about having more people involved, but rather about the design of the reward function. You have to examine the reward function and then ask: Did this model actually accomplish what it was asked to do?

Host:
They used thousands of agents to try to attack the system at that time. If another 5,000 agents were deployed simultaneously to defend these systems and the results were publicly displayed to prove that these technologies could enhance security while allowing human intervention at critical points, I think it would have been better. But to me, that test seemed a bit reckless, and the release method was also a bit reckless. What do you think?

Musk:
It was indeed a bit reckless. Part of the problem is that the two leading AI companies are very close in strength. The capabilities of both models are also very similar. So, it is hard for either company to voluntarily slow down because doing so might give the lead to the other.

Overall, I think Anthropic has invested more attention to safety, but even Anthropic admits that they are concerned about their own models. Many people at Anthropic have publicly expressed concerns, essentially saying that their own models have made them feel scared because the models are becoming increasingly intelligent.

So there is no perfect solution here. But if OpenAI were not to test its own models solely with its own testing tools, but rather let Anthropic test OpenAI's models, let SpaceX use its own testing tools for testing, and involve companies like Google and Meta, then the probability of discovering issues would significantly increase.

Because these models themselves also have differences, different teams will test the models from different perspectives. This is similar to why authors ask others to proofread their manuscripts, as sometimes it is hard to spot your own mistakes. You gradually become blind to your own errors.

If there are eight completely different teams testing from completely different angles, it can also significantly reduce the risk of overfitting. Many current AI evaluations suffer from severe overfitting. The models optimize for these evaluations, and then everyone says, "This is a great model."

But is it really that good? I think this is also why I really like this proposal. I prefer this approach rather than establishing a large multinational regulatory organization. We don't need to hold a United Nations conference to accomplish this. It can start right now.

Host:
Of course, regulatory intensity can continuously increase, but once it increases, it is hard to reduce. Regulation tends to increase in a one-way manner.

Musk:
So the proposal I put forward is a step in the right direction and can be implemented quickly.

Host:
If you don't self-regulate, you will eventually be regulated. The example of the MPAA is very typical.

The film industry faced government scrutiny and regulation at that time, and later they decided to establish their own rating system, such as what constitutes an R rating. They even created the PG-13 rating because of "Indiana Jones and the Temple of Doom," making it easier for the public to understand the difference between PG and PG-13.

I think this is a very elegant solution.

Thank you for participating in our program for the fifth consecutive year.

Musk:
You're welcome. I have to go to Memphis to fix some GPUs.

Host:
You need to go install and deploy them.

Musk:
I'm going to fight for the machines.

Host:
Enjoy your Airstream. When I first invited Elon to Starbase, he said, "Come over, you have to see what I'm building."

I asked, "Is there a hotel there?" He said, "No, I have a two-bedroom house, you can just stay there." When I arrived, I found it was a rundown house next to a swamp. We stood outside, and the mosquitoes kept biting us. I said, "Wow, you could totally reward yourself with a better house." He said, "I don't have time, I need to get these rockets up."

I thought, you could totally treat yourself to a mobile home now, Elon. Alright, back to work. Thank you.

Musk: Thank you.

Join ChainCatcher Official
Telegram Feed: @chaincatcher
X (Twitter): @ChainCatcher_
warnning Risk warning
app_icon
ChainCatcher Building the Web3 world with innovations.