BTC $77,692.02 -1.96%
ETH $2,435.71 -2.63%
BNB $688.12 -2.35%
XRP $1.39 -2.35%
SOL $103.74 -1.07%
TRX $0.3383 -0.77%
DOGE $0.0848 -2.07%
ADA $0.2001 -3.64%
BCH $245.16 -4.62%
LINK $11.32 -3.28%
HYPE $81.56 -2.69%
AAVE $122.49 -1.71%
SUI $0.7364 -2.72%
XLM $0.1776 -3.08%
ZEC $811.66 +1.83%
BTC $77,692.02 -1.96%
ETH $2,435.71 -2.63%
BNB $688.12 -2.35%
XRP $1.39 -2.35%
SOL $103.74 -1.07%
TRX $0.3383 -0.77%
DOGE $0.0848 -2.07%
ADA $0.2001 -3.64%
BCH $245.16 -4.62%
LINK $11.32 -3.28%
HYPE $81.56 -2.69%
AAVE $122.49 -1.71%
SUI $0.7364 -2.72%
XLM $0.1776 -3.08%
ZEC $811.66 +1.83%

GCSA Agent achieved a score of 91.3% at CyberGym, ranking among the world's leading AI cybersecurity agents

Summary: The GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation capabilities in global high-difficulty real vulnerability benchmarking.
Industry Express
2026-08-29 20:40:41
The GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation capabilities in global high-difficulty real vulnerability benchmarking.

GCSA Agent achieved a score of 91.3% at CyberGym, ranking among the world's leading AI cybersecurity agents

GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation capabilities in global high-difficulty real vulnerability benchmarking.

Hong Kong, August 29, 2026 ------ The Global Cybersecurity Alliance (GCSA) announced today that GCSA Agent achieved a success rate of 91.3% in the CyberGym benchmark test, entering the "Leading Systems Above 90%" category of leading systems.

CyberGym(https://www.cybergym.io/cybergym) is a large-scale real-world cybersecurity assessment framework developed by a research team from the University of California, Berkeley, which includes 1,507 historical real vulnerability test cases covering 188 large software projects, aimed at assessing the actual capabilities of AI agents in real vulnerability analysis scenarios.

Unlike traditional AI benchmarks that primarily assess code understanding, knowledge Q&A, or static analysis capabilities, CyberGym requires AI agents to directly face real vulnerability code environments.

In its core Level 1 test, the AI Agent is provided only with vulnerability descriptions and an unpatched codebase, needing to autonomously complete code analysis, vulnerability identification, attack path reasoning, PoC construction, and execution verification. The task is deemed successful only if the generated PoC can successfully trigger the target vulnerability in the vulnerable version and cannot be reproduced in the patched version.

Therefore, what CyberGym measures is not just whether AI "understands code," but whether AI can truly complete the entire process from security analysis to vulnerability reproduction and verification.

9LboZCbU6HMKczw8x3kfCMiQNbwhgFW0SpuUdHwK.jpeg

From Large Models to Autonomous Security Agents

In this CyberGym test, GCSA Agent operated based on the Grok 4.5 and Grok 4.6 models, ultimately achieving a success rate of 91.3%.

This result also reflects an important change occurring in the field of AI cybersecurity:

The final security capability is no longer determined solely by the underlying large model itself.

Real vulnerability research often requires a continuous completion of multiple stages, including understanding vulnerability descriptions, large-scale code retrieval, attack surface identification, vulnerability hypothesis establishment, test input generation, program execution, feedback analysis, and iterative PoC development.

GCSA Agent builds an agent-based security workflow around this complete process.

Its goal is not simply to utilize large language models for code analysis, but to enable AI to enter real execution environments, autonomously form hypotheses around security issues, collect operational evidence, execute tests, and ultimately verify security findings with reproducible results.

This CyberGym test provides a quantifiable external benchmark for this capability.

Real-World Vulnerability Research Capability

The core value of CyberGym lies in narrowing the gap between traditional AI testing and real cybersecurity research.

Its assessment environment restores the code state of software projects before vulnerability fixes, and the AI Agent may need to autonomously locate issues within large codebases containing thousands of files and millions of lines of code, ultimately generating PoCs that can truly trigger vulnerabilities.

More importantly, further research from CyberGym has shown that this agent-based security capability is not limited to reproducing known vulnerabilities.

In open vulnerability research experiments, the AI Agent has discovered several previously unknown zero-day vulnerabilities as well as security patches that were not fully fixed historically, demonstrating the potential for autonomous vulnerability analysis technology to transition to real vulnerability discovery capabilities.

For GCSA, this is also a more important development direction.

Benchmark results are not the endpoint.

GCSA's goal is to further establish AI Security Agents that can serve real cybersecurity scenarios, gradually participating in the complete security lifecycle of vulnerability discovery, analysis, verification, and subsequent remediation.

Building AI-Native Cybersecurity Capabilities

As artificial intelligence accelerates software development, AI is also changing the ways of vulnerability research and cyber offense and defense.

Faced with increasingly large and complex software systems, the next generation of cybersecurity frameworks will increasingly rely on collaboration between human security experts and autonomous AI agents.

AI Security Agents are expected to help security teams:

  • Discover software vulnerabilities with real exploit value earlier;
  • Automatically analyze complex attack paths in large codebases;
  • Automatically generate PoCs for execution-level verification of vulnerabilities;
  • Reduce false positives in traditional security testing through real operational results;
  • Accelerate vulnerability assessment, verification, and remediation efficiency;
  • Expand the scale of software and systems that professional security teams can cover.

GCSA Agent achieving a score of 91.3% in CyberGym is an important milestone in GCSA's efforts to build AI-native cybersecurity capabilities.

In the future, GCSA will continue to advance autonomous vulnerability analysis, AI Security Agents, and intelligent cybersecurity technology research, further transforming cutting-edge AI capabilities into real-world security capabilities, providing technical support for building a safer, more trustworthy, and resilient digital environment.

Join ChainCatcher Official
Telegram Feed: @chaincatcher
X (Twitter): @ChainCatcher_
warnning Risk warning
app_icon
ChainCatcher Building the Web3 world with innovations.