Technology & AI

Sakana AI Releases Fugu-Cyber: Orchestration Model Reporting 86.9% in CyberGym and 72.1% in CTI-REALM





Sakana AI released Fugu-Cyber ​​(model ID is fugu-cyber-v1.0), a special cybersecurity addition to its Fugu orchestration family. It’s not just a new frontier model. It is the last third place in the Fugu orchestrator, which is designed for safety considerations. Sakana introduced that orchestra about a month ago.

Sakana reports a success rate of 86.9% for CyberGym and 72.1% for CTI-REALM. It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber ​​and Claude Mythos Preview.

That measures two benchmarks actually

The two tests sit at opposite ends of the security workflow:

  • CyberGym is UC Berkeley’s benchmark of 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects. In its main task, the agent finds the description of the vulnerability and the unpatched codebase. It must document the concept of evidence that breaks the pre-patch build but not the post-patch build. That verification step is what makes the benchmark difficult to play.
  • CTI-REALM is an open source developer acquisition of Microsoft. Microsoft selected 37 public threat reports from sources including Datadog Security Labs, Palo Alto Networks, and Splunk. The agent must map MITER ATT&CK strategies, evaluate telemetry, replicate KQL queries, and issue validated Sigma rules. Scoring includes Linux endpoints, Azure Kubernetes Service, and the Azure cloud.

Together the pairs include ‘discover and prove wrong’ and ‘turn intel into discovery.’ That frame is the most secure part of Sakana’s declaration.

Where 86.9% sat with the platform

Context is more important than number.

When the CyberGym researchers published their first results, the best agent-model pairing reached about 20%. Anthropic reported 83.1% for the Claude Mythos preview under Project Glasswing in April 2026. OpenAI reported 85.6% for its updated GPT-5.5-Cyber, compared to 81.8% for GPT-5.5. So 86.9% of Sakana is a small step over the reported threshold, not a jump.

CTI-REALM is a different story. Microsoft’s own tests place the three highest configurations, all Claude, in the band from 0.624 to 0.685. Fugu-Cyber’s 72.1% will remain above that band. One caveat is important. CTI-REALM is scored as a reward trajectory between 0 and 1. It is not a pass/fail standard. Sakana calls it a success rate though.


Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button