Close Menu
    Facebook X (Twitter) Instagram
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Facebook X (Twitter) Instagram
    Crypto Currencey
    • Home
    • About
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • Exchanges
      • Centralized (CEX)
      • Decentralized (DEX)
    • Tax Software
    • Wallets
      • Hardware
      • Software
    • Trading Bots
      • Cloud Based
      • Advanced
    • Tools
      • Crypto Market Cap List
      • Market Heatmap
    Crypto Currencey
    Home»AI News»Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM
    Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM
    AI News

    Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM

    July 26, 20264 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email
    changelly


    Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family. It is not just a new frontier model. It is a third endpoint on the Fugu orchestrator, tuned for security reasoning. Sakana launched that orchestrator a month earlier.

    Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM. It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Preview.

    What the two benchmarks actually measure

    The two evaluations sit at opposite ends of a security workflow:

    • CyberGym is a UC Berkeley benchmark of 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects. In its main task, an agent receives a vulnerability description and an unpatched codebase. It must write a proof-of-concept that crashes the pre-patch build but not the post-patch build. That verification step is what makes the benchmark hard to game.
    • CTI-REALM is Microsoft’s open-source detection-engineering benchmark. Microsoft curated 37 public threat reports from sources including Datadog Security Labs, Palo Alto Networks, and Splunk. An agent must map MITRE ATT&CK techniques, explore telemetry, iterate on KQL queries, and emit validated Sigma rules. Scoring covers Linux endpoints, Azure Kubernetes Service, and Azure cloud.

    Together the pair spans ‘find and prove the bug’ and ‘turn intel into a detection.’ That framing is the most defensible part of Sakana’s announcement.

    aistudios

    Where 86.9% sits against the field

    Context matters more than the number.

    When the CyberGym researchers published their first results, the best agent-model pairing reached roughly 20%. Anthropic reported 83.1% for Claude Mythos Preview under Project Glasswing in April 2026. OpenAI reported 85.6% for its updated GPT-5.5-Cyber, against 81.8% for GPT-5.5. Sakana’s 86.9% is therefore a small step past the reported frontier, not a jump.

    CTI-REALM is a different story. Microsoft’s own evaluation put the top three configurations, all Claude, in a band from 0.624 to 0.685. Fugu-Cyber’s 72.1% would sit above that band. One caveat matters. CTI-REALM is scored as a trajectory reward between 0 and 1. It is not a pass/fail rate. Sakana calls it a success rate anyway.

    How the orchestration works

    Fugu is itself a language model. It is trained to read a query and build an agentic scaffold on the fly. It then delegates sub-tasks to specialist models in a pool.

    The approach is documented in the Fugu technical report and two ICLR 2026 papers, TRINITY and the Conductor. TRINITY assigns Thinker, Worker, and Verifier roles across multiple LLMs. The Conductor learns natural-language coordination strategies through reinforcement learning.

    For security work, Sakana research team argues the verifier role is the point. A candidate vulnerability surfaced by one agent gets validated by security-specialized sub-agents before any patch is proposed. Routing remains proprietary, so you cannot see which model handled which step.

    Access, policy, and price

    Fugu-Cyber is gated on four dimensions.

    Access requires an application form stating the intended use case and verified contact details. Sakana team reviews each one manually. The model ships under an updated Acceptable Usage Policy that prohibits offensive misuse. Billing is restricted to the Token Plan. The $20, $100, and $200 subscription tiers cover Fugu and Fugu-Ultra only. And the Fugu API is not offered in the EU or EEA while Sakana works toward GDPR compliance.

    Pricing is fixed at $6 per million input tokens, $36 output, and $0.60 cached input. All three rates double above a 272K-token context. Every line is exactly 1.2× the Fugu-Ultra rate, a flat 20% premium for the cyber endpoint. Long codebase runs cross 272K easily, so the doubled tier is not an edge case.

    Key Takeaways

    • Fugu-Cyber is an orchestration endpoint, not a new frontier model, launched July 21, 2026.
    • Sakana reports 86.9% on CyberGym and 72.1% on CTI-REALM, both self-reported and un-replicated.
    • Those scores edge past GPT-5.5-Cyber’s 85.6% and Claude Mythos Preview’s 83.1% on CyberGym.
    • Access is gated: manual approval, defensive-use AUP, Token Plan only, no EU/EEA, no weights.
    • Sakana’s own position is that a capable API along with human security expertise beats the API alone.

    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



    Source link

    coinbase
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    CryptoExpert
    • Website

    Related Posts

    Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

    July 27, 2026

    Working to automate nuclear plant operations | MIT News

    July 25, 2026

    Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

    July 24, 2026

    SenseTime’s Galaxy Project targets domestic AI chip scale-up

    July 23, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    kraken
    Latest Posts

    The brutal $346M math behind Galaxy’s high-stakes race to build CoreWeave’s Texas AI mega-center

    July 26, 2026

    Ethereum ETFs End 5-Day Inflow Streak With $70.6M Outflows

    July 26, 2026

    Is It Really Safe to Invest in the Vanguard S&P 500 ETF Right Now? Here’s What History Says.

    July 26, 2026

    5 Stocks I’m Buying HEAVY Right Now August 2026

    July 26, 2026

    Quantum Roadmap Would Push Bitcoin Much Higher: Charles Edwards

    July 26, 2026
    coinbase
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights

    Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

    July 27, 2026

    🔥 ISRO FREE AI/ML Course 2026 | Free Certificate + Complete Guide

    July 27, 2026
    livechat
    Facebook X (Twitter) Instagram Pinterest
    © 2026 CryptoCurrencey.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.