Anthropic's Mythos 5, Fable 5, OpenAI's GPT-5.6, and other advanced artificial intelligence (AI) models are starting to encounter regulatory measures from the U.S. government, highlighting AI "orchestration" approaches that combine several AI systems instead of using standalone versions. This change arises due to increasing worries over minimizing reliance on one specific AI provider.
Among these, the most notable is the AI orchestration model “Fugu” released 10 days after the U.S. government’s regulation of Anthropic by Japan’s unicorn (unlisted startup valued at over 1 billion dollars) “Sakana AI.” In Japanese, “sakana” means fish, and “fugu” refers to pufferfish. The company’s name reflects its vision of creating AI where smaller models collaborate and evolve like a school of fish in nature, rather than building a single massive AI model like Big Tech. The pufferfish metaphor symbolizes how small-scale models can expand their capabilities and leverage “toxins” (strengths) to survive threats, aiming to efficiently evolve compact models into powerful performance.
Founded in 2023 by Llion Jones, who was one of the authors of the influential AI research paper "Attention Is All You Need," along with David Ha, a previous researcher at Google Brain, Sakana AI reached a value of $2.7 billion (about 4.15 trillion South Korean won) in November of the prior year. Fugu, created by Sakana AI, functions like an "AI maestro" that manages and combines several artificial intelligence systems. Within the system, a smaller model containing roughly seven billion parameters works as the central command hub. Unlike current coordination models that choose just one AI to address a user's request, Fugu breaks down tasks into smaller components and utilizes various AI agents concurrently—adopting a strategy involving multiple agents working together. Several AI technologies operate side-by-side on a single task, splitting responsibilities and merging their contributions. According to Sakana AI, Fugu offers cutting-edge performance without concerns related to export restrictions.
Its performance stands out. As per Sakana AI, the leading model "Fugu Ultra" obtained a score of 93.2 on LiveCodeBench, an automated and competitive programming assessment, exceeding Anthropic’s Fable 5 (89.8). Additionally, it attained 95.5 on GPQA Diamond, a high-level scientific reasoning test, surpassing Mythos Preview (94.6). Nevertheless, on the extremely challenging academic examination known as "Human Last Exam (HLE)," it earned 50.0, falling short of Fable 5's 53.3. In the real-world coding benchmark SWEBench Pro, it received 73.7, which is lower than Fable 5's 86.0.
Fugu exploits the weakness of depending solely on one AI vendor, highlighting its ability to seamlessly redirect workloads to different models inside its network if a particular supplier encounters legal restrictions. This approach has resulted in Fugu being referred to as a "practical model for national AI." With nations facing challenges in building their own artificial intelligence systems, coordination provides an effective method to sustain efficiency by shifting between models when access is limited.
Opponents claim that Fugu functions as an "effective step down," offering just mid-level open-source models without including high-end controlled artificial intelligence systems. Assigning different jobs and managing several AIs simultaneously can raise token expenses and reduce efficiency when contrasted with using individual models.
However, the technology sector perceives Fugu's rise as important. It indicates a change in the AI race from focusing solely on power to emphasizing affordability and the capability to "execute, partition, validate, and integrate" functions. Although not an ideal answer, Fugu-like coordination is regarded as a strategy for reducing risks amid unexpected challenges such as trade restrictions, particularly as rules governing cutting-edge AI become stricter.