Skip to main content
Fugu Ultra v2.0 is Sakana AI’s model for work where quality, depth, and continuity matter more than a fast response. It appears as a single model, but coordinates different AI specialists to research, build, review, and verify parts of the same problem. In Tess, this approach is particularly useful when a task involves multiple sources, tools, and decisions: reproducing a scientific paper, investigating an entire repository, comparing patents, or conducting a technical analysis that must remain consistent across many steps.

A multi-agent system with a model experience

Instead of relying on one model to perform all the work, Fugu Ultra can activate one to three specialists and coordinate how they collaborate. One may explore possible paths, another may execute a technical step, and a third may verify the result. The composition changes with the task. For the user, however, the experience remains simple: select Fugu Ultra and describe the objective. Specialist selection, work distribution, and information exchange happen behind the scenes.
Fugu Ultra’s specialist pool is fixed and its internal routing is not exposed. This preserves the experience of a single model, but it means individual participating models cannot be selected or removed.

What version 2.0 prioritizes

Version 2.0 was designed to maximize quality in difficult, high-impact tasks. Its main strengths appear in:
  • multi-step reasoning: maintains hypotheses, decisions, and checks throughout extended work;
  • autonomous research: consults sources, connects evidence, and investigates points that require further analysis;
  • software engineering: navigates repositories, implements changes, runs tests, and reviews its own result;
  • visual and structured data: interprets charts, documents, images, and information organized in different formats;
  • consistency: remains among the strongest models across different categories of agentic tasks, not only on a single benchmark.
The model accepts text and image input, offers up to 1 million tokens of context, and supports high reasoning levels. In the provider implementation, high balances performance and response time, while xhigh deepens the analysis of complex problems.

What the benchmarks show

In results published by Sakana AI, Fugu Ultra v2.0 achieved the best or joint-best score on five of eight benchmarks and ranked in the top two on seven of eight. On Chartography, Sakana reports 48.3 points for Fugu Ultra v2.0, compared with 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5. On DeepSWE, the model reaches 74.3 points.
The most relevant signal is not winning one isolated test. Fugu Ultra v2.0 combines strong results in software, documents, tools, and visual interpretation—the same mix found in complex end-to-end work.
These benchmarks were published by Sakana AI and depend on the harness, tools, and parameters used. According to the provider, Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra are not part of the Fugu Ultra v2.0 pool. Validate the model in your own workflow before replacing a production configuration.

Where it creates the most value

Research that must become a decision

The work starts with papers, patents, reports, and open questions. Fugu Ultra can distribute the investigation among specialists, challenge conclusions, and deliver an analysis that connects sources, exposes gaps, and recommends the next step. Expected result: a decision document with traceable evidence, disagreements, and risks—not merely a summary of the sources.

A software change that must work

In a large repository, finding the right file is only the beginning. The model can map the flow, locate dependencies, implement the change, run tests, and review the diff before concluding. Expected result: changed code, validated behavior, executed tests, and documented remaining risks.

Analysis where mistakes are costly

In audits, security, and technical assessments, a first answer is not enough. Fugu Ultra can explore competing hypotheses, look for contradictory evidence, and add a verification step before consolidating the result. Expected result: evidence-backed conclusions, a confidence level, and points that still require human validation.

How to apply it in Tess

1

Choose work that justifies depth

Prefer tasks with multiple sources, chained decisions, or a need for validation. For simple rewrites and quick answers, a lighter model is usually more efficient.
2

Provide context and tool access

Include files, repositories, internal criteria, and the necessary tools. Do not treat 1M context as a target: provide only material relevant to execution.
3

Define what finished means

Specify tests that must pass, required sources, delivery format, and conditions that must be verified.
4

Adjust the depth

Start with High. Use XHigh when the task requires deeper investigation, comparison of hypotheses, or execution and verification cycles.
5

Keep control over critical actions

Require approval before publishing, deleting, sending, changing production systems, or taking any irreversible action.

Three playbooks to get started

1. Reproduce technical research

Input: paper, available data, execution environment, and expected result.
Tools: search, files, code, and terminal.
Done when: the method can be repeated, the results have been compared, and differences are explained.

2. Full-stack change with validation

Input: repository, current behavior, desired result, and project standards.
Tools: code, terminal, tests, and technical documentation.
Done when: the flow works end to end and the relevant validations pass.

3. Investigation with conflicting sources

Input: documents, spreadsheets, images, and decision criteria.
Tools: file search, web, and document generation.
Done when: every conclusion is traceable and relevant uncertainties are explicit.

Limitations and best practices

  • Reserve Fugu Ultra for tasks where quality justifies higher latency and cost.
  • The model may coordinate several internal calls; deep work tends to consume more tokens than a direct response.
  • You cannot see which specialists were used for each request.
  • Results in research, security, finance, and critical decisions require human review.
  • Set clear tool boundaries and require confirmation for irreversible actions.
  • Compare results with lighter models before standardizing Fugu Ultra for everyday tasks.
Captura De Tela 2026 09 17 Às 16 45 53
See also: Models and Costs · Sakana Fugu · Fugu Ultra v2.0 — technical announcement.