A multi-agent system with a model experience
Instead of relying on one model to perform all the work, Fugu Ultra can activate one to three specialists and coordinate how they collaborate. One may explore possible paths, another may execute a technical step, and a third may verify the result. The composition changes with the task. For the user, however, the experience remains simple: select Fugu Ultra and describe the objective. Specialist selection, work distribution, and information exchange happen behind the scenes.Fugu Ultra’s specialist pool is fixed and its internal routing is not exposed. This preserves the experience of a single model, but it means individual participating models cannot be selected or removed.
What version 2.0 prioritizes
Version 2.0 was designed to maximize quality in difficult, high-impact tasks. Its main strengths appear in:- multi-step reasoning: maintains hypotheses, decisions, and checks throughout extended work;
- autonomous research: consults sources, connects evidence, and investigates points that require further analysis;
- software engineering: navigates repositories, implements changes, runs tests, and reviews its own result;
- visual and structured data: interprets charts, documents, images, and information organized in different formats;
- consistency: remains among the strongest models across different categories of agentic tasks, not only on a single benchmark.
The model accepts text and image input, offers up to 1 million tokens of context, and supports high reasoning levels. In the provider implementation,
high balances performance and response time, while xhigh deepens the analysis of complex problems.What the benchmarks show
In results published by Sakana AI, Fugu Ultra v2.0 achieved the best or joint-best score on five of eight benchmarks and ranked in the top two on seven of eight.
On Chartography, Sakana reports 48.3 points for Fugu Ultra v2.0, compared with 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5. On DeepSWE, the model reaches 74.3 points.
These benchmarks were published by Sakana AI and depend on the harness, tools, and parameters used. According to the provider, Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra are not part of the Fugu Ultra v2.0 pool. Validate the model in your own workflow before replacing a production configuration.
Where it creates the most value
Research that must become a decision
The work starts with papers, patents, reports, and open questions. Fugu Ultra can distribute the investigation among specialists, challenge conclusions, and deliver an analysis that connects sources, exposes gaps, and recommends the next step. Expected result: a decision document with traceable evidence, disagreements, and risks—not merely a summary of the sources.A software change that must work
In a large repository, finding the right file is only the beginning. The model can map the flow, locate dependencies, implement the change, run tests, and review the diff before concluding. Expected result: changed code, validated behavior, executed tests, and documented remaining risks.Analysis where mistakes are costly
In audits, security, and technical assessments, a first answer is not enough. Fugu Ultra can explore competing hypotheses, look for contradictory evidence, and add a verification step before consolidating the result. Expected result: evidence-backed conclusions, a confidence level, and points that still require human validation.How to apply it in Tess
1
Choose work that justifies depth
Prefer tasks with multiple sources, chained decisions, or a need for validation. For simple rewrites and quick answers, a lighter model is usually more efficient.
2
Provide context and tool access
Include files, repositories, internal criteria, and the necessary tools. Do not treat 1M context as a target: provide only material relevant to execution.
3
Define what finished means
Specify tests that must pass, required sources, delivery format, and conditions that must be verified.
4
Adjust the depth
Start with High. Use XHigh when the task requires deeper investigation, comparison of hypotheses, or execution and verification cycles.
5
Keep control over critical actions
Require approval before publishing, deleting, sending, changing production systems, or taking any irreversible action.
Three playbooks to get started
1. Reproduce technical research
Input: paper, available data, execution environment, and expected result.Tools: search, files, code, and terminal.
2. Full-stack change with validation
Input: repository, current behavior, desired result, and project standards.Tools: code, terminal, tests, and technical documentation.
3. Investigation with conflicting sources
Input: documents, spreadsheets, images, and decision criteria.Tools: file search, web, and document generation.
Limitations and best practices
- Reserve Fugu Ultra for tasks where quality justifies higher latency and cost.
- The model may coordinate several internal calls; deep work tends to consume more tokens than a direct response.
- You cannot see which specialists were used for each request.
- Results in research, security, finance, and critical decisions require human review.
- Set clear tool boundaries and require confirmation for irreversible actions.
- Compare results with lighter models before standardizing Fugu Ultra for everyday tasks.


