Sakana Fugu ULTRA Review (Better than Fable 5?!)

Published
Jun 22, 2026
Duration
9:50
Click to load the YouTube player

Useful orchestration, punishing quota burn

  • Fugu Ultra is an orchestrator, not a new base model. It coordinates models such as GPT and Claude, then combines their work. The Fable 5 comparison is therefore about the quality of the assembled result, not one new “brain.” (source video zaLAonePx38, 01:09; source video zaLAonePx38, 01:58)
  • Sakana’s launch-day benchmarks put Ultra around Fable 5 and Mythos on hard coding and reasoning tasks, but Ron says those results were still mostly self-reported. This video does not independently reproduce the benchmark. (source video zaLAonePx38, 01:27; source video zaLAonePx38, 02:03)
  • Cost dominated Ron’s June 22, 2026 test: one Fugu Ultra prompt used 29% of his weekly allowance on the $20 monthly standard plan, and the directory-review run consumed 84% of a five-hour usage limit. These are plan-meter observations from that test, not current pricing or cost-per-token data. (source video zaLAonePx38, 00:31; source video zaLAonePx38, 03:14)
  • Ron judged xhigh excessive for a simple directory review. His practical split was high for reviews, system checks, questions, and feature drafts; xhigh for difficult refactors or long-horizon tasks. (source video zaLAonePx38, 05:20; source video zaLAonePx38, 05:40)
  • The tested value is convenience: automatic multi-model routing inside one task. If you can already route strong models manually, measure whether that convenience is worth Ultra’s higher latency and price. (source video zaLAonePx38, 02:40; source video zaLAonePx38, 04:19; source video zaLAonePx38, 04:29)

Fugu Ultra is smart, but “better than Fable 5” is the wrong buying question. Fable 5 is treated here as one premium model; Fugu Ultra is a learned multi-agent orchestrator, a control layer that routes parts of a task across other models and combines the result. Ron’s simple project review used xhigh, consumed most of the five-hour allowance shown in his test, and led him to call that effort choice a mistake. Start with a small, capped experiment, use high by default, and pay for xhigh only when the task is difficult enough that automatic routing can justify the cost. (source video zaLAonePx38, 00:47; source video zaLAonePx38, 03:14; source video zaLAonePx38, 05:40; source video zaLAonePx38, 08:40)

Watch the test

Ron in his own words

“Sakana Fugu is not a single frontier LLM.” — Ron, source video zaLAonePx38, 01:09

“So it’s really about convenience to be honest.” — Ron, source video zaLAonePx38, 04:29

“think of Fugu Ultra as the orchestrator helping route all of those other models.” — Ron, source video zaLAonePx38, 08:40

“Figure out a way to manage your cost first before building out your project.” — Ron, source video zaLAonePx38, 09:23

Fugu coordinates models

The default Fugu tier is described as balancing performance and latency for everyday coding, code review, and chat-style work. It also allows specific providers or models to be excluded from the pool for data privacy or compliance constraints. Ultra uses a deeper mixed-agent pool and more aggressive orchestration; Ron describes it as perhaps one to three agents per task, trading higher latency and a much higher per-token price for quality on complex work. (source video zaLAonePx38, 02:14; source video zaLAonePx38, 02:30)

That architecture explains the headline without endorsing it. Ron says the output shown in Codex was a compilation of GPT and Claude models examining the same task and reaching a conclusion. Later, he could see which models had been called for different parts of the job. Calling that “Fable 5 intelligence” is his shorthand for the combined result, not evidence that Sakana trained a standalone model equal to Fable 5. (source video zaLAonePx38, 02:55; source video zaLAonePx38, 08:29; source video zaLAonePx38, 08:52)

A routing decision table

The table organizes Ron’s launch-day observations by workload. It does not add evidence from a later Fugu test.

Your jobRoute suggested by the videoWhyEvidence boundary
Review a directory, check a system, ask questions, or draft a featureFugu Ultra on highRon says high still suits deep reasoning and would have been the better choice for his simple review. (source video zaLAonePx38, 05:31; source video zaLAonePx38, 05:55)The video does not run the same task on high, so it does not show the savings or result quality.
Difficult refactor or genuinely long-horizon taskConsider xhighRon reserves maximum reasoning effort for work difficult enough to need it. (source video zaLAonePx38, 05:36; source video zaLAonePx38, 05:49)No matched high versus xhigh test is included.
Routine work where you already know which model should do each stepRoute manuallyRon questions paying for full automation when strong models are already available and can be switched manually. (source video zaLAonePx38, 04:18)The video does not calculate the cost or operator time of manual routing.
Experimenting without a known budget profileUse a small capped trialRon recommends pay-as-you-go with $5 or less rather than immediately relying on the $20 subscription. (source video zaLAonePx38, 03:55; source video zaLAonePx38, 04:07)This was Ron’s June 22 recommendation; current plans and limits were not checked for this companion.

The project review found the tradeoff

Ron tested Fugu Ultra on his CBR TradingView-indicator repository. It flagged a strong indicator and blueprint layer alongside an unreliable backtesting and research layer. Ron explains why that diagnosis mattered: the project needed a real backtesting engine and paid access to order-flow history for his intended methodology, while he was trying to proceed with free inputs; he describes the resulting mismatch as specification drift. (source video zaLAonePx38, 06:51; source video zaLAonePx38, 07:34; source video zaLAonePx38, 08:14)

That is useful diagnostic evidence, but it is not a controlled Fugu-versus-Fable comparison. Fable 5 had built the indicator and blueprint earlier; several other models had continued the data and research work. The test shows Fugu finding known weaknesses in a mixed-history project. It does not show Fugu building the project, repairing the backtest, or beating Fable 5 on the same prompt. (source video zaLAonePx38, 07:06; source video zaLAonePx38, 07:23)

Freshness note

The video was published June 22, 2026. This companion was reviewed against the source video on July 18, 2026. No current Sakana documentation, pricing page, plan limits, benchmark, API behavior, or later creative-task test was added. The $5 pay-as-you-go suggestion, $20 standard plan, $100 Pro plan, 29% weekly usage, 84% five-hour usage, high and xhigh settings, API details, model pool, and benchmark position above are therefore a dated record of Ron’s June 22 test, not confirmation of Sakana’s state on July 18.

Continue learning