Something shifted this week in how AI requests get routed, and most people haven't noticed yet.
On September 11, Sakana AI — a Tokyo-based research lab founded by ex-Google DeepMind researchers, independent, not Big Tech — shipped two new products: Fugu Max and Fugu Ultra v2. Neither one is a language model. They're orchestrators — systems trained to decide which models in a large pool should handle which part of your task, then stitch the results back together behind a single API call.
The pricing got our attention immediately: $2 per million input tokens, $6 per million output tokens for Fugu Max. Sakana says that's 40-60% cheaper on output than Sonnet 5, GPT 5.6 Terra, and Kimi K3. And it's OpenAI-compatible — change a base URL and an API key, and your existing code runs.
That's a meaningful combination.
What an orchestrator actually does
If you've been routing all your AI traffic through one model — sending every request to the same frontier API — you've been paying top-dollar for tasks that don't need a top-dollar model.
Summarize a meeting transcript: probably doesn't need Opus 5. Extract structured data from a form submission: definitely doesn't. Write a thorough first draft of a grant proposal: probably does. Reply to a routine vendor email: absolutely doesn't.
An orchestrator knows this. You send one request. The orchestrator figures out which model in its pool is the cheapest one that can actually do the job well. It routes there, collects the output, synthesizes if needed, and returns a result. You pay less because you stopped using a fighter jet to drive to the grocery store.
Fugu does this at the API level, using a model that was trained specifically to coordinate other models — not a handwritten routing script, but something learned from actual multi-model task performance data. Two research papers out of ICLR 2026 (TRINITY and Conductor) describe the underlying architecture. The practical result is a system that assigns subtasks to whichever model in Sakana's pool is the lightest fit for that step.
From your application's perspective: one endpoint, one API key, one request format. Behind it, Sakana's orchestrator is making real-time routing decisions across a pool of open-weight and specialized models.
The benchmark numbers — and a caveat worth reading
Sakana reports Fugu Max at the top of the cost-performance frontier across six benchmarks including Terminal Bench 2.1, GPQAD, and AutomationBench. Fugu Ultra v2 claims 48.3 on Chartography (a visual reasoning benchmark) versus 27.3 for Opus 5.
These numbers have not been independently verified yet. The release post is transparent about that, and Sakana published its benchmarking methodology — community researchers are actively running their own checks.
Self-reported benchmarks are always worth treating with some skepticism until third-party validation lands. But the pricing is public and independently checkable right now. And the architecture is documented.
What this actually means for a 10-30 person team
If your org is running AI tooling across a team — Slack summaries, document drafting, data extraction, email processing — and you're paying API costs directly, Fugu Max is worth a week of testing. The switching cost is low (swap a URL and key). The potential savings are real if the quality holds in your workloads.
If you're building internal tools on top of an AI API, the orchestration question matters even more. Right now, most small teams either pick one model and commit to it everywhere, or build custom routing logic themselves — which is expensive to build and painful to maintain when model pricing changes. Fugu handles routing automatically. Your logic stops being "which model do we use for this" and starts being "what does the task need."
Sakana's API is OpenAI-compatible (Chat Completions, Responses, Models endpoints) plus an Anthropic-compatible Messages API. Subscription tiers run $20/month, $100/month, and $200/month, with pay-as-you-go available. For a team currently burning several hundred dollars a month in API costs, the math is worth doing.
The broader shift this signals
Fugu isn't a one-off. The AI pricing war has been fought at the model level for two years — everyone racing to ship the cheapest capable model. What's happening now is the battleground moving to the orchestration layer.
Independent labs are shipping systems that sit between you and the big API providers. They route intelligently, charge for the optimization, and hand you a lower bill. You trade some transparency into what's running under the hood for simpler integration and reduced cost.
For small orgs, the direction this moves is mostly good. AI tooling economics keep shifting your way. But there's a new vendor dependency to think about: you're now trusting Sakana to maintain a reliable model pool, negotiate continued access to the underlying models, and stay in business. Sakana is a credible team, but they're a small independent lab competing with companies that have considerably more capital. Don't build critical workflows on any new vendor without a fallback.
What to do right now
Test it on your highest-volume workflow. Pick the one generating the most API cost. Run it through Fugu Max for a week. Compare output quality and the actual bill.
Don't migrate everything at once. New orchestration layers can behave unexpectedly on edge cases. Test one workflow before you trust the second.
Keep your direct API keys active. If Sakana has an outage or reprices, you need to fall back quickly. Design your stack so the orchestration layer is swappable.
Watch the independent benchmarks when they land. The community is working through Sakana's methodology now. In a few weeks there should be third-party numbers. Those matter.
This is the kind of thing that seems obvious in retrospect — of course you'd want a layer that routes requests to the right model automatically. Of course the market would build it. It's here now, it's cheap enough to try, and the switching cost is an afternoon.
We've been testing orchestration layers with client teams throughout the year. Finding one stable enough to build on has been the consistent challenge. If you want help evaluating this against your actual stack and workloads, that's a conversation we're happy to have.