Rumor: NEW Claude Opus is Coming! (Marshmallow & Melon Leak)
Two new model IDs showed up in partner tools this week: Claude Marshmallow EAP and Claude Melon EAP. The rumor is these are early access previews of the next Opus and Sonnet, not a Haiku refresh. Internal testing reportedly shows Marshmallow beating Opus 5, so the smart money is on Marshmallow being the big model and Melon the mid-tier. We saw a similar leak last month with Honeycomb, which turned out to be Opus 5 two weeks later. So we could see these models within a few weeks.
What's wrong with Opus 5 right now
Anthropic engineers have quietly admitted Opus 5 has real problems. Boris Cherny, who works at Anthropic, said on social media that Opus is not perfect and fixing it is a big priority. The main complaint is verbosity. Opus 5 writes way too much for no reason, gets confusing on long runs, and burns tokens. There is a quick fix, though. You can run claude/config output style equals concise to force shorter responses. That helps, but it should not be necessary on a top-tier model.
Users have been more blunt. They call Opus 5 a verbose, assumption-laden token monster that expands scope, over-verifies its own work, and feels sloppier than Opus 4.6 for real debugging. Many people think Opus 4.6 back in February and March was the peak. The current Opus 5 is bench-maxed, scores well on agent arena, but breaks old workflows and demands a whole new prompt style to be usable.
Anthropic's spin is that Opus 5 is more agentic and less prompt injectable, which is true, but that's not what most users want. They want a model that follows instructions, not one that decides to go 180 degrees on its own. Fable 5, Anthropic's planning model, routes coding tasks to Opus anyway, so we've basically been using Sonnet the whole time.
Sonnet 5's hidden costs
Sonnet 5 has its own money problem. It looks like a mid-tier price, but reviews show it actually costs more per task than Sonnet 4.6. That's because it breaks jobs into more sub-agents and takes more steps. Sonnet 5 is smarter on paper but slower, fuzzier, and more expensive in practice. It also refuses more legitimate but sensitive tasks, getting stuck in loops of arguing instead of executing. Long chats lose context or misprioritize earlier details. So we end up with a smart Sonnet for Opus and a dumb Sonnet for Sonnet.
What Marshmallow and Melon need to fix
If Marshmallow is the next Opus, it needs to keep the deep reasoning and security gains while taming the over-agentic behavior. Less scope expansion, fewer self-initiated review marathons, better adherence to what you actually asked. If Melon is the next Sonnet, it needs to keep the reasoning boost but cut the hidden costs. Smarter defaults around when to spawn sub-agents, lower refusal rates, more predictable context in long sessions. Basically, the next-gen Claude should not force every user to rewrite their entire setup just to survive an upgrade.
Right now, 80 to 90% of our work runs on GPT-5.6 Sol and Terra. We use Fable 5 for really hard problem-solving, Grok 4.6 for speed, and Kimi K3 or DeepSeek V4 Pro as budget alternatives. If the new Opus and Sonnet are actually better than Sol and Terra in day-to-day use, we'll rotate back to Claude. But Anthropic has to ship a fix, not just a new model name.
