Hermes Agent Just Proved It's a SERIOUS Coding Harness
## The refactor that changed my mind News Research, the team behind Hermes, used Hermes to refactor its own codebase. That's a big deal. I usually reach for Hermes for easy tasks like research. If I need real coding, I use Codex or Claude Code. But this refactor makes me reconsider.
The repository had grown to over a million lines of Python, not counting tests. One file, gateway/run.py, was nearly 35,000 lines. That's a lot of surface for bugs. Tech Nium, the founder, wanted smaller files and less duplication. On September 2nd, he put a regular Hermes agent on the job. The goal: cut at least 30% of the code while keeping behavior the same.
## How the run worked Hermes split the repo into 36 work areas and dispatched 1,393 sub agents across about 19 hours. At peak, 218 ran at once. The main agent only coordinated. It read reports, merged branches, and ran checks. That division of labor is what catches my attention. Getting a patch from one agent is ordinary. Managing an army of them on a codebase this size is a different game.
Claude Opus 5.1 was the model underneath, so it deserves some credit. But Hermes was the harness moving the pieces. The result was concrete. That 35,000 line file shrank to about 5,500 lines. Non-test Python dropped 34.4%. Functions longer than 300 lines fell from 192 to just two. On paper, that's a clean win.
News also simulated symbol lookups. Average tokens returned dropped from around 2,200 to under 1,000. That means an agent can find and read a definition faster. The median lookup actually got slightly more expensive, so it's not a perfect improvement. But the cleanup gave a real benefit beyond fewer lines.
## The catch It wasn't fully automatic. Reviewers found public names workers had removed because nothing inside the repo called them. But external plugins could still rely on those names. Another rewrite changed exception handling at about 65 sites, and existing tests missed the regressions. Human review caught that. The run also needed a second Hermes session when a token expired. The same way I handle Codex: save a handoff file, boot a new session, and pick up the work.
There's also the bill. The main run cost around $19,300 in model time. With follow-ups, about $25,000. Human review adds more. For my own projects, I wouldn't drop that on day one. I'd run a contained refactor, check the diff, and ask an actual engineer before going all in. Refactoring without experience is dangerous. Our boss drilled that into us.
The refactor merged on September 4th after two rounds of community review. More fixes followed. So this isn't "agent writes perfect code." It's "agent writes large patches that humans still have to check." That's a fair trade for substantial work. I still use Codex for most coding. But after this, I'm giving Hermes a real coding task I'd normally send straight to Codex. If you've been saving Hermes for the easy stuff, this is a reason to try it on one hard thing. Tech Nium said he doesn't use Codex or Claude Code, just Hermes to fix Hermes. That's a committed founder, and it shows.
