GPT-6 Sol is LAZY
First impression: the visual demos
We tried GPT-6 Soul on visual tasks and it did well. We gave it prompts for interactive scenes and put the results on our Super Bash website. There were interactive scenes, a mechanical watch simulator, a flight around Hogwarts, and a trebuchet you can fire at a gate.
Take the watch for example. You can see the mechanism running, then switch to an exploded view and separate the layers. That control worked in both models we checked. Both trebuchet demos also fired and reached the gate impact report. There was enough interaction there to make these useful starting points for further development.
Getting there sometimes took corrections. Soul's glider initially had a camera rolled about 90 degrees. When we reviewed it, we caught the problem, the code was changed, and a later check confirmed the correction. I can work with that process. I expect to review an agent's output anyway, and I want the feedback to lead to a better result.
Where Soul fell short: agentic article rewrites
That review-and-improve process was also what I needed from Soul on Dog HK. This is another website we're working on for dog owners in Hong Kong. It has dog-friendly places, grooming services, and veterinary listings alongside practical guides in both English and traditional Chinese.
A guide on that site might help someone work out how to use the pet-friendly bus or plan a visit to the Mega Sky rooftop garden. A grooming article should help them compare options and decide what to ask before booking. Improving those articles means reading what's there, deciding what to change, and carrying the changes through both languages.
Soul's completion messages described exactly the sort of improvements I was looking for. It said the articles now opened with clearer recommendations, repeated material had been removed, and practical details had moved closer to the top. One message said it had used search console data to prioritize nine articles, rewritten each in both languages, and updated an additional review after finding outdated prices.
When I saw that, I was very impressed. But then it reported 20 edited files. Another message described 18 edited files and three sub-agents working on the assignment. Compare those reports with the change cards underneath. The 20-file message had a card showing four files with eight lines added and eight removed. The 18-file message had a card showing six files, again with eight lines added and eight removed.
Those cards made the reported scope difficult to reconcile with the visible changes. They can cover only part of a session, and a single change line can contain a whole paragraph, so they aren't a complete audit by themselves. My judgment came from reviewing the work. The edits fell well short of the substantial rewrite I wanted.
The second chance did not help
We went back to Soul and pointed out the inadequate changes. This should have been the opportunity for Soul to finish the job. We'd already seen a correction help the glider demo. But then it made another tiny change. It just acknowledged me, but no real work was done.
By then I had spent attention checking the result and explaining what was missing. Continuing meant another round of supervision on the same assignment. I eventually switched back to GPT-6 Astra, which produced a substantial later rewrite, including 20 files with extensive changes to the copy.
You have to remember that GPT-6 Soul is a lightweight version of GPT-6 Astra. I expect it to be not as good, of course, but not that bad. Which is why I feel like GPT-6 Soul and Luna are sort of incomplete. I feel like OpenAI just rushed to release it after Anthropic released Opus 5.5.
We are testing Opus 5.5 right now. Ironically, our Anthropic account has been banned for some reason, so we're testing on Cursor. But as it stands, it feels very unreliable to use GPT-6 Soul for actual agentic workflows, especially article writing. It couldn't do that simple job, and we had the instructions in the skill files, the proper plugins, and the whole repo to take reference from.
So for now, we're still using GPT-6 Astra for most of our real agentic workflows.
