Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5. More concurrency than …
Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG. https://tools.simonwillison.net/mark…
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son... 2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu... U…
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange. Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5%…
Unless you’re using frontier models like Astra, Sol, Fable, or Opus, I think you’re often better off using Chinese models for a fraction of the price. I’m not sure people realises just how competitive they’ve become. GLM and DeepSeek are g…
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but high…
pro tip: if you're on free tier, using Medium settings is far more intelligent than Max, and tokens don't run out so fast. Max is cranky and verbose, Medium is patient and somewhat goofy, but does the thing as expected. At least within Sep…