WTF Is Going On(Weights, Tools & Frameworks)

한국어
OpenAI 2 outlets

GPT-6 Astra: A new generation of intelligence

read it at the source — OpenAI →

“Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.”

who filed

Simon Willison 42h
GPT‑6 Astra

What was said about it

230 points · 141 comments
2195 points · 2013 comments
25 points · 16 comments
18 points · 6 comments
255 points · 163 comments
GPT 6 Astra, so good even OpenAI are worried
11,197 views · 139 comments
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since…
I posted this in the other Astra thread but it's just fallen off the homepage, so... Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p... I think this is a genuinely…
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage an…
I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a…
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actual…
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fab…
Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "re…
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how the…
That's some crazy SVG generation: https://aibenchy.com/compare/openai-gpt-6-astra-high/google-... It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach whe…
I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new fac…

Also on the board