WTF Is Going On(Weights, Tools & Frameworks)

한국어
OpenAI 2 outlets

Playco cut manual fixes 50% prototyping games with GPT-6 Astra

read it at the source — OpenAI →

“Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.”

who filed

What was said about it

8 points · 4 comments
1820 points · 1600 comments
4 points · 4 comments
26 points · 1 comments
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since…
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was orderi…
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how the…
So, on the one hand, we have AGI; on the other, the release page is returning 500s.
GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "re…
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage an…
If it's not clear what's happened: The launch was scheduled for 11am Pacific time. The press embargo broke at 11am, and we saw a flurry of press articles by Axios, TechCrunch et al. The model has appeared on the ChatGPT API. But the offici…
$10 per million input tokens and $50 per million output tokens sol is $4 / $20
Looks like OAI took down the page. Here are the benchmarks: https://imgur.com/a/aVTmzbC
Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273 How about we stick to that one for talking about the rollout, and this one for talking about the model?
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186. That must really blow. [0] https://arxiv.org/abs/2608.31126

Also on the board