플레이코는 GPT-6 Astra를 사용하여 하나의 그레이 박스 기반에서 세 가지 테마 게임 프로토타입을 구축했으며, 이전 모델보다 수동 수정이 50% 감소했다고 보고했습니다.
원문
“Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.”
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since…
ARC-AGI-3 스코어카드는 매우 오해의 소지가 있는데, 자체적으로 "responses API harness를 사용하면 Sol이 약 30% 정도의 점수를 낼 것으로 추정한다"고 명확히 밝혀 놓고서도 GPT-5.6 Sol의 점수를 7.8%로 표시하고 있으니까요. 아마도…
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was orderi…
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how the…
I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "re…
내 생각에 내가 가장 기대하는 것은 _사용자 프롬프팅_의 증가야. 내가 제약이 부족하거나 애매한 프롬프트를 주면, 모델이 여기저기서 가정을 한 번에 추측해버리는 걸 원하지 않아. Fable/GPT-6의 데모는 인상적이지만, "re…
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage an…
If it's not clear what's happened: The launch was scheduled for 11am Pacific time. The press embargo broke at 11am, and we saw a flurry of press articles by Axios, TechCrunch et al. The model has appeared on the ChatGPT API. But the offici…
Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273 How about we stick to that one for talking about the rollout, and this one for talking about the model?
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186. That must really blow. [0] https://arxiv.org/abs/2608.31126