WTF Is Going On(Weights, Tools & Frameworks)

English
OpenAI 2개 매체

GPT-6 Astra: 새로운 세대의 인텔리전스

GPT-6 Astra: A new generation of intelligence
원문 읽기 — OpenAI →

GPT-6 Astra를 소개합니다. 당사의 가장 지능적이고 정렬된 모델로서, 컴퓨터 사용, 코딩, 사이버 보안 및 과학 분야에서 최첨단 역량을 갖추고 있습니다.

원문

“Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.”

어디가 다뤘나

OpenAI 최초 3d
Safety overview: GPT-6 Astra
Simon Willison 42h
GPT‑6 Astra

어떤 얘기가 오갔나

230포인트 · 141댓글
2195포인트 · 2013댓글
25포인트 · 16댓글
18포인트 · 6댓글
255포인트 · 163댓글
GPT 6 Astra, so good even OpenAI are worried
11,197회 재생 · 139댓글
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since…
ARC-AGI-3 스코어카드는 매우 오해의 소지가 있는데, 자체적으로 "responses API harness를 사용하면 Sol이 약 30% 정도의 점수를 낼 것으로 추정한다"고 명확히 밝혀 놓고서도 GPT-5.6 Sol의 점수를 7.8%로 표시하고 있으니까요. 아마도…
I posted this in the other Astra thread but it's just fallen off the homepage, so... Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p... I think this is a genuinely…
다른 Astra 스레드에 이걸 올렸었는데 홈페이지에서 막 내려가서요... Astra의 Pelicans, 비교용으로 5.6 Sol, Terra, Luna도 함께: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p... 제 생각에 이건 진정으로...
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage an…
I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a…
드디어 접근 권한을 얻었어요. 여기 펠리컨들입니다! https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 맨 아래에 있는 "max"는 4분 2초가 걸렸고 비용은 63.206센트였습니다. 비교를 위해, 여기 그 새로운 Astra 펠리컨들이...
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actual…
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fab…
Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "re…
내 생각에 내가 가장 기대하는 것은 _사용자 프롬프팅_의 증가야. 내가 제약이 부족하거나 애매한 프롬프트를 주면, 모델이 여기저기서 가정을 한 번에 추측해버리는 걸 원하지 않아. Fable/GPT-6의 데모는 인상적이지만, "re…
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how the…
That's some crazy SVG generation: https://aibenchy.com/compare/openai-gpt-6-astra-high/google-... It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
그거 꽤 미친 SVG 생성이네요: https://aibenchy.com/compare/openai-gpt-6-astra-high/google-... 테스트하는 데 시간이 좀 걸렸는데, 처음에는 OpenRouter가 이 모델 ID에 대해 Not Found 오류를 띄우고 있었어요.
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach whe…
I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new fac…

상황판의 다른 소식