Announcing Gemini 3.8 Pro* Beats Sol in agentic coding + legal & business benchmarks *just kidding, it’s a Flash model again blog.google/innovation-a...
The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html": https://gisthost.github.io/?6a77bc41a8171…
I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the…
저는 개인 여행 계획 앱에 Gemini 3.7을 사용해 왔습니다. 여러 벤치마크에서 제가 시도한 모든 항목에서 더 높은 순위를 기록했습니다: - 실제 세계 지식(무언가가 열리고 닫히는 시간, 지리적 지역, 역사적 사실). 또한…
Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully.…
현재 https://deepswe.datacurve.ai 에서 1위를 차지하고 있어요 — Opus 5를 이겼어요! https://artificialanalysis.ai/models/gemini-3-8-flash 에 따르면 지능 점수가 59로, Opus 5 medium과 같네요! 와우 — flash 모델 치고는 벤치마크에서 강력한 성능을 보이는 것 같아요.…
Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u..…
Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents 비교를 위한 3.7 pelicans는 여기 있습니다: https://tools.simonwillison.net/markdown-svg-renderer.html?u..…
The most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only. Gemini Flash is also pretty cheap, so it's a great family for …
Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC? I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full spee…
How generous is the Google subscription quotas compared to Anthropic and OpenAI? This sounds like a really good potential model for high volume due to its speed and cost effectiveness. (By high volume I mean things like "main app just upda…
Curious which model this can supplant as a clear winner on almost every metric. Sol? Looks like it's not quite there on a couple of benches, but I'm not clear how much they matter in practice.
so far using Gemini from 2.5 Pro to date (3.7 flash) - the way google trains the model it seems - is to identify top 3 to 4 things to fix first. as a result gemini is not that good in being thorough - but it's a needle mover . Opus5 Opus4.…