@kunchenguid: day 1 observations for grok 4.7 ignore the reports that say “it’s terrible” and the only thing they reference is a publ…
Summary
The user shares day 1 observations on Grok 4.7, highlighting its close adherence to system prompts, stability, and conservative behavior, while noting it is slower and more costly than previous versions.
View Cached Full Text
Cached at: 09/22/26, 03:56 PM
day 1 observations for grok 4.7
ignore the reports that say “it’s terrible” and the only thing they reference is a public benchmark. the same benchmarks told us opus 5 was better that fable - they are useless
also ignore the reports that compare models with 3d games - that’s not real work. it’s made for attention on social media
i used grok 4.7 for a whole day as my firstmate, and it has been a really solid model with visible improvements over 4.5 (i’m ignoring 4.6 because 4.5 has been working better in my experience)
key differences with 4.7 -
- it follows system prompt very, very closely
i noticed firstmate showing many new behaviors that i’ve never seen before, such as asking me to name specific red CI checks that i’m ok with bypassing, and refuse a simple “yolo” instruction
i traced it and it’s indeed how i instructed it in firstmate’s system prompt, but none of the other models followed it closely enough to make this behavior visible - grok 4.7 is the first to pick that up
there were a few other similar examples as well. so to me this is a clear behavioral difference
- it’s very “stable”
if you’ve used astra then you know what a “spiky” model is. it can have some genius moments but you occasionally also wonder “how could it be so dumb and doesn’t get me”. grok 4.7 is the opposite of that
throughout the whole day so far, i’ll be honest i haven’t get a “wow this is absolutely genius” moment yet, but grok 4.7 has been very steady with no big surprises. its behavior feels predictable, which does help it gain trust from me quickly
- it’s a conservative model
it doesn’t like to take actions without asking, and would explicitly say so
this is a bit of a double edged sword, because it means i sometimes have to state the obvious “yes i do want that”, but in hindsight a lot of those cases are indeed a bit ambiguous and i may not have preferred the model to just move forward without my confirmation
- it’s a bit slower and costs more than 4.5, visibly
turns are taking a bit longer and my quota is draining at a visibly faster pace. i haven’t quantified exactly where this is coming from yet
so overall, i think it’s showing some clearly different traits, and i mostly like the changes. i’m going to keep it as my primary firstmate and observe more
if you’ve been using it, what qualitative insights have you gathered from real usage so far?
Similar Articles
@corbin_braun: I Spent 50 Hours Using Grok 4.7 (here is its best use case)
A user shares their 50-hour experience with Grok 4.7, highlighting its best use case.
@ericzakariasson: grok 4.7 is here, and its our best model so far! try it out in cursor, grok build, api or anywhere you get your tokens!…
Grok 4.7 is released as an improved version of Grok 4.6, offering better performance at the same price and speed, and can be accessed through various platforms.
Grok 4.7 first impressions
The article shares initial user impressions of Grok 4.7, an AI model, discussing frontend and backend work, comparisons to other models like 5.6 sol and fable 5.1, and raising questions about safeguards and censorship.
@stanine: Hmm. Today, we ran our 2100 scored runs with Grok 4.6. Versus 4.5, the pass rate regressed from 87.3% to 85.9%, and nea…
The article reports on performance regression in Grok 4.6 compared to 4.5, with lower pass rate and higher latency, affecting practical business tasks.
@ericzakariasson: what do you think of grok 4.6 so far? what can we improve in grok 4.7?
A user is asking for feedback on Grok 4.6 and suggestions for improvement in Grok 4.7, indicating a discussion about an AI model version.