Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
Grok 4.6 Day-One 'Sonnet-Class' Claims Need Receipts
On Grok 4.6 day one I replied that fantastic pricing is the only part of the victory lap I find credible.
Key takeaways
- On Grok 4.6 day one I replied that fantastic pricing is the only part of the victory lap I find credible.
- Calling a new model Sonnet-class, or 'definitely ahead of Gemini,' is a claim. Claims need evals, not vibes.
- I still use Grok. I still upgraded Cursor. I still refuse to launder a launch thread into a benchmark.
Local LLMs on NVIDIA Spark / ASUS GX10
I'm starting to think some of these launch-day accounts are trolling.
There is nothing about the statements — other than fantastic pricing — that is remotely credible based on any real metric. Calling Grok 4.6 a solid Sonnet-class model, and saying Grok is now definitely ahead of Gemini, is beyond a joke.
That was a quote-tweet. I am expanding it because launch-day AI Twitter is how bad buying advice gets into production.

What I am not doing
I am not running a hate campaign against xAI. The same week I:
- Used Grok Build to replace Linktree (writeup)
- Said I like the Grok app's Enhance your post button
- Kept Grok in the WikiWayne image pipeline
I also moved Cursor Pro+ → Ultra because I still pay for a serious coding seat. None of that requires me to agree that 4.6 is Sonnet-class at 14:00 UTC on launch day.
What "Sonnet-class" actually has to mean
Anthropic's Sonnet line earned the nickname the hard way: months of people putting it in agents, IDEs, and customer support, then complaining in public when it regressed.
To put a new Grok next to that bar you need, at minimum:
- Coding agents that finish a repo task without quietly deleting tests
- Long-context edits that do not invent files
- Refusal and instruction-following that match how you actually work
- A price — this part can be judged on day one
"We are already moving some Sonnet workloads to 4.6" is a migration anecdote. It is not a metric. I migrate workloads when a model is cheaper, too. That does not make it Sonnet-class.
"Ahead of Gemini" is worse. Ahead on what. Image? Video? Devrel blogging? Arena ELO with a cherry-picked system prompt? Say the task.
How I actually evaluate a new Grok
I keep a boring list. You can steal it.
- Price / token and rate limits — believe the price sheet first. It is the only number the vendor cannot hide.
- One private coding task I already know the answer to (WikiWayne chrome, a game bug, a broken MDX frontmatter).
- One agent loop that has to touch multiple files. Compare to whatever I currently pay for in Cursor.
- A week, not an afternoon. Day-one vibes are a press cycle.
Until those exist, I will use Grok where it already earns its keep (build, enhance, images) and I will not rewrite the Claude vs ChatGPT vs Gemini map because someone posted a bullet list.
Local vs cloud is a different axis
Easy: run it on your computer. Hard: frontier Grok. I posted a version of that ranking too.
The GX10 exists so I can keep DeepSeek Flash and Qwen honest on metal. Cloud Grok exists so I can move fast when I am not measuring a local box. Mixing those two conversations is how you get "Grok 4.6 is Sonnet-class and my 27B local model is AGI."
Pick a lane for the claim.
Receipts, then adjectives
If 4.6 is as good as the thread says, we will know. Independent benches will land. Agent users will either stop talking about Sonnet or they will come back.
Until then: fantastic pricing, maybe. Sonnet-class, not yet. Ahead of Gemini, show the task.
I will update this page when I have my own numbers — not when the next influencer posts fireworks. Follow the corrections on X (@wikiwayne).
Frequently asked questions
No. I am saying I will not repeat 'Sonnet-class' or 'ahead of Gemini' until I see tasks I care about: coding agents, long edits, local-vs-cloud tradeoffs. Pricing can be judged immediately. Quality cannot.
Yes. Grok Build, Grok App enhance-post, Grok for images on this site. Using a tool is not the same as rubber-stamping a peer's launch-day scoreboard.
Named evals, named prompts, independent reproductions, and a week of agent work that does not fall apart. Vibes are a starting point, not a result.
Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
