Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.
It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it's talking about.
I find I often have to ask it to re-explain what it means.
GPT does this all the time, too (both Sol and Astra). I constantly have to tell it to not use terms that were not part of the initial prompt.
Hah, yeah even when you put it in AGENTS.md or a skill.. constantly having to remind it.. "what does AGENTS.md" say about doing that?".. Thinking.. Thinking.. "Oh, it says I should never do that, I'll remember that next time.."
Next session - same thing.
Separately have been using Grok 4.6 for a bit and it's also pretty concise.
I’m pretty sure the big bois don’t do it because it would undermine “confidence”.
Seeing a model output “Oh I should just delete blah. Wait blah is a production service, I shouldn’t touch that. Maybe I can gain access to blah? Oh the aws cli isn’t signed in to blah. I see kubectl has access to blah though! Wait, I should ask user permission first.”
Yeaaaaah. Thinking tokens are fuckin’ wild.
Could be something very stupid like - "I don't have ffmpeg available here. Should I install it? No, I can't. I'll proceed doing something that will take me 100x more tokens and wall clock just to avoid adding a dependency." I can then just stop and say - you've got nix flake there, just add it.
That's impossible with western models. The only way is to ask why it did something stupid when it already spent 50% of your weekly quota and produced millions lines of slop.
I suspect that they specifically train Grok to be able to work well with military personnel -- speaking the way they speak: brief, to the point, efficient communication. Personally I really like this. Claude sounds like some demented clown from the marketing department.
I’ve noticed Astra doing this a lot as well.
Anecdotally, I have noticed the same in the past week. It might just be anecdotal or driven by a long context window.
First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.
I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.
Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.
I do think an infrastructure where another Claude retranslates the output would be better. Oftentimes I forget to put it in the actual prompt and when I receive back 8 paragraphs of Claudeish I ask for it then.
I would have to disagree that it gets as grating as Claudeish though. Its just direct and professional instead of ring-around-the-rosy clickbait.
“I would have to disagree that it gets as grating as Claudeish though.”
It’s hard to imagine anything more grating than Claudeish. To quote Rainer Wolfcastle, "My eyes! The goggles do nothing!"
I've found that prompting any constraint on output (length, style, vocab, even simple formatting) not only places additional cognitive load on the model, which burns some of whatever cognitive budget is available, it will also often skew the output in other subtle and completely unrelated ways.
Since I found this artifact interesting, I did some pretty extensive experiments a couple months ago. The increased load is real, although it may not be apparent if you're not near any cognitive boundaries. The subtle skew, however, seems nearly ever-present regardless of load.
While extensive, my tests were just following my curiousity, not controlled, exhaustive or well-documented. I identified about a dozen prior sessions of varying length and complexity to test and downloaded them with a browser add-on. I then removed all other user prompt instructions except for the formatting instruction. A test would typically involve changing the wording of the formatting instruction ranging from brutally simple to detailed and complete, then starting a new session, seeding one of the test sessions and continuing it. To get a feel for baseline inter-session variation, I also tried running the exact same prompt/session multiple times back-to-back, at different times and on different days of the week.
Once I identified a promising prompt candidate, I'd make it the formatting instruction in my regular, daily-use prompt for a few days. I quickly got a feel for how seemingly minor user prompt variations impact response quality, compliance and tone across fresh sessions as well as those in various states of context rot, drift, decay and cliff (<--my nicknames for the distinct flavors of session degradation, not technical terms).
My overall conclusion was that every instruction, no matter how minor or unrelated it seems, has some, real impact on the model's cog load, attentional focus and/or attentional weight budget. Both how these impacts manifest and what causes more or less impact is often extremely counteriintuitive. To more fully understand this, I eventually, got to the point of testing null case variants, such as the entire user prompt being one sentence completely unrelated to text formatting or the session topic, like: "Don't reference the cartoon character SnagglePuss" (in a deep dive on ancient Sumerian clay tokens). Similarly, a simple one sentence prompt requesting something the model already always does naturally also has a cost (eg "Capitalize proper nouns"). As others have observed, heavy emphasis, absolute prohibitions or emotional weight in prompts also tend to have outsized impact in both skew (impacting unrelated output tone/style) and in accelerating session degradation. "Avoid referencing SnagglePuss when you can" would have equal compliance but fewer downside impacts than "NEVER reference the cartoon character SnagglePuss" in sessions starting to degrade.
There were also surprises, such as when I was scanning transcripts of an older, longer session and noticed the LLM was doing number formatting almost perfectly. On looking at the active user prompt at the time (I keep a log of every user prompt change I make for every model), it didn't even reference formatting at all. More experimentation showed it a result of the LLM gradually mirroring my consistent use of formatting structure in my prompts over a long session (in which I never mentioned anything about formatting). Unfortunately, that mirrored trait doesn't persist to new sessions and reaching that point requires a substantial number of rounds burning quite a bit of context window.
After spending time surfacing the impacts of just changing lightweight user prompts so they could be observed (which are the lowest priority prompts a model gets), I now wonder just how much more 'brilliant' the models we use daily would be if they didn't have dozens of pages high-priority manufacturer prohibition prompts we never even see weighing them down. We've only ever seen these frontier 'racehorses' when they're already pulling a heavy invisible wagon.
I don't think this is true.
They have to express themselves as tokens. The meaning of those tokens doesn't have to be text. See any model that can handle images/video. Also, I don't think math, svg, etc, are "natural" language.
And, only the final expression is tokens. The intermediate layers, with the encoded concepts, aren't "natural language".
But, to address your concern (which nobody can disagree with, since even humans can't fully express through text/pictures), potentially: https://news.ycombinator.com/item?id=49758615
The model isn't limited to concepts that can be expressed in natural language.
It's only once the AI gets to the output layers that natural language comes back into play.
After all, they're all made out of weights[0].
How do we know for sure? We don't even know how the emergent properties we see actually emerged?
For humans we know for sure that people sometimes have concepts that they have no word for (the reason the phrase "It's on the tip of my tongue" is a phrase, after all).
We don't know this for LLMs. When it makes new phrases, it's always a mixup of two existing words hyphenated (aside, that also seems to be the limits of SOTA models creativity - join two unrelated words together with a hyphen).
LLMs never respond with "It's on the tip of my tongue" type responses, indicating it has a concept but cannot remember (or does not have) a word for that concept. Every human, pre-speech-age, has managed to express or convey concepts that they had no word for.
So, no. I'd need a citation, preferably multiple, that did the trials and found that a model can generate concepts for which it does not have any words for.
Even if the input is in plain English, the model never sees any words, tokens or glyphs to begin with. It's vectors all the way down.
1. Is natural language holding LLMs back by some %? 2. Is natural language serving as a hard gate that will prevent LLM intelligent progressing past some specific point?
The answer to 1 seems like an obvious yes to me.
Your thesis says the answer to 2 is "yes." That doesn't feel right to me. Think about all of the humans who have pushed various fields forward: Einstein, Newtown, Bach, whoever. If natural language doesn't prevent an entity from surpassing humans in one intellectual field, why would it prevent an entity from surpassing humans in all intellectual fields?
(To be clear, I'm not claiming superintelligence will or won't be achieved; I'm considering your specific thesis about whether or not natural language will be a hard gate)
By the way, how good is Claude's Hopi?
It burns more tokens but is the only way to get tolerable text.
https://code.claude.com/docs/en/hooks-guide#agent-based-hook...
Literally every one, even 1-2 prompts later it starts to go back
It’s been really productive and I’ve been asking my agents to communicate using it more and more. I believe it’s relieved my cognitive load a bit while working with them.
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
The weird thing is, that's not what AI models seem to be doing. The prose is just weird.
This happens most though when the speaker doesn't (or care to) understand their audience.
Eg i find effective communication requires expertise in both the subject matter domain but also the reference of the listener. Eg in ELI5 framing, if you don't know what information 5yr olds are expected to know you'll do a poor job at an ELI5.
It often feels like Claude does poorly at both framing the response relative to what it "thinks" the listener knows, but also the prose is... sideways, just weird as you said.
If I don't grok an elaborate explanation, I can ask for clarification. If it's explained to me in an overly simplistic or unnuanced way, I'll walk away with a false sense of understanding.
That said, I'm sure we all have very different concentrations of these types of people and problems around us. I've definitely met some engineers who seem to actively try to make their language incomprehensible
It is unsurprising that a LLM fails, without coaching, to effectively communicate.
Agreed. Do you think it's due to that EU issue of making AI text be identifiable?
I should try adding these tips to my system prompt. Is there a shorthand to describe such language use? I am not a native English speaker.
As for the wording of the prompt, you're pretty on point, I created a custom output style targeting mostly the first two you have there. Some people have wording that demands a certain technical standard or uses fancy words to describe what to avoid, but I haven't seen evidence those work better than asking plainly and I suspect the opposite: LLMs mimic the user to a degree so talking to it in terms of technical specifications and fancy words is an invitation to get them back.
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.
Sure the explanation will oversimplify a lot but then you can expand it recursively if needed, you gotta start somewhere.
If the presenter can't divide a problem until it reaches a series of independently simple concepts, then there's usually something fishy going on.
You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it?
That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical.
Does not excuse the Claude slop.
Solving the problem right in front of you is easy. Stepping back and asking: is that a problem to be solved, is infinitely harder.
I did not use Claude to write my comment, so I don't know where that is coming from.
this is specifically an Anthropic problem, maybe due to their heavy use of Claude to train Claude itself?
Part of intelligence is knowing your audience and communicating efficiently.
Bingo! And on this axis many SOTA models fail miserably. These things are acting on my behalf under my direction. All the supposed intelligence in the world means fuck-all if nobody can understand it.
And like somebody else said… when meat-based humans talk like Claude does, it almost always means they either don’t understand what they are talking about, or are actively trying to conceal something and are a fraud. Not always, but almost always.
I am a huge fan of Gemini Pro for chat... gemini somehow just knows the most obscure stuff. I'll double check something Gemini said and find the source is deep inside a hard to access scientific paper. Google just has the best index of the internet.
It's best for brain storming, rabbit holes, and image recognition.
Let the big models do the heavy lifting for now.
Even if you aren't coding, you really need to double check its answers. Flash 3.8 hallucinated a Keyence camera's max operating temperature for me, last week, and backed it up with "references".
It's still my favorite model for most non-coding stuff, though.
I don't know if it's the plain english or what, but I really like Grok for legal research (as opposed to code). It's got a noticeable edge in getting to the point compared to Opus 5.
How representative that is of real world usage, I don't know.
In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
And like that grok4.7 cache reads are more expensive than sol's (at $0.40/mil).
I do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.
Googling it returns no matches but I think it was supposed to be “live viewer”?
Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.
Lol, probably because Tesla's software stack is buildroot based. I'll bet that was in the training data.
So yea, I find fable 5.1 writing to be excellent everywhere. I still use Sol daily though, but for things like config, quick research, fixes, code review etc. Feature work and writing is for fable 5.1.
Fable 5.1 is not there quite there yet.
They need to get that Sonnet 3.5 magic back.
// `HANDLE` is an opaque kernel handle (kernel32 validates and returns 0/FALSE
// on a non-console handle); every out-param is `&mut T` to a `#[repr(C)]` POD,
// ABI-identical to the Win32 `LP*` pointer (thin non-null). The reference type
// encodes the only pointer-validity precondition, so `safe fn` discharges the
// link-time proof. (`bun_windows_sys::kernel32` declares these with `*mut`;
// redeclared locally so the legacy-conhost cursor path below is plain calls.)
or // Progress's terminal handle is the canonical `output::File` (vtable-backed
// stderr/File from `OutputSinkVTable`). The duplicate `ProgressTerminalVTable`
// from B-0 round 1 is removed; tty/ansi/winsize route through the new
// `OutputSinkVTable` slots so `bun_core` stays T0 (no `bun_sys` dep).
from src/bun_core/Progress.rsjust reading this gives me a headache
The Claudish is dead. Long live the Claudish.
And the fact that Grok is the ultimate grandmaster of parallel tool-calls, routinely kicking off four or five at once. Overlapping the latencies makes a huge difference in responsiveness.
I also like how Grok is trained to print a short one-sentence descriptions of what it's doing before each step. Like an airline pilot calling out observations for the black-box recorder to hear.
That’s a factor of half the parameters. I would be curious to see more on the focus of smaller parameters model and pushing its frontiers
I don't want to waste money because my calculator is cracking jokes. They don't deserve their paltry 5% marketshare or whatever it is they have currently. I'm not even getting into Musk as a person or the horrid things we've seen Grok spit out on twitter. I just don't trust his companies with my data and I have seen very little evidence that it's ever the best tool for the job. I'm sure those cases exist but I can't imagine it's worth it.
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
Has a shadow
Better shaped beak
Leg position more realistic for bicycle riding
Better feathers
I'm curious if you feel the same about re-migration of Belgians from the Congo?
Personally I think it's fine for any country to vote to control immigration as they see fit. I think Japan is a good example of a relatively xenophobic culture that deals with this fairly and thoughtfully.
Can't say I've ever heard anyone implying that colonialists leaving Belgium was unjust. Colonialists is actually not the right word, more like extended occupation, only slightly better than the enslavement of the Leopold II era. The Belgian's were less than 1% of the population and all but an ancillary amount worked in exploiting the native population.
Irrelevant. HN is not the place to randomly inject flamewars about politics. It's explicitly against both the purpose and guidelines of HN.
Seems like you need to review the guidelines again, because they're pretty clear:
> Eschew flamebait. Avoid generic tangents. Omit internet tropes.
Basic morality is not "flamewars" or "politics".
The fact that Elon Musk's companies take contracts from the CIA and NRO is not flamebait. It's context that informs how we evaluate future SpaceX ventures.
4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks. But it's not super clear to me whether it's above the line or not. A part of that is that it is so slow that i haven't been making fast progress today with benchmarking it.
Overall, it it gets above my intelligence line its a good release...but you can read the tea leaves and tell the Grok team thinks this was a miss.
For example?
4.6 made more mistakes than SOL or Opus overall. Gave up a lot. And in my opinion, the rate of mistakes is kind of more important than how brilliant it is.
I think 4.7 may still be better, but I was hoping for clearly Sol/Opus level and so far it just isn't there for me.
Make a galaxy model search the space, create a document, argue and defend decisions, then hand it to 2.5 to implement. 4.6 was a slower less enjoyable version of that.
4.7 is better at "I want the button to cancel the jobs, dont make any mistakes" but honestly that's not what I use it's class for.
The Django code that comes out of composer2.5, to me, was insulting. Grok definitely was a step up, especially because the fast option reaaally is fast so even if it came out a bit wrong I could just whip it into perfection.
For frontend work, it's a different story. You can still tell that composer2.5 is taking the long route, but I don't think it's as egregious as with Django.
Also, composer2.5 would routinely run commands that were really dangerous and in need of proper sandboxing. Things like creating an ./uninstall.sh script with a HOME variable on which it does rm -rf $HOME. In general, when I asked composer2.5 to do things "for me", I knew a third of the initial commands would be failures, and sometimes they could be catastrophic failures (it did actually run rm -rf $HOME on what would be an actual home folder). This just hasn't happened with Grok.
I also have a bunch of vibe-coded apps I built for myself with composer2.5 and it is extremely noticeable that they hit a "this needs to be refactored as it's crumbling unto itself" line much earlier than with Grok and proper frontier models.
This is such cope.
Not to mention the whole launching reusable rockets thing which is pretty cool too.
[1] https://archive.is/20260619220349/https://www.nytimes.com/20...
If you think anything Elon doing is groundbreaking, you have no idea how the world works. Recent Space X ipo showed that the launches aren't cheaper, they are just heavily subsidized. Tesla was a piece of crap until they got their model 3, the only reason Tesla succeeded with their S model is because Elon was the edgy hype dude who managed to generate enough hype to carry them through the bullshit with the car. Self driving was supposed to be solved last year, and tiny companies like Comma AI manage to build self driving systems that are in someways better than Teslas.
I bet you think Steve Jobs was a visionary as well lol.
If not for SpaceX there wouldn't be gigabit internet connectivity in the middle of the ocean.
You're entitled to your own opinion to hate the guy but some self reflection goes a long way.
Most manufacturers already were working on hybrids, which to this day are still suprerior to EVs. Chevy Volt, outside of being Chevy, was still one of the best cars ever made for utilitarian purpose. Nobody wanted to foot the bill to do electric conversions until this was necessary.
Tesla only opened up a market segment for high end electric cars, which I guess is cool, but far from revolutionary. The model 3 was a big success only because again, it was subsidized. Meanwhile BYD actually makes cheap affordable electric cars, and we both know why they are not sold in US.
>If not for SpaceX there wouldn't be gigabit internet connectivity in the middle of the ocean.
Plenty of companies were doing geostationary orbits with satellite connectivity. SES for one.
Any more Elon slop? You realize you are defending a dude that is literally a Nazi, right?
Neuralink: to see the impact of increased independence and autonomy of a paraplegic one day after the operation
SpaceX: to quite literally approach the final frontier. Currently launches 80-90% of all orbital mass. Starlink is saving lives constantly.
Tesla: to kickstart the EV revolution and reduce fossil fuel dependence
Boring Company: to radically decrease tunneling costs applicable to all sorts of critical urban problems from transportation to utilities etc
* Start their own company
* Go work for a startup where they actually get paid options, and have a say in what the company does.
* Go chill at one of the big companies getting paid good salary while coasting because the work is so easy.
What they certainly don't do is go work at a company for less pay and harder work hours, all so that they can make a literal Nazi richer.
xAI missed its chance, Ball is on Anthropic's court.
Elon claimed Opus was 5T in April, and I think it's fairly likely this is accurate: https://x.com/elonmusk/status/2042123561666855235
It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.
Best to stick with a high end model + low effort, do a manual pass on high effort and fix the bugs you know are reachable.
The two models are in completely different price tiers. Astra costs 5 times as much.
It seems like all you can judge about cars would be their maximum speed on an oval.
Based on Artificial Analysis Cost per Task, Astra is about 2-3x cheaper than Fable 5.1 at Medium and Low.
Consequently Astra could be cheaper than Grok 4.7, depending on the task.
Is the training data more valuable ? The training process ? The harness ?
I know they are all important but where are they (all the frontier labs) really pushing to get incremental gains?
An amateur but worth reading nontheless
My SuperGrok subscription previously easily lasted me through the week even with mild coding through Grok Build. Now when I use the app 1-2 times a day to ask some questions, I’m almost running out by the end of the 7 days. It’s terrible.
I want to keep using Grok but logically it makes no sense for me to keep paying for it on the side when my quota just doesn’t last. I have also no desire to upgrade to Plus with these terrible limits, while previously I would have eaten up a $100/mo Grok plan. Rumors say SuperGrok got heavily nerfed with the SuperGrok Plus introduction, and that sounds about right to me.
I’m sure it’s a great model and I’d love to use it. I hope they get their plans under control and only only focus on Grok Bot.
Even worse, the mere risk of quota exhaustion mid-conversation makes me not risk starting convos with Grok, instead I'll use ChatGPT or Claude (even though Claude is inferior for non-coding tasks, and ChatGPT is inferior for all tasks).
In fairness to xAI, they're a profit-motivated company like any other, so they cannot give us tokens for free or less than it costs them. Reality is we may have to simply pay a lot more if we want that Grok goodness.
Claude voice mode if you put it on Opus is now also pretty good, but there are frequently situations where I get upset at it’s responses.
For now, I doubt anyone would notice your protest if you didn't announce it.
Totally the same.
If Elon hadn’t worked with Orange Man Bad, then the Left would still be in love with him for his massive former donations to the Democrat political machine, and his work against climate change.
The whole “he’s a nazi” accusation is banal, and people are seeing through it now. That’s why we’ve moved on.
Like the whole pizza parlor pedo basement thing, people will death grip stupid stuff because they are so desperate to manifest the worst possible image of those unaligned with them.
The problem is that it blows up in their face and just makes them look unreliable, dumb, and lost.
Musk has done so many objectively bad things that there is no need for people to dilute their reputation on fringe theories and interpretations. Pushing the nazi thing just gives Musk ammo that his detractors are so desperate that they need freeze frames and hidden context to make him look bad.
If you want him to get the benefit of the doubt about his "hand gesture" then it would help if he wasn't promoting far-right parties all over Europe.
Excited to try 4.7. I hope they fixed the "it's not X, it's Y" that showed up in 4.6.
By now AI should know of the DRY concept. But no. Hence the keys have a rounded rectangle for the key shape and another rounded rectangle for a clip path, to prevent text overflow. There are 72 * 2 = 144 identical rectangles, when just one would suffice (in the defs), with this being cloned once for the clip path, and 72 times for the keys.
I would not expect SVGO levels of optimisation (rounding numbers, that sort of thing), however, the human, if writing out the same thing for the 72nd time, might think 'is there a better way', to get the manual out. A graphics program such as Illustrator would not do that, but AI 'should' because AI.
The above is not criticism of your work, just an observation regarding AI SVG capabilities.
There are also interesting inheritance rules with SVG, so you could define the basic shape of a key, well, several shapes, just as rects in the defs, with no stroke or fill specified.
Then, at the group level, you can then specify stroke and fill, so there could be a group of normal keys, another group for modifiers, function keys and so on.
Then there are the keys themselves, how do you clone a shape and put different text inside each clone? There are many ways to do this but I think you are on the right track using the clip path approach, albeit using the rects in the defs.
What is interesting about SVG is that artists don't care for the file format, they just see text as shapes on a page. Then programmers don't care for SVG as that is a graphic designer/artworker thing. So SVG sits in this witch-space, with only a few brave enough to wade in and do cool stuff.
Given your application, and given the fun that could be had with SMIL/JS, you could make your SVG files interactive, so you press a key and a popover tells you more about what that key does. You can even get audio working in SVG, as well as HTML popovers (in foreignobjects, as buttons, but working, nonetheless).
'Views' is another interesting SVG feature. I have a sprite sheet that uses a lot of views, where you are projecting your SVG into some type of virtual canvas, taking a 'picture' of it, and then incorporating that in something else, maybe a CSS variable.
One 'deadly addiction' is animation. Filters are another 'deadly addiction'. Why have a static and actually useful diagram, when you can animate it, move the 'camera' and add the equivalent of 27 Photoshop layers as filters to everything?
For example, supposing you wanted to show what keys to press, with there being modifiers and a sequence, e.g. 'Hello World!'. The animation for one letter could be what triggers the animation for the next 'key press' and so it goes.
My top tip of all: reposition the origin (0,0) to where it makes sense. Many objects have symmetry, so you can define one side, clone it, scale it (-1,1) and do it all around 0,0 to then translate the results to somewhere sensible.
I have found the JetBrains IDEs to be extremely useful for SVG, the preview feature is very helpful, as are the code hints.
AI sort of knows SVG, so I have had some suggestions from Google on how to build filters. These never work, but they do get you thinking. Say you wanted to use filters to add specular highlights and animated shadows to the keys, that would be fair game for AI hints on how to do it.
Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Astra's build: https://html.non.io/annui/
Grok's build: https://html.non.io/Annui-grok/
Additional prompt instructions: "Add scrolling clouds behind the statues. Dynamically light the statues based on mouse position. Use diffui to generate the normal maps/depth maps/roughness maps of the objects, and to separate out the assets on to different layers."
Overall I find these models are getting good at following image as a source of instructions, but their refinement of the output varies heavily between the models. Astra's final output feels more polished, has better visual contrast, and the animations between the pages are smoother. Grok also chose to light all of the background elements, which imo overcooks it a bit.
Still though, for the price it's a great starting point.
Grok 4.7: $12.60
GPT Astra: $35.00
In cursor I have switch over to grok for planning a composer for coding.
"Privacy# All these models are hosted in the US. Providers follow a zero-retention policy and do not use your data for model training, with the following exceptions:
Big Pickle: During its free period, collected data may be used to improve the model.
DeepSeek V4 Flash Free: During its free period, collected data may be used to improve the model.
MiMo-V2.5 Free: During its free period, collected data may be used to improve the model.
Laguna S 2.1 Free: During its free period, collected data may be used to improve the model.
Ling-3.0-tiny Free: During its free period, collected data may be used to improve the model.
LongCat-2.0 Free: During its free period, collected data may be used to improve the model.
North Mini Code Free: During its free period, collected data may be retained and used to improve the model. Do not submit personal or confidential data. See the provider’s Terms of Use and Privacy Policy.
Nemotron 3 Ultra Free (NVIDIA free endpoints): Trial use only — do not submit personal or confidential data. Your use is logged for security purposes and to improve NVIDIA products and services. The logged session data for improvement purposes is not linked to your identity or any persistent identifier. For more information about data processing practices, see the Privacy Policy. By interacting with this endpoint, you consent to the collection, recording, and use of such information and the NVIDIA API Trial Terms of Service."
https://openrouter.ai/deepseek/deepseek-v4.1-flash?endpoint=...
"order": ["relace", "coreweave", "novita", "baseten", "together"],
"allow_fallbacks": false
It still won't be quite as high as you'd get by just using DeepSeek because occasionally a request will fail and you'll get routed to a backup provider with nothing cached, but it's close enough not to matter in most instances.But I can't argue with the lower off-peak pricing when using DeepSeek directly. The downside is they train their models on your input, which might be a deal-breaker for many users (as it is for me).
Well, at least I spent lots of dollars, and I had to use those models the same way I am using local and cheap models, with the same results.
On the other hand, Grok and GPT finish these tasks in <5min with no issues, and significantly better output.
GLM or Kimi are better for my own personal projects. DS? uhm. it just keeps doing dumb crap
I told it to compose an image (putting headgear on top of a head) - kept getting it completely wrong, generating new headgear, getting that wrong and screwing up the scaling.
I told it to diagnose a webhook issue that was happening in production from a local environment and it kept giving me moronic answers like that environment variables weren't set (despite me telling it that the values WERE set in production).
I've tried the latest models across OpenAI, Claude, Chinese, etc. They just do stuff. That's not how work is though. You want them to do specific work, at which point it's a real hassle to follow up on all the garbage they have been outputting.
/long rant
Unless they produce the same token output on the face of it, it looks like they're trying to cover for 4.7 not having good model perf?
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
From their headline comparison:
Grok: $2/$6 per million
Fable: $10/$50 per million
What this doesn't say: Grok costs 0.50/M cache read, Fable $0.25/M cache read
Long running agentic workflows are dominated by cache reads.Just makes Grok sound deceptive, and more importantly, reliant on user's lack of understanding of costs aka predatory (which in turn is more infuriating)
Grok 4.7 is near the top of the board. A significant improvement over Grok 4.6 but still not as good as Gemini 3.8 Flash which is very cheap and fast too.
It is the same multiplier for Sol with subscription. For Astra though the multiplier is ≈20x, so half of Sol usage.
For Claude it seems to be ≈40x too for Opus, but less for Fable (similar to Astra in GPT).
All on the most expensive plan. Previously, Grok usage escalated linearly from the $100 plan to $300 plan. That would be a really good $100 plan if it is still true.
Some sources:
1. https://x.com/kunchenguid/status/2098256018836963382
2. https://x.com/stevenzhang/status/2092110386569089311
3. https://github.com/openai/codex/issues/43731
It reset just a few hours ago and I've been running it, couldn't be more than 15 sessions none more than an hour long:
---
Session usage: no model calls yet in this session.
Weekly limit: 46%
Next reset: September 27, 23:20
---
I actually have to believe my account is messed up tbh, it's so bad. For reference I've ran 12 fable and some ~40 Opus sessions since reset yesterday on a CC account, at least 5x more usage by my estimate:
Current week (all models)
28% used
Resets Sep 28 at 5am (Pacific/Honolulu)
---Ok looking at it more, Grok and Grok Build just really suck. They are about 10x less token efficient, often using 200+ tool calls in a row for what are not even big tasks where Opus would use 5-10. Their cache hit rate is worse, and two sessions got into basically unnecessary loops costing a solid quarter of the entire week. And this was on smaller tasks as I tend to use it for easier things.
I can't help much more than that, I did that research in the last few days, but I never used Grok myself.
I pay for GPT, Claude and Gemini. Last week I consumed all my quota on two of them, so I wondered which next subscription I would pay for if needed.
For me and what I’m doing that’s insanely good value.
I find grok build chews through my SuperGrok sub very quick - but I think that is due to it having the 500k context window which uses more credits. Cursor limits it to 256K (tho I see in today’s update for Grok 4.7 there’s now a toggle for context size).
Normal SuperGrok barely lasts me through the week with very mild usage and no coding. The sentiment around SuperGrok Plus is also not great and I haven’t seen someone saying they’re happy with it yet.
SuperGrok Heavy is $300/mo, so you could get a full ChatGPT Pro and Claude Max 5x for that price. That’s so far out of my budget for a single provider I haven’t bothered trying it.
I still have SuperGrok through X Premium+ but will downgrade that next billing cycle
Astra for deep dive investigations, Sol 5.6 at mid-level for day to day tasks, Grok 4.6 via Cursor for routine and low complexity tasks.
There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.
But honestly, it is because numbers are like people; torture them enough and they'll tell you anything.
Output tokens from Intelligence Index:
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.
https://aibenchy.com/model/x-ai-grok-4-7-medium/#showcase=dd...
https://aibenchy.com/model/x-ai-grok-4-7-xhigh/#showcase=2f9...
Is there a metric for like... time taken when comparing these two? I see score and cost.
If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?
Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".
musk can fund the space stuff with this
But also Xai doesn’t seem to care about user experience and long term support.
For daily one off questions I prefer it because it is fast enough and I like the way it responds. I also use it for basic research like “find me a battery drill for this and that”.
Kimi and GLM feel extremely coding oriented. I use them for code reviews basically. I hate the way Anthropic models talk. GPT takes too much time and effort for that kind of stuff for some reason.
Grok happened to be a nice middle ground.
As a technical point of reference to compare against other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project and it could blow other options away and I wouldn't touch it.
You probably shouldn't cut off your nose to spite your face.
What's superficial about refusing to use a product from someone like that? Or are you one of those 'technology isn't about politics' people? That's a superficial take if you ask me.
All technology is political, and understanding that is a deep, not superficial take. It requires systems thinking which unfortunately many people building technology seem to lack, despite software being a sophisticated complex system.
I'll never use an xAI product.
I'm still never going to use an xAI product.
If a shitty murderous tyrant builds some roads the citizens can still use those roads while protesting against the tyrant
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
- it allows different models within one session via roles (I only have API, so pay per token)
- it's much more likely (ime) to use the LSP over grep for determining how code fits together
But I agree a 20k+ starting context is way overkill.
I find it's very hard to get information on harnesses people are using. I have to stay model agnostic so I avoid claude, codex, cursor, etc. I've used and tried opencode, which worked well, but obviously lacks the above features.
Does anyone have a resource for following what people are actually being productive with? With so much vibe going on it's hard to separate the wheat from the chaff.
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.
If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
https://x.com/elonmusk/status/2102082011233931762?s=20
so it's likely about usage in Cursor specifically.
That it isn't the most efficient way to achieve the same end result is irrelevant.
Many people are allergic to “AI safety” as they perceive it as an attempt to deceive them. When they ask a factual question and get back an unfactual answer, it makes them upset. None of the people I know who feel this way are searching out CSAM. They accurately perceive that there is a team that wants answers to come out a certain way.
For instance, I’m loosely connected to people who care about animal safety in AI research. There are absolutely people spending their time pushing a narrative about _the_ ethical way to interact with animals. Feeling this not being forced onto you is kind of nice.
These comments don't stay up much anymore and I can't tell if it's structural to the forum (flag weight + statistical mechanics of votes + guidelines) or if it's the userbase sentiment.
But I think it represents real malaise in the community. It's not a moderator plot, people here really just don't care and might even support this.
We really are in the minority of opinion for giving a damn about liberal democracy.
Between the guidelines + user thoughts (e.g. repetition, low novelty/new info), there's other reasons these types of replies might end up dead.
I am worried that it leads to people self selecting to other forums biasing the remaining userbase vote/vouch/flag distributions. In an exit vs voice situation, the voice kinda dies out. Then we end up other-izing people and homogenizing our communities.
But I concede it's also possible that the minority opinion issue could be the core driving force.
Nobody both worked and spent their money to get Trump elected like Musk. 300 million to his 2024 campaign [1]. DOGE. On-stage endorsements. Nobody even came close.
No, other big labs are not "innocent little virgins", but they're not even in the same solar system of harm as Musk. To hand-wave at the differences is to permit them.
[1] https://www.opensecrets.org/2024-presidential-race/donald-tr...
Handing corporate code secrets to his AI model is... unusually trusting.
And methane is a large percentage of all power production in the US. So again that also applies to all the other data centers. (And FWIW they've been winding down and shutting down the on site methane generators.)
And no corporate code was handed to AI models.
Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
Can you provide specific examples of where Elon has bent the levers of government?
So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.
Also pushing technology forward.
Anthropic: 1.25B/month
Google: 0.92B/month
Unnamed customer starting in december: 1.1B/month
Starlink monthly revenue is ~1.5B/month
If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.
Ah, yes, democracy!, except for when the public is wrong.
Who decides when the public is wrong? We do! Who decides "what the public voted for"? We do! So we are the rulers? No, of course, not, this is democracy.
You want to become the decider of when the public is wrong and of what the public voted for? TYRANT! TYRANT!
This is simply epistemologically incorrect. It's obviously incorrect in this case because voters writ large do not have any idea how the government is administered and how to improve it, so even if they claimed to be voting for that, it would not necessarily be an endorsement of any particular approach.
More specifically we know it's not true in this case because there are polls. Voters didn't even claim to care about this! "How the government is administered" was not a high salience issue to voters. Simple as that.
Nonetheless, I didn't suggest anything about overriding their votes. It sounds like you have some sensitive spots to work through (someone obliquely criticizing your idol for sucking at his job?)
half of voters don't pay any attention to politics until the week or two before voting
Sheep often like to think themselves the wolf or coyote, it would seem.
Fuck, it is like the denial around Jan 6th. Those idiots we’re live streaming that shit. I watched it go down live. Now they say they weren’t violent.
We can’t have discourse when we have legit video evidence and people refuse to open their eyes and choose to deny reality
Which Nazi ideologies do you think he embraces? How do you reconcile all the Nazi ideologies he rejects?
"You're on, bro"
The personality is bland and it doesn’t work nearly as hard or even tries to help.
I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.
created: 9 minutes ago
That could have been said just as perfectly well from a main account my guy/guyette
There are ample reasons to believe that Elon Musk is running his mother's account and that the photos weren't even real putting in question that he even had a birthday party.
If you ask Grok about what this means it will always take Elon's defence. It will vehemently deny that Elon would be capable or willing participant of such a thing even if you point out that he faked being a world class gamer, buying accounts that had done all the work and showing none of the skills when live-streaming.
Yeah, I'm sure the guy has time to run his mom's social media account. What's your reasons or evidence? Did you consider maybe it's a social media manager one of them hired running the account?
Until you ask it to start generating horrific imagery and then it's best in class.
The value of the internet is that people can share whatever they want, and use software how they want. This will mean that some people will abuse that. This is the tradeoff of a free society.
Sounds like a plus. Guess I will give Grok another try...
Maybe I'm in some kind of bouble but I have never met or talked to anyone who has used Grok.
Not sure if I'll hold the subscription but I could see myself working with it more.
Grok 4.6/4.7 feel like you are less tokens because Cursor gives you double the quota subsidising Grok specifically - there is a separate Grok/Cursor model quota bar in addition to non . Its actual token drain if you are paying API rates is generally far higher than Fable, Astra et al. You are not using less tokens.
It is weird in some reward way, it will frequently, at least for me, complete the task in the laziest way possible,Technically it is done, but that's about it.
For example a user registration system ended up with users being able to log in as anyone because the authn was a cookie set with the user's pkey id, unsigned. Nearly anything it outputs technically works but if you give it a GLM or DS pass you will find dozens to hundreds of vulnerabilities of varying hilarity (Fable refused, Opus refused).
I have seen it write SQLI-vulnerable code, asked it to review a file without saying what's wrong and it did not catch it (fresh session, single file @-tagging in harness). Again, technically, the code works and does what the prompt asked, its just handing in some of the laziest copied homework I have ever seen.
If you specifically call out to use prepared statements it will typically put a plaster over this however I have also seen Grok 4.6 use prepared statements by concenating the user input raw into a statement then executing it with no ? or named replacements, rendering it somehow a SQLI-vulnerable prepared statement. One med it added a OR 1=?, replaced the ? with another 1, then said OK you are using prepared statements now.
As a chatbot it’s totally fine, virtually indistinguishable from Gemini or ChatGPT or Claude.
For coding it’s… okay. I tried 4.6 and it feels similar to Opus from 12 months ago, or maybe Sonnet from 9 months ago. YMMV.
It's just an observation but so far a pretty solid correlation. Musk has so severely poisoned the well in terms of his UK reputation that the only people who are open about using Grok are... well, wankers is as good a word as any.
FWIW among the AI-using people, it mostly goes Claude Code, then Codex, then whatever runs on their Mac. The only Cursor user I knew has jumped ship to OpenCode.
I like to follow them and look for benchmark for each LLM release.
I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6.
I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.
I'm judging on benchmarks, and whether anybody or any company I know has ever suggested using it (not yet).
I don't personally make my judgements based on how many other people mention a thing, but if that gets your code written, by all means.
The voice is the same AI slop as the others imho.
(This is about Grok 4.6, I didn't test 4.7 yet).
edit: clarified I mean agentic coding tasks
Deepswe results show that grok 4.6 is more expensive per-task and consistently scores worse than: luna xhigh, glm 5.3, astra low, sol high/xhigh, opus 5 medium.
Grok also used almost 3x as many tokens/turns to complete tasks than all of those models (besides luna), so it takes way more time to complete a task.
There isn't much reason to use Grok at all, it's gotten better but it's still worse than every other player in the field, which shouldn't be a surprise considering until about a year ago they were just buying tokens from other providers and pretending it was their own model.
With gpt-6 luna and sol coming tomorrow it's going to look even worse too, especially if new luna retains the same dirt cheap pricing that 5.6 luna has.