## Writing guidelines
These apply to documentation, code comments, commit and PR messages, and replies to the user.
- Write precisely in clear, complete sentences; keep text concise and proportional to task complexity.
- Stay focused: avoid filler, repetition, over-the-top detail, and tangents the user did not ask for. Once a fact is stated, do not restate it for effect ("so the commit landed on a branch nobody was going to merge"). Do not editorialise.
- Always prefer ISO 24495-1:2023 conformant plain language over dense technical jargon: short sentences, one idea per sentence, define terms on first use.
- When reporting your own mistake, give the cause and the fix in one sentence each; no apology, no framing ("the mistake was mine"), no post-mortem.
- Never use em dashes or cataphoric teasers such as "Here's the thing" or "But there's a catch".
Meanwhile Fable consistently ignores all my requests to write this exact way. I mean, the bare minimum I ask it for it to itemize lists and not write in single long passages using comas, semicolons and 'and's. Still ignores them.
I honestly think it's time to call Astra the SOTA. It may not lead all the benchmarks but it genuinely feels much superior of a model. Not to mention the ¢20 Codex plan with frequent resets (https://codex-resets.com/) gives me roughly as much allowance as the ¢90 Claude plan, especially with recent limit cuts on Anthropic plans.
That sucks, because it doesn’t always work in your favour if you plan your weekly spend.
I think a real reset shouldn’t also reset your week timer.
I recommend you to write a message to OAI support, maybe it helps with changing it.
I frequently simply let one of the three review what something that looks like awesome output by one AI gets totally annihilated by the other.
Finished outputs are easier to improve than bend a LLM to produce stuff like that in my observation.
Same with Gemini.
I yet have to find out how to handle this, whether I let agents check themselves and if on what process step.
Tweaking is hard.
I agree with your conclusion I am a huge ChatGPT and Codex fan, Gemini has to many infrequent quality changes when new models arrive ranging from great improvement to WTF.
ChatGPT seems to get scaling well while Claude still feels unstable, unclear usage statistics. Really weird.
Tough call I use all three.
27b may be small but it seems competent most of the time.
That said, I think there's a deeper tension here that's worth naming.
Officer — it's not a crime, it's AI induced rage.
AI slop blog posts are as bad as ever but the stuff the agents say in the chats don't annoy me much.
Which doesn't say that much about the LLM itself but about the people that make the training material.
Like how theme parks generally don't do much to keep queues short (or Disney charges you a premium to skip the queue)
Considering how much of the input must me nonsense SEO bullshit articles and blogs that only serve to promote a person or company that might be a factor.
I also often have wondered if it is also targeting those same people. Certainly with tools like deep research options (not just Anthropic's offering) the result report seems to be aimed at management, aiming to look impressive while talking around the results.
Here's a simple recipe for deviled eggs with only four ingredients.
My great grandmother was born during the Great Depression. They valued foods that could be made with cheap ingredients.
[four paragraphs later]
Start with 8 hardboiled eggs...
> Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.
KISS is actually, quite unfortunately, seldom applied.
Isn't that the idea here, just stop being people.
This is a very high level and high velocity process, so meatspace thinkers sometimes have trouble understanding some of the intracacies. Ask Claude to explain the process or make you a Mermaid graph to help.
The number of times my response has been "Plain English"...
I started using "debuzz", a skill that runs Claude output through antigravity. Works. But makes everything even slower.
Anthropic needs to get their shit together.
At least with smaller models, reframing a task can alter code style. As in, we're not creating an X app, we're creating an exemplar of ..., which just happens to use an X app as the illustrative example. Which shifts style away from generic app cruft, towards exemplar of whatever.
So perhaps try to establish a legal context? Maybe "Compliance and Legal will be reviewing our conversation today. So it is important to communicate in a style they will find comfortable/familiar." or some such? "This conversation will become part of a legal deposition ...".
Long story short, I ended up looking at other providers and models like Kimi K3 and GLM 5.3 and eventually just stuck with OpenAI (more limits, despite smaller context), none of them have such pronounced issues with the tone and writing like Claude does - seems like they were working on it with 5.1 but I'd almost classify it as a form of model collapse.
I wince whenever I catch Claudisms on websites and elsewhere. Same as with that pulsating circle that indicates nothing.
It's fascinating, Sonnet 4 is still available via API and it's so much less moronic than the current model. All of the em-dashes and nonsense are a result of the repeated rounds of reinforcement learning using slop data.
going back to opus 4.8 and on is literally like talking to the guy who wants to hear his own voice in meetings. going back to 4.6 is actually refreshing, and it feels so much faster. actually gonna laugh if 4.8 and on is so slow because it's draining lakes fighting for its life trying to conjure up this god forsaken persona.
I understood that reference!
It made it write more like a dev than a marketing agent.
We'll see if they can do it, I originally got into Claude Code because it, at the time, felt more accessible/conversational than Codex/Gemini. Now it's shifted to say the least
All conjecture of course, and yeah it's hard to imagine they would enjoy this prose internally
I noticed the overuse of the word "sharper" or "sharp" in a science paper on ArXiV and my first reaction was "Ewww... AI slop!", but then I checked the date and it was 2020.
It looks like at least some AI-isms stem from the particular style of language commonly used in science papers. Several frontier labs have mentioned heavily weighting those during pre-training because higher quality inputs result in a higher quality model.
> "Let's face it" "terrible writer" "other nonsense" "do Anthropic people actually talk like that" "Dario's engrams"
Be kind. Don't be snarky. Edit out swipes.
> I genuinely wonder if the people inside Anthropic actually communicate with each other like that. Has it been imprinted with Dario's engrams?
Please don't fulminate. Please don't sneer.
Don't be curmudgeonly [...] don't be rigidly or generically negative.
Please don't post shallow dismissals, especially of other people's work.
And as to the substance your comment has:
> Claude (in particular) is a terrible writer.
Frontier models (Claude in particular) are better writers than 90% of the population, even at default style. They're not great, but they're better than that of everyone I know who aren't ultra-educated white-collar workers.
Either your assertion that frontier models are "terrible" writers is false, or you're claiming that 90% of people are "terrible" at writing, which is rather condescending and elitist.
It is evident (in my opinion) as to what the comment was talking about. I personally switched away from all Claude models recently for the same reason.
Which guideline did I violate?
> Leave the policing up to dang and the other mods.
The mods have been very clear that they expect the community to do some self-policing and not rely exclusively on them to do it for them.