undefined

points

[-]

To me it feels like they're basically tweaking these things around the edges. I'm not seeing any difference in capability just preference. This has been the case for a while.

by _heimdall5 hours ago|

parent|

[-]

That makes sense, its seemed to me for a while now the competing product is the harness not the model itself.

by kingkongjaffa11 hours ago|

parent|

prev|

[-]

Most people thought Fable had more 'taste' than Opus, there was certainly a better quality of writing that felt more 'smart human' and not 'stochastic parrot stringing sentences together'.

by clickety_clack4 hours ago|

prev|

[-]

I think that Obama-esque, GMAT essay format is the AI flavor that turns me off AI-written articles. It used to be good writing, but because AI locked onto it as such, it's become the watermark of AI generated content.

by scottyah3 hours ago|

parent|

[-]

Oh boy, people are really going to lean into avoiding proper grammar now.

by Lerc6 hours ago|

prev|

[-]

>2. Review of a SPICE model. Models had different comments, none substantial. Both missed important issues that were simulated inadequately. Clearly a valley where they are undertrained.

When models miss things, there is always the possibility that it has the capability to identify the issues but it is misevaluating the level of analysis that you want it to do. The fine tuning will have them targeting a balance of subjective opinions of what is appropriate. To go beyond broad demographic guessing the model really needs to 'get to know you' to know what it means when you specifically request an action. Without that information about you it has to weigh your words against the level of sophistication it expects a standard user is able to express.

by kgwgk6 hours ago|

parent|

[-]

> has the capability to identify the issues but it is misevaluating the level of analysis that you want it to do.

I guess OP should have told it more explicitly to “find all errors without missing anything.”

by BlobberSnobber6 hours ago|

parent|

[-]

> Thinking. I know this user well, they don't actually want me to find all errors.

> Thinking.. But I found a smoking gun of an error with this SPICE model, maybe I should inform the user.

> Thinking... Hm, but again, I know this human well, they likely don't care about this error. That's absolutely right - it's not an assistant's job to decide this, it's the user's.

by Lerc4 hours ago|

parent|

prev|

[-]

Well if you want it go go off and try and validate the spice simulator and the kernel of the operating system that it's running on then that might be an approach to use.

by 6 hours ago|

parent|

prev|

[-]

deleted

by kolinko4 hours ago|

prev|

[-]

Did you use their native harnesses, or a generic one?

by varjag3 hours ago|

parent|

[-]

Native for both.