Am I missing something or the evals do not compare it to the baseline deepseek-v4-flash? Without a baseline comparison, it is hard to tell what works well and what doesn't
Look, if you're going to use AI to write your posts, use it to de-slop them. There are plenty of tools out there, AI can find them and run them on its own output. Reading the same mannerisms over and over and over is painful for your audience and it's literally one command to fix. Please do us the courtesy of running that command, or better yet, making it part of your permanent workflow.
In my view, it's not so much the writing style itself as the lack of 'taste'. Text that clearly seems AI-written has a flat level of exuberance that's just exhausting, kind of like a written version of the 'loudness war'.
Without some kind of dynamic range, I find myself having to do a lot of work to infer what points are truly important versus what are at best interesting implementation details.
Cool idea. I didn't understand what causes the bad response on the first query. Does it mean the first response in every new conversation, or just the first served response after startup?
The content is unreadable. No comparison to the underlying model. Massive text expansion. Hard to tell if the numbers are real or entirely hallucinated SEO slop.
The article is AI-written as well, with all the "honest problems" and "real weak spots". The code must be AI-generated too, so without a human properly verifying it, I can't take the project seriously. It may just be AI hallucinations congratulating themselves on imaginary achievements.
Without some kind of dynamic range, I find myself having to do a lot of work to infer what points are truly important versus what are at best interesting implementation details.