v4.4 | minor findings
deriving patterns from the results
this post has been recovered from a place that used image compression. excused. if you want the full experience and see patterns in the benchmark, try using the website yourself:
- click_me
- or download the results as json
some patterns i’ve observed
structure: analysis-header is followed by its picture
cost and score.
word-count: near zero correlation
n1: lexical diversity/jargon
n4: conceptual diversity/ideation
these copy or repeat entire common phrases (top right), making texts conceptually boring to read.
these come up with unique n4 (4 word) groups (bottom left), throwing quirky or novel concepts around.
findings
currently leading with: kimi-k2-0905 > o3 > ling-2.6 > grok 4.3
architecture-overlap
ling-2.6 and kimi-k2, which score very high across the board, follow the pattern of being extremely sparse non-reasoning models, having mla attention and swiglu activation. kimi k2 was historically praised for being extremely good at writing which is probably due to its rl-pipeline.
any “best” model?
o3 is the highest scoring reasoning-model with better word-constraint adherence. for writing tasks that involve constraints but require lexical sophistication, i would score it along kimi k2. in the majority of models, reasoning harder doesn’t do much, which i think is great, considering that models have an inherent bias they cant reason-away.