Best LLMs for Writing in 2026

Aggregated benchmark data across EQ-Bench Creative Writing, LMArena Text, and Artificial Analysis — covering 44 models, updated weekly.

Last updated:  ·  44 models tracked  ·  3 tiers: Premium · Mid-Range · Budget

The short answer

Updated August 2026

The best LLM for writing right now is Claude Opus 5 (Anthropic), which tops the EQ-Bench Creative Writing leaderboard ahead of Kimi K3 and GPT-5.6 Sol. For the best quality per dollar, Muse Spark 1.1 (Meta) delivers near-frontier writing at a fraction of the cost. Full ranking across 44 models below.

  1. 1 Claude Opus 5 EQ Creative 2105
  2. 2 Kimi K3 EQ Creative 2060
  3. 3 GPT-5.6 Sol EQ Creative 1959

How we rank models for writing

This ranking combines three independent data sources to give the most complete picture of writing quality across frontier LLMs. No single benchmark captures the full picture — so we aggregate:

Show benchmark details
  • EQ Creative EQ-Bench Creative Writing — specialist benchmark using trained raters to assess narrative quality, emotional depth, prose style, and character voice. Elo scale ~1300–2100. The most relevant signal for marketing copy, long-form content, and creative work.
  • Arena Text LMArena Text — crowd-sourced human preference leaderboard. Broad signal across all text tasks: a model that consistently wins votes is generally pleasant, clear, and useful to read. Elo scale ~1346–1506.
  • EQ General EQ-Bench General — measures emotional intelligence in roleplay scenarios. A proxy for character voice quality and tonal control — useful for brand voice work. Note: high EQ-General does not automatically mean strong creative writing; interpret alongside EQ Creative.
  • Speed Artificial Analysis — median output tokens per second across providers. Matters for iterative draft workflows where waiting costs time. ~75 tokens ≈ 55 words.

Prices are per 1M tokens (input / output) and reflect standard API pricing. A dash (—) means the model has not yet appeared on that leaderboard — never an estimated or interpolated value.

EVY Writing for LinkedIn or X? Four minutes of voice, a week of posts.

The best AI models for copywriting

Updated weekly  ·  Aug 17, 2026
Model EQ Creative Arena Text EQ General Speed Price / 1M
Claude Opus 5 Premium
Anthropic 📋 Consensus
2,104.9 1,493
$5.00 $25.00
Kimi K3 Premium
Moonshot AI 📋 Consensus
2,060.3 1,485
$3.00 $15.00
GPT-5.6 Sol Premium
OpenAI 📋 Consensus
1,959.1 1,481
$5.00 $30.00
Claude Fable 5 Premium
Anthropic 📋 Consensus
1,931.6 1,506 2,049.7
$10.00 $50.00
Muse Spark 1.1 Mid-Range
Meta 📋 Consensus
1,910.8 1,489
$1.25 $4.25
Claude Opus 4.7 Premium
Anthropic 📋 Consensus
1,904.2 1,502 1,884.3
$5.00 $25.00
GPT-5.6 Terra Premium
OpenAI 📋 Consensus
1,846.9 1,464
$2.00 $12.00
GPT-5.5 Premium
OpenAI 📋 Consensus
1,841.4 1,482 1,576.9
$5.00 $30.00
GPT-5.4 Premium
OpenAI 📋 Consensus
1,834.2 1,476 1,562.9
Claude Opus 4.8 Premium
Anthropic 📋 Consensus
1,833.8 1,481 2,029.8
$5.00 $25.00
GPT-5.6 Luna Budget
OpenAI 📋 Consensus
1,825.7 1,450
$0.20 $1.20
Claude Sonnet 4.6 Premium
Anthropic 📋 Consensus
1,803.1 1,472 1,714.1 50 t/s
$3.00 $15.00
Claude Opus 4.6 Premium
Anthropic 📋 Consensus
1,802 1,505 1,717.4 45 t/s
$5.00 $25.00
Claude Sonnet 5 Premium
Anthropic 📋 Consensus
1,786.9 1,462
$3.00 $15.00
GLM-5.2 Mid-Range
Zhipu AI 🧠 EQ-Bench
1,749.5 1,575.1
$1.40 $4.40
Kimi K2.6 Mid-Range
Moonshot AI 📋 Consensus
1,723 1,461 1,561.2
$0.60 $2.50
Gemini 3.7 Flash Mid-Range
Google 📋 Consensus
1,721.6 1,490
$1.50 $7.50
GPT-5.2 Premium
OpenAI 📋 Consensus
1,695.1 1,437 1,558.8
$1.25 $10.00
GPT-5.3 Chat Mid-Range
OpenAI 🧠 EQ-Bench
1,688.6 1,393
Claude Opus 4.5 Premium
Anthropic 📋 Consensus
1,683.1 1,473 1,545.7
$5.00 $25.00
Claude Sonnet 4.5 Premium
Anthropic 📋 Consensus
1,676.3 1,456 1,511
$3.00 $15.00
O3 Mid-Range
OpenAI 📋 Consensus
1,670.9 1,431 1,500
$2.00 $8.00
Kimi K2 Mid-Range
Moonshot AI 🧠 EQ-Bench
1,663.3 1,562.2 44 t/s
$0.55 $2.20
Horizon Alpha Mid-Range
Unknown 🧠 EQ-Bench
1,619.5 1,554.4
Gemini 3.6 Flash Mid-Range
Google 📋 Consensus
1,599.4 1,484
$1.50 $7.50
GLM-5 Budget
Zhipu AI 📋 Consensus
1,597.2 1,457 1,526 80 t/s
$0.80 $2.50
GLM-5.1 Mid-Range
Zhipu AI 📋 Consensus
1,588.9 1,467 1,566.5
Claude Opus 4 Premium
Anthropic 📋 Consensus
1,576.8 1,425 1,399.8
$15.00 $75.00
Kimi K2.5 Mid-Range
Moonshot AI 📋 Consensus
1,575.6 1,450 1,544.7
DeepSeek V4 Pro Budget
DeepSeek 📋 Consensus
1,552.5 1,458 1,570.2
$0.43 $0.87
Gemini 3 Pro Mid-Range
Google 📋 Consensus
1,519.9 1,485 1,559.4 80 t/s
$2.00 $12.00
DeepSeek V3.2 Budget
DeepSeek 📋 Consensus
1,511.2 1,425
$0.28 $0.42
Gemini 3.1 Pro Mid-Range
Google 📋 Consensus
1,488.8 1,486 1,537.7
$2.00 $12.00
Mistral Medium 3 Budget
Mistral AI 🧠 EQ-Bench
1,473.9
$0.40 $2.00
Qwen3-235B Budget
Alibaba 📋 Consensus
1,459 1,375 1,213.6
$0.18 $0.54
GPT-4o Premium
OpenAI 📋 Consensus
1,443 1,346 1,393 185 t/s
$2.50 $10.00
Gemini 2.5 Pro Premium
Google 📋 Consensus
1,418.7 1,445 1,545
$1.25 $10.00
GLM-4.7 Budget
Zhipu AI 📋 Consensus
1,410.4 1,442 1,428.3
$0.38 $1.70
MiniMax M2.5 Budget
MiniMax 📋 Consensus
1,358 1,390 395 t/s
$0.30 $1.20
Gemini 3 Flash Mid-Range
Google 🏟️ Arena
1,473 250 t/s
$0.50 $3.00
Gemini 3.1 Flash-Lite Budget
Google 🏟️ Arena
1,432
$0.25 $1.50
Grok 4.1 Mid-Range
xAI 🏟️ Arena
1,466 163 t/s
$0.20 $0.50
MiMo-V2.5 Budget
Xiaomi 🏟️ Arena
1,434
$0.40 $2.00
GPT-5.1 Premium
OpenAI 📋 Consensus
1,455 1,551
$1.25 $10.00

← Scroll to see all columns →

EQ Creative & EQ General: EQ-Bench  ·  Arena Text: LMArena  ·  Speed: Artificial Analysis  ·  Prices per 1M tokens  ·  — = not yet on leaderboard  ·  Click any row for sources

The right model depends on the task

Benchmark leaderboards rank models globally — but the best model for a 2,000-word thought leadership article is not necessarily the best model for a 15-word social media headline. Here's how the leading models split across common writing tasks:

Narrative & long-form

Thought leadership, case studies, email newsletters, ghostwriting. Requires emotional depth, tonal consistency, and the ability to sustain voice across thousands of words.

Best picks: Claude Opus 5 · Claude Sonnet 4.6

Structured commercial copy

Product descriptions, landing pages, ad copy, LinkedIn posts. Requires clarity, persuasion structure, and format adherence more than creative flair.

Best picks: GPT-5.5 · Claude Sonnet 4.6

High-volume / fast drafts

Social media scheduling, meta descriptions, bulk content variation. Speed and cost matter more than peak quality; fast iteration wins here.

Best picks: Gemini 3 Flash · Grok 4.1 · Kimi K2

Brand voice & consistency

Any content where staying on-brand is non-negotiable. Requires strong instruction-following, tonal control, and memory of brand guidelines.

Best picks: Claude Opus 4.8 · Claude Sonnet 4.6

Managing this by hand means juggling several API keys, pricing tiers, and a decision tree for every task type. The table above shows where each model wins, so you can match the model to the job instead of forcing one model onto everything.

Idea engine

You picked the model. Now you need something to say.

Up to five post ideas reach you every morning, built around your industry and profile. Pick one, talk it through, and it comes back written in your voice. Your phrasing, your rhythm, your point of view.

Join the EVY waitlist → Founder-led marketing took Lovable to $100M ARR.

Frequently asked questions

Which LLM is best for creative writing in 2026?

Anthropic's new Claude Opus 5 leads the EQ-Bench Creative Writing leaderboard at an Elo of 2105, with Kimi K3 (Moonshot AI) second at 2060 and GPT-5.6 Sol (OpenAI) third at 1959 as of August 2026. Claude Opus 5 tops the creative benchmark, but Claude Fable 5 still leads the broader boards: LMArena Text at 1506 and EQ General at 2050, where Claude Opus 4.8 is a close second at 2030. These models excel at narrative quality, emotional depth, and character voice — the core skills that separate great writing from generic AI output.

What is EQ-Bench and why does it matter for writing?

EQ-Bench is an independent benchmark that evaluates large language models on emotional intelligence and narrative quality, using a panel of human raters. Its Creative Writing sub-leaderboard specifically measures story quality, emotional resonance, and prose style — making it the most relevant benchmark for marketing copy, long-form content, and creative work. Scores are on an Elo scale where higher is better, typically ranging from ~1300 to ~2100.

What is LMArena Text and how is it different from EQ-Bench?

LMArena Text (formerly LMSYS Chatbot Arena) measures human preference through head-to-head votes: two anonymous models answer the same prompt, and users pick the better response. It's a broad preference signal across all text tasks, not just writing. EQ-Bench Creative Writing is narrower and more specialist — it specifically evaluates narrative and emotional writing quality with trained raters rather than crowd votes.

Which LLM is the best value for writing tasks?

Meta's Muse Spark 1.1 is the standout value pick: an EQ-Bench Creative score of 1911 at just $1.25 input / $4.25 output per 1M tokens — near-frontier writing for a fraction of the premium-tier price. GLM-5.2 (Zhipu AI) is another strong performance-per-dollar option at 1750 Creative for $1.40 input / $4.40 output — roughly 3.5× cheaper than Claude Opus 4.7 with ~90% of its creative writing performance. OpenAI's GPT-5.6 Luna is the deepest budget pick at 1826 Creative for just $0.20 input / $1.20 output — cheaper than Kimi K2 on both sides of the meter and 160 Elo ahead of it. Kimi K2 by Moonshot AI remains a solid low-cost option at 1663 Creative for $0.55 input / $2.20 output, and Kimi K2.6 scores even higher at 1723 Creative. GLM-5 (Zhipu AI) is another strong value option at $0.80/$2.50 with scores of 1597 EQ Creative and 1526 EQ General.

How often is this ranking updated?

Scores are updated weekly via an automated scraper that fetches the latest data from EQ-Bench and LMArena. Prices are reviewed manually and updated when providers announce changes. The 'Updated weekly' badge in the table header shows the date of the last successful update.

What does 'tokens per second' mean for writing?

Tokens per second (t/s) measures how fast a model outputs text — roughly, 75 tokens equals about 55 words. For writing workflows, speed matters when you need rapid iteration on drafts or real-time dictation-to-copy conversion. MiniMax M2.5 is the fastest tracked model at 395 t/s; Gemini 3 Flash at 250 t/s offers the best speed-to-cost ratio among paid models.

Does the best LLM for writing change depending on the task?

Yes — significantly. Claude Opus 5 tops EQ Creative (2105) with Kimi K3 right behind (2060), while Claude Fable 5 leads EQ General (2050) — making it especially strong for character voice and emotional tone, with Claude Opus 4.8 close behind at 2030. Claude Sonnet 4.6 remains excellent for structured commercial copy where format consistency matters at a lower price. GPT-5.5 is a strong contender for creative work. Faster models like Gemini 3 Flash or Grok 4.1 suit high-volume, lower-stakes content.

Does picking the best LLM guarantee good writing?

No. The model sets the ceiling on quality, but it does not decide how the writing sounds: your voice, your angle, your point of view. Output from even the top-ranked models often reads as generic and gets flagged as AI, which costs reach. The fix is to give the model your voice and a system for shipping. EVY does both: you speak for four minutes, and it returns a week of LinkedIn and X posts written the way you talk, running a combination of the models ranked on this page underneath.