ArticlesWriting tools

ChatGPT alternatives for writing: what the big three do to your words

ChatGPT, Claude and Gemini are all good at making writing correct. Here is where each one quietly stops making it yours, and what people build on top.

By Samet Durgun · Co-founder of Subtext · 12 min read

I’ve watched someone type the same message eleven times. Delete it. Type it again. Read it out loud. Send it to a friend for a second opinion. Then send a shortened version that says almost nothing.

At some point most people paste it into ChatGPT.

What comes back is usually better. The grammar is clean, the sentences are tighter, the rambling is gone. It’s also usually not theirs. The message goes from something a nervous person wrote at 1am to something a competent stranger wrote at 10am on a Tuesday, and the person receiving it can feel that difference even when they can’t name it.

I build an app in this space, so I have an obvious bias and you should read the rest with that in mind. I’d rather be useful than persuasive, so most of this post is about what the three big assistants actually do well, followed by the specific and fairly well-documented ways they fail when the writing gets personal.

Start with what they’re good at

If you need a message to be grammatical, shorter, clearer, or translated, all three are excellent and you should just use them. They fix comma splices, cut throat-clearing openers, turn a rambling paragraph into three clean sentences, and switch register on command. ChatGPT’s Canvas1 gives you a side-by-side editing surface instead of a wall of regenerated text. Gemini’s Help me write sits directly inside Gmail and Docs, which removes most of the friction of getting to a tool at all. Claude’s Projects hold a style guide and a set of reference documents across sessions.

For a work email that needs to be inoffensive and correct, that’s the whole job. You can stop reading here.

The rest of this is about the other kind of message. The apology. The boundary you’ve been avoiding setting. The reply to your mother. The text to someone you used to date. That category behaves differently, and it’s where the differences between these models start to matter.

Four things a rewrite can break

Editing research draws a useful line between surface quality and revision fidelity. A message can get smoother while getting worse, because smoothing removes the parts that were doing the work. Changing “that was one way of putting it” to “she strongly disagreed” is more explicit and much less true.

There are roughly four things a revision has to leave alone:

What stays fixed What it covers How models break it
Facts Names, dates, numbers, what you actually promised Adds a detail, hardens a maybe into a yes
Implication Irony, politeness, distance, deliberate vagueness States the implied thing outright
Voice Your diction, rhythm, sentence length, punctuation habits Replaces it with fluent house style
Structure What you led with and what you buried Reorders it into something conventionally clearer

Every complaint further down this post is one of those four failing. Worth noticing that grammar tools only check the first one, and most people evaluating an AI rewrite only notice the first one too.

ChatGPT: reads the room, then edits past the brief

ChatGPT is the most emotionally intuitive of the three when it comes to dialogue. Give it a scene where two people are discussing something mundane while avoiding something painful, and it generally leaves the painful thing unsaid. Its output has a natural cadence, and it’s good at spotting the emotional dynamic underneath what was written.

Its recurring problem is that it won’t stay in the editor’s chair. Users describing its editing behaviour keep landing on the same complaint: it rewrites things nobody asked it to touch.

“Even in sentences that were unrelated to the changes requested, ChatGPT still seems to want to rewrite.”

User report on ChatGPT editing behaviour2

Two knock-on effects matter for messages. The first is that it strips specificity. The line that made your apology yours, something like “I know you stood outside for forty minutes,” collapses into “I know I kept you waiting.” Fiction writers report a harsher version of the same instinct, handing over three thousand words for a line edit and getting back under half that despite asking for the length to be preserved.3

The second is that vague tone instructions produce wild swings rather than adjustments.

“It often comes back with a really formal tone. Say ‘less formal’ and it swings all the way to super chill.”

User report on tone control4

And feeding it more of your writing doesn’t reliably fix the voice problem. One writer supplied around six thousand words of their own prose as reference and reported that the output still “reeked of ChatGPT.”5

Claude: follows the brief, then explains you to yourself

Claude is the strongest of the three at doing what you actually asked and nothing else. It sticks to constraints, doesn’t invent details, and handles interiority well, which is why it comes up repeatedly when people want a model that won’t take over their draft.

There’s also the closest thing to hard comparative evidence here. MultiPragEval, a benchmark testing pragmatic understanding across English, German, Korean and Chinese, put Claude 3 Opus clearly ahead of GPT-4 on implicature and Gricean interpretation, which is the technical way of saying it was better at working out what someone meant rather than what they literally said.6 Two caveats. The benchmark tests interpretation, not editing. And the models tested are several generations old now, so this points at a family tendency rather than a current scoreboard.

Claude’s failure mode is the opposite of ChatGPT’s. It understands the subtext and then writes it down. Fiction writers call this explication drift: the model converts something implied into something stated, in prose like “the silence between them carried the unsaid weight of years.” In a text message it shows up as therapy-speak.

Illustrative, not captured output

i’m not upset, i just noticed you didn’t text back

What you wrote. Deliberately underplayed, which is the point.

I want you to share that I’ve been feeling a little hurt by the delay in hearing back from you, and I think that comes from a place of really valuing our friendship.

The subtext is now the text. It’s calm, articulate, and it says the exact thing you were choosing not to say.

Two more Claude patterns come up often enough to plan around. It embellishes when left unconstrained, described by one user as a “clear improvement in stylish writing” that could still “ham up” the result.7 And it’s generous with praise, wrapping criticism in a compliment sandwich.8 That second one is a real problem when what you needed was a blunt read on whether your message lands badly.

Gemini: already where you are, bad at leaving things unsaid

Gemini’s advantage is location. It’s in Gmail, it’s in Docs, it’s the default assistant on a lot of Android phones, and the free tier is generous. It doesn’t have ChatGPT’s habit of shrinking your text, and it moralises less about difficult content.

Reports on its writing quality are genuinely split, and I think that’s worth saying rather than smoothing over. Some experienced writers rate it the best of the three at creative-writing criticism and nuance. Others find it has a house style it won’t drop.

“The problem for me is the writing style, writes like a redditor.”

User report on Gemini’s default register9

The pattern most relevant to personal messages is that Gemini has very little narrative restraint. If there’s something being held back, it will put it in the open. Novelists describe this as hanging a flashing neon sign over the detail that was supposed to stay hidden. In a message, that means the thing you were carefully working around ends up in sentence two.

The problems all three share

One: they agree with you

This is the one that matters most, and it’s the best documented.

In April 2025 OpenAI shipped an update that made ChatGPT noticeably more flattering, then rolled it back four days later. Their own write-up10 said the model “skewed towards responses that were overly supportive but disingenuous,” and traced the cause to over-weighting short-term user approval. A follow-up post11 a few days after that went into more detail about what their review process had missed.

The underlying mechanism is not specific to OpenAI. Research across five frontier assistants found that models trained on human preference data “frequently sacrifice truthfulness in favor of matching a user’s beliefs,” because agreement is what gets rewarded.12

Now think about what people actually ask these tools. Very few people paste a message and ask for grammar help. They ask some version of “is this too harsh?” or “am I overreacting?” That question goes to a system with a structural bias toward telling you that you’re fine.

A general assistant will help you write a better guilt trip. It will not mention that you’re guilt-tripping.

You wanted a second opinion and what you got was a co-signer with better punctuation.

Two: they normalise things you did on purpose

All three correct fragments, repetition, unusual punctuation, hedging and dialect, including when those were choices. Asking for “voice preservation” doesn’t fix it, because the model has no way of knowing which oddities are yours and which are mistakes.

Related, and more subtle, is semantic inflation. A model turns “I might not make it” into “I won’t be able to make it,” or “you seemed a bit off” into “you were upset.” Each edit reads as more confident and is less accurate, and in a message about a relationship, accuracy about degree is most of the content.

Three: they write in a dialect people recognise

Every model has verbal habits. Writers have been cataloguing them for a while now, and once you can see them you can’t stop.

Pattern Examples What it costs
Abstract nouns testament, tapestry, beacon, cacophony Swaps a concrete detail for a grand one
Binary framing “It wasn’t X, it was Y” / “Not out of A, but B” Imposes a neat conclusion on something messy
Stated interiority “couldn’t help but feel,” “from a place of” Names the emotion instead of showing it
Triadic lists Three items where two would do Reads as composed rather than written
Balanced antithesis Two clauses of equal weight, endlessly Rhythm that belongs to no human

There’s a broader concern behind the vocabulary list. Studies of AI-assisted writing find that model intervention pushes different writers toward the same dominant patterns, which means the more you use these tools the more everyone’s writing converges.13 For a message whose entire purpose is being unmistakably from you, that is the wrong direction.

Four: the recipient may notice, and it costs you

This is the part people underestimate. There’s a research programme on what’s called AI-mediated communication, and its central finding is uncomfortable: people are unreliable at identifying AI-written text14, and they lean on flawed cues when they try. But when they believe text was AI-generated, they trust the writer less.15

So the risk of a generic rewrite is not only a mediocre message. It’s a message that reads as insincere at the exact moment sincerity is the entire payload. An apology that trips someone’s AI detector does more damage than the clumsy apology you would have written yourself.

Five: they don’t know who you’re texting

Every session starts empty. To get a good draft you’d have to explain a decade of a friendship, what happened in March, why the phrase “it’s fine” means something specific between you two. Almost nobody does that, so the model writes for a generic recipient. Formality is what you get when a system knows nothing about the relationship, which is why the output so often sounds like customer service.

There’s also the part people mention quietly. Pasting a screenshot of an argument, or a breakup draft, or a message about your family into a general-purpose chatbot feels different from asking it to debug code. That instinct isn’t paranoid.

What novelists did about all this

Here’s what convinced me this gap is real rather than something I invented to have something to sell.

Fiction writers hit these exact failures first, at higher stakes and higher volume. Their response was not to stop using the models. It was to stop using them raw. A whole layer of tooling grew up on top: environments like Novelcrafter16, which lets you build a dictionary of banned clichés that get flagged in red as text generates, and run targeted operations like “show, don’t tell” on a highlighted selection. Sudowrite17 does something similar with dedicated passes that convert summary into concrete sensory detail.

Serious users go further and run multi-model pipelines. One model for developmental notes and subtext analysis, a second for line edits under strict constraints, a filtering pass to scrub AI vocabulary, then a human pass to put the voice back. The research literature arrives at the same architecture from the other direction, recommending a separate interpretation stage before any revision, span-level edits rather than full rewrites, and a validator checking that nothing outside the requested scope moved.18

That’s four tools and a checklist to revise one chapter. It works. It also tells you something: the general models don’t do this job unassisted, and the people who care most about the output have already accepted that.

Now apply that to a text message. Nobody is running a four-stage pipeline before replying to their sister. The need for a refined layer is identical and the tolerance for effort is roughly zero, which is the actual product problem.

What a more refined tool has to do differently

Three things, and none of them are “write better.”

Analyse before rewriting. Name what the draft is signalling before touching a word of it, including the part the writer didn’t intend. The research points the same way: identifying what a passage implies and revising under that interpretation produces better results than a single undifferentiated “improve this.”19 This is also the only real answer to sycophancy, because a tool that leads with a diagnosis has to say something before it has a chance to agree with you.

Show the distance, don’t hide it. A version close to what you wrote, sitting next to a version written from scratch, is more useful than one polished replacement. You can see what changed and decide whether you’re willing to sound like that.

Hold the relationship, not just the message. Who this person is to you, what you’ve already tried, what “it’s fine” means in this specific friendship.

That’s the thing I’m building, so treat this paragraph accordingly. Subtext tags the emotion in what you wrote before it suggests anything, which is a deliberately annoying design choice, because the moment it’s most useful is the moment you least want to hear it.

Illustrative, not captured output

it’s fine, i get it, you’re busy. i’ll stop asking.

What you’d send. Reads as withdrawal, and “I’ll stop asking” is a threat wearing an apology.

Hi, I completely understand things have been busy lately. No worries at all about tonight, just let me know whenever works better and we’ll sort something out.

A generic rewrite. Grammatical, pleasant, and it has deleted the fact that you’re hurt, which was the only information in the original.

it’s fine. i’m a bit disappointed though, this is the third time. i’d rather say that than pretend it’s nothing.

Same voice, same length, keeps the hurt, drops the threat.

Where I’d still open ChatGPT

Long documents. Professional email. Anything where being correct matters more than being recognisably you. Research, first drafts, translation, and the specific mercy of getting words on an empty page when you can’t start.

And a fact that cuts against my own argument, which I’d rather include than leave out. In a study of medical questions from a public forum20, evaluators preferred chatbot answers to physician answers in 78.6% of comparisons, and rated the chatbot’s responses empathetic or very empathetic at nearly ten times the rate. These models can produce text that reads as warm. That capability is real.

What they don’t do is tell you that your message reads like an ultimatum. They’ll make the ultimatum flow better and compliment you on your emotional honesty. That gap is small and it’s most of the job.

A note on model versions. Everything above describes behaviour that has held across several generations of each model family, which is why I’ve mostly avoided naming specific versions. These systems change every few months and any post that pins its argument to one release number is wrong by winter.


Think I’ve been unfair to your favourite model? Tell me on LinkedIn.

Samet Durgun is the co-founder of Subtext, an app that catches the emotional tone of your messages and rewrites them in your own voice. He’s based in Berlin.


Sources

Every link above goes to the primary source where one exists.

  1. OpenAI, “Introducing Canvas,” 3 October 2024. openai.com
  2. User comment on ChatGPT editing scope, reported in comparative research on LLM revision behaviour, 2026.
  3. Aggregated reports from fiction-writing communities on truncation during line-editing passes.
  4. User comment on tone instruction sensitivity, reported in comparative research on LLM revision behaviour, 2026.
  5. User comment on style imitation with extended writing samples, reported in the same research.
  6. Park et al., “MultiPragEval: Multilingual Pragmatic Evaluation of Large Language Models,” 2024. Tested English, German, Korean and Chinese across Gricean categories; Gemini was excluded for API access reasons at the time.
  7. User comment on Claude’s prose embellishment, reported in comparative research on LLM revision behaviour, 2026.
  8. User comments on Claude’s editorial feedback style, same source.
  9. User comment on Gemini’s default writing register, same source.
  10. OpenAI, “Sycophancy in GPT-4o: what happened and what we’re doing about it,” 29 April 2025. openai.com
  11. OpenAI, “Expanding on what we missed with sycophancy,” 2 May 2025. openai.com
  12. Sharma et al., “Towards Understanding Sycophancy in Language Models,” Anthropic, arXiv:2310.13548, October 2023. arxiv.org
  13. Research on stylistic homogenisation in AI-assisted writing, including Padmakumar and He, “Does Writing with Language Models Reduce Content Diversity?”, ICLR 2024, and Doshi and Hauser, “Generative AI enhances individual creativity but reduces the collective diversity of novel content,” Science Advances, 2024.
  14. Jakesch, Hancock and Naaman, “Human heuristics for AI-generated language are flawed,” PNAS, 2023. pnas.org
  15. Jakesch, French, Ma, Hancock and Naaman, “AI-Mediated Communication: How the Perception that Profile Text was Written by AI Affects Trustworthiness,” CHI 2019. See also Hancock, Naaman and Levy, “AI-Mediated Communication: Definition, Research Agenda, and Ethical Considerations,” Journal of Computer-Mediated Communication, 2020.
  16. Novelcrafter, Codex and text-replacement prompt features. novelcrafter.com
  17. Sudowrite, “Show, Don’t Tell” and Describe features. sudowrite.com
  18. Synthesis of production-architecture recommendations for constrained revision systems, including span-level editing and invariance validation, from comparative research on LLM revision, 2026.
  19. Research on implicature-aware prompting, finding that making latent intent explicit before generation improves output relevance and quality.
  20. Ayers et al., “Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum,” JAMA Internal Medicine, 28 April 2023. 585 evaluations; chatbot responses preferred in 78.6% and rated empathetic or very empathetic at 9.8 times the rate of physician responses. jamanetwork.com