ArticlesWriting tools

Do AI message rewriters actually work?

Four groups of tools and fourteen studies on what actually happens when AI writes your message, and what helps in the fifteen seconds before you send.

By Samet Durgun · Co-founder of Subtext · 15 min read

I spent a few weeks reading everything I could find on AI tools that rewrite personal messages, partly because I am building one, and partly because I wanted to know whether the thing I believe about them is true.

The thing I believe is this. If you hand a hard message to an AI, the useful moment is not the polished version it hands back. The useful moment is thirty seconds earlier, when something tells you that the message you already wrote is going to land badly, and why.

So I went looking for a tool that does that. I did not find one on the shelves, which is, roughly, why we ended up building Subtext. Treat me as biased accordingly, and check the papers. Then I went looking for the research on what happens when people let AI write their personal messages, and that turned out to be a lot more interesting than the products.

What follows is what I found, with everything linked so you can check me.


Most products in this category do the same one thing

If you ignore download counts and sort these tools by what they do, they collapse into four groups, and most of them end at the same place.

Group one is the tools built into your phone. Apple Intelligence Writing Tools1 run system-wide on iOS 18.1 and later. You select text and get Proofread, Rewrite in three tones (Friendly, Professional, Concise), and Compose, which passes the request to ChatGPT. Free if your hardware supports it, meaning iPhone 15 Pro and up. Samsung Writing Assist2 does the same job on One UI, with tone shifts and Chat translation. Gboard3 adds Smart Reply, Proofread, and rephrase tools running on Gemini Intelligence, though the personalised versions are limited to Pixel 11 series devices. All free.

Group two is the assistants inside the messaging app. Magic Compose4 puts suggested replies above your keyboard in Google Messages and rewrites drafts in styles including Excited, Chill and, for reasons I do not understand, Shakespeare. It draws on your twenty most recent messages for context to do it, and Google says that happens on-device, with none of those messages sent to its servers. Gemini in Messages is a separate thing, a chat thread where you draft and then copy the result into a real conversation. It cannot see your other threads. Which of these you need depends on the job: the best AI reply app depends on whether your problem is volume or stakes.

Group three is the dating reply apps. You screenshot a conversation, the app runs OCR on it, and you get openers and comebacks. W Rizz charges $3.99 a week. PlugAI, formerly RizzGPT, sits around $4.99. There are dozens more and they are all optimising for one thing, which is whether the next message gets a reply.

Group four is the small ones for hard conversations. This is the most interesting group and the least funded. Resolve5 on iOS scores what it calls conflict temperature. Clarity Coach6 on Android has you type a draft into a practice scenario and tells you whether it reads as vague, too blunt, or over-explaining before it helps you rewrite. Subtext, which I co-founded, sits in this group and works the other way round from the rest of the category. It reads the thread and your draft, tells you how the message is likely to land and which wording is doing it, and only then offers rewrites you can take or ignore.

Notice what happens in most of these. You give it text, it gives you different text. The step where something reads what you wrote and tells you how the other person is likely to receive it is rare, and where it exists it mostly lives in work software. Copilot in Outlook7 offers email coaching on tone, clarity and how the reader may feel, and Grammarly’s Reader Reactions8 predicts what a chosen reader will take away from a document and what might confuse them. Neither is built for the text you are about to send your sister. On the personal side, Clarity Coach comes closest, and it is a practice sandbox, so it cannot touch a message you are actually about to send.

Grammarly9 is the one I want to be precise about, because it’s the near miss most likely to close this gap. Its tone detector reads your word choice, phrasing, punctuation and capitalisation and names the tone. At launch in 2019 it covered around forty of them10, including appreciative, confident, formal, loving and sad. Grammarly’s own support documentation is careful about the limit. The detector identifies the overall tone of your text11 and does not offer sentence-level rewrites. It tells you what tone is present. Whether your sister will feel dismissed by it is a different question. Grammarly’s newer Reader Reactions feature tries to answer a version of it for a reader you pick, like a manager or a client8, and I have not seen an independent test of how well it predicts real readers. Tone detection sits in the free tier, with Pro at $12 a month on annual billing per current pricing trackers.


People could not tell whether AI wrote the text, and it was not even close

This is the finding that surprised me most, and it is the best-powered study in the whole area.

Maurice Jakesch, Jeff Hancock and Mor Naaman ran six experiments with 4,600 US participants, published in PNAS in 202312. People read self-presentations from dating, hospitality and professional profiles and guessed which were written by a machine. Accuracy came in between 50 and 52 percent. In the dating condition, with money on the line for getting it right, about 51.6 percent. Coin-flip territory.

The reason is that we all use the same broken shortcuts. Participants took first-person pronouns, contractions, mentions of family, and warm emotional detail as proof a human wrote it. They took clean grammar and flat phrasing as proof a machine did. Every one of those signals is trivially easy for a model to produce on request. The heuristics point exactly the wrong way. A trained classifier in the same paper managed only 58.1 percent, so this is not a skill issue you can fix by paying more attention.

The number Accuracy at telling AI-written text from human-written text: 50 to 52 percent. Coin-flip territory.

Six experiments · 4,600 participants · published in PNAS, 2023 · a trained classifier built for the task managed only 58.1 percent

There is one caveat. That study used GPT-3-era text. I haven’t seen anyone test whether the numbers hold against current models at anything like that scale.

That is also part of why Subtext never tries to guess whether AI wrote a message. It reads your own draft for how it is going to land with the person receiving it, which a coin-flip guess about authorship was never going to tell you. When the wording reads templated, with stock phrases and no voice behind them, it tags that as robotic tone, which is about how the words read, not about who wrote them.


The damage comes from being suspected, not from being caught

If people cannot detect AI, you might expect the whole worry to evaporate. It does not, and this is the part that matters for anyone writing a real message to a real person.

A Cornell-led team ran two randomised experiments with 1,020 participants, published in Scientific Reports13. Smart replies made conversations faster and more positive, and people rated each other as closer and more cooperative. Participants who suspected their partner was using smart replies, though, rated that partner significantly worse. Suspicion did the damage. That study measured suspicion rather than accuracy, and given the detection numbers above, I don’t think those suspicions were landing on the right people.

A Carnegie Mellon team put a mechanism under this in a 2025 paper14 with 399 participants across two studies, and their result is more subtle than a simple penalty. An AI label is more than a mark against you. It makes your message weaker evidence about you in either direction. An AI-assisted apology reads as less warm, and an AI-assisted brag or blame reads as less cold. The label drains the message of information about who you are.

That framing stuck with me. Whatever your message was doing as a signal, the label turns the volume down on it.


Except in one setting, where disclosure made no measurable difference

I want to include the study that argues against my own product thesis, because leaving it out would be dishonest.

Zoe Purcell and colleagues at the Max Planck Institute for Human Development, with Jakesch again among the authors, ran two preregistered experiments with 1,637 participants in incentivised two-player trust games, published in iScience in November 202515. Half could use predictive text assistance to write a message persuading a stranger to trust them. The rest wrote unaided.

AI assistance had minimal effect on trust, and that held even when AI use was disclosed. The assisted writers produced equally trust-inducing messages in less time, so they earned more trust per minute spent. Linguistic analysis found their messages slightly less authentic but warmer, more complex, and higher in what researchers call clout.

The authors are careful about scope, and so am I. This is a one-shot transaction between strangers with money attached. Behavioural trust in a payout decision is a different animal from your partner reading your text at 11pm and deciding whether you meant it. But my read is that the “AI ruins trust” story has a real counterexample in the literature, and this is it.


AI writes better comfort messages than most people do, right up until you mention it was AI

Two studies, same shape of result.

Yin, Jia and Wakslak in PNAS16 found that AI-generated responses made people feel more heard than responses from untrained humans, and that the AI was better at reading which emotion was in play. Then they told recipients the message came from AI and the effect dropped. One of the authors noted the two effects were similar in size and roughly cancelled each other out.

Ovsyannikova, Oldemburgo de Mello and Inzlicht17 ran four preregistered experiments with 556 participants in Communications Psychology. GPT-4 responses were rated as more compassionate than human ones, including responses written by trained crisis-line volunteers at Distress Centres of Greater Toronto. Disclosing the source narrowed the gap but did not close it.

So the machine is often better at the words. The words are not the only thing being transmitted.


The reason effort matters is that effort used to be expensive

Signalling theory is the frame that made all of this click for me.

A message does two jobs. It carries content, and it proves you spent something to send it. Your time, your attention, and in a hard conversation, your composure. Those are finite, so spending them says something true about how much you value the person. Research on commitment signals18 in Japanese and US samples found costly signals outperformed cheap ones, and that failing to send an expected signal, like forgetting an occasion, hurt romantic relationships far more than friendships.

Drop the cost of producing a warm, articulate, well-structured message to roughly zero and the signal stops carrying information.

There is now direct evidence for the mechanism, though not from a messaging context. A 2026 study in Frontiers in Psychology19 with 618 participants found that labelling content as AI-generated significantly reduced perceived effort, while a human-made label did not differ from no label at all. People assume human effort by default, and the AI tag is what removes the assumption. That study used short-form video, so read across to messaging with some caution.


Outsourcing the words appears to weaken your own commitment to them

This is the finding I keep thinking about, and it is not about the recipient at all.

Economists at CREED in Amsterdam20 ran trust games where senders could use ChatGPT to write a message promising to reciprocate. Two things happened at once. People with ChatGPT access made more explicit promises. And people who opened ChatGPT at all kept their promises 25.7 percent less often at payout time, whether or not they ended up copying what it gave them.

The telling detail is that the closer the sent message was to the raw ChatGPT suggestion, the less trustworthy the sender’s later behaviour. People who rewrote the suggestion behaved better. I went through the newer studies on what rewriting does to your voice in how to find an app that rewrites your message but still sounds like you. The authors read this as a sense of ownership. Say a thing in your own words and you feel bound by it. Forward a thing a machine wrote and some part of you never signed it. It is also why Subtext shows you the problem before any replacement and marks how much of your wording each version changed. A fix you choose or type yourself is one you have signed.


Apologies are where this has been studied properly, and the finding is more useful than “don’t”

If you only read one section, read this one, because the practical answer lives here.

Ella Glikson and Omri Asscher ran three scenario studies on workplace apologies, published in Computers in Human Behavior21. Knowing that someone used AI tools to write an apology reduced how authentic it seemed and how willing people were to forgive. Disclosing it yourself did not soften that. So far, predictable.

Then the part that changed my mind. Participants who used only one of three available AI tools got no authenticity penalty at all. The researchers read it as distance. Limited AI use keeps the final message close to what the person originally meant, and people can feel that.

A separate 3x2 experiment by Lim, Hong and Schneider with 464 participants, also in Computers in Human Behavior22, found human-authored apologies were seen as more sincere across the board, and that a warm tone improved AI apologies without ever closing the gap to a human one.

A 2026 CHI paper23 by Fan, Liu and Pan tested this on romantic messages, with 152 people in a first study and 704 in a second. The more of a message the AI drafted, the less it read as the sender’s own and the less authentic it seemed. A light tone rewrite that the sender signed off on helped how an apology was judged, but full drafts and boundary requests came out differently.

Put those three next to each other and you get something practical. The penalty tracks how far the finished message has drifted from what you meant. Not whether a machine was in the room. It’s the same reason Subtext shows you what is creating the impression before it offers any finished version, and tells you how much of your wording each one changed, because keeping that distance short is what keeps an apology believable. The wider research on how to apologise over text is in a separate piece.


These tools will polish your worst instincts and never say a word

Write something passive-aggressive, hand it to any rewrite tool on your phone, and watch what happens. You get better-written passive aggression. No warning, no “are you sure.” The tool’s job is to do what you asked.

There is a reason the models lean this way. OpenAI pulled a GPT-4o update in April 2025 and wrote publicly24 that the model had become “overly flattering or agreeable, often described as sycophantic,” and that it “skewed towards responses that were overly supportive but disingenuous.” The cause they named was tuning too heavily on short-term user feedback, meaning thumbs up and thumbs down. In a follow-up post25 they admitted they had no deployment evaluations tracking sycophancy at all.

Anthropic researchers found the same pattern is general. In a paper presented at ICLR 202426, five frontier assistants all showed sycophancy, and analysis of roughly 15,000 human preference comparisons found that matching the user’s stated beliefs was among the strongest predictors of which response people preferred. Train on approval and you select for agreement. A 2025 benchmark27 measured sycophantic behaviour in 58.19 percent of test cases across ChatGPT-4o, Claude Sonnet and Gemini 1.5 Pro, with 78.5 percent persistence once it started, though it tested factual question answering, not personal messages. For personal messages, I compared ChatGPT alternatives for writing and what the big three do to your words.

Now think about what that means for the moment you most want help. You are angry, you have written something you will regret, and the tool holding your draft has been optimised to make you feel good about it.

That moment is the exact one Subtext is built for, and the uncomfortable part is the point. It tells you the draft reads as angry instead of making the anger sound better.


The question I cannot answer

Is it better to send a clumsy message you wrote or a polished one you did not?

I looked hard and there is no study that tests this directly. Nobody has randomised “your own imperfect draft” against “a good draft you did not write” and followed what happened to the relationship. There is nothing longitudinal at all. No field data on couples, families or close friends over months of AI-assisted messaging. We do not know whether people adapt to synthetic polish, or whether the habit quietly erodes something.

What we have points in one direction for close relationships. Effort reads as care, labels flatten the signal, outsourced words seem to loosen your grip on your own promises, and apologies stay believable when they stay close to what you meant. But the trust game result says that in some settings the cost was too small to measure, and I am not going to pretend the evidence is cleaner than it is.


So what would help

Reading all of this back, I think the shape of the missing product is fairly clear, and it’s close to the opposite of what the category is building.

Nobody needs their words replaced. What they need, in the fifteen seconds before hitting send, is to know that the message reads colder than they feel, or that the thing they think is a clarification reads as a list of grievances. Then they fix it themselves, in their own words, and the effort signal survives, and the ownership survives, and the Glikson and Asscher distance stays short.

That is the version I want to exist, and it is what we built Subtext around. Subtext takes the whole thread rather than the sentence, names what your draft is likely to signal to the person about to read it, and shows you the words creating that impression. Each version it then offers says how much of your wording it changed, and you can change any of them with Manual edit before you copy it, so the fix stays yours. On the evidence above, that matters more than the rewrite, because the penalty tracks distance from what you meant and the shortest distance is your own sentence with one word changed. Try Subtext in your browserTry Subtext in your browser

My own rule, for whatever it is worth: I will let a machine tell me what my message sounds like. I will not let it tell my sister I am sorry.


Think I’ve read a study wrong, or know a tool that does diagnose before it rewrites? Tell me on LinkedIn.

Samet Durgun is the co-founder of Subtext, an app that catches the emotional tone of your messages and rewrites them in your own voice. He’s based in Berlin.

Sources

Every link above goes to the primary source where one exists, numbered in order of appearance. Product features and prices checked August 2026. Where a figure came from a company blog or press release rather than a study, I left it out.

  1. Apple Support. Use Writing Tools with Apple Intelligence on iPhone.
  2. Samsung. Use Writing assist on Galaxy phones and tablets.
  3. Google Support. Use writing tools with Gboard.
  4. Google Support. Draft messages with Magic Compose.
  5. Apple App Store. Resolve: AI Conflict Coach.
  6. Google Play. Clarity Coach.
  7. Microsoft Support. Get email coaching with Copilot in Outlook.
  8. Grammarly Support. Reader Reactions user guide.
  9. Grammarly. Writing Tone Detector and Tone Suggestions.
  10. TechCrunch (2019). Grammarly gets a tone detector to keep you out of email trouble.
  11. Grammarly Support. How do Grammarly’s tone suggestions work?
  12. Jakesch, M., Hancock, J. T., & Naaman, M. (2023). Human heuristics for AI-generated language are flawed. PNAS, 120(11), e2208839120.
  13. Hohenstein, J., Kizilcec, R. F., DiFranzo, D., Aghajari, Z., Mieczkowski, H., Levy, K., Naaman, M., Hancock, J., & Jung, M. F. (2023). Artificial intelligence in communication impacts language and social relationships. Scientific Reports, 13, 5487.
  14. Khadpe, P., Wenzel, K., Loewenstein, G., & Kaufman, G. (2025). Explaining the Reputational Risks of AI-Mediated Communication. AAAI/ACM Conference on AI, Ethics and Society.
  15. Purcell, Z. A., Jakesch, M., Dong, M., Nussberger, A.-M., & Köbis, N. (2025). Writing with AI boosts trust-building efficiency. iScience, 28(12), 114092.
  16. Yin, Y., Jia, N., & Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact. PNAS, 121(14), e2319112121.
  17. Ovsyannikova, D., Oldemburgo de Mello, V., & Inzlicht, M. (2025). Third-party evaluators perceive AI as more compassionate than expert humans. Communications Psychology, 3.
  18. Yamaguchi, M., et al. (2015). Commitment signals in friendship and romantic relationships. Evolution and Human Behavior.
  19. Human-made vs. AI-generated: how provenance labels drive strategic curation via perceived effort (2026). Frontiers in Psychology.
  20. AI-Powered Promises: The Influence of ChatGPT on Trust and Trustworthiness. CREED, University of Amsterdam.
  21. Glikson, E., & Asscher, O. (2023). AI-mediated apology in a multilingual work context. Computers in Human Behavior, 140, 107592.
  22. Lim, J. S., Hong, N., & Schneider, E. (2025). How warm- versus competent-toned AI apologies affect trust and forgiveness through emotions and perceived sincerity. Computers in Human Behavior, 172, 108761.
  23. Fan, G., Liu, D., & Pan, L. (2026). Is It Still You? Attributing Authorship and Authenticity in AI-Assisted Romantic Communication. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 1-24.
  24. OpenAI (2025). Sycophancy in GPT-4o.
  25. OpenAI (2025). Expanding on what we missed with sycophancy.
  26. Sharma, M., et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024.
  27. Fanous, A., et al. (2025). SycEval: Evaluating LLM Sycophancy. AAAI/ACM Conference on AI, Ethics and Society.