ArticlesWriting tools
Every AI message tool rewrites your draft. Not one tells you how it will land.
Four groups of tools and twelve studies on what actually happens when AI writes your message, and what helps in the fifteen seconds before you send.
By Samet Durgun · Co-founder of Subtext · 12 min read
I spent a few weeks reading everything I could find on AI tools that rewrite personal messages, partly because I am building one, and partly because I wanted to know whether the thing I believe about them is actually true.
The thing I believe is this. If you hand a hard message to an AI, the useful moment is not the polished version it hands back. The useful moment is thirty seconds earlier, when something tells you that the message you already wrote is going to land badly, and why.
So I went looking for a tool that does that. I did not find one on the shelves, which is the short version of why we ended up building Subtext. Treat me as biased accordingly, and check the papers. Then I went looking for the research on what actually happens when people let AI write their personal messages, and that turned out to be a lot more interesting than the products.
What follows is what I found, with everything linked so you can check me.
Every product in this category does the same one thing
If you sort these tools by what they actually do rather than by how many downloads they have, they collapse into four groups, and all four end at the same place.
Group one: the tools built into your phone. Apple Intelligence Writing Tools1 run system-wide on iOS 18.1 and later. You select text and get Proofread, Rewrite in three tones (Friendly, Professional, Concise), and Compose, which passes the request to ChatGPT. Free if your hardware supports it, meaning iPhone 15 Pro and up. Samsung Writing Assist2 does the same job on One UI, with tone shifts, suggested replies and Live Translate. Gboard3 adds Smart Reply, Proofread, and rephrase tools running on Gemini Nano on newer devices. All free.
Group two: assistants inside the messaging app. Magic Compose4 puts suggested replies above your keyboard in Google Messages and rewrites drafts in styles including Excited, Chill and, for reasons I do not understand, Shakespeare. It ships up to twenty of your recent messages to Google’s servers to do it, which means those messages are not end-to-end encrypted anymore. Google says it does not store them. Gemini in Messages is a separate thing, a chat thread where you draft and then copy the result into a real conversation. It cannot see your other threads.
Group three: the dating reply apps. You screenshot a conversation, the app runs OCR on it, and you get openers and comebacks. W Rizz charges $3.99 a week. PlugAI, formerly RizzGPT, sits around $4.99. There are dozens more and they are all optimising for one thing, which is whether the next message gets a reply.
Group four: the small ones for hard conversations. This is the most interesting group and the least funded. Resolve5 on iOS scores what it calls conflict temperature. Clarity Coach6 on Android has you type a draft into a practice scenario and tells you whether it reads as vague, too blunt, or over-explaining before it helps you rewrite. Subtext, which I co-founded, sits in this group and works the other way round from the rest of the category: it reads the thread and your draft, tells you how the message is likely to land and which wording is doing it, and only then offers rewrites you can take or ignore.
Notice what happens in every single one of these. You give it text, it gives you different text. The diagnosis step, where something reads what you wrote and tells you how the other person is likely to receive it, does not exist in any tool that works inside a real conversation. Clarity Coach comes closest and it is a practice sandbox, so it cannot touch a message you are actually about to send.
Grammarly7 is the near miss worth being precise about, because it is the one most likely to close this gap. Its tone detector reads your word choice, phrasing, punctuation and capitalisation and names the tone. At launch in 2019 it covered around forty of them8, including appreciative, confident, formal, loving and sad. Grammarly’s own support documentation is careful about the limit: the detector identifies the overall tone of your text9 rather than offering sentence-level rewrites. It tells you what tone is present. Whether your sister will feel dismissed by it is a different question, and Grammarly does not claim to answer it. Tone detection sits in the free tier, with Pro at $12 a month on annual billing per current pricing trackers.
That distinction is the whole gap.
Nobody can tell whether a message was written by AI, and it is not even close
This is the finding that surprised me most, and it is the best-powered study in the whole area.
Maurice Jakesch, Jeff Hancock and Mor Naaman ran six experiments with 4,600 US participants, published in PNAS in 202310. People read self-presentations from dating, hospitality and professional profiles and guessed which were written by a machine. Accuracy came in between 50 and 52 percent. In the dating condition, with money on the line for getting it right, about 51.6 percent.
Coin-flip territory.
The reason is that we all use the same broken shortcuts. Participants took first-person pronouns, contractions, mentions of family, and warm emotional detail as proof a human wrote it. They took clean grammar and flat phrasing as proof a machine did. Every one of those signals is trivially easy for a model to produce on request. The heuristics point exactly the wrong way. A trained classifier in the same paper managed only 58.1 percent, so this is not a skill issue you can fix by paying more attention.
One caveat I want to be honest about. That study used GPT-3-era text. Whether the numbers hold against current models has not been tested at anything like that scale.
The damage comes from being suspected, not from being caught
If people cannot detect AI, you might expect the whole worry to evaporate. It does not, and this is the part that matters for anyone writing a real message to a real person.
A Cornell-led team ran two randomised experiments with 1,036 participants, published in Scientific Reports11. Smart replies made conversations faster and more positive, and people rated each other as closer and more cooperative. Then the twist. Participants who suspected their partner was using smart replies rated that partner significantly worse. Suspicion did the damage. That study measured suspicion rather than accuracy, and given the detection numbers above, there is not much reason to think those suspicions were landing on the right people.
A Carnegie Mellon team put a mechanism under this in a 2025 paper12 with 399 participants across two studies, and their result is more subtle than a simple penalty. An AI label does not just make you look bad. It makes your message weaker evidence about you in either direction. An AI-assisted apology reads as less warm, and an AI-assisted brag or blame reads as less cold. The label drains the message of information about who you are.
That framing stuck with me. Whatever your message was doing as a signal, the label turns the volume down on it.
Except in one setting, where disclosure costs you nothing at all
I want to include the study that argues against my own product thesis, because leaving it out would be dishonest.
Zoe Purcell and colleagues at the Max Planck Institute for Human Development, with Jakesch again among the authors, ran two preregistered experiments with 1,637 participants in incentivised two-player trust games, published in iScience in November 202513. Half could use predictive text assistance to write a message persuading a stranger to trust them. The rest wrote unaided.
AI assistance had minimal effect on trust, and that held even when AI use was disclosed. The assisted writers produced equally trust-inducing messages in less time, so they earned more trust per minute spent. Linguistic analysis found their messages slightly less authentic but warmer, more complex, and higher in what researchers call clout.
The authors are careful about scope, and so am I. This is a one-shot transaction between strangers with money attached. Behavioural trust in a payout decision is a different animal from your partner reading your text at 11pm and deciding whether you meant it. But the honest version of this article says that the “AI ruins trust” story has a real counterexample in the literature, and this is it.
AI writes better comfort messages than most people do, right up until you mention it was AI
Two studies, same shape of result.
Yin, Jia and Wakslak in PNAS14 found that AI-generated responses made people feel more heard than responses from untrained humans, and that the AI was better at reading which emotion was in play. Then they told recipients the message came from AI and the effect dropped. One of the authors noted the two effects were similar in size and roughly cancelled each other out.
Ovsyannikova, Oldemburgo de Mello and Inzlicht15 ran four preregistered experiments with 556 participants in Communications Psychology. GPT-4 responses were rated as more compassionate than human ones, including responses written by trained crisis-line volunteers at Distress Centres of Greater Toronto. Disclosing the source narrowed the gap but did not close it.
So the machine is often better at the words. The words are not the only thing being transmitted.
The reason effort matters is that effort used to be expensive
Signalling theory is the frame that made all of this click for me.
A message does two jobs. It carries content, and it proves you spent something to send it. Your time, your attention, and in a hard conversation, your composure. Those are finite, so spending them says something true about how much you value the person. Research on commitment signals16 in Japanese and US samples found costly signals outperformed cheap ones, and that failing to send an expected signal, like forgetting an occasion, hurt romantic relationships far more than friendships.
Drop the cost of producing a warm, articulate, well-structured message to roughly zero and the signal stops carrying information.
There is now direct evidence for the mechanism, though not from a messaging context. A 2026 study in Frontiers in Psychology17 with 618 participants found that labelling content as AI-generated significantly reduced perceived effort, while a human-made label did not differ from no label at all. People assume human effort by default, and the AI tag is what removes the assumption. That study used short-form video, so read across to messaging with some caution.
Outsourcing the words appears to weaken your own commitment to them
This is the finding I keep thinking about, and it is not about the recipient at all.
Economists at CREED in Amsterdam18 ran trust games where senders could use ChatGPT to write a message promising to reciprocate. Two things happened at once. People with ChatGPT access made more explicit promises. And people who opened ChatGPT at all kept their promises 25.7 percent less often at payout time, whether or not they ended up copying what it gave them.
The detail that makes it: the closer the sent message was to the raw ChatGPT suggestion, the less trustworthy the sender’s later behaviour. People who rewrote the suggestion behaved better. The authors read this as a sense of ownership. Say a thing in your own words and you feel bound by it. Forward a thing a machine wrote and some part of you never signed it.
Apologies are where this has been studied properly, and the finding is more useful than “don’t”
If you only read one section, read this one, because the practical answer lives here.
Ella Glikson and Omri Asscher ran three scenario studies on workplace apologies, published in Computers in Human Behavior19. Knowing that someone used AI tools to write an apology reduced how authentic it seemed and how willing people were to forgive. Disclosing it yourself did not soften that. So far, predictable.
Then the part that changed my mind. Participants who used only one of three available AI tools got no authenticity penalty at all. The researchers read it as distance. Limited AI use keeps the final message close to what the person originally meant, and people can feel that.
A separate 3x2 experiment by Lim, Hong and Schneider with 464 participants, also in Computers in Human Behavior20, found human-authored apologies were seen as more sincere across the board, and that a warm tone improved AI apologies without ever closing the gap to a human one.
Put those two next to each other and you get something practical. The penalty tracks how far the finished message has drifted from what you actually meant. Not whether a machine was in the room.
These tools will polish your worst instincts and never say a word
Write something passive-aggressive, hand it to any rewrite tool on your phone, and watch what happens. You get better-written passive aggression. No warning, no “are you sure.” The tool’s job is to do what you asked.
There is a reason the models lean this way. OpenAI pulled a GPT-4o update in April 2025 and wrote publicly21 that the model had become “overly flattering or agreeable, often described as sycophantic,” and that it “skewed towards responses that were overly supportive but disingenuous.” The cause they named was tuning too heavily on short-term user feedback, meaning thumbs up and thumbs down. In a follow-up post22 they admitted they had no deployment evaluations tracking sycophancy at all.
Anthropic researchers found the same pattern is general. In a paper presented at ICLR 202423, five frontier assistants all showed sycophancy, and analysis of roughly 15,000 human preference comparisons found that matching the user’s stated beliefs was among the strongest predictors of which response people preferred. Train on approval and you select for agreement. A 2025 benchmark24 measured sycophantic behaviour in 58.19 percent of test cases across ChatGPT-4o, Claude Sonnet and Gemini 1.5 Pro, with 78.5 percent persistence once it started, though that was factual question answering rather than personal messages.
Now think about what that means for the moment you most want help. You are angry, you have written something you will regret, and the tool holding your draft has been optimised to make you feel good about it.
The question I cannot answer
Is it better to send a clumsy message you wrote or a polished one you did not?
I looked hard and there is no study that tests this directly. Nobody has randomised “your own imperfect draft” against “a good draft you did not write” and followed what happened to the relationship. There is nothing longitudinal at all. No field data on couples, families or close friends over months of AI-assisted messaging. We do not know whether people adapt to synthetic polish, or whether the habit quietly erodes something.
What we have points in one direction for close relationships. Effort reads as care, labels flatten the signal, outsourced words seem to loosen your grip on your own promises, and apologies stay believable when they stay close to what you meant. But the trust-game result says that in some contexts none of this costs you anything, and I am not going to pretend the evidence is cleaner than it is.
So what would actually help
Reading all of this back, the shape of the missing product is fairly clear, and it is close to the opposite of what the category is building.
Nobody needs their words replaced. What they need, in the fifteen seconds before hitting send, is to know that the message reads colder than they feel, or that the thing they think is a clarification reads as a list of grievances. Then they fix it themselves, in their own words, and the effort signal survives, and the ownership survives, and the Glikson and Asscher distance stays short.
That is the version I want to exist, and it is the one we built. Subtext takes the whole thread rather than the sentence, names what your draft is likely to signal to the person about to read it, and shows you the words creating that impression, so the fix stays yours. On the evidence above that matters more than the rewrite: the penalty tracks distance from what you meant, and the shortest distance is your own sentence with one word changed. Free to start, on your phone or in the browser.
My own rule, for whatever it is worth: I will let a machine tell me what my message sounds like. I will not let it tell my sister I am sorry.
Think I’ve read a study wrong, or know a tool that does diagnose before it rewrites? Tell me on LinkedIn.
Samet Durgun is the co-founder of Subtext, an app that catches the emotional tone of your messages and rewrites them in your own voice. He’s based in Berlin.
Sources
Every link above goes to the primary source where one exists, numbered in order of appearance. Product features and prices checked August 2026. Where a figure came from a company blog or press release rather than a study, I left it out.
- Apple Support. Use Writing Tools with Apple Intelligence on iPhone. https://support.apple.com/guide/iphone/find-the-right-words-with-writing-tools-iph6f08da1d2/ios
- Samsung. Use Writing assist on Galaxy phones and tablets. https://www.samsung.com/us/support/answer/ANS10000943/
- Google Support. Use writing tools with Gboard. https://support.google.com/gboard/answer/17470061
- Google Support. Draft messages with Magic Compose. https://support.google.com/messages/answer/13632636
- Apple App Store. Resolve: AI Conflict Coach. https://apps.apple.com/us/app/resolve-ai-conflict-coach/id6759972332
- Google Play. Clarity Coach. https://play.google.com/store/apps/details?id=com.zamrudila.claritycoach
- Grammarly. Writing Tone Detector and Tone Suggestions. https://www.grammarly.com/tone
- TechCrunch (2019). Grammarly gets a tone detector to keep you out of email trouble. https://techcrunch.com/2019/09/24/grammarly-gets-a-tone-detector-to-keep-you-out-of-email-trouble/
- Grammarly Support. How do Grammarly’s tone suggestions work? https://support.grammarly.com/hc/en-us/articles/10674801783309
- Jakesch, M., Hancock, J. T., & Naaman, M. (2023). Human heuristics for AI-generated language are flawed. PNAS, 120(11), e2208839120. https://www.pnas.org/doi/10.1073/pnas.2208839120
- Hohenstein, J., Kizilcec, R. F., DiFranzo, D., Aghajari, Z., Mieczkowski, H., Levy, K., Naaman, M., Hancock, J., & Jung, M. F. (2023). Artificial intelligence in communication impacts language and social relationships. Scientific Reports, 13, 5487. https://www.nature.com/articles/s41598-023-30938-9
- Khadpe, P., Wenzel, K., Loewenstein, G., & Kaufman, G. (2025). Explaining the Reputational Risks of AI-Mediated Communication. AAAI/ACM Conference on AI, Ethics and Society. https://arxiv.org/abs/2509.09645
- Purcell, Z. A., Jakesch, M., Dong, M., Nussberger, A.-M., & Köbis, N. (2025). Writing with AI boosts trust-building efficiency. iScience, 28(12), 114092. https://www.cell.com/iscience/fulltext/S2589-0042(25)02353-3
- Yin, Y., Jia, N., & Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact. PNAS, 121(14), e2319112121. https://www.pnas.org/doi/10.1073/pnas.2319112121
- Ovsyannikova, D., Oldemburgo de Mello, V., & Inzlicht, M. (2025). Third-party evaluators perceive AI as more compassionate than expert humans. Communications Psychology, 3. https://www.nature.com/articles/s44271-024-00182-6
- Yamaguchi, M., et al. (2015). Commitment signals in friendship and romantic relationships. Evolution and Human Behavior. https://www.sciencedirect.com/science/article/abs/pii/S1090513815000525
- Human-made vs. AI-generated: how provenance labels drive strategic curation via perceived effort (2026). Frontiers in Psychology. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1840483/full
- AI-Powered Promises: The Influence of ChatGPT on Trust and Trustworthiness. CREED, University of Amsterdam. https://www.creedexperiment.nl/creed/pdffiles/chatGPT.pdf
- Glikson, E., & Asscher, O. (2023). AI-mediated apology in a multilingual work context. Computers in Human Behavior, 140, 107592. https://www.sciencedirect.com/science/article/abs/pii/S0747563222004125
- Lim, J. S., Hong, N., & Schneider, E. (2025). How warm- versus competent-toned AI apologies affect trust and forgiveness through emotions and perceived sincerity. Computers in Human Behavior, 172, 108761. https://www.sciencedirect.com/science/article/pii/S0747563225002080
- OpenAI (2025). Sycophancy in GPT-4o. https://openai.com/index/sycophancy-in-gpt-4o/
- OpenAI (2025). Expanding on what we missed with sycophancy. https://openai.com/index/expanding-on-sycophancy/
- Sharma, M., et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. https://arxiv.org/abs/2310.13548
- Fanous, A., et al. (2025). SycEval: Evaluating LLM Sycophancy. AAAI/ACM Conference on AI, Ethics and Society. https://arxiv.org/abs/2502.08177
- Hancock, J. T., Naaman, M., & Levy, K. (2020). AI-Mediated Communication: Definition, Research Agenda, and Ethical Considerations. Journal of Computer-Mediated Communication, 25(1), 89-100. https://academic.oup.com/jcmc/article/25/1/89/5714020