ArticlesWriting tools

What an AI message checker can check, and how to do it yourself

An AI message checker gives you a second read on a message you already wrote. What it can tell you, the line it should not cross, and the seven-step version you can run yourself in under a minute.

By Samet Durgun · Co-founder of Subtext · 15 min read

If you searched for an AI message checker, you probably already have the message. It is sitting in the text box, you have read it six times, and nothing is obviously wrong with it. It just feels off. Too cold, or too much, or defensive in a way you cannot point at.

An AI message checker reads a message you have already written and shows you how the wording may come across before you send it. Some rewrite afterwards. The checking is the part that matters, and it is the part most tools skip.

I co-founded Subtext, which is built around this exact moment, so I am not neutral about the category. What I can be useful about is the narrower question: what can software actually tell you about a message, and what do you still have to decide yourself. Every research claim below links to its source.

What a message checker can check, and the line it should not cross

A checker can analyse the language in front of it. Whether the wording reads formal or casual, direct or indirect, warm or distant. Whether a phrase states the other person’s motive as a fact. Whether there are three apologies inside a two-line refusal.

What it cannot do is verify what anybody privately feels.

Take a two-word draft: “Fine. Do whatever you want.” A checker can reasonably say that “Fine” is clipped and that “whatever you want” reads as withdrawal. It cannot establish that the person who wrote it is angry. Those are different claims, and a lot of tone-detection marketing crosses between them without saying so. I went through the evidence on that in a separate piece on what AI does and does not know about tone.

For a message you are about to send, mind-reading is not what you need anyway. You need a second reader, because you cannot be one for your own draft.

Why rereading your own draft does not work

Kruger, Epley, Parker and Ng ran five experiments on this in 20051, and the paper gets quoted far more often than it gets read.

The largest of the five had 154 pairs of students write statements meant to convey sarcasm, seriousness, anger or sadness, then predict whether their partner would identify the intended tone. Senders expected to land 88.8 percent of them. They landed 70.4 percent.

The gap is the finding, and the authors are careful about what it does not mean. Accuracy in their first experiment was 84 percent, and they write that it would be “misleading to suggest from these data that people are poor at communicating sarcasm over e-mail.” Their summary of what it does mean: “however able people are, they are not as able as they believe.”

Three details make this harder for a sender. Friends were no more accurate than strangers, and just as overconfident, so writing to someone who knows you well does not close the gap. Emoticons do not close it either: a follow-up reported inside the same paper manipulated whether senders could use them and found that “overconfidence did not differ between those who were using emoticons and those who were not.” And the widely circulated claim that people read tone correctly 56 percent of the time is a reading of a graph rather than a number in the text. What the paper prints for that condition is an accuracy rate indistinguishable from chance, among 29 pairs. The 56 is not worth repeating.

One honest limit before the steps. This is 2005 research on email. The authors expected the mechanism to extend to chat and instant messaging and said so, but nobody has rerun it on modern texting. Treat the mechanism as well established and the exact percentages as belonging to email.

What your message reliably carries, and what it loses

Holtgraves2 gives the most useful modern numbers. Senders wrote text messages conveying one of 22 emotions without naming them. Readers identified the emotion by free response 20.1 percent of the time, and 46.3 percent with multiple choice. When they got it wrong, they still landed in the correct positive or negative valence 84.5 percent of the time in the free-response study and 90.6 percent in the multiple-choice study.

That produces one instruction for a sender. The reader will almost certainly register that something is off. They will probably get the reason wrong. So the check is aimed at the ambiguity that lets them fill in a worse reason, rather than at hiding the feeling.

Byron’s framework3 predicts the direction of the loss: positive emails tend to arrive as neutral, and neutral ones as negative. It is a theory paper rather than an experiment, so weigh it accordingly. Write one notch warmer than feels natural and it arrives at natural.

The seven-step check

There is no validated scientific test for whether a text sounds all right. What follows is a way of creating distance from your own draft. It takes under a minute and needs no app.

1. Read it without supplying what you meant

Take the sentence carrying the most weight and read it as though you do not know who wrote it.

“I guess just do whatever works for you.” You know whether that was sincere. Your reader does not. “I guess” can signal reluctance. “Just” can add impatience. “Whatever” can read as withdrawing from the decision.

None of those readings is automatically the right one. They are available to the reader, which is the whole point of checking.

2. Count the apologies and the hedges

This is the most common defect I see, and both halves have evidence behind them.

Freedman, Burgoon, Ferrell, Pennebaker and Beer4 found across studies totalling roughly 1,880 participants that apologies attached to rejections increased hurt feelings and created pressure to say “I forgive you” without increasing actual forgiveness. Apologetic rejections produced more aggression in a behavioural task. About 39 percent of people spontaneously included an apology when asked to write a good rejection, so the counterproductive move is also the popular one. The evidence is specific to rejections and declines, and I will not stretch it further.

On hedges, Givi, Kirk, Grossman and Sedikides5 ran six experiments, five preregistered, on replying to an invitation with a tentative “maybe” instead of a straight no. Invitees overestimated how much the inviter would prefer the maybe, because they underestimated how much more disrespected a maybe makes the inviter feel. Their explanation is the part worth sitting with: the maybe serves the person sending it more than the person receiving it, and motivated reasoning does the rest. That evidence is about invitations, so applying it to any decision you have already made and are dressing as uncertainty is my extension rather than their finding.

Practically: count every sorry, every “sorry to bother you”, every “I hope this isn’t annoying”. Keep at most one, and only where you are apologising for something you did. Then search for maybe, possibly, we’ll see, let me check and get back to you. Some are honest uncertainty and stay. The ones to cut are the ones you felt relieved to write.

This count is also the first thing Subtext runs on a draft, for the reason the Freedman numbers suggest: the apology you added instinctively is the one you are least likely to notice on a reread.

3. Find the sentence that is about you rather than them

Read each sentence and ask who it serves. “I’ve been meaning to message for ages and I feel terrible about it” is you managing your own guilt in front of someone who now has to reassure you about it. It is a request wearing the clothes of an admission.

This step is reasoning rather than a finding. Nobody has run the experiment. I include it because it is the most reliable pattern I have found reading thousands of drafts, and because it sits next to something the research does document: an apology that asks for absolution and an apology that acknowledges a cost are different acts, and only the second does anything for the recipient.

The fix is usually deletion. Occasionally it is moving the sentence to the end, where it reads as a footnote rather than as the point.

4. Check whether the request also carries a judgement

Compare “Can you let me know by Thursday?” with “Can you please actually let me know this time?”

Both make a request. The second also comments on the person’s past behaviour. That may be deliberate, and sometimes it should be. More often it arrives because you were frustrated while typing and the frustration found a place to sit.

The word doing the work is usually small. Actually. Again. Finally. This time. Look for those specifically.

5. Check whether the explanation has taken over

Explanations are useful and they quietly turn into defences. It happens most in apologies, cancellations and disagreements. You start with what happened, then why it happened, then why your reaction to it made sense. By the end the reader has received a case for the defence when you thought you were sending an apology.

Delete the explanation for a moment and look at what is left. Is the request still clear? Is the acknowledgement still there? Is the no still a no? If the message collapses without the explanation, the explanation was carrying something it should not have been.

6. Ask whether it adds anything or only asks for a reply

Lew, Walther, Pang and Shin6 crossed response latency with conversational contingency, meaning how directly a message engages with what came before. Contingent replies outperformed generic ones overall, but the effect of speed depended on contingency: a fast reply helped only when it was also contingent, while a fast but generic, scripted-sounding reply was rated the worst of all four conditions, worse than a slow generic reply. Contingency didn’t cushion slowness. It was a precondition for speed to pay off at all.

So a message that engages with the conversation is evaluated on its content, and a message that only prompts for a response is evaluated as a demand for something you feel owed. “?” and “hello?” are the pure form. “Just checking you saw this” does identical work in a better coat.

Fast test: delete the last line. If the message still says something, the last line was a prompt and it can stay deleted.

7. Read it out loud in the wrong tone

This is the only step here with an experiment directly behind it, and the result is more specific than the advice you normally hear.

In the fourth experiment of the Kruger paper1, 54 students typed out sarcastic and serious statements to send by email, then read each one aloud into a tape recorder before predicting how well their partner would decode it. Half read each statement in the tone they had intended. Half read it in the opposite tone, saying the sarcastic ones seriously and the serious ones sarcastically.

The second group stopped being overconfident. In the paper’s words, “the phenomenology manipulation completely erased participants’ overconfidence.” The first group, reading their own words in their own intended tone, stayed exactly as miscalibrated as everyone in the earlier studies.

So reading your draft aloud the way you meant it does nothing, which is awkward for the standard version of this advice. Read it aloud in the tone you are afraid of instead. Say it flatly. Say it coldly. Say it as though you were annoyed with the person. If it survives that, send it. If it becomes a different message in your mouth, that reading was sitting in the text the whole time and you could not see it.

One experiment, 54 people, on email. It is still the best-evidenced item on this page. It is also the closest description of what Subtext does when it reads a draft: a pass through your words in a voice that is not the one in your head.

Then read it once as the person receiving it

The reading a message gets depends on who is holding it, and they are overconfident too. In the same Kruger study, email readers believed they had identified the sender’s intended tone 89.3 percent of the time and were right 62.8 percent of the time1. Whatever they decide your message meant, they will not be checking.

Who they are shifts the odds. Sillars and Zorn7 documented a negative intensification bias in workplace email, where receivers rated messages more negatively than uninvolved observers did and were only weakly anchored to the message’s actual features. The effect was stronger in poor communication climates and among people lower in the hierarchy. Kingsbury and Coplan8 found socially anxious readers interpret ambiguous texts more negatively, across samples of 215 and 353.

So if you are messaging someone junior to you, or someone you have been tense with, or someone anxious, the sharpest available reading of your draft is a likely one rather than a paranoid one. Read the message as them, in that state, and see whether it survives.

Worth saying plainly, because the popular version overshoots: the evidence does not show everyone defaulting to the worst reading of everything. It shows negative interpretation being predicted by who is reading and what the relationship is like at the time.

Message checker or message rewriter

The two can run on similar models underneath. The useful moment happens at a different point.

Message checker Message rewriter
Starts with A message you already wrote A draft or a prompt
Main job Show how the current wording may read Produce different wording
Useful when You cannot name what feels off You already know what you want changed
Main risk Treating one interpretation as fact Replacing more of your voice than needed
Good outcome You understand your own draft You get a better alternative

Back to “Fine. Do whatever you want.” A rewriter might hand you “Sounds good, go with whichever option works best for you.” That is smoother, and it is a different message. Maybe you are frustrated and want some of that to survive. Checking first lets you see the effect before you decide whether to delete it.

Can ChatGPT check a message before you send it

Yes, if you ask it the right way. The failure mode is that a general assistant will solve the easier problem and hand you a rewrite before it has told you anything.

Separate the two explicitly:

Read this as the person receiving it. Do not rewrite it yet. Tell me which phrases could read as colder, more defensive or more intense than I probably intend, and explain why each one does.

Then paste the draft, and add the relationship and the previous message if either changes the meaning.

For an occasional difficult text that works well. What a dedicated checker removes is the setup, which matters mainly because you are least able to construct a careful prompt at the exact moment you need one.

What your phone already does

Apple Intelligence Writing Tools rewrite selected text and offer Friendly, Professional and Concise, plus a “Describe your change” box for anything more specific9. Google’s Magic Compose suggests replies and rewrites drafts in different styles inside Google Messages on supported devices10. Samsung’s Writing Assist generates text, changes writing style and checks spelling and grammar on supported Galaxy phones11. Grammarly goes furthest into detection, and its tone checker states that it analyses word choice, phrasing, punctuation and capitalisation12.

All of these are good when you know what you want. “This is too formal, make it more casual” is a solved problem and it is free on hardware you already own.

The harder version is “I do not know what is wrong with this, does it sound all right.” That question is where checking and changing stop being the same job.

How Subtext runs the check

Subtext starts from the message you are writing and shows how it may come across before you decide whether to change anything. You can write from scratch, paste a draft, or add a screenshot of the conversation for context.

The analysis reads your draft’s tone and clarity and flags wording that could land as unclear, harsh, passive or easy to misread. From there you can take alternative wording, or tap shorter, warmer, more casual, or add confidence, and it rewrites while keeping your voice13.

A screenshot does a different job. It supplies conversation context so the reply fits the exchange. That is worth being precise about: the trigger and emotion analysis runs on your own draft. A screenshot of what somebody else sent is context for your reply, not proof of what they privately feel, and no tool including this one can establish that from a screenshot.

The order is the part I care about. Read the draft. See what may be landing differently from what you meant. Then decide whether anything needs changing. Sometimes the useful result is that the message was already fine.

Subtext is free to start, on iPhone and Android.

What none of this catches

Three limits, and they matter more than the steps do.

The check is structural. It finds apologies, hedges, prompts, judgements and self-serving sentences. It has no opinion on whether the thing you are saying is correct, fair or worth saying, and a well-worded message can be wrong in every way that counts.

The reader brings their own weather. A 2025 preregistered study of 51 couples14 found people partly read their own current mood into their partner’s messages. Nothing in your draft controls that.

And the check assumes the problem is the wording. Sometimes it is not. If you are rewriting the same message for the ninth time, or sending three follow-ups to find out whether someone is annoyed, the wording has stopped being the thing under strain. I have written separately about the research on texting anxiety and on reading anger into ambiguous messages, and both are more use than a better draft. If it is persistent enough to be shaping your days, a doctor or therapist is a better resource than an article by someone who makes an app.

Common questions

What is an AI message checker? It analyses a message you have already written and shows how the wording may come across before you send it. Depending on the tool that can cover tone, clarity, formality, ambiguity, or phrases likely to be misread. Some also offer rewrites afterwards.

Can AI tell if my message sounds rude? It can identify wording that may reasonably be read as blunt, dismissive or accusatory. It cannot guarantee that a particular person will find it rude. A useful checker names the specific wording that produced the reading rather than handing you one label.

Is a message checker the same as a tone checker? They overlap. A tone checker typically names qualities in the writing, such as formal, confident or accusatory. A message checker can also look at clarity, specific wording and conversation context.

Can AI know how someone feels from a text? It can analyse language and offer plausible readings. A text alone does not establish the writer’s private emotional state, and confident claims about what somebody really feels go beyond what the words support.

What is the difference between a message checker and an AI texting app? “AI texting app” is a broad term covering dating reply generators, writing keyboards, chatbots and automatic reply tools. A message checker has a narrower job: analysing something you already wrote, before you send it.


Think I have read a study wrong, or know research on this I have missed? Tell me on LinkedIn.

Samet Durgun is the co-founder of Subtext, an app that catches the emotional tone of your messages and rewrites them in your own voice. He’s based in Berlin.

Sources

Every link goes to the primary source where one exists, numbered in order of appearance. Where a step rests on reasoning rather than evidence, this post says so instead of borrowing authority from a nearby study. Quotations are verbatim.

  1. Kruger, J., Epley, N., Parker, J., & Ng, Z.-W. (2005). Egocentrism over e-mail: Can we communicate as well as we think? Journal of Personality and Social Psychology, 89(6), 925-936. Five experiments. Study 1 N = 12, Study 2 N = 60, Study 3 N = 308 in 154 pairs, Study 4 N = 54, Study 5 N = 58. Funded by the University of Illinois Board of Trustees and NSF grant SES-0241544. The widely quoted 56 percent figure is a reading of Figure 1 in Study 2 rather than a number in the text. doi.org
  2. Holtgraves, T. (2022). Implicit communication of emotions via written text messages. Computers in Human Behavior Reports, 7, 100219. N = 136 and N = 167. doi.org
  3. Byron, K. (2008). Carrying too heavy a load? The communication and miscommunication of emotion by email. Academy of Management Review, 33(2), 309-327. Theory paper, no sample. doi.org
  4. Freedman, G., Burgoon, E. M., Ferrell, J. D., Pennebaker, J. W., & Beer, J. S. (2017). When Saying Sorry May Not Help: The Impact of Apologies on Social Rejections. Frontiers in Psychology, 8, 1375. Combined N approximately 1,880; findings specific to rejections. doi.org
  5. Givi, J., Kirk, C. P., Grossman, D. M., & Sedikides, C. (2025). Maybe don’t say “maybe”: How and why invitees fail to realize that they should not respond to invitations with a “maybe”. Journal of Experimental Social Psychology, 121, 104814. Six experiments, five preregistered. Findings specific to invitations. doi.org
  6. Lew, Z., Walther, J. B., Pang, A., & Shin, W. (2018). Interactivity in Online Chat: Conversational Contingency and Response Latency in Computer-Mediated Communication. Journal of Computer-Mediated Communication, 23(4), 201-221. N = 131, 2×2 between-subjects design. doi.org
  7. Sillars, A., & Zorn, T. E. (2021). Hypernegative Interpretation of Negatively Perceived Email at Work. Management Communication Quarterly, 35(2), 171-200. doi.org
  8. Kingsbury, M., & Coplan, R. J. (2016). RU mad @ me? Social anxiety and interpretation of ambiguous text messages. Computers in Human Behavior, 54, 368-379. N = 215 and N = 353. doi.org
  9. Apple Support. Use Writing Tools with Apple Intelligence on iPhone. support.apple.com
  10. Google Messages Help. Draft messages with Magic Compose. support.google.com
  11. Samsung Support. Use Writing Assist on Galaxy phones and tablets. samsung.com
  12. Grammarly. Writing Tone Detector and Tone Suggestions. grammarly.com
  13. Subtext. subtext.it, checked August 2026.
  14. Steinebach, P., Stein, M., & Schnell, K. (2025). Messenger-based assessment of empathic accuracy in couples’ smartphone communication. BMC Psychology, 13, 147. N = 102 in 51 couples; recruitment fell short of the preregistered target. doi.org