ArticlesMessages

What makes a reply land well?

A reply lands when it shows you engaged with this specific message from this specific person. Speed matters far less than people think, and a fast generic reply tested worse than a slow one.

By Samet Durgun · Co-founder of Subtext · 12 min read

Most advice about replying is advice about speed. Answer within five minutes, answer within an hour, do not leave it three days. I co-founded Subtext and spend most of my time looking at messages people are stuck on, and speed is almost never the thing that is wrong with them.

The research points somewhere else and it points there consistently. What makes a reply work is whether it demonstrates that you engaged with this particular message from this particular person. Psychologists call it perceived responsiveness, meaning the sense of having been understood, validated and cared about. It is the measured mechanism in the question-asking research, the mechanism in the concreteness research, the thing that decides whether support helps or hurts, and the thing that suspicion of AI destroys. Everything else in this article is a different way of solving that one problem, and it is the same problem Subtext is built to help you solve.

Speed matters less than you think, except where it doesn’t

Two different situations, and blending them produces bad advice.

With strangers, in a transaction, speed is real money. Hart, VanEpps, Sezer and Amir combined transaction data from a freelance marketplace covering roughly 11.66 million service requests with experiments totalling several thousand participants1. Delays of even five to ten minutes measurably reduced the chance of winning the contract. The mechanism they identify is inferential: a fast reply signals you will probably be responsive later too. Keep the two halves separate, though. The marketplace data is correlational field data. The mechanism is the experimental part.

With colleagues, the error runs the other way. Giurge and Bohns documented an email urgency bias across eight preregistered studies with 4,004 working adults2. Receivers overestimate how fast senders expect a reply to non-urgent out-of-hours email, and pay for it in stress. A single line from the sender saying there is no rush reduced the bias. Related work defines workplace telepressure, the urge to answer immediately, and links it to worse recovery, burnout and sleep3.

And silence means whatever the norm between you was. Kalman and Rafaeli tested reply latency as an expectancy violation and found the effect of delay depended on how the sender was already regarded4. A slow reply was forgiven for a promising job candidate and damaging for a weaker one. Two hours from someone who always answers in five minutes is a different signal from two hours from someone who answers twice a day, and no clock-based rule captures that, which is exactly why Subtext does not try to give you one.

The finding that reframes the whole speed question

Lew, Walther, Pang and Shin ran the experiment everyone should know about and almost nobody does5.

They varied two things independently: how fast a reply arrived, eight seconds against forty, and whether the reply was contingent, meaning tailored to what the person actually asked, against a generic canned response. 131 US adults watched a simulated customer service chat and rated it on four outcomes.

Speed produced no significant main effect on any of the four. Contingency produced significant effects on all four, with effect sizes from 0.48 to 0.66.

The authors had predicted that slow generic replies would be worst. They were wrong, and they say so. Fast generic replies scored lowest on every outcome. Their explanation is that a fast reply that does not engage seems robotic, and without a delay to suggest the person is busy, it reads as inattentive at best.

That is the practical headline of this entire literature. Rushing to send something empty is worse than taking your time and sending something real.

This is also the exact failure Subtext is built to catch, because a fast empty reply does not look wrong on the screen. It looks efficient. The problem only becomes visible when you ask what it does for the person reading it, which is the question the product asks on your behalf.

The one tactic that generalises

Ask a follow-up question about what they just said.

Huang, Yeomans, Brooks, Minson and Gino ran three studies of live conversations6. In an online chat study with 199 dyads, people instructed to ask many questions were significantly better liked, and the effect was mediated by perceived responsiveness. A second study replicated it and showed the work was being done specifically by follow-up questions, meaning questions that could only come from having read the answer. A speed-dating component found follow-up question rate predicted second-date offers.

Three honest notes on that paper, because it is contested and it matters. There is a published dispute: Kluger and Malloy reanalysed the speed-dating study and argued the effect disappears once you account for who was asking, and the original authors replied that the reanalysis used the wrong measure and an unsuitable model7. The conversation experiments are the stronger support and the speed-dating component is the contested part. The paper carries a published correction from March 2025, and it carries no expression of concern and no retraction. And a third-party study within the same paper found observers reading transcripts did not share the preference, which is a real boundary condition.

Separately, people avoid asking sensitive questions because they fear giving offence, and they significantly overestimate that cost8. Whether it is the obvious follow-up or the sensitive one, the question you are not asking is probably safe to ask, and it is the kind of question Subtext will nudge you towards asking rather than talking yourself out of.

Name the specific thing

Packard and Berger isolated linguistic concreteness as a signal of listening9. Five studies, including text analysis of over 1,000 real customer interactions across two field settings. Customers were more satisfied, more willing to buy, and spent more when the person replying used concrete language, and the mechanism was that they inferred the employee was actually paying attention. A one standard deviation increase in concreteness improved satisfaction by 9 per cent and spending by at least 13.

The translation is unusually clean. “That grey t-shirt” rather than “your issue”. “The Thursday meeting” rather than “the thing you mentioned”. Naming the specific object is what reads as having listened, and it costs nothing. It is also the easiest thing to check mechanically, and one of the few places where Subtext can give you a straight answer rather than a judgement.

If it is a difficult conversation, phone them

This is the most quotable result in the whole set and it will annoy you.

Bevis, Schroeder and Yeomans ran randomised experiments with 1,576 conversation partners across 1,842 conversations, plus 1,432 observers10. Spoken conversations produced greater perceived understanding and lower perceived conflict than written ones, along with higher liking, more enjoyment and greater attitude change. Pooled across their first three studies, speaking raised understanding by a standardised beta of 0.45 and lowered conflict by 0.23.

Then they asked people what they would choose. 83.9 per cent preferred to write when they expected disagreement, and predicted that writing would produce less conflict and more understanding. Their experiments found the opposite.

There is a twist that saves this article from being pointless. Receptive language, meaning hedging your claims, acknowledging the other side before countering, and naming genuine points of agreement, predicted constructive outcomes more strongly in writing than in speech11. The authors call it a puzzle: people use less receptive language in the medium where it appears to matter most.

So if you must write it, the receptiveness moves matter more than they would on a call. If you have the option, the call goes better than you expect it to. Since Subtext only ever works on what you write, this is the section I would take most seriously: the tool can help you hedge, acknowledge and name agreement well, and it cannot make the case for picking up the phone instead. That case belongs to you.

Caveats the authors state: participants were mostly strangers, conflict was low across all conditions, and most were students. Two people who know each other well arguing about something central to their relationship might behave differently.

Holding replies, and why the obvious advice needs a condition

There is no experiment I could find that directly tests a short “got it, will reply properly” against a delayed full reply with nothing in between. The mechanism is well supported, though, so I would treat this as reasoning rather than as a finding.

The uncertainty about whether you will respond at all is what the latency research identifies as the damaging part1. A holding reply addresses exactly that. Giurge and Bohns showed a brief expectation-setting note reduced a documented misreading2.

But Lew and colleagues put a condition on it5. A generic holding reply is a fast non-contingent reply, which is the combination that scored worst of all four in their study. Both of these buy you the same few days. Only one of them names anything.

Generic “Thanks for your email, we’ll be in touch”

Fast and non-contingent, the exact combination that scored worst on every outcome in Lew and colleagues’ study5

Specific “I’ve read the bit about the timeline and I need until Thursday to answer properly”

Subtext will tell you which version of a holding reply you have actually written, since the two look similar and score nothing alike.

Where AI comes into it

Worth being precise here, since I sell an AI product and the evidence is more interesting than either the boosters or the critics say.

Rubin and colleagues ran nine studies with 6,282 participants12. Every response was AI-generated. What varied was the label: some were presented as written by a human, some as written by AI. Human-labelled responses were rated more empathic and supportive and produced more positive emotion. Participants’ own unprompted suspicion that AI had helped write a human-labelled response reduced perceived empathy too. It replicated across different lengths, delays and models, and was driven by responses emphasising emotional sharing rather than practical understanding.

Read the design carefully, because it decides what the finding means. This is evidence about attribution and belief. It is not evidence that AI writes worse messages, because the AI wrote all of them.

A similar pattern shows up elsewhere. In two randomised experiments with 1,036 participants, smart replies increased efficiency and positive language, and participants who believed their partner was using them rated that partner as less cooperative, after controlling for actual use13. Belief carried the penalty. And an earlier study of AI-written profiles found trust was similar when everything was AI-written or everything was human-written, and dropped only when people had to guess14.

There is a genuine counter-finding that should not be smoothed over. Ovsyannikova, Oldemburgo de Mello and Inzlicht found third-party evaluators perceived AI as more compassionate than expert humans, and that it held even when evaluators knew15. The two results are not straightforwardly reconcilable. The difference may sit in evaluating someone else’s exchange against receiving a message yourself.

A separate study on the smart-reply penalty above is worth setting against it directly, because it found the opposite. Mieczkowski and colleagues had participants exchange messages where one side used suggested smart replies, and the senders who used them were not rated lower on warmth or competence16. Put next to the belief-driven penalty in source 13, the two studies point at the same conclusion from different sides: actual AI assistance is not obviously punished. The suspicion that a partner is outsourcing effort is.

Liu, Mittal, Yang and Bruckman tested that suspicion directly, with three kinds of email, a product enquiry, a party invitation, and consoling a friend whose pet had died, each described as AI-written17. Trust fell as attributed AI involvement rose, and it fell hardest on the emotional email, which is what the authors expected going in. What they had not expected: trust rose, not fell, when the interpersonal stakes were higher, so the AI-labelled condolence still landed more trusted than the AI-labelled product enquiry. Their own hypothesis predicted the opposite ordering. Interviewed afterwards, all ten participants said flatly they would not want AI writing anything emotional for them, which did not match what their own ratings had just shown. What people say they will tolerate and what actually moves their trust are not the same thing here.

Khadpe and colleagues found the mechanism underneath both results. An AI-assisted label does not just mark a message down. It dampens what the message tells the reader about the sender’s character at all, in both directions: a warm message reads as less warm once labelled, and a cold one reads as less cold18. The label does not cost you credit for warmth so much as it makes the message stop counting as evidence of who you are.

What I take from it, running a product in this category: the useful thing is help writing something that is actually yours, and the damaging thing is a reply that reads as outsourced. Which lands in the same place as everything else here. The reply works when it demonstrates you engaged. Nothing that skips the engaging step survives contact with the person reading it.

Subtext works from what you wrote rather than replacing it, and shows you how much it changed. That design choice is not a marketing position, it is the only one the evidence supports.

Free to start, on your phone or in the browser.

What it looks like by context

The underlying problem is the same everywhere. How you solve it differs enough to mislead if you generalise.

  • Work and commercial. With strangers in a transaction, speed carries real weight. With colleagues, the documented error is overestimating urgency. Concrete specifics and a clear next action carry throughout.
  • Dating. Timing genuinely signals something here in a way it does not elsewhere, with a curvilinear optimum rather than a rule: a moderately delayed text after a first date outperformed both an immediate one and a long delay, and the same study found no support for a playing-hard-to-get mechanism19.
  • Support. Responsiveness is the whole game, and unresponsive support can be worse than none. Maisel and Gable found in a study of 67 cohabiting couples that support helped only when it matched what the person actually needed20.
  • Conflict. Hedge, acknowledge, name real agreement, and if you can, pick up the phone.

This is why Subtext reads the message in front of it rather than applying one rule everywhere. What “engaged” looks like in a work thread and in a first-date text are not the same thing, even though the underlying problem is.

Rules with nothing behind them

The five-minute lead response rule. The claim that replying within five minutes raises conversion 21-fold traces to vendor whitepapers. The underlying research on reply speed and contracts is real1. The multiplier is not.

The three-hour dating text delay. No empirical support, and the study closest to it explicitly failed to find the scarcity mechanism it would need19.

Professional email benchmarks and customer service expectation figures. Widely circulated, sourced to business news releases and vendor reports with no disclosed samples.

“31 per cent of people see texting as a daily source of anxiety.” Commercial blog, no traceable study.

Sources

Numbered in the order they appear above. Where a figure could not be verified against the primary source, the text says so. Checked 28 August 2026.

  1. Hart, E., VanEpps, E. M., Sezer, O., and Amir, O. (2026). Speed Is a Signal: When Faster Replies Increase Hiring Likelihood. Management Science, published online 3 June 2026. https://pubsonline.informs.org/doi/abs/10.1287/mnsc.2024.06185
  2. Giurge, L. M., and Bohns, V. K. (2021). “You don’t need to answer right away!” Organizational Behavior and Human Decision Processes, 167, 114 to 128. https://www.sciencedirect.com/science/article/pii/S0749597821000480
  3. Barber, L. K., and Santuzzi, A. M. (2015). Please respond ASAP: Workplace telepressure and employee recovery. Journal of Occupational Health Psychology, 20(2), 172 to 189. https://psycnet.apa.org/record/2014-49215-001
  4. Kalman, Y. M., and Rafaeli, S. (2011). Online Pauses and Silence: Chronemic Expectancy Violations in Written Computer-Mediated Communication. Communication Research, 38(1), 54 to 69. https://journals.sagepub.com/doi/10.1177/0093650210378229
  5. Lew, Z., Walther, J. B., Pang, A., and Shin, W. (2018). Interactivity in Online Chat: Conversational Contingency and Response Latency in Computer-Mediated Communication. Journal of Computer-Mediated Communication, 23(4), 201 to 221. 131 US adults, mean age 51.32. https://doi.org/10.1093/jcmc/zmy009
  6. Huang, K., Yeomans, M., Brooks, A. W., Minson, J., and Gino, F. (2017). It Doesn’t Hurt to Ask: Question-Asking Increases Liking. Journal of Personality and Social Psychology, 113(3), 430 to 452. Carries a published correction from March 2025, and it carries no expression of concern and no retraction. Data at osf.io/8k7rf. https://psycnet.apa.org/record/2017-39236-002
  7. Kluger, A. N., and Malloy, T. E. (2019). Journal of Personality and Social Psychology, 117(6), 1132 to 1138, with the authors’ reply at Yeomans, Brooks, Huang, Minson and Gino (2019), 117(6), 1139 to 1144. https://psycnet.apa.org/record/2019-70290-011
  8. Hart, E., VanEpps, E. M., and Schweitzer, M. E. (2021). The (better than expected) consequences of asking sensitive questions. Organizational Behavior and Human Decision Processes, 162, 136 to 154. https://www.sciencedirect.com/science/article/abs/pii/S0749597820303733
  9. Packard, G., and Berger, J. (2021). How Concrete Language Shapes Customer Satisfaction. Journal of Consumer Research, 47(5), 787 to 806. https://doi.org/10.1093/jcr/ucaa038
  10. Bevis, M., Schroeder, J., and Yeomans, M. (2026). Spoken disagreement is more constructive than written disagreement. Nature Communications, 17, article 5792. Open access; preregistrations and data at OSF. https://doi.org/10.1038/s41467-026-71669-5
  11. Yeomans, M., Minson, J., Collins, H., Chen, F., and Gino, F. (2020). Conversational receptiveness: Improving engagement with opposing views. Organizational Behavior and Human Decision Processes, 160, 131 to 148. https://www.sciencedirect.com/science/article/abs/pii/S0749597820303423
  12. Rubin, M., Li, J., Zimmerman, F., Ong, D. C., Goldenberg, A., and Perry, A. (2025). Comparing the value of perceived human versus AI-generated empathy. Nature Human Behaviour, 9(11), 2345 to 2359. Nine studies, 6,282 participants; all responses were AI-generated and only the label varied. https://www.nature.com/nathumbehav/
  13. Hohenstein, J., Kizilcec, R. F., DiFranzo, D., and colleagues (2023). Artificial intelligence in communication impacts language and social relationships. Scientific Reports. Two randomised experiments, 1,036 participants. https://www.nature.com/srep/
  14. Jakesch, M., French, M., Ma, X., Hancock, J. T., and Naaman, M. (2019). AI-Mediated Communication: How the Perception that Profile Text was Written by AI Affects Trustworthiness. CHI 2019. 527 participants. https://dl.acm.org/doi/10.1145/3290605.3300469
  15. Ovsyannikova, D., Oldemburgo de Mello, V., and Inzlicht, M. (2025). Third-party evaluators perceive AI as more compassionate than expert humans. Communications Psychology, 3(1), 4. Runs against the direction of source 12; both are named here rather than reconciled. https://www.nature.com/commspsychol/
  16. Mieczkowski, H., Hancock, J. T., Naaman, M., and Jung, M. F. (2021). AI-Mediated Communication: Language Use and Interpersonal Effects in a Referential Communication Task. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), article 17. https://doi.org/10.1145/3449091
  17. Liu, Y., Mittal, A., Yang, D., and Bruckman, A. (2022). Will AI Console Me when I Lose my Pet? Understanding Perceptions of AI-Mediated Email Writing. CHI 2022. Three email scenarios of increasing interpersonal stakes: a product enquiry, a party invitation, and consoling a friend over a pet’s death; ten follow-up interviews. https://doi.org/10.1145/3491102.3517731
  18. Khadpe, P., Wenzel, D., Loewenstein, G., and Kaufman, G. (2025). Explaining the Reputational Risks of AI-Mediated Communication: Messages Labeled as AI-Assisted Are Seen as Less Diagnostic of the Sender’s Character. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8, 1413 to 1423. https://doi.org/10.1609/aies.v8i2.36641
  19. Teichmann, F., Petrowsky, H. M., Boecker, L., Soliman, M., and Loschelder, D. D. (2026). Journal of Social and Personal Relationships. Curvilinear timing effect after a first date; no support found for a scarcity mechanism. https://journals.sagepub.com/home/spr
  20. Maisel, N. C., and Gable, S. L. (2009). The Paradox of Received Social Support: The Importance of Responsiveness. Psychological Science, 20(8), 928 to 932. 67 cohabiting couples. https://journals.sagepub.com/doi/10.1111/j.1467-9280.2009.02388.x