ArticlesWriting tools
AI can read a screenshot of your texts. It doesn't always read it right.
The published error rates for reading small, dark mode and non-Latin text in a screenshot, what three screenshot apps say about your upload, and whether stripping metadata does anything measurable.
By Samet Durgun · Co-founder of Subtext · 13 min read
Paste a screenshot of a chat into an AI tool and it usually just works. It tells you who said what, gets the joke, and drafts a reply that makes sense. “Usually” is the part nobody puts a number on, and it is worth separating from a different question this site already covers elsewhere. Once a model has the words in front of it, working out who is speaking across turns, tracking a question that never got answered, reading sarcasm, is the subject of the page on reading the message you received and the research behind reading a whole thread before replying. This piece is about the earlier step, whether the model reads the picture correctly in the first place.
I co-build Subtext, a tone-checking and reply-drafting app that takes a chat screenshot as one of its inputs, so I have reason to want a favorable answer here. I do not have one. What follows is what I could verify on a narrower, less flattering question, how good today’s vision-capable models are at reading a phone screenshot as an image, not a transcript, and what happens to a screenshot once you hand it to one of these apps. Every source below is one I opened myself.
Reading a screenshot is an image problem before it is a language problem
A chat screenshot is a picture of light, not a block of text a computer can already parse. Before a model can reason about what a message says, it has to turn pixels into characters correctly, in whatever font, size, contrast and layout the messaging app happened to use. That step is optical character recognition, riding on the same vision system that reads any other picture. Most accuracy claims made about AI and messaging concern the reasoning that happens after this step. Few concern the step itself.
The numbers that exist
The most direct test is a 2023 evaluation of GPT-4V’s OCR ability across several benchmarks1. On English scene text it scored 88.0% word accuracy on the CUTE80 benchmark and between 62.0% and 66.0% on three harder English sets, against a specialized OCR system’s 68.2% to 98.6% on the same four benchmarks. On Chinese scene text, the ReCTS benchmark, GPT-4V scored zero. In a separate multilingual test in the same paper, its text-spotting F1 score reached 82.49% for English and 83.42% for French, against 16.55% for Arabic and 1.36% for Chinese. The paper states plainly that “despite its versatility in handling diverse OCR tasks, GPT-4V does not outperform existing state-of-the-art OCR models,” and that it “showed limitations when dealing with non-Latin languages”1.
That is one model generation, tested on curated scene-text datasets, not messaging screenshots, so read it as a floor under the general problem, not a verdict on any app available today. It is still the most direct, checkable number against the plain claim that AI reads non-Latin scripts inside an image as well as it reads Latin ones. On this evidence, it does not.
Resolution matters too. The same paper found “a positive correlation between the input image resolution and the recognition performance”1, which works against a compressed or heavily downscaled screenshot, exactly what many phones produce when an image gets re-saved or reshared, before the model reads a single word. A narrower 2025 preprint tested this more precisely, though only on single Japanese kanji characters rendered at controlled sizes, not on chat screenshots. Multimodal models matched a dedicated OCR method at roughly 300 pixels per inch, and “their performance deteriorates significantly below 150 ppi”2. Whether that threshold holds for small text inside a compressed phone screenshot is untested. The direction, smaller and lower-resolution text degrading faster for these models than for purpose-built OCR, is the part two separate papers agree on.
On dark mode specifically, and on low-contrast text more generally, I could not find a published evaluation measuring multimodal model accuracy against light-mode equivalents of the same screenshots. Classic OCR pipelines routinely invert dark images before processing, because contrast reversal is a known trigger for misreads in that older technology, which suggests the underlying concern is reasonable. Nobody has run the equivalent test on the multimodal models people now use to read a chat screenshot, at least not in anything I could find, so I am not going to state a number I cannot support.
Emoji sit inside the same image, and the research on them answers a different question than the one people usually ask. It is less about whether a model can tell that a shape is an emoji, and more about which emoji it is looking at, because the same character does not draw the same picture everywhere. The Unicode Consortium standardizes the code point and its meaning, not the artwork. Apple, Google, Samsung and every other vendor draw their own glyph for it3. In a survey of 710 Twitter users who had just posted an emoji-bearing tweet, at least 25% did not know their emoji could look different on someone else’s phone, and after being shown how one of their own tweets actually rendered on another platform, 20% said they would have edited it or not sent it at all3. That gap exists between two people looking at the same character on two different phones. A screenshot narrows it further. Whatever the sender’s platform drew is flattened into a still image, and a model reading that image cannot recover which underlying emoji it was without the platform’s own font already baked into the pixels. Whether that combination, cross-platform ambiguity plus screenshot compression, changes how often a model names the wrong emotion behind an emoji is a test nobody appears to have run yet.
Failure modes the workflow pages here do not cover
A handful of specific chat features make correct reading harder in their own way, beyond the pixel-level problem above. Some of what follows rests on a study. Some is a plain description of how the interface works, with no dedicated research behind it, and I have tried to keep the two apart so one does not borrow the other’s weight.
Group chats and multi-participant threads. A one-to-one screenshot gives a model two visual groups to sort speech into, one side and the other. A group thread breaks that pattern. Three or more people can share one bubble color, a sender’s name may appear only above the first message in a run and not on every line after it, and a cropped avatar can be the only visual cue for who is speaking. I did not find a study testing model accuracy on attributing lines correctly inside a group-chat screenshot specifically. The closest research measures a related but different problem, speaker attribution in written meeting transcripts, not screenshotted bubbles, and that ground is already covered on the page about reading a whole conversation thread. What is here is a mechanical description of the screenshot version of the same problem, not a tested finding. More speakers sharing fewer visual cues per line is a harder attribution task on its face, whether or not anyone has measured how much harder.
Tap reactions and emoji reactions. A small emoji pinned to the corner of someone else’s bubble, a heart, a laugh, a thumbs up, carries two separate risks. The first is whether a model can correctly identify the small, sometimes overlapping glyph in the first place, which I found no direct test of. The second is what the reaction means once it is correctly read, and that part has been studied. An analysis of more than 650,000 reaction-bearing Telegram messages found that positive reactions dominated the responses regardless of whether the underlying message read as neutral or negative, concluding that emoji reactions “do not reliably function as indicators of emotional mirroring or resonance of the content”4. That is a large sample from one platform and one topic area, crypto-related channels, and it remains a preprint, so treat the precise figures as provisional. The point that matters for a screenshot reader holds regardless of those caveats. Naming the reaction emoji correctly does not tell anyone, model or person, what the reaction actually signals.
Quoted replies and threading UI. Most messaging apps let you reply to a specific earlier message, rendering the quoted snippet as a smaller, dimmer block above the new one. In a screenshot, that visual treatment, reduced size, lighter color, sometimes a vertical rule, is the only signal that a line is quoted context rather than a new message from whoever’s turn it visually looks like. Crop the image a few pixels differently and the distinction can disappear. I found no study measuring how often a model mistakes a quoted snippet for a fresh message or the reverse. This is a plausible mechanical failure mode built into how the interface represents threading, not a tested one.
Disappearing and ephemeral messages. An ephemeral message is designed to leave no record once viewed, which is exactly what a screenshot defeats. Before any AI question enters the picture, forensic researchers had already shown the underlying premise is shakier than users assume. Testing WhatsApp, Snapchat and Telegram’s disappearing-message features directly, one study recovered supposedly deleted content from device data and cloud backups in several of the scenarios it tried5. That is a finding about the platforms, not about AI reading a screenshot of one, and it measures nothing about a model’s accuracy. What it does establish, independent of any AI, is that a screenshot is not the only way ephemeral content outlives its intended deletion. Whether an app’s own on-screen blur or a “screenshot taken” banner introduces visual artifacts a model could mistake for content is, again, a plausible concern with no dedicated study behind it.
Stickers. A sticker carries meaning through pose and expression more than through words, which makes it a harder read for a machine and, it turns out, for people too. A study of five student discussion groups using stickers in real project chats, drawing on interviews with seven of the participants about their own sticker use, found a mismatch between what the sender meant and what the receiver understood in 34.7% of the stickers examined, chiefly because of ambiguous facial and bodily expressions in the artwork itself6. That is a small, qualitative study of one population, university students at one Asian institution, and it measures human misreading, not a model’s. It is still real evidence that a sticker is an ambiguous signal even between people who already know each other, which sets a low ceiling for how well any reader, human or machine, should be expected to do from the image alone.
Code-switching mid-conversation. Switching languages inside one message, or between messages in the same thread, is common in bilingual and multilingual texting and a documented hard case for language models generally. A 2025 survey of the field states plainly that “most LLMs still struggle with mixed-language inputs”7, without a number that would transfer cleanly to a screenshot with two languages inside one speech bubble. I could not find a study isolating code-switching as a screenshot-reading problem specifically, apart from the broader multilingual-text problem the survey describes. Treat this as a real, documented difficulty in the underlying technology, applied here to a case nobody has tested directly yet.
What three screenshot apps say about your upload
Subtext is not the only app built to read a chat screenshot, and the useful comparison is what each one’s own current documentation says happens to an upload, not what its marketing implies. The wider field of AI writing tools is mapped separately; here I checked three products named there, against their own privacy pages, for one narrow question: retention, deletion, and use in training a model.
Keys AI Texting Coach, built by Charmed Inc. around reading screenshots for dating and messaging advice, appears to no longer be live. Apple’s own iTunes lookup service returns zero results for its App Store listing, app id 1510154956, as of 27 September 20268, and its known web domain redirects every path, including the one that used to hold its privacy policy, to a domain-parking page. I could not locate a current privacy policy for this app through any channel. Whatever it once said about screenshot handling, I cannot verify it today, and I am not reconstructing a claim from a page that no longer loads.
Mei, an Android default-messaging app with an optional AI assistant, states in its current terms, last modified 3 October 2023 and checked 27 September 20269, that message content including “images, photos, audio or videos” is stored on its servers “for the sole purpose of communication” and “used by us in no other way.” Separately, describing what its optional AI assistant actually receives, the same document says the last 20 messages sent to the AI as context include “emojis, reactions, and URLs,” but that “messages with attachments, voice messages, and images aren’t collected” for that purpose. Read together, Mei’s own documentation says its AI suggestion feature does not process image content at all, screenshots included. The storage clause about photos concerns ordinary message delivery within the app, not anything sent to the AI. A separate, opt-in AI assistant feature works differently: when a user turns it on, Mei uploads the device’s whole SMS and MMS message database, including the text of each message, plus hashed phone numbers and contacts’ names, and the policy states that data is “used for training AI models.” That clause does not name screenshots or images specifically, but it is about training on message text broadly, not only on contact metadata.
ConfiText’s privacy policy, effective 10 February 2026 and checked 27 September 202610, addresses only “text you enter in the App” and does not mention photos, images or screenshots anywhere across its nine sections. Its own marketing site describes exactly one input method, pasting a message the user already wrote, and no screenshot or photo upload feature appears anywhere on the page. As things stand, the retention question does not appear to apply to ConfiText. The product, per its own current site, does not take a screenshot as input in the first place.
Across the three, the count is one app that no longer appears to exist, one whose own documentation excludes images from what its AI actually reads but does train on message text when its assistant is switched on, and one that does not take a screenshot at all. That does not settle how a screenshot-reading app should handle retention. It does mean that comparing three named competitors’ screenshot policies produced less agreement, and less applicability, than the premise assumed going in.
Stripping the metadata is common sense, not something measured
Every photo a phone camera takes can carry EXIF data, device model, timestamp, sometimes location, embedded in the file. Advice to strip this data before sharing a photo is everywhere online. I looked for a controlled study measuring whether removing that data, or manually redacting a screenshot’s visible contents, produces a measurable drop in real-world privacy harm, as against a theoretical exposure that a redaction merely closes on paper. I did not find one. What exists is descriptive, security writing that explains which fields exist and how to remove them, not an experiment comparing outcomes for people who stripped metadata against people who did not. Treat the advice as reasonable and untested, not as a measured protection.
On what Subtext itself does with a screenshot’s metadata, I read the backend’s system prompts and tool definitions directly, the same file this site’s own rules require checking before any product claim (functions/src/subtext_prompts.ts). Nothing in that code reads, strips, inspects or otherwise touches a screenshot’s file metadata. Images are handed to the underlying vision-capable model as visual input, the same way any other image would be, and no separate metadata-processing step appears anywhere in the prompts or tool definitions I read. I am stating that as an absence, not a feature. I found no evidence Subtext does anything with a screenshot’s metadata, in either direction, and I am not claiming a capability the code does not show. No independent evaluation of how accurately Subtext itself reads a screenshot exists either, the same gap that runs through every app named above. What the privacy policy does confirm, and what the two workflow pages on this site already state, is the general rule that applies to every attachment including a screenshot. A conversation, and anything attached to it, deletes itself five days after last use or ninety days if pinned, and nothing uploaded is used to train an AI model.
What this adds up to
Put the four pieces together and the picture is uneven by design, not by accident. The one part with a real, checkable accuracy number, general OCR performance inside a vision model, is mixed: strong on English, measurably weaker on non-Latin scripts and on low-resolution text, tested on curated benchmarks, not actual chat screenshots. Four of the six extra failure modes above cite real research on an adjacent question: what a reaction actually signals once it is read correctly, how reliably ephemeral apps erase content in the first place, how often people misread each other’s stickers, and how badly language models handle mixed-language text in general. None of those four studies tested the exact case of a model reading it off a screenshot. The other two, group-chat attribution and quoted-reply confusion, have no research behind them at all, adjacent or otherwise, just a mechanical description of how the interface works. The competitor check found less to compare than the premise assumed: one app gone, one that excludes images from its own AI feature by its own account, one that never accepted screenshots at all. And the metadata question turned out to be advice everyone repeats and nobody appears to have measured.
None of this means a model reading your screenshot is unreliable in general use. The largest, most direct study here still measured a model reading most English scene text correctly. It means the specific claim, that a vision model reads a chat screenshot as reliably as it reads plain typed text, has real, quantified exceptions, some of them large, and several hard cases nobody has measured at all.
That is where things stand today. Subtext is one more app doing this, no more and no less verified than the rest above.
Try Subtext in your browserScan it with your phone camera to install.Try Subtext in your browser
Checked 27 September 2026.
Sources
- Shi, Peng, Liao, Lin, Chen, Liu, Zhang and Jin (2023). Exploring OCR Capabilities of GPT-4V(ision): A Quantitative and In-depth Evaluation. arXiv. Scene-text and document OCR benchmarks against one 2023 model generation, not messaging screenshots.
- Inoue (2025). Context-Independent OCR with Multimodal LLMs: Effects of Image Resolution and Visual Complexity. arXiv preprint, single author. Tests 100 isolated Japanese kanji characters at controlled font size and resolution, not chat screenshots.
- Miller Hillberg, Levonian, Kluver, Terveen and Hecht (2018). What I See is What You Don’t Get: The Effects of (Not) Seeing Emoji Rendering Differences across Platforms. Proceedings of the ACM on Human-Computer Interaction, CSCW. Survey of 710 Twitter users about their own emoji-bearing tweets, not a test of a model reading them.
- Tardelli, Alvisi, Cima, Cresci and Tesconi (2025). Emoji Reactions on Telegram: Unreliable Indicators of Emotional Resonance. arXiv preprint. Over 650,000 reactions on crypto-related Telegram messages, one platform and one topic area.
- Heath, MacDermott and Akinbi (2023). Forensic analysis of ephemeral messaging applications: Disappearing messages or evidential data? Forensic Science International: Digital Investigation. Tests WhatsApp, Snapchat and Telegram’s own deletion behavior, not AI reading of a screenshot.
- Tang, Hew, Herring and Chen (2021). (Mis)communication through stickers in online group discussions: A multiple-case study. Discourse & Communication. Five discussion groups and seven interviewees at one university; measures human misreading, not a model’s.
- Sheth, Sinha, Patil, Beniwal and Singh (2025). Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities. arXiv preprint survey, no screenshot-specific test.
- Apple. iTunes Lookup API result for Keys AI Texting Coach, app id 1510154956, checked 27 September 2026. Zero results returned; the app no longer appears to be listed.
- Mei. Terms of Service and Privacy Policy, last modified 3 October 2023, checked 27 September 2026.
- ConfiText. Privacy Policy, effective 10 February 2026, checked 27 September 2026.