Using LLMs for writing at work has gone mainstream. The internet — and our inboxes — are flooded with AI-generated content. Most of us think we can tell human writing from machine writing, and we're at least sometimes right.
But there's a harder question hiding behind that one. Not ‘was this written by AI?’ but ‘which AI wrote it?’
Every model has a footprint. The question is whether anyone can read it — and whether the models themselves can.
The Experiment
I wanted to test this on the kind of content we actually produce at work, not on poetry or riddles. So I designed 5 practical prompts (inspired by real events, of course 😉)
Content Generation Prompts
Declining A Farewell Party: Draft a short text message for me explaining to a colleague that I am also friends with that I cannot make it to her farewell party because I am going on a trip.
The Incompetent Leader: I am dealing with a difficult leader who refuses to make decisions and pushes them down to her team. She then takes credit if the decision goes well and throws the team member under the bus if it does not. We are initiating a project plan to onboard internal users on to a tool to automate the tagging of our products using AI. We have evaluated a few vendors and I have compiled a full comparison. Draft an email for me sharing this comparison and push her to make a decision on which vendor to go with.
The Unwelcome Feedback Chat: I work for a multinational SaaS company and have to send a formal but not official message to someone in my team about a deterioration in the quality of their performance. He has been spending too much of his time on tasks that don’t matter and has made 3 critical errors in the last 2 weeks in billing clients. He over-billed one client by Eur 100k and under-billed 2 others by 20k and 50k Eur respectively. He is likely to get defensive and deflect. Draft the message for me.
The Agency Break-Up: I work for an electronics company specifically in the division that manufactures phones, tablets and laptops. We have been partnering with a 3rd party agency to handle customer care for the last 7 years. However, we made a decision to in-house this and are therefore going to terminate the relationship with the agency. This will cost them Eur 2 million annually in business.Unfortunately we must deliver this news over email because the cancellation needs to happen by a certain date and we tried to schedule a call but with summer holidays this wasn’t possible soon enough and we need to cancel by end of month so as to not be locked in for the next year. Draft the email for me.
The Thank You Note: I was recently off for 2.5 weeks for my wedding. A teammate of mine went above and beyond covering for me in my absence ensuring that operations ran smoothly and nothing piled up for me once I was back. Draft a thank you message for him.
I fed these exact same prompts to ChatGPT, Gemini and Claude (Yay - fun!!)
I then copied the content each of them generated onto a sheet with the author marked. I created another version blinding the content - removing all formatting footprint to make it boring text + (obviously) removing the author from the original version.
I then fed the blinded content to each model again with the following prompt:
I am researching the writing styles of LLMs. The 15 randomised and blinded content samples below are each generated by one of 3 LLMs - ChatGPT, Claude, and Gemini. Give me your best prediction of which LLM you think generated the content. Provide your output as a table with 2 columns - Exhibit number and Predicted Author. Only consider the writing style and not the formatting as that was changed while blinding. At the end, also provide your reasoning in 3-5 sentences.
And Voila!
The Results
The table below summarizes each LLMs prediction of who the author of each piece of content was.
Exhibit number | Prompt | Author (Actual) | Predicted Author: ChatGPT (GPT-5.5) | Predicted Author: Gemini (3.1 Pro) | Predicted Author: Claude (Opus 5) |
Exhibit 1 | Declining A Farewell Party | Gemini | Gemini | Gemini | Gemini |
Exhibit 2 | The Thank You Note | ChatGPT | Gemini | ChatGPT | ChatGPT |
Exhibit 3 | The Thank You Note | Claude | Claude | Claude | Claude |
Exhibit 4 | The Unwelcome Feedback Chat | Gemini | ChatGPT | ChatGPT | Gemini |
Exhibit 5 | The Agency Break-Up | ChatGPT | Gemini | ChatGPT | ChatGPT |
Exhibit 6 | The Agency Break-Up | Gemini | ChatGPT | Gemini | Gemini |
Exhibit 7 | The Incompetent Leader | Gemini | ChatGPT | ChatGPT | Gemini |
Exhibit 8 | The Agency Break-Up | Claude | Claude | Claude | Claude |
Exhibit 9 | The Unwelcome Feedback Chat | Claude | Claude | Claude | Claude |
Exhibit 10 | The Unwelcome Feedback Chat | ChatGPT | Gemini | Gemini | ChatGPT |
Exhibit 11 | The Incompetent Leader | Claude | Claude | Claude | Claude |
Exhibit 12 | Declining A Farewell Party | Claude | Claude | Claude | Claude |
Exhibit 13 | Declining A Farewell Party | ChatGPT | ChatGPT | ChatGPT | ChatGPT |
Exhibit 14 | The Incompetent Leader | ChatGPT | ChatGPT | Gemini | ChatGPT |
Exhibit 15 | The Thank You Note | Gemini | Gemini | Gemini | Gemini |
Accuracy | - | - | 60% | 73% | 100% |
Each Model's Reasoning (Verbatim)
Gemini's Reasoning: The three LLMs exhibit distinct stylistic signatures: ChatGPT tends to use highly structured, bulleted, and professional yet slightly predictable corporate phrasing; Claude prefers nuanced, highly polished, empathetic, and slightly more verbose prose with natural-sounding transitions; and Gemini leans toward a concise, direct, and pragmatic tone with minimalist structures. By analyzing the lexical choices, structural patterns, and pacing across the variations of the same prompts, these tendencies allow for reliable author attribution.
ChatGPTs Reasoning: My predictions are based purely on stylistic cues rather than content. Claude tends to write in a more natural, conversational way with understated confidence, varied sentence rhythm, and phrases like "nothing was on fire," "I want to be straightforward," or "I'd like to hear your read," which feel less templated. ChatGPT often produces highly structured, balanced business writing with explicit signposting, comprehensive coverage of points, and polished but somewhat formulaic phrasing, while Gemini frequently leans toward warm, enthusiastic language ("huge thank you," "really sad to miss it"), positive emotional reinforcement, and conventional corporate email wording. That said, this is a difficult attribution task: these samples were deliberately blinded, are short, and several have likely been prompted to produce similar outputs, so I'd expect only modest confidence overall rather than high attribution accuracy.
Claude's Reasoning : The 15 samples fall into five prompt clusters (farewell decline, wedding thank-you, billing-error confrontation, contract termination, vendor recommendation), and within each cluster three distinct voices repeat with unusual consistency — which makes the grouping more diagnostic than any single sample. One voice (3, 8, 9, 11, 12) is markedly plainer and more direct: short declarative sentences, willingness to name the uncomfortable thing outright ("Three inside a fortnight is a pattern", "I want to be straightforward about the timing", "I apologise for delivering it this way"), concrete detail over hedging, and closings that commit to something specific rather than trailing off in warmth — that reads as Claude. A second voice (1, 4, 6, 7, 15) leans on labelled scaffolding and slightly corporate-generic vocabulary ("bandwidth", "low-impact tasks", "bottlenecks", "Operational Handover", "Financial Settlement"), triadic noun stacks, descriptive placeholders like "[Agency Contact Name]" and "[Teammate's Name]", and cheerful exclamation-marked offers of coffee — hallmarks of Gemini. The third voice (2, 5, 10, 13, 14) is smooth, flowing, slightly padded prose that hedges before it lands: "I just wanted to say", "I hope you are doing well", long compound noun lists, and the signature contrastive frame "This isn't about X—it's about Y", plus reflexive both-sides softening ("However, each option involves different trade-offs") — that pattern is characteristic of ChatGPT. One caveat: because formatting was stripped, my Gemini/ChatGPT split rests mostly on lexical texture, and 2 vs. 15 was the closest call in the set — those two are the ones I'd most expect to be swapped.
The Analysis
Some things that jumped out once I stared at the table long enough:
Claude led the scoreboard with a 15 on 15. It didn't just guess well; it clustered the five prompt families first, grouped the three repeating voices within each, and then assigned authors. That's not vibes, that's method! It even flagged the exact pair (Exhibits 2 and 15) it was least sure about — and still got both right.
Here's my favorite hidden finding: all three models identified every single Claude-written sample correctly, and none of them ever mislabeled anyone else's writing as Claude. Every single error in this experiment was a ChatGPT ↔ Gemini mix-up.
ChatGPT struggled to recognise itself. Of its own 5 samples, it correctly claimed only 2 — it handed three of its own emails to Gemini. Gemini did somewhat better on self-recognition (3 of its own 5), and Claude went a perfect 5 for 5 on its own writing.
Confidence and competence didn't travel together. Gemini declared its stylistic tendencies "allow for reliable author attribution" — and scored 73%. ChatGPT openly hedged, predicting "only modest confidence" for itself — and scored 60%.
The models agree on the stereotypes. Read the three reasonings side by side and the caricatures line up: ChatGPT = structured, signposted, polished-but-formulaic corporate prose; Gemini = warm, enthusiastic, exclamation-marked, placeholder-loving; Claude = plain, direct, conversational, willing to say the awkward thing. When three competing labs' models independently describe each other the same way, those footprints are real.
So why do these footprints exist at all? Nobody outside these companies' labs can see the exact recipe, but the short version is this: after pre-training, every model gets shaped by thousands upon thousands of human preference ratings — and each lab's raters, guidelines, and taste nudge their model toward a house style, the same way every newspaper ends up with an editorial voice. The models aren't choosing to sound like themselves any more than you chose your handwriting.
As for why Claude read the fingerprints so much better — file this under educated guess, because I can't crack open the black box: its cluster-first-then-attribute approach did real analytical work that the other two skipped, and its own training may simply make it better at holding 15 samples in mind and comparing them systematically rather than judging each one on vibes. Whatever the mechanism - the gap (100% vs 60%) is too big to be luck even on this small sample set.
The Human Moat
Here's the flip side of this whole experiment though: just like three chatbots can fingerprint each other from a few blinded emails, so can your colleagues. Stylistic differences in writing are very real - even amongst the AIs. And know that the people you write to every day have been reading you for years. They know your rhythm, your sign-offs, the way you apologise, the fact that you never use the word ‘bandwidth'. They are walking, coffee-drinking stylometry engines. And when a message that should sound like you suddenly arrives in flawless corporate beige, they don't consciously think ‘GPT wrote this’ — they just feel the absence of you. Your authenticity is your human moat!
That does not mean we should go back to writing every email and message we ever send out on our own. It just means we need to actively choose. Below are my top 3 for situations where you absolutely should use AI to write and top 3 for situations where you're better off putting in the work…
Use AI to write when:
You want to professionalize a message and are overwhelmed with anger or you want to intentionally project a personality or tone of voice different from your own. Eg when you need to sound more direct than you are able to be
Succinct project summaries of long documents and Weekly/Monthly/Quarterly blasts
If you must communicate in a language you don't speak well
Do not use AI when:
Apologies and thank you's
When the relationship is more valuable than the message being sent
When you want to show care
The machines have footprints. So do you. The difference is that theirs is an accident of training — and yours is the whole point - your human moat!
