I graduated college 4 months ago with a Computer Science degree. And I am grateful that I did not have to take a single English class in my 4 years in college. Why? In high school, English Language Arts (ELA) was my worst class and I wanted no part of it.
I did though have to write extensive essays in my small stint as a Communications minor, I dropped the minor my junior year lol. My favorite essay I wrote was about my favorite movie of all time, Get Out from the great Jordan Peele. And my assignment was to write an intercultural media analysis essay on a movie or TV show where intercultural communication is a central part of the narrative. And as a Black person I loovveeee writing essays about Black things.
It is important to note to you dear reader, I submitted this essay on April 25, 2023, which is before the use of ChatGPT became rampant in everyday use in schools.
This past August, Anthropic, the creators of Claude introduced a text watermarking feature to comply with transparency requirements as stated in documents such as EU's AI Act. You can read more about it here in How Claude's Watermarking Works. Another thing we have seen in masse, is schools trying to crackdown on AI-generated papers, a favorite tool at my university my junior and senior year was the infamous, Turnitin (which accused me of turning in AI-generated papers I completely wrote myself).
So let's put a human and AI head-to-head.
Teachers swear up and down they can tell the difference between an essay written by their own student vs generated by AI. They either read it and know or used tools like Turnitin's AI checker. But what does the data say about the difference?
Methodology
I took my COMM 4323 essay as the human baseline and had Claude completely generate the same essay given the same prompt and essay structure requirements.



Prologue: The Raw Numbers
After cleaning up both essays: removing the title page, headers, footers, and anything outside of the essay's contents what are the differences?
| human.txt | ai.txt | |
|---|---|---|
| Word count | 2,542 | 2,659 |
| Sentence count | 109 | 81 |
| Avg. sentence length | 23.4 words | 32.8 words |
| Reading grade level (Flesch-Kincaid) | 9.3 | 17.7 |
| Movie citations (parenthetical) | 10 | 8 |
The biggest thing that stands out is Claude's lower sentence count over mine, but it's significantly higher average sentence length.
According to my fiance, the biggest thing that stood out to her was that I was a sophomore in college writing at a 9th grade reading level.
Act 1: Writing Style
Chapter 1: The Stylistic Fingerprint
I have been told by so many teachers that they "know what my writing looks like" or in terms of my computer science classes, they "know what my code looks like". They call this a fingerprint, everyone has a specific way and quirk on things they write, as it stems from how we learned and developed our writing skills.
What does my fingerprint look like in comparison to Claude's?

My essay and Claude's don't have subtle differences but instead they occupy completely different grounds on five out of eight axes. Claude pushed more out on sentence length, reading grade level, hedge language, and citation-verb use. I, on the other hand, seem to excel in sentence and paragraph variance.
In CoNcLuSiOn, Claude's essay is not "more" of anything but instead it is more even in structure.
Chapter 2: Words Used
Words are the powerhouse of an essay, and a huge red flag that always makes teachers think I used AI is when while writing my essays I pull out the good ole thesaurus (I just hated using the same simple words over and over).
Apparently, big words = AI.
This comment I would always get honestly diminished my skill set and made me feel my teacher thought I was dumber than I was only because I used a thesaurus. But what does the data say?

Filtering to words that appear at least three times across both essays and ranking by rate difference surfaces a register split. My essay's most distinctive words are concrete nouns tied directly to Get Out's plot: black, culture, Chris, cult, Dean, this is the vocabulary of someone narrating what happens on screen. While Claude (who probably DID NOT watch the move 1,000 times), used distinctive words that are abstract and theory-naming: cultural, exploitation, historical, racial, value, the vocabulary of someone naming a framework and then applying it.
Chapter 3: Habits of Phrasing
Since the essays are not the same size, normalizing by length was the best way to compare both essays contents. The biggest gap is between citation verbs. Something I noticed when reading Claude's essay is it would lean more on using the scaffolding sentence of "Samovar et al. argue/explain/emphasize" to introduce almost every textbook's theoretical claim, it did this at a rate of 5.6 per 1,000 words. Compared to me, where I directly integrated the textbook's text in literal quotes instead of just indirectly referencing it.

Hedge words and transitions both run higher in Claude's essay too (4.5 vs. 2.8, and 2.6 vs. 1.6 per 1,000 words), consistent with a writing process that leans on connective tissue to make each paragraph's logic explicit.
Chapter 4: Sentence-Length Rhythm
For as long as I can remember, I hated writing long sentences in my essays. I was never a fan of using semicolons (;) to continue a thought, I would just write a new sentence. This was the biggest difference between my essay and Claude's

My sentences cluster in the 8–23 word range with a real tail on both sides, six sentences under 8 words, four sentences over 48. While Claude's essay has almost no sentences under 16 words and its mode sits a full bin higher, at 32–39 words. Fewer short sentences means less rhythmic contrast, no one-line assertions dropped in for emphasis, no fragments, no rhetorical pause. Every sentence in Claude's essay is doing roughly the same amount of syntactic work.
NOTE: a mode is a statistical number that represents the value that appears most frequently in a dataset.
Chapter 5: Paragraph Shape
Average paragraph length is nearly identical: 127.25 words (mine) vs. 126.57 (Claude).

Based on the prompt itself, this makes sense, since both essays were built around the same "roughly one page per section" structure. But my essay's paragraph-to-paragraph standard deviation is 59.2 versus 42.0 for Claude, 41% higher. Look at the shape of the two lines: my essay swings from a 61-word paragraph up to a 273-word paragraph and back down within a few beats. While Claude's line stays inside a much narrower band throughout.
Act 2: Semantic Meaning
Get Out is my favorite movie of all time, and I relate with a lot of the race related scenes in the movie, because of this, the true intented meaning between my essay and Claude's is different.
Text analysis about the meaning of words is called semantic meaning.
Chapter 6: Semantic Map of Every Sentence
NOTE: Principal Component Analysis (PCA) is a statistical method that reduces the amount of variables in a large dataset but keeps the important patterns.
To visualize sentence distribution, every sentence from both essays was converted into a *300-dimensional embedding and reduced to 2D via PCA (capturing ~35% of total variance). Two main observations stand out:
-
Substantial overlap: The two sentence clouds share heavy overlap, as expected given that both essays analyze the same film using identical analytical categories. Rather than signaling plagiarism or copying, this overlap indicates basic adherence to the prompt.
-
Distinct edge cases: The highest cross-essay similarity (0.98) links one of my sentences ("Family is the building block of culture...") and Claude's sentence ("Family is one of the primary institutions through which culture is transmitted...") that virtually paraphrase each other. Conversely, the most isolated sentence (0.48 similarity) is a basic foundational definition which comes from my essay: "Intercultural communication concerns the communication between two different cultures." It was interesting that Claude avoids this kind of redundant baseline definition, driving the separation at the edge.
*What does the 300 mean? Each sentence gets a numerical 'meaning fingerprint' made up of 300 numbers. That's impossible to visualize, so I used PCA to flattens those 300 numbers down to just 2 (like squashing a globe into a flat map) so the sentences could be plotted on a simple chart. You lose some precision in the flattening, but you keep the big picture: which sentences are talking about similar things, and which ones stand apart.

| Closest Pair | Most Isolated Sentence | ||
|---|---|---|---|
| Mine: | Family is the building block of culture, and can best be described as an important “social unit [that] forms the basic cooperative structure that ensures an individual’s primary needs and provides the necessary care for children to develop as healthy and productive members of the group and thereby ensure its future.” (Samovar et al., 2017, p.73). | Mine: | Intercultural communication concerns the communication between two different cultures. |
| Claude: | Family is one of the primary institutions through which culture is transmitted from one generation to the next. | Claude: | NA |
Chapter 7: Semantic Redundancy
Here is where I see a difference between the word intent between my essay and Claude's. While I read Claude's essay, it seemed repetitive in so many areas, and I was right. Claude's sentences were 4.2% more self-similar than mine.
The essay was more repetitive in its own sentences meaning Claude just kept on returning to the same semantic territory rather than ranging across new ground sentences, i.e. it kept using the same citation verbs to introduce a textbook theory.

The chart itself covers semantic meaning but it asks, "how similar is each essay to itself" instead of Chapter 6's "where do the two essays overlap".
My essay covers way more distinct semantic ground per sentence, while Claude paraphrases and reinforces a narrower set of ideas a lot more tightly. This is that "evenness" fingerprint covered in Chapter 1.
Conclusion
So, who writes better, me or Claude?
Honestly, after all this, I don't think that's even the right question anymore. "Better" implies one of us is doing it wrong, and the data doesn't really back that up, it just shows two completely different processes landing on the same assignment.
My essay is uneven on purpose (even if I didn't know it at the time). Long paragraph, short paragraph. Eight-word sentence, forty-word sentence. A direct quote instead of a paraphrase because I wanted you to hear Larry A. Samovar et al. in their own words, not mine. That unevenness, the swings in Chapter 4 and Chapter 5, is what a teacher is actually picking up on when they say they "know my writing." It's not vocabulary, it's rhythm.
Claude's essay is even on purpose too, just for a different reason. It's not trying to have a fingerprint, it's trying to be consistently legible: same sentence length, same hedge words doing the same scaffolding work, same citation-verb pattern every time a theory shows up. That's not a flaw, it's just what "optimize for clarity across every paragraph" looks like when you zoom out to 2,659 words.
The one place the data genuinely surprised me was Chapter 7. I expected Claude's essay to cover more ground semantically since it's longer and uses more abstract, theory-naming language. Instead it was more self-similar, more repetitive, circling the same ideas with slightly different words. My essay, messier on the surface, actually ranges across more distinct semantic territory sentence to sentence. Turns out "sounds smarter" and "says more" aren't the same thing.
So if a teacher tells you they can "just tell," they're not wrong, but they're probably not reading for vocabulary or citation count like Turnitin is. They're reading for the swing, the surprise short sentence after three long ones, the paragraph that runs long because I couldn't shut up about the Sunken Place. That's the part that's hard to fake, and it's the part this whole breakdown ended up proving mattered most.
Methodology & data footnotes
How this came together
I took my best (non-AI) written analysis essay I did in college and used text analysis to compare a completely AI-generated essay on the same prompt and essay structure.
Browse the analysis ↗