ScribeLens

← Back to all articles

What the Gettysburg Address looks like to an AI detector

By the ScribeLens Team ·

Two hundred and seventy-two words, delivered on a Pennsylvania battlefield in November 1863 — the Gettysburg Address is held up in every writing class as a model of economy and rhythm. We wondered what a modern, pattern-based AI detector would make of it. So we ran it, alongside Lincoln's Second Inaugural Address, through ScribeLens exactly like any guest scan, and we're publishing everything it reported: the verdict, the score, the sentence-level evidence, and the one sentence that actually got flagged.

What we actually ran

We combined the full text of the Gettysburg Address (1863) and the Second Inaugural Address (1865) into a single 983-word excerpt, sourced verbatim from Project Gutenberg's editions of Lincoln's speeches (ebooks #4 and #925), and submitted it to ScribeLens as a guest scan — no special handling, no cherry-picked passage, the same pipeline any visitor uses.

This is one experiment on one document, not a benchmark or a study across a text corpus. We're not claiming it tells you how AI detectors handle "classic literature" in general, and we're not publishing a pass rate. It's a look at what the evidence view actually shows on one famous, heavily rhetorical text — which is genuinely worth seeing, given how often "formal writing gets flagged" comes up as a concern.

Two honest caveats before the numbers. First, Project Gutenberg's edition of the Gettysburg Address uses non-standard elliptical punctuation between clauses, which split several sentences into short fragments — including the "government of the people" line, more on that below. We kept the text verbatim rather than editing it to look cleaner. Second, and more interesting: the Gettysburg Address and the Second Inaugural are almost certainly present, more than once, in the training data behind large language models — they're two of the most reproduced passages of English prose that exist. Memorized text tends to read as unusually predictable to a statistical model, which can bias a detector's perplexity-based signals toward "AI-like," not away from it. If anything, that confound should make the result below harder to earn, not easier.

The verdict: likely human, high confidence

ScribeLens read the combined text as 78% human-like / 22% AI-like, with a "likely human" verdict at high confidence — its top confidence tier. The report's own confidence statement reads: "We are highly confident this text was written by a human."

That verdict came from weighing six signal categories together — formulaic structure, stylometric markers, editorial AI-writing tells, GPT-2 perplexity analysis, GPT-2 token-rank distribution, and RoBERTa neural classification — not from any one of them alone. The report also broke the text into 62 individual prose sentences and classified each one separately: 58 read as human-like, 3 as mixed, and exactly 1 as AI-like. That sentence-level breakdown is what makes a document verdict more than a single number — and it's where the interesting part of this experiment actually lives.

The one sentence that got flagged — and why

Only one sentence out of 62 crossed into "likely-ai" territory, and it's from the Second Inaugural, not the more famous Gettysburg text:

“While the inaugural address was being delivered from this place, devoted altogether to saving the Union without war, urgent agents were in the city seeking to destroy it without war--seeking to dissolve the Union and divide effects by negotiation.”

ScribeLens scored that single sentence 47% AI-like / 53% human-like — the highest AI score anywhere in either speech — and flagged it for three reasons: “Balanced caveat frame,” matching the phrase “from this place, devoted altogether to”; “False range,” matching the longer clause “from this place, devoted altogether to saving the Union without war, urgent agents”; and “Formulaic structure.”

Read the sentence again with those labels in mind and the pattern is genuinely there: it's a long, hedged, syntactically balanced construction — a clause of context, a qualifier, then a parallel “seeking to … seeking to” repetition — built the way careful antithesis is built, which happens to sit structurally close to how AI-generated text often pads a hedge. That's exactly the kind of evidence a bare percentage never shows: not “22% AI,” but this specific sentence, these specific reasons, out of 62 sentences total.

The rule-of-three that tripped a rule-of-three signal — and the one that didn't

Three sentences landed in “mixed” territory: 31%, 26%, and 25% AI-like. The middle one is the most interesting, also from the Second Inaugural:

“To strengthen, perpetuate, and extend this interest was the object for which the insurgents would rend the Union even by war, while the Government claimed no right to do more than to restrict the territorial enlargement of it.”

It was flagged for “Rule-of-three phrasing” — the report matched the exact clause “strengthen, perpetuate, and extend” — plus “Mechanical rhythm” and “Known AI-writing pattern.” That's a real Lincoln triad, tripping a real detection signal that also fires on the three-part lists modern AI writing leans on constantly. Great oratory and machine-generated prose share more structural DNA than either side of that comparison usually likes to admit, and sentence-level evidence is what lets you actually see that happening instead of just distrusting the number. (The other two mixed sentences, both long and symmetrically built, were flagged for the same family of reasons: “Formulaic structure” and “Mechanical rhythm” on one, “Mechanical rhythm” and “Generic advisory phrasing” on the other.)

Here's the twist, though: American rhetoric's single most famous rule-of-three — “government of the people, by the people, for the people” — did not get flagged at all. Project Gutenberg's edition splits that clause into three separate short “sentences” at the ellipsis marks in the source text, and each fragment scored 94% human-like / 6% AI-like individually, with zero reason codes attached. The most quoted triad in American history slipped past the exact pattern it should have tripped, because of how the source file happened to be punctuated — not because the detector treated it specially. We're disclosing that rather than smoothing it over: it's exactly the kind of formatting quirk an evidence-first report should surface, not hide.

What this means for your formal writing

None of this means detection is unreliable, and that's not the point of publishing it. The point is what it shows: formal register and parallel structure — the deliberate repetition, the balanced clauses, the measured antithesis that make a speech or an essay sound authoritative — carry some of the same surface statistics that AI-generated writing produces by default. That overlap is real and structural, and a detector worth trusting shouldn't pretend otherwise.

It's also exactly why ScribeLens reports sentence-level evidence and a graded verdict instead of stopping at a single score. A document doesn't get convicted because one sentence out of 62 leans on a rhetorical triad — 58 of Lincoln's 62 sentences came back cleanly human-like, and the one flagged sentence didn't move the document past “likely human, high confidence.” If your own essay, cover letter, or research introduction leans on parallel structure and formal phrasing, some sentences may carry real signal for the same structural reasons Lincoln's did. What you want from a detector at that point isn't a single number to argue with — it's the sentence, the reason keyword, and the room to look at both and judge for yourself.

Bottom line

If you take one thing from this:

Frequently asked questions

Did the detector think Lincoln used AI?

No. The overall verdict was likely human, high confidence — 78% human-like against 22% AI-like across the combined 983-word text, with the report's own confidence statement reading “We are highly confident this text was written by a human.” One sentence out of 62 individually scored higher on AI-like patterns, but it wasn't enough to move the document verdict.

Why did any sentence get flagged at all, then?

Because pattern-based detection reads structure, not authorship. The flagged sentence is long, hedged, and built from parallel clauses (“seeking to … seeking to”) — a genuinely formal, balanced construction that shares surface statistics with how AI-generated text often pads a qualified statement. That's real evidence worth seeing at the sentence level, not a false alarm to bury inside a single score.

What was the “rule-of-three” flag actually about?

It matched a genuine Lincoln triad — “strengthen, perpetuate, and extend” — in the Second Inaugural. Rule-of-three phrasing is a real detection signal because three-part lists show up constantly in AI-generated writing, and Lincoln's own oratory happens to use the same structure. His most famous triad, “of the people, by the people, for the people,” wasn't flagged at all — it was split into fragments by the source file's punctuation, not treated differently by the detector.

Would my formal essay read the same way?

Possibly, in part. Parallel structure, measured transitions, and consistently even sentence rhythm are stylistic choices a careful writer makes on purpose, and they're also patterns detection is built to notice. That's not a reason to avoid writing formally — it's why a fair detector shows you which sentence carries the signal and why, instead of leaving you with a single number to either trust blindly or dismiss.

Can I run this experiment myself?

Yes. Both speeches are public domain and available from Project Gutenberg; paste the same combined text into ScribeLens and you'll get the same kind of sentence-level breakdown described here. It's free and doesn't require an account for a quick scan.

See the sentence-level evidence for yourself

Paste any text — a speech, an essay, a draft — and see the same kind of sentence-by-sentence breakdown and reason keywords this investigation is built on. Free, no signup required for a quick check.

← Back to all articles