Images won the feed.
Then the feed grew subtitles.
A theory worth testing says that attention pays for pictures, so writing retreats into a caption while meaning moves into images, video and memes. The structure of the argument is coherent and the economics behind it are well documented. The argument is also sixty-four years old, it has failed every time it has been made, and the largest measurable change in how video is consumed runs against it.
Four observations in this report
- 70%Of Gen Z watch video with subtitles on most of the time (Preply, 2022)
- 1962McLuhan first published the argument that electronic media would end print culture
- 0 of 4Populations where written post length declined, measured across 2012 to 2024
- 0Function words emoji produced in fifteen years: no verbs, no prepositions, no determiners
The measured question
More pictures.
Fewer words?
Four archives, 2012 to 2024. Fitted change in median post length, per year.
No estimate clears |t| > 2.31, the study’s threshold for a trend. This does not measure people who left writing altogether.
Read the measurement and its limitsFigure 01 — The claim, on repeat
What was predicted, and what showed up
Each column holds one announcement that pictures were about to displace writing. Under the rail is the medium that arrived instead. The fourth column is drawn open because its arrival is still being measured.
Predicted
Electronic media will dissolve print culture and return the West to an oral, acoustic, tribal condition.
Predicted
Television is producing a public that no longer reads, and public argument will follow it into images.
Predicted
A templated picture with two lines of Impact type does the work of a paragraph, and will replace it.
Predicted
Attention pays for images and video, so language reduces to a caption and meaning moves into pictures.
Arrived
Email in 1971, Usenet in 1980, the web in 1991. Every one of them a text protocol.
Arrived
SMS arrived in 1992, then chat, forums and comment threads. Ordinary life was re-textualised.
Arrived
The format died. Advice animals are catalogued as a late-2000s fad, and platforms now throttle static jokes.
Arriving
Video acquired a compulsory text layer. Most of Gen Z reads its video rather than only watching it.
Figure 02 — The case for
Four reasons to take the theory seriously
One: Egypt solved this and ran the solution for three thousand years.
Egyptian scribes wrote unpronounced signs at the ends of words to mark what class of thing the word named. Orly Goldwasser calls them semantic classifiers and maps each one as the head of a conceptual category. They were silent, visible, and combined with phonetic spelling rather than competing with it. Anyone who appends a picture to a sentence today is doing that job with different pictures.
Two: a picture has already signed a contract.
In March 2021 an employee of South West Terminal texted a photo of a flax contract to Chris Achter with the words “Please confirm flax contract.” Achter answered with a thumbs-up emoji and then delivered no flax. The Court of King’s Bench treated the emoji as acceptance and awarded about $82,200 in damages. Saskatchewan’s Court of Appeal upheld the decision two to one, finding the emoji a valid electronic signature on the facts.
Three: arrangement does the grammatical work that sequence cannot.
Neil Cohn and colleagues tested whether emoji strings behave like sentences. No word order beat chance, verbs came in below chance, and participants produced no prepositions or determiners at all. Their explanation is more useful than their result. Linear sequencing suits speech, and graphic communication was never linear. Templates encode their operations spatially, and once you look at layout the operations are all present.
Four: published linguistics treats memes as constructions.
Barbara Dancygier and Lieven Vandelanotte analysed image macros as multimodal constructions in Cognitive Linguistics in 2017. Their account gives the image a constructional slot, filling a role the text leaves open. One of their examples reads “Uses shopping cart / Leaves it in empty parking space.” Nothing in that text supplies a subject. The picture supplies it.
Grammatical operations, realised as layout
Four templates reduced to geometry. Every box is empty because the content varies from post to post and the arrangement does not. The arrangement performs the operation named above each card.
Polarity and preference
Two stacked panels
Vertical position marks rejection above and endorsement below. A negation with no phonetic form.
Ordinal degree
Four ascending tiers
Height maps to magnitude across an ordered scale, so the axis itself does the quantifying.
Three-argument predicate
Three figures in one frame
Spatial relation assigns the roles: one committed, one attractive, one disapproving.
Exclusive disjunction
Two adjacent options
Adjacency under a single actor forces the choice and marks it as costly.
Figure 03 — The case against
Six ways the theory comes apart
The argument is sixty-four years old.
Marshall McLuhan published it in 1962. The Gutenberg Galaxy predicted that electronic media would dissolve print culture and restore something closer to oral tribal life. Walter Ong’s secondary orality descends from the same reasoning. What arrived after McLuhan was email in 1971, Usenet in 1980 and the web in 1991, three text protocols in a row.
The format with the most grammar lost.
The construction-grammar evidence comes from image macros, and image macros are gone. Advice animals are catalogued as a fad of the late 2000s, and the format is described as bled dry of its novelty. Platforms now favour original short video and longer watch time, which throttles static jokes. Compositional structure did not turn out to be the adaptive trait.
Video did not shed its language. It grew a text layer that most young viewers now depend on.
Subtitles turned video into a reading medium.
Preply found 70% of Gen Z watching with subtitles most of the time, against 53% of millennials, and the BBC put viewers aged 18 to 35 near 80%. Hearing loss explains little of it. Respondents cited muddled audio at 72% and hard accents at 61%, and about 74% watch in public where the sound is off. Viewers also report using captions to move through content faster than it plays. Short-form video now ships with burned-in type as standard.
Memes are parasitic on syntax rather than a substitute for it.
Construction grammar is a theory of language, so filing memes as constructions places them inside language. The shopping-cart example works because English has a subject position available to leave empty. A picture can fill a gap that grammar opens. Filling that gap is a different achievement from replacing the grammar that opened it.
Memes and writing solve opposite problems.
Research on meme competence describes it as the boundary between members in the know and outsiders, where a failed reading can exclude. Density comes from knowledge both parties already share. Writing was invented to reach strangers across distances nobody could verify. A system engineered to be opaque to outsiders cannot take over from a system engineered to reach them.
Speed disqualifies the orality frame.
Ong lists oral cultures as conservative, and they must be, because memory is their only storage. Homeric formulae are stable for that reason. Meme popularity windows have run the other way, from roughly two years in 2008 to about four months in 2023 on indicative figures. Volatility on that scale describes fashion rather than tradition. Writing exists to defeat time, and cuneiform stays legible after four thousand years.
The frame itself is contested.
Secondary orality inherits the great-divide problem, the sharp split between oral and literate cultures that anthropologists have attacked as ethnocentric and unsupported. Ong draws standing charges of technological determinism. Most of what gets called digital orality is written, in chat windows and comment threads. Calling text oral because it feels informal names a mood rather than a mechanism.
Figure 04 — Measured for the report
Four populations, thirteen years
Published work supplies every other indicator. Nobody appears to have tested the length prediction directly, so it was tested: 58,453 records across 169 cells, four populations, the same week of June every year from 2012 to 2024.
Ordering the populations by how much room a picture has turns a yes-or-no question into a directional one. If pictures crowd out prose, the fall should be steepest on 4chan, where a picture starts every thread, and absent on Hacker News, which hosts no images at all.
Hacker News
No images exist on the platform
median 40 words
Stack Exchange
Images possible but rare
median 97 words/answer
Mixed link and discussion
median 17 words
4chan
An imageboard: a picture starts every thread
median 13 words
Left to right: increasing room for a picture
Each panel spans ±40% of its own median. Band is the fit’s residual scatter. Dashed line is the fitted trend. A slope needs |t| above 2.31 to count.
Reading a result
Two points can tell the wrong story.
4chan’s first and last years suggest a fall. The years between them change how that fall should be read.
Same median values in both views · axis starts at zero
Endpoint difference
−15.6%
16.7 words in 2012, 14.1 in 2024. Subtraction makes this look like a simple decline.
Two observations leave the intervening variation out.Reported fit · all years
−0.48% / year
t = -0.54. Below the study’s |t| > 2.31 threshold. The reported result is no detectable trend.
The first year is unusually high; later values fluctuate. A fitted trend asks a different question from subtracting endpoints.
Read the 13 annual medians
201216.7 words
201311.5 words
201412.8 words
201513.8 words
201613.5 words
201713.5 words
201810.5 words
201913.9 words
202011.5 words
202112.7 words
202213.7 words
202313.0 words
202414.1 words
No population declined.
A slope needs a t above 2.31 across thirteen fitted years to count as a trend. None of the four approaches it. Reddit fits at −0.17% a year with a t of −0.34, Hacker News at 0.40% and 0.71, Stack Exchange at 1.75% and 2.00. The imageboard fits at −0.48% a year with a t of −0.54.
The directional prediction fails harder than the flat one.
4chan cannot open a thread without a picture. Hacker News cannot host one. Their slopes are −0.48% and +0.40% a year, each indistinguishable from zero and from the other. The platform where images dominate by construction behaves like the platform where images are impossible.
4chan carries a second series nobody else can supply.
Its posts arrived with an image 24.5% of the time in 2012 and 26.9% in 2024, fitting at 1.41% a year with a t of 1.48. On an actual imageboard, across thirteen years, posts are not increasingly bringing a picture with them.
Three things the measurement cannot carry
The Stack Exchange verdict moves with the window. Starting in 2011 rather than 2012 returns a t of 2.61, which clears the threshold and reads as rising. The conservative reading is used and the other one is recorded rather than dropped.
On 4chan the endpoints alone suggest a 16% decline, from 16.7 words to 14.1. The fit says flat, because 2012 is a high outlier and every year since sits between 10.5 and 14.1 against 1.6 words of scatter. Subtracting endpoints would have produced a finding on the one source where a finding would have been most welcome.
All four populations measure people who are still writing somewhere. Anyone who abandoned writing altogether for video appears in none of them. That remains the strongest surviving form of the theory, and no public archive can reach it.
Figure 05 — Scoreboard
What would have to be true
8 indicators would move if writing were reverting to pictures. Of those, 7 run against the theory and 1 supports it. Every row states what was observed and where the observation comes from, so a reader can weigh the source rather than the verdict. The first two rows are measured; the rest are other people’s work.
| Would have to be true | Observed | Direction |
|---|---|---|
| Written posts getting shorter as language retreats into a caption | Measured, and flat on all four populations sampled. Reddit fits at −0.17% a year, Hacker News at 0.40%, Stack Exchange at 1.75%, 4chan at −0.48%. None clears the threshold a slope needs to count as a trend.Original measurement, 58,453 records, 2012 to 2024 | AGAINST |
| The fall running steepest where a picture has the most room | The opposite of a gradient. 4chan, where a picture starts every thread, and Hacker News, which hosts no images at all, return slopes of −0.48% and +0.40% a year. Neither is distinguishable from zero or the other.Original measurement, four boards and one forum | AGAINST |
| A meme format conveying a claim the audience does not already hold | None documented. Meme density is described in the literature as compression over knowledge both parties already share.Information, Communication & Society, 2022 | AGAINST |
| Meme literacy transferring across communities rather than marking their edges | Runs the other way. Competence marks the boundary between members in the know and outsiders, and failed decoding can exclude.Meme Studies Research Network; JCMC, 2018 | AGAINST |
| Video shedding its text layer as it matures | Runs the other way. Caption use spans about 40% to about 80% depending on the question, and rises among the youngest viewers.Preply, 2022; BBC | AGAINST |
| Meme lifespan lengthening toward the stability a writing system needs | Runs the other way. Popularity windows have shortened by roughly a factor of five since 2008, on indicative figures.Reddit corpus analyses; Humor 2.0, ch. 16 | AGAINST |
| Pictorial units producing function words of their own | Emoji-only messages stayed at one-unit patterns. Verbs came in below chance, and no prepositions or determiners appeared at all.Cohn et al., 2019 | AGAINST |
| A pictorial unit performing a binding act in a domain that punishes ambiguity | Established. A thumbs-up emoji was held to be a valid signature on a grain contract, and the ruling survived appeal.South West Terminal v Achter Land & Cattle | SUPPORTS |
Figure 06 — Reading
What survives the objections
A register, working inside literate culture.
Memes compress at extraordinary rates over shared context, encode their operations in layout, and coin new units without asking permission from a standards body. Emoji cannot say the same. Unicode added eight base characters in 2025, against thirty to sixty a release before 2020, and its own membership has spent a decade arguing over whether emoji distract from the work of encoding historical scripts. What memes cannot do is inform somebody who does not already share the reference. That limit is structural, and it keeps them alongside writing rather than ahead of it.
The economics moved while the theory was being written.
The attention economy paid for capture, and capture favoured pictures. That payment rail is thinning. Small publishers lost about 60% of search referral traffic over two years, and chatbot referrals still account for under 1% of publisher page views. Work on generative engine optimisation finds citation rates rising by up to 40% when content adds statistics, quotations and authoritative language. Whatever replaces the feed rewards structure and propositional text.
The pattern worth keeping
Sixty-four years of this argument have produced one reliable regularity. Each time somebody announces that pictures are taking over from writing, the reply arrives in writing.
Sources
Figure 04 is original measurement, collected 2026-08-24 by scripts/text-corpus-sample.mjs and reduced by scripts/text-corpus-measure.mjs, both in the repository. 58,453 records across 169 cells, sampled from 1 to 8 June each year, UTC so that seasonal swings in traffic and tone cannot be read as trend. Medians rather than means, because post length is severely right-skewed and one long essay can move an average. Sources are never pooled: median length runs from about 13 words on 4chan to about 96 on Stack Exchange, so a combined figure would track sampling yield instead of writing. Deleted, bot and empty records are removed at measurement rather than at collection, and the removal rates are kept, because moderation intensity rose over the period and each year’s survivors are a differently-selected population. 18 cells holding fewer than 40 records were excluded from their fits, all of them Stack Exchange sites in years when those sites barely existed. Hacker News and Stack Exchange bodies arrive as HTML and were stripped of code blocks and blockquotes before counting; 4chan quotelinks were removed and greentext kept.
Compiled 23 August 2026, measurement added 24 August 2026. Source quality is uneven and the text says so where it matters. The meme lifespan figures of 23.6 months in 2008 and about four months in 2023 trace to an informal analysis of Reddit popularity that was then recirculated by marketing blogs; the direction replicates in peer-reviewed work on meme competition dynamics, the two numbers have no primary citation this project could verify, and no curve is drawn between them. Caption rates come from three surveys asking different questions of different age bands, running from about 40% of 18-to-44s at “always or often” to about 80% of 18-to-35s; the lowest of the three anchors the headline figure. Emoji counts are base-character additions from Unicode’s own version charts and exclude skin-tone and sequence variants, which is why they fall well below the totals Emojipedia publishes for the same releases. Widely circulated statistics claiming that 71% of social images are AI-generated, or that six trillion emoji messages are sent monthly, trace only to content farms and are left out. Figure 02 draws template geometry rather than any template, because the argument concerns arrangement and reproducing the images would add nothing to it.