Why do we still treat video captions like a niche favor for the few?
80 %
Watched on total mute
Eighty percent of social media videos are consumed without sound-a number that has redefined digital engagement.
It is a flat, unvarnished number that should, by all rights, have revolutionized the way we produce digital media a decade ago. Instead, we treat the absence of sound as a temporary glitch or a specialized requirement for a tiny demographic. We are still designing for an era of desktop speakers and private offices, ignoring the fact that the modern internet is mostly consumed in the fragile, silent gaps of public life.
The Library of Glowing Rectangles
On the commuter train heading into Zurich, the carriage is packed, yet the only sound is the rhythmic, metallic clatter of wheels on tracks and the occasional rustle of a heavy coat. Every second person has a glowing rectangle in their hand. They are scrolling through a frantic, high-definition world of cooking tutorials, political commentary, and late-night talk show clips. But the carriage remains a library.
A woman near the window pauses on an explainer video about the recent energy price caps. It looks well-produced, with crisp graphics and a narrator who is clearly speaking with urgency. She watches for exactly . There are no captions. There is no text overlay to tell her what the urgency is about.
She doesn’t reach for her bag to find headphones, and she certainly doesn’t turn on the volume to subject forty strangers to a lecture on fiscal policy. She simply swipes upward. The video vanishes, relegated to the graveyard of “content that required too much effort.”
The Curb Cut Effect in a Digital World
This isn’t an accessibility failure in the traditional, legalistic sense of the word. The woman isn’t deaf. She doesn’t have a diagnosed hearing impairment. But in this specific context-on this train, in this social contract of silence-she is effectively excluded from the information.
We have created a world where captions are framed as an “accommodation,” a polite nod toward inclusivity for a minority. In reality, captions have become shared public infrastructure, as essential to the flow of digital information as curb cuts are to a city sidewalk.
Designed for wheelchair users to transition safely to the street.
Benefits strollers, delivery workers, and travelers with suitcases.
I spent this morning clearing my browser cache in a fit of desperate, tech-induced rage because my livestream dashboard kept flickering like a haunted Victorian lamp. It’s the kind of invisible friction that makes you realize how much we rely on things just *working* without our intervention. When the technology fails to meet us where we actually live-whether that’s in a glitchy browser or a silent train car-we don’t blame ourselves. We just move on.
The “Curb Cut Effect” is a well-documented phenomenon in urban planning. When you cut a wedge out of a sidewalk to allow a wheelchair to transition to the street, you aren’t just helping people with disabilities. You’re helping the parent with a stroller, the delivery worker with a heavy dolly, the teenager on a skateboard, and the traveler dragging a suitcase. By designing for the edge case, you improve the experience for the entire population.
The Distributed Benefits vs. Centralized Costs
Captions were pioneered for the deaf and hard-of-hearing community, but they are now the primary way the majority of people consume video in transit, in bed next to a sleeping partner, or in an open-plan office. Yet, we still treat them as an optional extra. We treat them as a “cost” rather than a “utility.”
The reason for this disconnect is a classic problem of distributed benefits and centralized costs. The benefit of a captioned video is spread across every single viewer who doesn’t want to turn their sound on. The cost of those captions-the time spent transcribing, the money spent on a service, the tedious task of timing SRT files-falls entirely on the individual creator or the small marketing team.
To understand why this is such a persistent hurdle, you have to look at how this process actually works under the hood. Audio transcription isn’t a linear “listening” process for a computer; it’s a complex mathematical translation.
The software takes a continuous stream of audio and performs what’s called a Fast Fourier Transform (FFT). This breaks the sound waves down from a messy, wobbling line into a “spectrogram”-a visual map of frequencies over time. The AI then looks at this map and identifies “features” that correlate to specific phonemes, the smallest units of speech. It isn’t just listening for words; it’s calculating the probability that a specific spike in the 2kHz range followed by a soft hiss in the 5kHz range represents the letter “s.”
The End of the Cloud Tax
Once those phonemes are identified, a language model kicks in to guess the most likely word based on context. It’s a computationally expensive dance of physics and linguistics. For a long time, if you wanted high accuracy, you had to ship that audio file off to a massive server farm, wait for the “brain” in the cloud to crunch the numbers, and pay a per-minute fee for the privilege. That friction-the cost, the privacy concerns of sending your data away, and the sheer time involved-is why so many videos remain silent and unreadable.
We are currently seeing a shift where that computational “tax” is being eliminated. Tools like SpeechPulse are moving that entire mathematical dance onto the user’s local machine. By using optimized models that run on your own processor, the “cost” of creating that public infrastructure drops to near zero.
The Advantage Shift
By processing transcription locally, creators gain speed and privacy while losing the per-minute fees. The creator of that energy price cap video would have gained a viewer if transcription had taken just of local math.
When you lower the barrier to entry for accessibility, you stop seeing it as a moral obligation and start seeing it as a competitive advantage. If that energy price cap video had been processed through a local transcription tool in thirty seconds, the woman on the Zurich train would have watched it. The creator would have gained a viewer, the viewer would have gained information, and the “silent carriage” contract would have remained unbroken.
The Tragedy of the Silent Commons
The current state of the internet is a tragedy of the commons. We are drowning in “content,” yet so much of it is locked behind an audio wall that we refuse to climb. We have more information at our fingertips than any generation in human history, but we are losing it to the simple fact that we are often in places where we cannot-or will not-make noise.
I’ve moderated livestreams for years, and I’ve seen the same pattern over and over. A creator spends forty hours editing a masterpiece. They obsess over the color grade. They buy a $900 microphone to ensure the “warmth” of their voice is captured. Then they upload it without a single line of text on the screen.
Within minutes, the analytics show a massive “drop-off” in the first . They assume the “hook” wasn’t good enough. They think the audience is fickle. They rarely realize that half their audience was just standing in a grocery store line without their AirPods.
We need to stop asking “Who needs captions?” and start asking “Who are we excluding by not having them?” Treating accessibility as a niche requirement is a failure of imagination. It assumes that “users” are static entities who always exist in a vacuum. But users are humans, and humans move through environments.
“The Zurich train is a library where the books refuse to be read unless you bring your own silence to the page.”
– Observations on the 8:12 to Zurich
Reclaiming Attention in Public Spaces
We are moving toward a post-audio web, not because we are losing our hearing, but because we are reclaiming our attention in public spaces. The rise of the “silent scroll” isn’t a fad; it’s a fundamental shift in how humans interact with machines. We want the information, but we don’t want the noise. We want the connection, but we don’t want the social awkwardness of a blaring phone in a crowded lift.
Providing captions is no longer about “helping people.” It’s about ensuring your voice can be heard in a world that has collectively decided to turn the volume down. When we make it easy to generate these digital curb cuts-when we remove the fees, the cloud-based hurdles, and the technical complexity-we aren’t just being inclusive.
We are making sure the to Zurich isn’t just a carriage full of people swiping past things they might have loved, if only they could have read them. The technology to close this gap exists. The models are accurate, the hardware is fast enough, and the need is universal.
The only thing left is to stop seeing the “CC” button as an afterthought and start seeing it as the primary interface of the modern age. If 80% of the room isn’t listening, maybe it’s time we gave them something to read.