Why are short videos bad for learning?

An essay version of the video is below – not just a transcript but a text that I’ve rewritten thoroughly to work as text. As is an audio version. You can subscribe to the audio-only unlisted podcast feed here if audio is more your speed.

Overview

In this opinion piece – very much my opinion based on my experience and practice, not on academic studies – I explain one reason why short videos, such as what you get on TikTok, are bad for learning and information.

Did I mention that this is my opinion? Well it is. Don’t go all “well, that’s just your opinion, man”. I already told you so. And the same goes for the academics who suddenly find themselves filled with the urge to debate me on how montage or Tarkovsky have been understood in academic writing over the years.

In. My. Opinion.

Video

Audio


The TikTok Problem

TikTok and its copycats have come to define how much of our societies consume information. The information that reaches us is bounded by the nature of the media that convey it.

By emphasising dubious theories that misrepresent neuroplasticity and “brain changes” we’ve managed to downplay that these platforms are doing exactly what they’re designed to do. Their form and design direct how we experience “facts” and how we learn.

I like to blame the Russians.

Not the current lot. Older Russians.

Though, if I’m to be fair, they more described the problem than caused it directly.

Specifically, I tend to blame Eisenstein and Tarkovsky.

Eisenstein was one of the early pioneers of cinema and was the first major filmmaker to come out of Russia in the immediate aftermath of the Soviet revolution.

His contribution was the concept of the “montage”. As legend would have it – probably untrue and apocryphal, but it’s a good story – early Soviet filmmakers struggled with shortages. That meant that they had a harder time of constructing films using the method that had been the norm in early cinema: a series of long takes spliced together, as you would if you were recording a play. They had to work with scraps of film.

So, as legend has it, they resorted to splicing together short bits of film and discovered that spliced end result worked very differently.

As he described the concept in his book (I think, it’s been twenty years since I read it, so I’m probably both misremembering and paraphrasing it badly):

If you record a long take of a man walking to a dinner table, looking down at the food, and licking his lips, you are very directly telling the viewer that the man is hungry and is about to eat.

But if you edit together a shot of the man entering a room, a close-up of food on the table, and a close-up of his face as he licks his lips, you are giving the viewer the impression of the same without actually showing it. You never actually visually show the man responding to the food one the screen at the same time. And if you change the shot of the food to the shot of another person, you transform the impression into a sexual one. Same core elements and shots, but because you aren’t directly telling the story visually – you’re implying it – you can transform the impression just by changing one shot.

The meaning, the scene, and the whole is constructed in the mind of the viewer without having to go through the effort of doing the thing itself on screen.

This is labour-saving, but it also means you can tell the story without having to show the story in detail.

The counterpoint to this approach came from another Russian – Tarkovsky – who made Solaris and Stalker and quite a few other movies. He believed that the montage was a bit of a deception. (Again, paraphrasing as it’s been almost two decades since I read his book as well.)

He saw the strength of film as being its time-based nature. It could show you time unfold, actions taking place over time, all on a large screen with detail and resolution that gave you the feeling of being there – gave you the sensation of a reality that impressionistic montage couldn’t and wouldn’t give you.

This is the point where, as a teacher, you’ve lost your students. As soon as you say “Russians” in any sort of cultural context, their eyes glaze over.

So, I always preferred to use action movies as examples to explain these two ideas. Back in the day I would contrast the Jason Bourne movies with Jackie Chan’s, but if I were teaching today I would probably instead use Marvel movies versus John Wick.

The Bourne movies are a stronger example, though, of the use of montage to deliver the impression of a fight without actually showing you the fight. They show punches, throws, and grimaces but at the same time they make it hard for you to follow the action in detail. You get the highlights – “he went out the window!” – but the exact scene itself simply isn’t shown. You get an impression of it through short edits. This is logistically quite convenient because it means you don’t have to actually shoot the scene you want to convey. You can conceal the fact that the actors don’t know how to fight, that fewer things are broken than you’d think, that the entire thing was much slower than portrayed, and the space was too constrained for an actual fight to happen.

You are given the impression of an experience that never took place.

Whereas in the Jackie Chan movies – or the John Wick movies if you aren’t familiar with early Jackie Chan – you get a broader view that often shows the full bodies of all the participants. You see the scene unfold as it happened in front of the camera. Every motion and action is shown in full detail.

It’s visual exposition, and it gives you a much stronger sense of being there, of experiencing the action that you’re watching, and it also gives you enough of an experience of the fight scene to be able to describe it to somebody else because you saw it unfold.

Try to describe a fight scene from one of the Bourne movies and, even if you’ve just watched it, you’ll only be able to convey the highlights of the scene because that’s all you were given.

Try to describe a fight scene from a Jackie Chan movie or John Wick, even a decade later, and you can often outline much of what happened because you witnessed it, and it felt real. It’s a richer and more vivid experience, especially when experienced on a large screen.

These two approaches are important because TikTok and YouTube Shorts – and many longer videos – are impressionistic and not expository.

Because they make you “lean forward” as you’re using the apps the expectation of a shorter duration is baked into the design of the software. You interact with streams of video – you don’t browse or forage – and are constantly ready to swipe to the next.

These services may let you upload longer videos, but the design itself enforces brevity.

This means the shorts themselves tend to resort to the economy of montage to deliver their message. Each short – brief already – is often composed of a sequence of even shorter clips. The stream itself sutures the many shorts together to create an algorithmic montage that’s specifically tailored to your tastes, constantly creating unfounded and even outright false impressions of what the stream is pretending to document.

They give you the feeling of having listened to somebody explain something in detail, but you haven’t actually – you just got convinced by an impression. You didn’t hear an argument. You haven’t actually heard through the reasoning of what you’re being told.

It becomes a mechanism of convincing people of political ideas, of news, of sequences of events out in the world that never actually happened, or are severely misrepresented through editing to create an impression.

This is the most effective medium to deliver a condensed message. Montage lets you create a dense message that is concise enough to fit within the duration parameters enforced by TikTok’s design.

These de facto duration limits our ability to use the Tarkovsky or the Jackie Chan approach of unfolding an argument or story over time.

You can’t deliver complex visual action or complex visual arguments as easily on TikTok using edited media as you can using a larger screen with linear time-based media.

One way to counteract this is through simplicity: the talking head video. A person, looking into the camera, outlining their argument verbally, doesn’t require much in terms of visual information and so isn’t constrained by the tiny screen size. Audio isn’t compromised by screen real-estate. It’s effectively just radio with a face on it, and it’s limited to the specific kind of information you could deliver on radio, but it suffers less from the limitations of modern short-form video.

They may not know it themselves, but this is probably the reason why many educators on TikTok tend to use this approach, or a similar approach with minimal editing, to deliver their work.

My own videos would be more convincing if they were two-minute videos full of short clips from various movies giving you the impression of my argument. (Might not pass content-id, but that’s a separate issue.)

This approach would be more convincing over a short duration, but I would not have the same time to explain the argument, and you would have a harder time integrating it and later on explaining it somebody else.

Because that’s the issue with the impressions created through montage – like the Jason Bourne fight scenes. If you’re sitting with a friend in a pub and trying to tell them about some of the scenes in a movie you just saw, you’d have a hard time giving them the details of a specific fight beyond a single highlight from the scene and how it concluded. You couldn’t describe the actual sequence of events because you never witnessed the sequence of events.

You were only given an impression of it.

Whereas the Jackie Chan or John Wick movies exhibited the event visually. That makes it easier for you to describe in your own words.

This tends to be missing from TikTok and the many YouTube videos that often to emphasize editing over explaining, over outlining the argument.

There are exceptions like the TikToker that goes by the handle OddprideAstrid Lundberg. She’s excellent at the genre of video you could describe as “sit down in front of a camera and explain a single concise idea directly to the audience”. She uses basically the talking-head approach. She sits down. Explains things in sequence. No cuts. Maybe an image or two is overlaid. Just a single camera; single shot.

It delivers the message.

When you’re trying to teach, or educate, or convince people of an argument the edited approach works, but it’s one of the reasons why misinformation is on the rise.

It’s so easy to deliver a false message that doesn’t have an argument but gives people the impression of an argument that they never actually watched unfold.

They come away with ideas that have no basis in factuality.

That isn’t to say the two approaches don’t have their place.

If you’re keeping with the action movie comparison, John Woo demonstrates how it can work. In his movie Hard Boiled there are action scenes with rapid editing, that build up a visceral impression that keeps the details intentionally unclear, before cutting to a wide and detailed shot that shows you where things ended up after the rapid sequence.

The impressionistic edits are used to strengthen and reinforce the emotions of the wider shot that follows.

In other scenes John Woo switches to extreme long takes – literally the time-based storytelling that Tarkovsky advocated – that show you the entire action sequence with two characters alternating and going through a hospital, fighting their way to their goal. This gives you the feeling of being there – of experiencing it with them. It makes the story feel much more real. He mixed the two approaches together to deliver a really effective action movie.

This is something you can do when you’re creating a longer piece of work, like a documentary, a movie, a series, or a longer YouTube essay.

Especially if you have the luxury of delivering for a larger screen, even if it’s just the larger screen of people’s living room TV.

This mix-and-match approach isn’t something that works on the phone. It doesn’t really scale down nor does it work with the constraints created by TikTok’s app design.

Talking head videos, like the ones I’m trying to make, kind of scale down. Sort of. For the most part. It’s one of the reasons why I’m using this approach to video.

This format is more likely to be able to deliver an argument that people will understand and integrate and then apply their own life.

Because they are more likely to be able to explain what they learned to others.

If I delivered this as a highly edited video, it would be more convincing, but it would rob you as a viewer of the tools you need to convince others.

The second order effect would be gone. The only way you could spread the message would be to share the video. The first order effect of manipulating you – the viewer – would be there, but if we’re trying to truly reach people and reach an audience we need to give you and ourselves the tools of explaining the things we learn ourselves.

That requires more time. Both in the duration and length of the work we make and in the experience that we demand of the reader or viewer.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论