AIuthority

Google's Liz Reid: LLMs Unlock Advanced Audio and Video Indexing for SEO

By Charles Ryder

I’ve been watching Google’s search evolution for years, and Liz Reid’s latest comments make one thing clear: SEO is no longer a text-only game.

In a recent appearance on the ACCESS podcast, Reid—Google’s VP and Head of Search—said large language models are finally giving Google the ability to understand audio and video in ways that weren’t possible before. She wasn’t only talking about better transcription. Her point was bigger: multimodal LLMs can interpret what a video is about, how it delivers its message, and even the style of the content itself.

That’s a significant shift for search.

Google's Liz Reid discussing LLMs unlocking advanced audio and video indexing for SEO on a podcast, emphasizing multimodal understanding. Why This Matters Now

For years, audio and video SEO leaned heavily on metadata, titles, descriptions, and whatever transcript quality creators could provide. If Google couldn’t parse the spoken content clearly, publishers had to make up for it with text-based signals. That created obvious limits, especially for podcasts, interviews, tutorials, and content with regional or conversational speech.

Google had already run into these issues in earlier systems. Speech-to-text often stumbled over proper nouns, accents, and context-heavy phrasing. So even strong content could be misunderstood because the machine only picked up fragments of what was said.

Reid’s comments suggest those barriers are starting to come down.

When Google describes LLMs as multimodal, it means audio and video are no longer being treated like secondary formats that must be flattened into text first. The system can interpret meaning across formats more directly. For SEO professionals, that changes the optimization equation in a very real way.

From Transcripts to True Understanding

What stood out most in Reid’s remarks was the emphasis on depth. She explained that Google can now understand not just the words in a video, but what the video is actually about and what kind of style it uses.

That distinction matters.

A transcript may tell Google the literal words spoken in a cooking tutorial, product review, or podcast discussion. Deeper AI understanding can go further:

  • identify the central topic
  • detect the format and intent
  • understand whether the style is instructional, conversational, or opinion-driven
  • connect visual and spoken signals
  • infer whether the content is practical, authoritative, or entertainment-led

That opens the door to more accurate indexing for podcasts, webinars, interviews, explainer videos, and short-form clips.

For creators, this is encouraging. Strong multimedia content may no longer need perfect manual transcript cleanup to earn visibility. For marketers, it means substance and delivery are becoming searchable assets in their own right.

Google’s Broader Search Direction

Reid’s comments also fit neatly into Google’s broader search transformation over the past few years.

We’ve already seen Google test and expand AI-powered features like Audio Overviews in Search Labs. We’ve also seen ranking changes that surface more short-form video, forums, and user-generated content as user behavior shifts. On top of that, Google has introduced features like Preferred Sources, designed to surface content from publishers and creators users already trust or pay for.

Taken together, the direction is pretty clear: search is moving toward richer, more personalized, and more format-flexible discovery.

People don’t just want links. They want answers, context, trusted voices, and content in the format that fits the moment. Sometimes that’s a traditional article. Sometimes it’s a short video. Sometimes it’s a podcast clip or an AI-generated spoken summary.

LLMs are what make that possible at scale.

Diagram illustrating the shift from text-only SEO to multimodal AI understanding of audio and video content, showing deeper interpretation. What This Means for SEO Strategy

If I were advising brands right now, I’d say this is the moment to stop treating audio and video as supporting content. They’re becoming core search assets.

That means SEO strategy should include:

1. Multimedia content with real depth

Thin videos and low-value podcasts won’t suddenly rank just because Google understands media better. If anything, stronger understanding should make it easier to separate useful content from filler.

2. Strong topical consistency

If your brand publishes blog posts, videos, interviews, and audio episodes around the same themes, you build a broader authority footprint. Google’s multimodal understanding is likely to reward that consistency.

3. Better structure and context

Titles, chapters, descriptions, speaker labels, timestamps, and surrounding page content still matter. Even with better AI interpretation, clear context helps search engines classify content faster and more accurately.

4. Style as a ranking signal

This may be one of the most interesting implications. Reid’s comments suggest Google may increasingly understand not just subject matter, but presentation style. That could influence how educational, entertaining, technical, or conversational content is surfaced for different queries and audiences.

5. Multilingual opportunity

Reid also pointed to the ability of LLMs to bridge language gaps. That creates real opportunity for creators and publishers targeting multilingual audiences, especially in underserved markets where high-quality localized content has been limited.

The Competitive Impact

I think this will widen the gap between brands creating genuinely useful multimedia content and those still relying on outdated, text-only SEO habits.

Podcasters, educators, YouTubers, niche experts, and founder-led brands may benefit most. If Google gets better at understanding nuance, expertise, and style, authentic voices become more discoverable. That’s a major advantage for creators who communicate authority more naturally in speech or video than in polished article form.

At the same time, it raises the stakes for everyone else. If search becomes better at evaluating multimedia meaning, low-effort AI slop in audio and video should face tougher scrutiny as well. Better indexing doesn’t reward more content. It rewards clearer value.

FAQ

How are LLMs changing audio and video SEO?

They allow Google to understand more than a transcript. Multimodal LLMs can interpret topic, intent, style, and signals across both spoken and visual content.

Do transcripts still matter for SEO?

Yes. Improved AI understanding reduces dependence on transcripts alone, but transcripts, titles, chapters, descriptions, and timestamps still provide valuable context.

What types of content stand to benefit most?

Podcasts, webinars, interviews, tutorials, explainer videos, short-form clips, and multilingual content all stand to benefit as Google improves media understanding.

Does this mean low-quality video content will rank more easily?

No. Better understanding should make it easier for Google to separate useful, original content from shallow or low-value material.

Conclusion

Liz Reid’s comments point to a real turning point. Search is shifting from reading pages to understanding media. For SEO, that means optimization is becoming less about isolated keywords and more about total content comprehension across text, voice, and video.

I see this as both a challenge and an opportunity. Brands that invest in quality multimedia now will be in a much stronger position as Google expands multimodal search. And for teams looking to stay ahead of that shift with smarter AI-powered publishing support, it’s worth exploring AIuthority as part of a modern content strategy.