MikhbarMIKHBAR
Artificial Intelligence

Condé Nast Builds Multimodal Video Discovery with AWS

By moving away from manual, title-based keyword scrubbing, Condé Nast transformed its video search workflows and drastically cut content discovery times across its major media brands.

Condé Nast Builds Multimodal Video Discovery with AWS

The Operational Challenge of Legacy Video Search

For prominent publishing brands like Vogue, GQ, Vanity Fair, and Wired, editorial teams historically lacked a fast, scalable method for conducting multimodal video discovery. Staff members spent an average of 250 minutes per content discovery task by manually scrubbing through a vast asset library containing more than 140,000 videos. Because teams relied exclusively on text titles and brief descriptions, finding relevant video clips became an inefficient operational bottleneck.

This reliance on human-authored metadata and institutional knowledge created single points of failure whenever specific staff members were unavailable. Furthermore, large volumes of valuable underutilized content remained locked away in the archive, entirely undiscoverable because standard keywords failed to connect them to the specific search intents of editors.

Partnering for Generative AI Solutions

To overcome these persistent structural limitations, Condé Nast partnered directly with the AWS Generative AI Innovation Center (GenAIIC) to architect and build a modern, AI-powered video discovery platform.

The collaborative effort focused on designing a scalable platform capable of running intent-based semantic searches across video transcripts, visual elements, and audio tracks simultaneously. By deploying the new technical framework, Condé Nast successfully reduced asset discovery time from 250 minutes down to under 2 minutes per task.

Core Architectural Decisions and Model Selection

When scoping the project, engineering and editorial teams noted that basic keyword searches were fundamentally insufficient for creative workflows. Editorial staff rarely look for strict alphanumeric strings, preferring instead to search for conceptual ideas like beginner-friendly tutorials or behind-the-scenes footage with specific background atmospheres. This requirement pointed directly toward semantic vector embeddings.

To handle the massive scale of more than 140,000 video assets, the technical architecture explicitly decoupled the compute-heavy ingestion pipeline from the low-latency query tier. The development team selected the TwelveLabs Marengo embedding model to jointly encode visual, audio, and transcript signals. This specialized model was integrated via Amazon Bedrock, allowing the organization to utilize advanced foundation models through a unified API without managing dedicated model-serving infrastructure.

Two-plane architecture: an asynchronous ingestion and indexing pipeline and a synchronous query and serving tier
Image related to the report from AWS Machine Learning Blog · Source: AWS Machine Learning Blog

Infrastructure Security and Managed Vector Search

Operating inside Amazon Bedrock granted the project team access to robust enterprise security features. These controls include AWS Identity and Access Management (IAM) for secure access protocols, Amazon Virtual Private Cloud (Amazon VPC) for strict network isolation, and AWS CloudTrail for comprehensive auditability.

For maintaining the vector index, the platform relies on Amazon OpenSearch Service. This managed service provides multi-AZ replication, metadata filtering, and k-nearest neighbor similarity search capabilities, allowing engineers to query large media catalogs efficiently.

Expanded Editorial Capabilities and Event-Driven Pipelines

The deployed platform fundamentally transforms how staff interact with media archives by providing intent-based search, multimodal understanding, image-based queries, typo tolerance, and precise timestamp markers that point directly to exact moments inside video files.

Under the hood, the ingestion framework operates as an asynchronous, event-driven pipeline. New media uploads land in storage buckets, trigger orchestration events, get preprocessed, and are split into chunks running on Amazon Elastic Container Service with AWS Fargate. These chunks undergo parallel embedding generation before getting indexed for vector searches, ensuring high reliability and scalability across the entire archive.

Sources

Continue chronologically

Related entity coverage