Twelve Labs Launches Marengo 2.7, Introducing New Multi-Vector Approach to Video Understanding
Summary
With Marengo 2.7, Twelve Labs deploys multi-vector representation for the first time to address the complexities inherent in video. Each vector independently captures distinct aspects of the video content - from visual appearance and motion dynamics to OCR text and speech patterns. For example, one vector might capture what things look like (e.g., "a man in a black shirt"), another tracks movement (e.g., "waving his hand"), and another remembers what was said (e.g., "video foundation model is fun"). Marengo 2.7 demonstrates particular strength in detecting small objects while maintaining exceptional performance in general text-based search tasks. Now, with Marengo 2.7, users can search complex visual scenes, find specific brand appearances, locate exact audio moments, match images to video segments, and more.