Analyzes video frames, speech, and sound together with multimodal AI models. Processes visual and audio simultaneously for search.