Journey into Learning - 4:00 PM
Interactive Demo and Multimodal RAG System Architecture
A multimodal AI system should be able to understand both text and video content.
Step 1 - Learn Gradio (UI) (30 mins)
Gradio is a powerful Python library for quickly building browser-based UIs. It supports hot reloading for fast development.
Key Concepts:
- fn: The function wrapped by the UI.
- inputs: The Gradio components used for input (should match function arguments).
- outputs: The Gradio components used for output (should match return values).
π Gradio Documentation
Gradio includes 30+ built-in components.
π‘ Tip: For inputs and outputs, you can pass either:
- The component name as a string (e.g.,
"textbox") - An instance of the component class (e.g.,
gr.Textbox())
Sharing Your Demo
demo.launch(share=True) # Share your demo with just one extra parameter.
Why Didnβt Hot Reloading Work?
(Investigate potential caching issues, missing dependencies, or incorrect function signatures.)
Gradio Advanced Features
Gradio.Blocks
Gradio provides gr.Blocks, a flexible way to design web apps with custom layouts and complex interactions:
- Arrange components freely on the page.
- Handle multiple data flows.
- Use outputs as inputs for other components.
- Dynamically update components based on user interaction.
Gradio.ChatInterface
- Always set
type="messages"ingr.ChatInterface. - The default (
type="tuples") is deprecated and will be removed in future versions. - For more UI flexibility, use
gr.ChatBot. gr.ChatInterfacesupports Markdown (not tested yet).
Step 2 - Learn Bridge Tower Embedding Model (Multimodal Learning) (15 mins)
Developed in collaboration with Intel, this model maps image-caption pairs into 512-dimensional vectors.
Measuring Similarity
- Cosine Similarity β Measures how close images are in vector space (efficient & commonly used).
- Euclidean Distance β Uses
cv2.NORM_L2to compute similarity between two images.
Converting to 2D for Visualization
- UMAP reduces 512D embeddings to 2D for display purposes.
Preprocessing Videos for Multimodal RAG
Case 1: WEBVTT β Extracting Text Segments from Video
- Converts video + text into structured metadata.
- Splits content into multiple segments.
Case 2: Whisper (Small) β Video Only
- Extracts **audio** β `model.transcribe()`.
- Applies `getSubs()` helper function to retrieve **WEBVTT** subtitles.
- Uses **Case 1** processing.
Case 3: LvLM β Video + Silent/Music Extraction
- Uses **Llava (LvLM model)** for **frame-based captioning**.
- Encodes each frame as a **Base64 image**.
- Extracts context and captions from video frames.
- Uses **Case 1** processing.