File size: 3,204 Bytes
04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde 9201f50 04d7dde f97d782 04d7dde f97d782 04d7dde f97d782 04d7dde | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 | # Journey into Learning - 4:00 PM
[**Interactive Demo and Multimodal RAG System Architecture**](https://learn.deeplearning.ai/courses/multimodal-rag-chat-with-videos/lesson/2/interactive-demo-and-multimodal-rag-system-architecture)
### A multimodal AI system should be able to understand both text and video content.
---
## Step 1 - Learn Gradio (UI) (30 mins)
Gradio is a powerful Python library for quickly building browser-based UIs. It supports hot reloading for fast development.
### Key Concepts:
- **fn**: The function wrapped by the UI.
- **inputs**: The Gradio components used for input (should match function arguments).
- **outputs**: The Gradio components used for output (should match return values).
📖 [**Gradio Documentation**](https://www.gradio.app/docs/gradio/introduction)
Gradio includes **30+ built-in components**.
💡 **Tip**: For `inputs` and `outputs`, you can pass either:
- The **component name** as a string (e.g., `"textbox"`)
- An **instance of the component class** (e.g., `gr.Textbox()`)
### Sharing Your Demo
```python
demo.launch(share=True) # Share your demo with just one extra parameter.
```
### Why Didn’t Hot Reloading Work?
(Investigate potential caching issues, missing dependencies, or incorrect function signatures.)
---
## Gradio Advanced Features
### **Gradio.Blocks**
Gradio provides `gr.Blocks`, a flexible way to design web apps with **custom layouts and complex interactions**:
- Arrange components freely on the page.
- Handle multiple data flows.
- Use outputs as inputs for other components.
- Dynamically update components based on user interaction.
### **Gradio.ChatInterface**
- Always set `type="messages"` in `gr.ChatInterface`.
- The default (`type="tuples"`) is **deprecated** and will be removed in future versions.
- For more UI flexibility, use `gr.ChatBot`.
- `gr.ChatInterface` supports **Markdown** (not tested yet).
---
## Step 2 - Learn Bridge Tower Embedding Model (Multimodal Learning) (15 mins)
Developed in collaboration with Intel, this model maps image-caption pairs into **512-dimensional vectors**.
### Measuring Similarity
- **Cosine Similarity** → Measures how close images are in vector space (**efficient & commonly used**).
- **Euclidean Distance** → Uses `cv2.NORM_L2` to compute similarity between two images.
### Converting to 2D for Visualization
- **UMAP** reduces 512D embeddings to **2D for display purposes**.
## Preprocessing Videos for Multimodal RAG
### **Case 1: WEBVTT → Extracting Text Segments from Video**
- Converts video + text into structured metadata.
- Splits content into multiple segments.
### **Case 2: Whisper (Small) → Video Only**
- Extracts **audio** → `model.transcribe()`.
- Applies `getSubs()` helper function to retrieve **WEBVTT** subtitles.
- Uses **Case 1** processing.
### **Case 3: LvLM → Video + Silent/Music Extraction**
- Uses **Llava (LvLM model)** for **frame-based captioning**.
- Encodes each frame as a **Base64 image**.
- Extracts context and captions from video frames.
- Uses **Case 1** processing.
|