Instructions to use jinaai/jina-embeddings-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jinaai/jina-embeddings-v4 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("jinaai/jina-embeddings-v4", trust_remote_code=True, device_map="auto") - ColPali
How to use jinaai/jina-embeddings-v4 with ColPali:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- sentence-transformers
How to use jinaai/jina-embeddings-v4 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("jinaai/jina-embeddings-v4", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
How to perform retrieval using fused [image, text] as the input query?
Hi All!
Could you please advise: what would be the best option for user query which is a combination of [text, image]? Generally, how can I generate this "fused" embedding for [text, image] which works best with jina-v4?
For example, the user wanted to retrieve document using this query ["can you identify the mechanical tool type in this image and how should I operate this tool?" + img_of_tool]. In this case, both the image and text are important, I want jina-v4 to return documents discussing both the tool and the operation procedure.
Thank you!
Hi @ququwowo ,
We haven’t trained or tested the model on the "fused" embeddings, this is why the model class does not support them. However, if you still want to try it, you can modify the prompt here when encoding an image and pass the desired text, for example:
<|im_start|>user\n Can you identify the mechanical tool type in this image and how should I operate this tool?<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n