Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

nvidia
/
llama-nemotron-rerank-vl-1b-v2-fp8

Text Ranking
Transformers
Safetensors
sentence-transformers
multilingual
llama_nemotron_vl_rerank
feature-extraction
reranker
cross-encoder
visual-document-retrieval
question-answering retrieval
multimodal reranking
semantic-search
rag
custom_code
modelopt
Model card Files Files and versions
xet
Community
1

Instructions to use nvidia/llama-nemotron-rerank-vl-1b-v2-fp8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • Transformers

    How to use nvidia/llama-nemotron-rerank-vl-1b-v2-fp8 with Transformers:

    # Load model directly
    from transformers import AutoModel
    model = AutoModel.from_pretrained("nvidia/llama-nemotron-rerank-vl-1b-v2-fp8", trust_remote_code=True, device_map="auto")
  • sentence-transformers

    How to use nvidia/llama-nemotron-rerank-vl-1b-v2-fp8 with sentence-transformers:

    from sentence_transformers import CrossEncoder
    
    model = CrossEncoder("nvidia/llama-nemotron-rerank-vl-1b-v2-fp8", trust_remote_code=True)
    
    query = "Which planet is known as the Red Planet?"
    passages = [
    	"Venus is often called Earth's twin because of its similar size and proximity.",
    	"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
    	"Jupiter, the largest planet in our solar system, has a prominent red spot.",
    	"Saturn, famous for its rings, is sometimes mistaken for the Red Planet."
    ]
    
    scores = model.predict([(query, passage) for passage in passages])
    print(scores)
  • Notebooks
  • Google Colab
  • Kaggle
llama-nemotron-rerank-vl-1b-v2-fp8
2.4 GB
Ctrl+K
Ctrl+K
  • 2 contributors
History: 2 commits
dovchinnikov's picture
dovchinnikov
Add quantized checkpoint weights (#1)
a306061 3 days ago
  • .gitattributes
    1.57 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • LICENSE
    21.8 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • README.md
    22 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • THIRD_PARTY_NOTICES.md
    16.5 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • chat_template.jinja
    3.83 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • config.json
    4.63 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • configuration_llama_nemotron_vl.py
    5.89 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • hf_quant_config.json
    759 Bytes
    Add quantized checkpoint weights (#1) 3 days ago
  • model.safetensors
    2.38 GB
    xet
    Add quantized checkpoint weights (#1) 3 days ago
  • modeling_llama_nemotron_vl.py
    27.8 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • processing_llama_nemotron_vl.py
    21.5 kB
    Add quantized checkpoint weights (#1) 3 days ago
  • processor_config.json
    419 Bytes
    Add quantized checkpoint weights (#1) 3 days ago
  • special_tokens_map.json
    454 Bytes
    Add quantized checkpoint weights (#1) 3 days ago
  • tokenizer.json
    17.2 MB
    xet
    Add quantized checkpoint weights (#1) 3 days ago
  • tokenizer_config.json
    52.9 kB
    Add quantized checkpoint weights (#1) 3 days ago