--- tags: - sentence-transformers - cross-encoder - reranker - generated_from_trainer - dataset_size:5749 - loss:BinaryCrossEntropyLoss pipeline_tag: text-ranking library_name: sentence-transformers metrics: - pearson - spearman model-index: - name: CrossEncoder results: - task: type: cross-encoder-correlation name: Cross Encoder Correlation dataset: name: sts validation type: sts-validation metrics: - type: pearson value: 0.9194308725342207 name: Pearson - type: spearman value: 0.9162169914711082 name: Spearman --- # CrossEncoder This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model trained using the [sentence-transformers](https://www.SBERT.net) library. It computes scores for pairs of texts, which can be used for text reranking and semantic search. ## Model Details ### Model Description - **Model Type:** Cross Encoder - **Maximum Sequence Length:** 512 tokens - **Number of Output Labels:** 1 label ### Model Sources - **Documentation:** [Sentence Transformers Documentation](https://sbert.net) - **Documentation:** [Cross Encoder Documentation](https://www.sbert.net/docs/cross_encoder/usage/usage.html) - **Repository:** [Sentence Transformers on GitHub](https://github.com/UKPLab/sentence-transformers) - **Hugging Face:** [Cross Encoders on Hugging Face](https://huggingface.co/models?library=sentence-transformers&other=cross-encoder) ## Usage ### Direct Usage (Sentence Transformers) First install the Sentence Transformers library: ```bash pip install -U sentence-transformers ``` Then you can load this model and run inference. ```python from sentence_transformers import CrossEncoder # Download from the 🤗 Hub model = CrossEncoder("cross_encoder_model_id") # Get scores for pairs of texts pairs = [ ['Ah ha, ha, ha, ha, ha!', 'Ha, ha, ha, ha, ha, ha!'], ['Besides battling its sales slump, Siebel also has been sparring with some investors upset about huge stock option windfalls company managers have pocketed.', 'Besides a sales slump, Siebel is sparring with some shareholders over management stock option windfalls.'], ['Rosenthal declined comment on the Garrett situation Tuesday but said in a statement: "We had a big contract negotiation.', 'The show\'s creator and executive producer, Phil Rosenthal, quipped in a statement: "We had a big contract negotiation.'], ['The army said the raid, which comes just days after Israeli troops shot and killed Abdullah Kawasme, the Hamas leader in the city, targeted militants in Hamas.', "The arrests came just days after Israeli troops shot and killed Abdullah Kawasme, the militant group's leader in Hebron."], ['Saudi to give Lebanese army $3 billion', 'Saudi Arabia to grant Lebanese army $3 billion'], ] scores = model.predict(pairs) print(scores.shape) # (5,) # Or rank different texts based on similarity to a single text ranks = model.rank( 'Ah ha, ha, ha, ha, ha!', [ 'Ha, ha, ha, ha, ha, ha!', 'Besides a sales slump, Siebel is sparring with some shareholders over management stock option windfalls.', 'The show\'s creator and executive producer, Phil Rosenthal, quipped in a statement: "We had a big contract negotiation.', "The arrests came just days after Israeli troops shot and killed Abdullah Kawasme, the militant group's leader in Hebron.", 'Saudi Arabia to grant Lebanese army $3 billion', ] ) # [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...] ``` ## Evaluation ### Metrics #### Cross Encoder Correlation * Dataset: `sts-validation` * Evaluated with [CECorrelationEvaluator](https://sbert.net/docs/package_reference/cross_encoder/evaluation.html#sentence_transformers.cross_encoder.evaluation.CECorrelationEvaluator) | Metric | Value | |:-------------|:-----------| | pearson | 0.9194 | | **spearman** | **0.9162** | ## Training Details ### Training Dataset #### Unnamed Dataset * Size: 5,749 training samples * Columns: sentence_0, sentence_1, and label * Approximate statistics based on the first 1000 samples: | | sentence_0 | sentence_1 | label | |:--------|:------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------|:---------------------------------------------------------------| | type | string | string | float | | details | | | | * Samples: | sentence_0 | sentence_1 | label | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------|:------------------| | Ah ha, ha, ha, ha, ha! | Ha, ha, ha, ha, ha, ha! | 0.9 | | Besides battling its sales slump, Siebel also has been sparring with some investors upset about huge stock option windfalls company managers have pocketed. | Besides a sales slump, Siebel is sparring with some shareholders over management stock option windfalls. | 0.76 | | Rosenthal declined comment on the Garrett situation Tuesday but said in a statement: "We had a big contract negotiation. | The show's creator and executive producer, Phil Rosenthal, quipped in a statement: "We had a big contract negotiation. | 0.75 | * Loss: [BinaryCrossEntropyLoss](https://sbert.net/docs/package_reference/cross_encoder/losses.html#binarycrossentropyloss) with these parameters: ```json { "activation_fn": "torch.nn.modules.linear.Identity", "pos_weight": null } ``` ### Training Hyperparameters #### Non-Default Hyperparameters - `eval_strategy`: steps - `per_device_train_batch_size`: 96 - `per_device_eval_batch_size`: 96 - `fp16`: True #### All Hyperparameters
Click to expand - `overwrite_output_dir`: False - `do_predict`: False - `eval_strategy`: steps - `prediction_loss_only`: True - `per_device_train_batch_size`: 96 - `per_device_eval_batch_size`: 96 - `per_gpu_train_batch_size`: None - `per_gpu_eval_batch_size`: None - `gradient_accumulation_steps`: 1 - `eval_accumulation_steps`: None - `torch_empty_cache_steps`: None - `learning_rate`: 5e-05 - `weight_decay`: 0.0 - `adam_beta1`: 0.9 - `adam_beta2`: 0.999 - `adam_epsilon`: 1e-08 - `max_grad_norm`: 1 - `num_train_epochs`: 3 - `max_steps`: -1 - `lr_scheduler_type`: linear - `lr_scheduler_kwargs`: {} - `warmup_ratio`: 0.0 - `warmup_steps`: 0 - `log_level`: passive - `log_level_replica`: warning - `log_on_each_node`: True - `logging_nan_inf_filter`: True - `save_safetensors`: True - `save_on_each_node`: False - `save_only_model`: False - `restore_callback_states_from_checkpoint`: False - `no_cuda`: False - `use_cpu`: False - `use_mps_device`: False - `seed`: 42 - `data_seed`: None - `jit_mode_eval`: False - `use_ipex`: False - `bf16`: False - `fp16`: True - `fp16_opt_level`: O1 - `half_precision_backend`: auto - `bf16_full_eval`: False - `fp16_full_eval`: False - `tf32`: None - `local_rank`: 0 - `ddp_backend`: None - `tpu_num_cores`: None - `tpu_metrics_debug`: False - `debug`: [] - `dataloader_drop_last`: False - `dataloader_num_workers`: 0 - `dataloader_prefetch_factor`: None - `past_index`: -1 - `disable_tqdm`: False - `remove_unused_columns`: True - `label_names`: None - `load_best_model_at_end`: False - `ignore_data_skip`: False - `fsdp`: [] - `fsdp_min_num_params`: 0 - `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False} - `tp_size`: 0 - `fsdp_transformer_layer_cls_to_wrap`: None - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None} - `deepspeed`: None - `label_smoothing_factor`: 0.0 - `optim`: adamw_torch - `optim_args`: None - `adafactor`: False - `group_by_length`: False - `length_column_name`: length - `ddp_find_unused_parameters`: None - `ddp_bucket_cap_mb`: None - `ddp_broadcast_buffers`: False - `dataloader_pin_memory`: True - `dataloader_persistent_workers`: False - `skip_memory_metrics`: True - `use_legacy_prediction_loop`: False - `push_to_hub`: False - `resume_from_checkpoint`: None - `hub_model_id`: None - `hub_strategy`: every_save - `hub_private_repo`: None - `hub_always_push`: False - `gradient_checkpointing`: False - `gradient_checkpointing_kwargs`: None - `include_inputs_for_metrics`: False - `include_for_metrics`: [] - `eval_do_concat_batches`: True - `fp16_backend`: auto - `push_to_hub_model_id`: None - `push_to_hub_organization`: None - `mp_parameters`: - `auto_find_batch_size`: False - `full_determinism`: False - `torchdynamo`: None - `ray_scope`: last - `ddp_timeout`: 1800 - `torch_compile`: False - `torch_compile_backend`: None - `torch_compile_mode`: None - `include_tokens_per_second`: False - `include_num_input_tokens_seen`: False - `neftune_noise_alpha`: None - `optim_target_modules`: None - `batch_eval_metrics`: False - `eval_on_start`: False - `use_liger_kernel`: False - `eval_use_gather_object`: False - `average_tokens_across_devices`: False - `prompts`: None - `batch_sampler`: batch_sampler - `multi_dataset_batch_sampler`: proportional - `router_mapping`: {} - `learning_rate_mapping`: {}
### Training Logs | Epoch | Step | sts-validation_spearman | |:------:|:----:|:-----------------------:| | 0.3333 | 20 | 0.9100 | | 0.6667 | 40 | 0.9133 | | 1.0 | 60 | 0.9128 | | 1.3333 | 80 | 0.9152 | | 1.6667 | 100 | 0.9159 | | 2.0 | 120 | 0.9162 | ### Framework Versions - Python: 3.12.2 - Sentence Transformers: 5.0.0 - Transformers: 4.51.3 - PyTorch: 2.7.1+cu126 - Accelerate: 1.9.0 - Datasets: 4.0.0 - Tokenizers: 0.21.2 ## Citation ### BibTeX #### Sentence Transformers ```bibtex @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } ```