SentenceTransformer based on Mihaiii/gte-micro-v4

This is a sentence-transformers model finetuned from Mihaiii/gte-micro-v4. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: Mihaiii/gte-micro-v4
  • Maximum Sequence Length: 256 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("philipp-zettl/gte-micro-v4-mtg-v2")
# Run inference
sentences = [
    'creature — drake',
    'Title: Fighting Drake\nCost: {2}{U}{U}\nColors: U\nType: Creature — Drake\nDesc: Flying',
    "Title: Big Game Hunter\nCost: {1}{B}{B}\nColors: B\nType: Creature — Human Rebel Assassin\nDesc: When this creature enters, destroy target creature with power 4 or greater. It can't be regenerated.\nMadness {B} (If you discard this card, discard it into exile. When you do, cast it for its madness cost or put it into your graveyard.)",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.6089,  0.0016],
#         [ 0.6089,  1.0000, -0.1129],
#         [ 0.0016, -0.1129,  1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 2,132,917 training samples
  • Columns: sentence_0 and sentence_1
  • Approximate statistics based on the first 100 samples:
    sentence_0 sentence_1
    type string string
    modality text text
    details
    • min: 3 tokens
    • mean: 10.1 tokens
    • max: 41 tokens
    • min: 24 tokens
    • mean: 62.76 tokens
    • max: 146 tokens
  • Samples:
    sentence_0 sentence_1
    magnifying glass Title: Magnifying Glass
    Cost: {3}
    Type: Artifact
    Desc: {T}: Add {C}.
    {4}, {T}: Investigate. (Create a Clue token. It's an artifact with "{2}, Sacrifice this token: Draw a card.")
    sorcery Title: Graveyard Shift
    Cost: {4}{B}
    Colors: B
    Type: Sorcery
    Desc: This spell has flash as long as there are five or more mana values among cards in your graveyard.
    Return target creature card from your graveyard to the battlefield.
    beacon of unrest Title: Beacon of Unrest
    Cost: {3}{B}{B}
    Colors: B
    Type: Sorcery
    Desc: Put target artifact or creature card from a graveyard onto the battlefield under your control. Shuffle Beacon of Unrest into its owner's library.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 96
  • fp16: True
  • per_device_eval_batch_size: 96
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 96
  • num_train_epochs: 3
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 96
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Click to expand
Epoch Step Training Loss
0.0225 500 1.8730
0.0450 1000 0.5353
0.0675 1500 0.4567
0.0900 2000 0.4285
0.1125 2500 0.4066
0.1350 3000 0.3939
0.1575 3500 0.3833
0.1800 4000 0.3718
0.2025 4500 0.3707
0.2250 5000 0.3703
0.2475 5500 0.3630
0.2701 6000 0.3537
0.2926 6500 0.3575
0.3151 7000 0.3548
0.3376 7500 0.3565
0.3601 8000 0.3476
0.3826 8500 0.3422
0.4051 9000 0.3423
0.4276 9500 0.3414
0.4501 10000 0.3410
0.4726 10500 0.3480
0.4951 11000 0.3406
0.5176 11500 0.3293
0.5401 12000 0.3354
0.5626 12500 0.3337
0.5851 13000 0.3371
0.6076 13500 0.3376
0.6301 14000 0.3347
0.6526 14500 0.3297
0.6751 15000 0.3326
0.6976 15500 0.3297
0.7201 16000 0.3213
0.7426 16500 0.3270
0.7651 17000 0.3291
0.7876 17500 0.3291
0.8102 18000 0.3235
0.8327 18500 0.3293
0.8552 19000 0.3212
0.8777 19500 0.3261
0.9002 20000 0.3191
0.9227 20500 0.3176
0.9452 21000 0.3145
0.9677 21500 0.3267
0.9902 22000 0.3197
1.0127 22500 0.3160
1.0352 23000 0.3230
1.0577 23500 0.3182
1.0802 24000 0.3221
1.1027 24500 0.3169
1.1252 25000 0.3125
1.1477 25500 0.3135
1.1702 26000 0.3149
1.1927 26500 0.3182
1.2152 27000 0.3246
1.2377 27500 0.3161
1.2602 28000 0.3173
1.2827 28500 0.3137
1.3052 29000 0.3147
1.3278 29500 0.3110
1.3503 30000 0.3129
1.3728 30500 0.3111
1.3953 31000 0.3132
1.4178 31500 0.3176
1.4403 32000 0.3092
1.4628 32500 0.3177
1.4853 33000 0.3031
1.5078 33500 0.3126
1.5303 34000 0.3144
1.5528 34500 0.3061
1.5753 35000 0.3120
1.5978 35500 0.3083
1.6203 36000 0.3087
1.6428 36500 0.3131
1.6653 37000 0.3108
1.6878 37500 0.3131
1.7103 38000 0.3092
1.7328 38500 0.3099
1.7553 39000 0.3104
1.7778 39500 0.3049
1.8003 40000 0.3061
1.8228 40500 0.3105
1.8454 41000 0.3031
1.8679 41500 0.3008
1.8904 42000 0.3108
1.9129 42500 0.3071
1.9354 43000 0.3067
1.9579 43500 0.3077
1.9804 44000 0.3094
2.0029 44500 0.3031
2.0254 45000 0.3045
2.0479 45500 0.3056
2.0704 46000 0.3075
2.0929 46500 0.3054
2.1154 47000 0.2982
2.1379 47500 0.3003
2.1604 48000 0.3077
2.1829 48500 0.3012
2.2054 49000 0.3060
2.2279 49500 0.2995
2.2504 50000 0.3060
2.2729 50500 0.3098
2.2954 51000 0.3002
2.3179 51500 0.3004
2.3404 52000 0.3095
2.3629 52500 0.3028
2.3855 53000 0.3040
2.4080 53500 0.3056
2.4305 54000 0.3066
2.4530 54500 0.3013
2.4755 55000 0.3074
2.4980 55500 0.3054
2.5205 56000 0.3053
2.5430 56500 0.3001
2.5655 57000 0.2987
2.5880 57500 0.3104
2.6105 58000 0.3044
2.6330 58500 0.3015
2.6555 59000 0.3076
2.6780 59500 0.3012
2.7005 60000 0.3022
2.7230 60500 0.3084
2.7455 61000 0.3004
2.7680 61500 0.3056
2.7905 62000 0.3057
2.8130 62500 0.2993
2.8355 63000 0.3010
2.8580 63500 0.3019
2.8805 64000 0.2993
2.9031 64500 0.3040
2.9256 65000 0.3003
2.9481 65500 0.3000
2.9706 66000 0.2982
2.9931 66500 0.3003

Training Time

  • Training: 1.7 hours

Framework Versions

  • Python: 3.13.11
  • Sentence Transformers: 5.6.0
  • Transformers: 5.14.1
  • PyTorch: 2.13.0+cu130
  • Accelerate: 1.14.0
  • Datasets: 5.0.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
43
Safetensors
Model size
19.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for philipp-zettl/gte-micro-v4-mtg-v2

Finetuned
(2)
this model

Collection including philipp-zettl/gte-micro-v4-mtg-v2

Papers for philipp-zettl/gte-micro-v4-mtg-v2

Free AI Image Generator No sign-up. Instant results. Open Now