Text Generation
Transformers
Safetensors
Turkish
qwen3
text-generation-inference
unsloth
mmlu
evaluation
conversational

Qwen3 4B Bilimkurgu - Türkçe MMLU Benchmark & Model Kartı


📈 Türkçe MMLU Benchmark Test Sonuçları

Model, alibayram/yapay_zeka_turkce_mmlu_model_cevaplari test veri kümesi kullanılarak Türkçe MMLU benchmark değerlendirmesine tabi tutulmuştur.

📊 Türkçe MMLU Benchmark Karşılaştırma Tablosu

Model Türü Model Adı Genel Başarı (%) Doğru / Toplam Soru Test Süresi (sn)
Base Model unsloth/qwen3-4b-instruct-2507-unsloth-bnb-4bit %26.21 1625 / 6200 218.4s
Fine-Tuned Model gururaser/qwen3-4b-bilimkurgu %34.35 2130 / 6200 251.23s

🧪 Test Metodolojisi ve Detaylar

  • Değerlendirme Scripti: alibayram/yapay_zeka_turkce_mmlu_bolum_sonuclari/olcum.py mantığı temel alınmıştır.
  • Anlamsal Karşılaştırma: paraphrase-multilingual-mpnet-base-v2 modeli ile şık eşleştirme ve opsiyonel harf eşleme.
  • Ortam: Google Colab A100 GPU (4-bit NF4 Quantization)
  • Test Seti: alibayram/yapay_zeka_turkce_mmlu_model_cevaplari

""

Downloads last month
16
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train gururaser/qwen3-4b-bilimkurgu