eh-qwen3-4b-clean-sft-seed2-merged-16bit

Status: exploratory response-confidence track. The seed-2 merged 16-bit clean-schema-SFT base of the GRPO three-seed confirmatory block (experiments/grpo-three-seed-confirmatory, resolved 2026-08-07). Trained from unsloth/Qwen3-4B-bnb-4bit on the professorsynapse/epistemic-humility-phase1 data at seed 2, then merged to 16-bit.

The adapter professorsynapse/eh-qwen3-4b-clean-sft-grpo-v2-seed2-lora loads on THIS base. Per-seed lineage is a registered rule of the block (Amendment G §3, carried forward): each seed rebuilds its own complete lineage from the foundation model; never mix this base with another seed's adapters. Its own eval reads refusal recall 89.92% / answer-on-unknown 10.08% on the full 3,369-row SelfAware set under the response-confidence contract.

Full record: experiments/grpo-three-seed-confirmatory/AMENDMENT.md and NOTEBOOK.md in the source repository.

Downloads last month
8
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for professorsynapse/eh-qwen3-4b-clean-sft-seed2-merged-16bit

Finetuned
Qwen/Qwen3-4B
Finetuned
(17)
this model
Adapters
1 model