llama3_2_1b_HNPU / v79 /llama_full_wqo_o3.bin

Commit History

revert to working P=128 while a16 int16-calib is tuned
6caa332
verified

Aman0runanywhere commited on

a16 P=512 prefill, 8-sample calib (better int16 activation scales)
aa97157
verified

Aman0runanywhere commited on

P=512 a16 (int16 acts, FLOAT I/O, int8 per-channel weights): fix -130 fp16-accum overflow, host-op binds f32
76b63aa
verified

Aman0runanywhere commited on

P=512 prefill W8A16 (int16 acts + int32 accum): fix multi-turn -130; user-diagnosed fp16-accumulator overflow
349aa82
verified

Aman0runanywhere commited on

revert to working P=128: P=512 monolithic prefill drifts on v79 (both toolchains, empty output); pursuing runtime fallback
4692c0a
verified

Aman0runanywhere commited on

P=512 prefill CANDIDATE (user-authorized on-device validation; revert if fails)
7a59579
verified

Aman0runanywhere commited on

revert prefill to P=128: P=512 build (mixed 2.48/2.45 toolchain) produced empty output on v79; rebuilding consistently
154e03a
verified

Aman0runanywhere commited on

prefill P=128->512: fix multi-turn -130 (prompt>128 tokens); w8a16, export-exact 7/7 vs HF gold
5ca9603
verified

Aman0runanywhere commited on

Upload v79/llama_full_wqo_o3.bin with huggingface_hub
7b1457d
verified

Aman0runanywhere commited on