aurick commited on
Commit
bf42454
·
verified ·
1 Parent(s): dd9e1f2

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +7 -0
README.md CHANGED
@@ -756,6 +756,13 @@ The training data curation process includes cleaning, processing, and modifying
756
  </tbody>
757
  </table>
758
 
 
 
 
 
 
 
 
759
  ## 6. Safety
760
 
761
  We conducted safety evaluations ahead of release, spanning both everyday human-AI interaction and dangerous-capability testing. Because Inkling-Small is multimodal, we paid attention to whether safety behavior held consistently across text, audio, and image inputs. We applied mitigations to reduce risks before release.
 
756
  </tbody>
757
  </table>
758
 
759
+ - <small>Inkling-Small against open-and closed-weights models across the full eval suite. Activated and total parameters are given for scale; a dash means the score was not available at the time of writing.</small>
760
+ - <small>SWEBench Verified: Inkling and Inkling-Small’s numbers are reported using a bash-only harness. We use self-reported numbers for external models.</small>
761
+ - <small>Terminal Bench 2.1: Inkling and Inkling-Small’s numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. We use self-reported numbers for external models where available. Otherwise, we report performance using our internal harness.</small>
762
+ - <small>Audio MC: Other models were evaluated internally since they are not on the official leaderboard.</small>
763
+ - <small>VoiceBench: VoiceBench uses rule-based, hard-coded string matching for grading, making the evaluation sensitive to output-formatting differences. We therefore added a system message instructing models to follow the expected answer format.</small>
764
+ - <small>HLE with tools: We benchmarked Minimax M2.7, Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna using our internal harness.</small>
765
+
766
  ## 6. Safety
767
 
768
  We conducted safety evaluations ahead of release, spanning both everyday human-AI interaction and dangerous-capability testing. Because Inkling-Small is multimodal, we paid attention to whether safety behavior held consistently across text, audio, and image inputs. We applied mitigations to reduce risks before release.