Benchmark run complete: 595/600 evaluations, grounded cuts misstatements 56% vs naive 1a8d5dc ashe0042 commited on Jun 23
Phase 3 complete: two-layer judge with deterministic citation checker, LLM entailment judge, Cohen's kappa — smoke test passed e793b82 ashe0042 commited on Jun 23
Phase 2 complete: all five configs + eval harness working end to end 8337957 ashe0042 commited on Jun 18
Config 4: KG-augmented RAG with cross-reference graph expansion (KG working) d53dec4 ashe0042 commited on Jun 18
Configs 1-3: naive, hybrid, rerank RAG pipelines + citation format fix b293818 ashe0042 commited on Jun 18
Phase 1 complete: retrieval layer (dense+BM25+RRF) + naive RAG config end-to-end b7b7b49 ashe0042 commited on Jun 18
Retrieval layer: dense search + BM25 + RRF combiner (smoke test verified) a02cc42 ashe0042 commited on Jun 18
Embed all corpus chunks with text-embedding-3-large (11613 chunks, /bin/zsh.23) 177a6eb ashe0042 commited on Jun 17
Clause-aware chunker: paragraph ID preservation, header fix, SCHEDULE tagging, APRA artifact filter 97459e2 ashe0042 commited on Jun 17
Clause-aware chunker: paragraph ID preservation, header fix, SCHEDULE tagging, APRA artifact filter e949834 ashe0042 commited on Jun 17