researchit-reranker-phase6 / EVAL_V2_DESIGN.md
siddhm11's picture
Add eval v2 design document: explains why old eval was weak and how new one works
4c07339 verified
|
Raw
History Blame
6.41 kB

Evaluation Framework V2 — Survey Reading List Benchmark

Why the Old Eval Was Wrong

The Problem

The V1/V2 evaluation tested: "Can the model rank papers that were cited above papers that weren't?"

Labels: cited = 2, co-cited = 1, not cited = 0

This produced inflated metrics (nDCG@10 = 0.919, MRR = 1.0) because:

  1. Easy negatives: Random ANN results that happened not to be cited. Many of these are actually good recommendations.
  2. Binary signal: All cited papers are equally "good" — no distinction between essential reading and passing mentions.
  3. Wrong task: Predicting citations ≠ predicting what a user should read next.
  4. Model learned popularity: The top feature (14,986 importance) is candidate_num_cited_by. The model is a citation counter.

What This Means

The 0.919 nDCG@10 doesn't tell you "the model makes good recommendations." It tells you "highly-cited papers tend to be cited by other papers." That's trivially true and not useful.


The New Eval: Survey Reading List Benchmark

The Insight

Survey papers are expert-curated reading lists. When a survey author includes 200 citations organized into sections, they're telling you:

  • Which papers are essential background (Related Work)
  • Which papers provided tools/methods (Methods)
  • Which papers are just context (Introduction)
  • Which papers they chose NOT to include (the hard negatives)

This is the closest approximation to "what should a researcher read?" without having real user data.

Data Source

ines-besrour/unarxive_2024 — 2.28M full-text arXiv papers with:

  • Section-level structure (Introduction, Related Work, Methods, etc.)
  • Citation spans annotated with position and reference ID
  • Bibliography entries with DOIs, arXiv IDs, and raw text
  • Paper metadata (categories, dates, authors)

Label Scheme (5-tier)

Label Meaning Source Signal
4 Essential reading Cited ≥3 times or ≥2× in Related Work section
3 Relevant work Cited in Related Work section
2 Tool/method Cited in Methods or Experiments section
1 Background Cited only in Introduction
0 Hard negative Same arXiv category, cited by OTHER surveys, but NOT this one

Why Hard Negatives Matter

A "hard negative" is a paper that:

  • Is in the same field (same arXiv category)
  • Is good enough that OTHER survey authors in the field cited it
  • Was NOT included by THIS survey's author

This represents an expert exclusion decision. The author was likely aware of this paper and chose to leave it out. This is a much harder signal than "random paper from Qdrant that happened not to be cited."

Time Split

  • Train surveys: Published before 2023 (for training better models)
  • Eval surveys: Published 2023+ (for honest evaluation)

This prevents temporal leakage: eval surveys reference newer papers.


Metrics

Metric What It Tests Old Eval Equivalent
nDCG@10 Overall ranking quality with graded relevance nDCG@10 (but harder)
nDCG@20 Ranking quality at deeper positions nDCG@20
Recall@10 (tier≥3) How many important papers are in top-10? HR@10 (but stricter)
MRR (tier≥3) Where does the first important paper appear? MRR
MAP Average precision across all positions — (new)
Hard Negative AUC Can model separate cited from expert-excluded? — (new, most honest)
Essential nDCG@10 Can model find the must-read papers? — (new, most important)
Relevant nDCG@10 Can model find Related Work papers? — (new)

Key Metric: Hard Negative AUC

This is the single most honest metric. It answers: "If you give the model a paper that was cited by a survey expert and one that was excluded, can it tell which is which?"

  • 0.5 = random (model has no signal)
  • 0.7 = decent (better than popularity)
  • 0.8+ = strong (model understands relevance beyond citations)
  • 1.0 = perfect

Your current model's hard_neg_auc will likely be 0.55-0.65 because it's essentially a citation counter.


Expected Results

Model nDCG@10 Hard Neg AUC Essential nDCG@10
Random ~0.30 0.50 ~0.25
Citation Count (popularity) ~0.45 ~0.60 ~0.40
LightGBM V2 (current) ~0.50 ~0.62 ~0.45
LightGBM V3 (reading-path labels) ~0.65+ ~0.75+ ~0.60+
Cross-encoder ~0.70+ ~0.80+ ~0.65+
Oracle (perfect) 1.00 1.00 1.00

These are estimates. The actual numbers from running the eval will be the ground truth.


Pipeline

scripts/04_extract_survey_reading_lists.py
  ↓ (processes unarXive 2024, finds CS surveys, extracts section-annotated citations)
eval_v2/eval_survey_reading_lists.parquet + train_survey_reading_lists.parquet
  ↓ (used by evaluation script)
scripts/05_evaluate_on_surveys.py
  ↓ (runs models against eval data, reports metrics)
eval_v2/eval_results.json

Files

File Location Purpose
scripts/04_extract_survey_reading_lists.py Model repo Extract eval data from unarXive
scripts/05_evaluate_on_surveys.py Model repo Run evaluation
eval_v2/eval_survey_reading_lists.parquet Dataset repo Eval data (2023+ surveys)
eval_v2/train_survey_reading_lists.parquet Dataset repo Training data (pre-2023 surveys)
eval_v2/eval_metadata.json Dataset repo Stats and documentation
EVAL_V2_DESIGN.md Model repo This document

Running

# Step 1: Extract (run as HF job, needs ~50GB disk, 2-3hr CPU)
python scripts/04_extract_survey_reading_lists.py \
  --max-surveys 1000 --min-citations 20 --push-to-hub

# Step 2: Evaluate (quick, <5min)
python scripts/05_evaluate_on_surveys.py \
  --eval-file eval_v2_data/eval_survey_reading_lists.parquet \
  --baselines-only

# Step 3: Evaluate your model (needs Qdrant + Turso)
python scripts/05_evaluate_on_surveys.py \
  --eval-file eval_v2_data/eval_survey_reading_lists.parquet \
  --model-file production_model/reranker_v2.txt \
  --qdrant-url $QDRANT_URL --qdrant-api-key $QDRANT_API_KEY \
  --turso-url $TURSO_URL --turso-token $TURSO_DB_TOKEN