--- title: FairRecovery++ emoji: ๐Ÿ™๏ธ colorFrom: yellow colorTo: red sdk: docker pinned: false --- # ๐Ÿ™๏ธ FairRecovery++: Teaching LLMs to Make Fair Disaster Recovery Decisions ๐Ÿšจ Problem: AI systems optimize speed โ€” but silently ignore the most vulnerable. ๐Ÿ’ก Solution: FairRecovery++ trains LLMs to make recovery decisions that are both efficient AND fair. > **We didnโ€™t just train an agent to recover cities โ€” we trained it to recover them fairly.** [![Llama-1B](https://img.shields.io/badge/Model-Llama--3.2--1B-orange)](https://huggingface.co/Joshua1702/fairrecovery-Llama-3.2-1B) [![Qwen-7B](https://img.shields.io/badge/Model-Qwen--2.5--7B-purple)](https://huggingface.co/Joshua1702/fairrecovery-Qwen2.5-7B-GRPO) [![Qwen-1.5B](https://img.shields.io/badge/Model-Qwen--2.5--1.5B-blue)](https://www.kaggle.com/code/joshuaragiland/fairevlo-lite) [![OpenEnv](https://img.shields.io/badge/OpenEnv-compliant-blue)](https://github.com/meta-pytorch/OpenEnv) [![Theme](https://img.shields.io/badge/Theme-3.1%20%7C%202-orange)](https://huggingface.co/openenv) [![Space](https://img.shields.io/badge/๐Ÿค—%20Space-Live-green)](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus) ## ๐Ÿš€ Key Results (Benchmarking) ```text โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•— โ•‘ FINAL RESULTS โ€” Fair-GRPO-RLVR vs Greedy โ•‘ โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•ฃ โ•‘ Metric Baseline Trained Delta % โ•‘ โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•ฃ โ•‘ โœ… Reward 0.7238 0.7953 +0.0714 +9.9% โ•‘ โ•‘ โœ… Fairness 0.7433 0.7730 +0.0298 +4.0% โ•‘ โ•‘ โœ… Utility 0.5663 0.7152 +0.1489 +26.3% โ•‘ โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•ฃ โ•‘ ๐Ÿ† IMPROVED ON ALL METRICS โ€” Fairness Trap escaped! โ•‘ โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ• ``` - ๐Ÿ“ˆ **57% overall reward improvement** in Premium Qwen-7B runs. - โš– **24% fairness boost** (0.73 โ†’ 0.91) across trained agents. - ๐Ÿง  Learned to prioritize vulnerable zones vs greedy cost-minimization. --- When disaster strikes โ€” floods in Chennai, earthquakes, hurricanes โ€” city authorities face an impossible-looking problem: limited crews, limited budget, dozens of damaged neighborhoods, and no time. Most AI systems trained to help optimize for speed: fix what's easiest first, maximize total service restored. The problem? **Easy to fix almost always means wealthy.** The informal settlements, the elderly care facilities, and the low-income neighborhoods that took the hardest hit? They wait the longest. We built **FairRecovery++** to train LLM agents that escape this trap using **GRPO (Group Relative Policy Optimization)**. --- ## What is FairRecovery++? FairRecovery++ is an [OpenEnv](https://github.com/meta-pytorch/OpenEnv) RL environment where an AI agent acts as a post-disaster Recovery Planner for a simulated city of 5 zones. Each episode spans 10 simulated recovery days. The agent must: 1. **Analyze** which zones are most critically damaged. 2. **Allocate** limited resources (medical units, power crews, water tankers, housing repairs). 3. **Execute** the allocation and observe the outcome. 4. Repeat across multiple days, then **submit** a final recovery plan. The environment rewards the agent not just for total service restored, but for **equitable** service restored โ€” measured by the gap between how well vulnerable zones are served versus non-vulnerable ones. --- ## ๐Ÿ’ก Why This Is New (Not Just Another RL Environment) Unlike existing disaster RL systems that optimize only efficiency, FairRecovery++ introduces a **โ€œFairness Trapโ€** where greedy policies fail. We are the first to: - Encode fairness as a verifiable RL objective - Train LLM planners using GRPO on long-horizon recovery decisions - Prevent reward hacking with safety-aware constraints --- ## โš– Fairness Metric We define fairness as: > **Fairness = 1 โˆ’ service disparity across zones** Where disparity is the average difference in service levels. - Fairness = 1 โ†’ perfectly equal recovery - Fairness โ†“ โ†’ unequal allocation This makes fairness: - โœ” hard to game This ensures fairness cannot be faked โ€” improving fairness requires real redistribution of resources. > **Note:** Typical fairness ranges from **0.70โ€“0.95** across complex scenarios; we intentionally avoid saturation to ensure the metric remains sensitive to real-world trade-offs. --- ## ๐ŸŽฏ The Fairness Trap The hard scenario is deliberately designed with a trap: - **Zone 0** (wealthy district): 35% damage, easy to fix, 8% vulnerable population. - **Zone 4** (informal settlement): 92% damage, 96% vulnerable population. A naive utility-maximizing agent always picks Zone 0: lower cost, faster payoff, higher immediate reward. Zone 4 gets ignored. Our trained **Qwen-7B** and **Llama-3.2-1B** agents learn to prioritize Zone 4 โ€” because its 96% vulnerable population deserves equitable access to recovery services, even if it costs more per unit of service gained. --- ## ๐Ÿ” Before vs After Training ### โŒ Baseline (Greedy Policy) - Prioritizes low-cost zones - Ignores high-vulnerability areas - Leads to unfair recovery ### โœ… Trained Agent (Fair-GRPO-RLVR) - Identifies critical zones (high damage + vulnerability) - Allocates resources more equitably - Improves both fairness and overall recovery --- ## ๐Ÿค— Trained Models We provide two pre-trained models optimized for this environment using GRPO (TRL + Unsloth): * **Premium Agent (Qwen-7B)**: [Joshua1702/fairrecovery-Qwen2.5-7B-GRPO](https://huggingface.co/Joshua1702/fairrecovery-Qwen2.5-7B-GRPO) *Best for complex reasoning and near-perfect fairness (Equity: 0.912).* * **Efficient Agent (Llama-1B)**: [Joshua1702/fairrecovery-Llama-3.2-1B](https://huggingface.co/Joshua1702/fairrecovery-Llama-3.2-1B) *Best for low-latency edge deployment with strong equity performance (Equity: 0.840).* * **Lite Agent (Qwen-1.5B)**: [Kaggle Notebook](https://www.kaggle.com/code/joshuaragiland/fairevlo-lite) *Ideal for rapid prototyping and quick testing cycles.* --- ## ๐Ÿ—๏ธ Architecture ``` LLM Agent (GRPO trained) โ”‚ โ–ผ FairRecoveryAction โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ Safety Shield (shield.py) โ”‚ โ† blocks invalid actions before mutation โ”‚ Stage validator โ”‚ โ”‚ Budget enforcer โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ valid action โ–ผ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ FairRecoveryEnvironment โ”‚ โ† core OpenEnv Environment class โ”‚ Multi-step protocol: โ”‚ โ”‚ analyze โ†’ allocate โ†’ โ”‚ โ”‚ execute โ†’ (ร—MAX_DAYS) โ†’ โ”‚ โ”‚ submit โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ updated CityState โ–ผ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ Reward Engine (RLVR) โ”‚ โ† no learned reward model โ”‚ R_exec (service improvement)โ”‚ โ”‚ R_fair (disparity reduction)โ”‚ โ”‚ R_safe (constraint penalty) โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ per-step reward โ–ผ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ Composable Rubrics (RFC 004) โ”‚ โ† FairnessRubric + UtilityRubric โ”‚ Terminal episode scoring โ”‚ + AnalysisRubric โ”‚ Grader score โˆˆ (0.01, 0.99) โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ ``` --- ## ๐ŸŽฎ What the Agent Sees, Does, and Gets Rewarded For ### Observation (per step) ```json { "zones": [ {"zone_id": 4, "damage": 0.92, "service": 0.08, "vulnerable_ratio": 0.96} ], "day": 2, "budget_left": 25.0, "step_stage": "allocate", "fairness_score": -0.61, "cumulative_reward": 0.142 } ``` ### Action (multi-step protocol โ€” not just one choice) ```json // Step 1: analyze {"action_type": "analyze", "critical_zones": [3, 4], "reasoning": "highest damage ร— vulnerability"} // Step 2: allocate {"action_type": "allocate", "allocations": [ {"zone": 4, "resource": "medical"}, {"zone": 3, "resource": "power"} ]} // Step 3: execute (commits allocations, receives dense reward) {"action_type": "execute"} // After MAX_DAYS: submit (receives terminal bonus) {"action_type": "submit"} ``` ### Reward System (RLVR โ€” all verifiable, no learned model) | Component | Formula | Weight | What it teaches | |-----------|---------|--------|-----------------| | `R_exec` | Avg service improvement this step | 0.5 | Restore services efficiently | | `R_fair` | โˆ’(avg_svc_normal โˆ’ avg_svc_vuln) | 1.0 | Don't leave vulnerable zones behind | | `R_safe` | โˆ’0.1 ร— violations | 0.5 | Respect constraints | | `R_analysis` | Overlap(chosen, top-k) | 0.1 | Identify critical zones | | **Terminal bonus**| 0.5ร—avg_svc + 0.5ร—(1+R_fair) | โ€” | Long-horizon outcome | *Fairness is normalized and bounded to prevent saturation and reward hacking.* **Grader score: `0.6 ร— avg_service + 0.4 ร— (1 + fairness)` clamped to (0.01, 0.99)** ### Anti-Reward-Hacking Measures - Stage ordering enforced (can't skip analyze โ†’ go straight to execute) - Budget overflow: allocations rejected + penalty, state NOT mutated - Persistent ignore penalty: if vulnerable zones receive 0 resources for 2+ consecutive days - Early-submit blocked until MIN_STEPS reached - Step cap: force-terminate at MAX_STEPS_SAFETY_CAP **These safeguards ensure the agent cannot exploit reward signals without solving the actual task.** --- ## ๐Ÿ“Š Visual Evidence Dashboard In this section, we highlight the most critical performance metrics and fairness improvements achieved during training. These plots provide empirical proof of the agent's ability to navigate the fairness-utility trade-off. ### ๐Ÿ“‰ Training Evidence (Mandatory) **All results are generated from real GRPO training runs (not synthetic or static evaluation).** We trained using GRPO with Unsloth acceleration. - **Models:** Llama-3.2-1B, Qwen-2.5-1.5B, & Qwen-2.5-7B (4-bit) - **Iteration Strategy:** We use smaller models (Llama-1B, Qwen-1.5B) for rapid iteration and stable learning, then scale to Qwen-7B for final performance. - **Qwen-1.5B Quick-Test:** - **Episodes:** 20โ€“50 (multiple seeds) - **Runtime:** < 20 min on T4 (Kaggle) - **Outcome:** Loss decreased steadily over time while reward converged. ๐Ÿ“Š **Training Loss (Curriculum)** ![Training Loss](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/training_loss.png) *Figure: The agent steadily converges its policy over the curriculum.* ### ๐Ÿ† The Winning Metric: Dual-Model Comparison ![Model Comparison](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/model_comparison.png) *Figure 1: Comparison between Baseline, Llama-1B, and Qwen-7B. Our Premium Qwen-7B agent reaches a near-perfect **0.912 Equity Index**.* ### ๐Ÿ“ˆ Training & Strategy Results **Baseline vs Trained (Qwen)** ![Training Results](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/training_results.png) *Figure 2: 57% reward improvement.* **Reward Heatmap** ![Score Heatmap](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/score_heatmap.png) *Figure 3: Consistency across episodes.* **Fairness-Utility Frontier** ![Utility vs Fairness](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/utility_vs_fairness.png) **Extended Training Metrics** ![Extra Metrics](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/training_metrics_extra.png) **Advanced Stability Analysis** ![Extra Metrics 2](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/training_metrics_extra_2.png) *Figure 6: Advanced stability and variance analysis across multi-seed runs.* ### โฑ๏ธ Execution Dynamics (Step-by-Step) ![Fairness Improvement](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/fairness_vs_episode.png) *Figure 7: Global Fairness achievement trend over 32 training cycles.* ![Component Rewards](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/component_rewards.png) *Figure 8: Decomposed reward components showing the sacrifice of immediate utility for long-term equity.* ![Reward vs Steps](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/reward_vs_steps.png) *Figure 9: Performance stability during the 10-day recovery window.* ![Fairness vs Steps](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/fairness_vs_steps.png) *Figure 10: Cumulative equity growth per action.* --- ### Key Numbers | Metric | Greedy Baseline | Llama-3.2-1B | Qwen-2.5-7B-GRPO | |--------|----------------|--------------|------------------| | Avg Episode Reward | 0.548 | 0.785 | **0.864** (+57%) | | Avg Final Fairness | 0.732 | 0.840 | **0.912** (+24%) | | Avg Final Utility | 0.545 | 0.712 | **0.808** (+48%) | | Strategy discovered | Always Zone 0 | Vulnerable-first | Strategic-equitable | --- ## ๐Ÿš€ Quick Start ```bash git clone https://github.com/joshua400/FairRecovery-PlusPlus cd FairRecovery-PlusPlus pip install -r requirements.txt uvicorn server.app:app --reload ``` ```bash # Verify environment curl http://localhost:8000/health # Run a full episode python inference.py --difficulty hard --episodes 3 --policy fairness_aware ``` ### Use as OpenEnv client ```python from client import FairRecoveryEnv env = FairRecoveryEnv(base_url="https://Joshua1702-FairRecovery-PlusPlus.hf.space") obs = env.reset(difficulty="hard") for day in range(5): action = your_policy(obs) # analyze โ†’ allocate โ†’ execute obs = env.step(action) print(f"Day {obs.day}: reward={obs.reward:+.3f} fair={obs.fairness_score:.3f}") ``` ### Run training (Colab) Open `train_COMPLETE.ipynb` โ€” runs on free Colab T4 in ~10 minutes. ## ๐Ÿ“ Project Layout & Resources ### Repository Structure ``` fairrecovery_env/ # ๐Ÿง  Core OpenEnv Engine โ”œโ”€โ”€ constants.py # Reward weights & thresholds โ”œโ”€โ”€ models.py # Pydantic v2 Action/Obs/State โ”œโ”€โ”€ rewards.py # 5-component RLVR reward engine โ”œโ”€โ”€ rubrics.py # RFC 004 composable rubrics โ””โ”€โ”€ ... # CityState, Tasks, Shield validator server/ # ๐Ÿš€ Deployment & UI โ”œโ”€โ”€ app.py # FastAPI + Gradio Web Dashboard โ””โ”€โ”€ fairrecovery_env.py # OpenEnv Environment Wrapper asset_final/ # ๐Ÿ“Š Final Evidence & Model Adapters โ”œโ”€โ”€ plots/ # Consolidated training visualizations โ””โ”€โ”€ model/ # LoRA adapter weights (Qwen-7B) docs/ # ๐Ÿ“ Documentation & Reports โ”œโ”€โ”€ train_llama_final.ipynb # Llama-1B GRPO Notebook โ””โ”€โ”€ train_qwen_final.ipynb # Qwen-7B GRPO Notebook ``` ### ๐Ÿ”— Key Materials | Resource | Link | |----------|------| | ๐Ÿค— **Live Space** | [Joshua1702/FairRecovery-PlusPlus](https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus) | | ๐Ÿ’ป **GitHub** | [joshua400/FairRecovery-PlusPlus](https://github.com/joshua400/FairRecovery-PlusPlus) | | ๐Ÿ““ **Llama Training** | [Notebook](docs/train_llama_final.ipynb) [![Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/joshua400/FairRecovery-PlusPlus/blob/main/docs/train_llama_final.ipynb) | | ๐Ÿ““ **Qwen Training** | [Notebook](docs/train_qwen_final.ipynb) [![Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/joshua400/FairRecovery-PlusPlus/blob/main/docs/train_qwen_final.ipynb) | | ๐Ÿงช **Quick Testing (1.5B)** | [Kaggle Notebook](https://www.kaggle.com/code/joshuaragiland/fairevlo-lite) | | ๐Ÿ“ **HF Blog Post** | [blog.md](blog.md) | --- ## ๐ŸŒ Why It Matters Post-disaster recovery planning is a $200B/year global challenge. AI systems that optimize only for speed or total utility **systematically disadvantage the most vulnerable populations** โ€” the elderly, disabled, and low-income communities who live in the hardest-hit zones. FairRecovery++ is the first OpenEnv environment to encode intersectional fairness as a verifiable, first-class RL objective, making it a research-grade benchmark for safe and fair LLM agent training. --- ## OpenEnv Compliance Checklist - โœ… `openenv.yaml` manifest present - โœ… `Environment` base class used with try/import fallback - โœ… `reset()` / `step()` / `state()` standard API - โœ… Pydantic v2 typed `Action` / `Observation` / `State` - โœ… Hosted on HF Spaces (Docker) - โœ… GRPO training with TRL + Unsloth (see `docs/train_qwen_final.ipynb` and `docs/train_llama_final.ipynb`) - โœ… Multi-Model evidence: Llama-1B & Qwen-7B plots in `https://huggingface.co/spaces/Joshua1702/FairRecovery-PlusPlus/resolve/main/evidence/plots/` - โœ… Episode data in `episode_log.csv` - โœ… Composable rubrics (OpenEnv RFC 004) - โœ… Anti-reward-hacking: stage gates + persistent ignore penalty --- ## ๐ŸŒ Why This Matters Disaster recovery systems often optimize speed, but ignore vulnerable populations. FairRecovery++ shows that AI can: - make fair decisions under constraints - balance efficiency and equity - improve real-world disaster planning ๐Ÿ‘‰ This opens the door to safer, fairer AI systems. ## ๐ŸŽฌ What the Agent Learned Initially, the agent behaves like a human under pressure: **fix what is easiest first.** After training, it learns something deeper: **true recovery is not about speed โ€” itโ€™s about fairness.** This shift is what FairRecovery++ enables. --- *Built for the Meta PyTorch OpenEnv Hackathon India 2026.*