File size: 6,716 Bytes
75df548
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
<div align="center">
    <a href="https://arxiv.org/abs/2603.03907"><img src="https://img.shields.io/badge/Arxiv-preprint-red"></a>
    <a href="https://yzc-ippl.github.io/FG-IAA/"><img src="https://img.shields.io/badge/Homepage-green"></a>
    <a href='https://github.com/yzc-ippl/FG-IAA/stargazers'><img src='https://img.shields.io/github/stars/yzc-ippl/FG-IAA.svg?style=social'></a>
</div>

<h1 align="center">Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks</h1>

<div align="center">
    Zhichao Yang<sup>1†</sup>,
    Jianjie Wang<sup>1†</sup>,
    Zhixianhe Zhang<sup>1</sup>,
    Pangu Xie<sup>1</sup>,
    Xiangfei Sheng<sup>1</sup>,
    Pengfei Chen<sup>1</sup>,
    Leida Li<sup>1,2*</sup>
</div>

<div align="center">
  <sup>1</sup>School of Artificial Intelligence,
  <sup>2</sup>State Key Laboratory of EMIM, Xidian University
</div>

<div align="center">
<sup>†</sup>Equal contribution &nbsp;&nbsp; <sup>*</sup>Corresponding author
</div>

<br>

<div align="center">
  <img src="FGAesthetics+Q.png" width="900"/>
</div>

<div style="font-family: sans-serif; margin-bottom: 2em;">
    <h2 style="border-bottom: 1px solid #eaecef; padding-bottom: 0.3em; margin-bottom: 1em;">News</h2>
    <ul style="list-style-type: none; padding-left: 0;">
        <li style="margin-bottom: 0.8em;">
            <strong>[2026-04-10]</strong> ✨</span>✨</span> The <strong>Inference Code</strong> and <strong>Pre-trained Weights</strong>, are now publicly available. A demo video demonstrating FGAesQ's application in <strong>LivePhoto Cover Recommendation</strong> is also provided.
        </li>
        <li style="margin-bottom: 0.8em;">
            <strong> [2026-04-09]</strong> πŸŽ‰</span>πŸŽ‰</span> Congratulations! Our paper has been accepted for an <strong>Oral Presentation</strong> at CVPR 2026.
        </li>
        <li style="margin-bottom: 0.8em;">
            <strong>[2026-02-21]</strong> πŸŽ‰</span>πŸŽ‰</span>  Our paper, "Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks", has been accepted to <strong>CVPR 2026</strong>!
        </li>
    </ul>
</div>

## Applicatons (More scenarios will be uncovered)

<div align="center">
  <video src="https://github.com/yzc-ippl/FG-IAA/releases/download/v1.0/demo_2.mp4" width="900" controls></video>
</div>

## Quick Start

This guide will help you get started with FGAesQ inference in minutes.

### 1. Installation

Clone the repository and install the required dependencies:

```bash
git clone https://github.com/yzc-ippl/FG-IAA.git
cd FG-IAA
pip install -r requirements.txt
```

> **Note:** The CLIP dependency is installed directly from the official OpenAI repository and will be fetched automatically via `pip install -r requirements.txt`.

### 2. Download Pre-trained Weights

Download the pre-trained model weights from: [**(Hugging Face)**](https://huggingface.co/yzc002/FGAesQ) &nbsp;|&nbsp; [**(Baidu Netdisk)**](#)

Place the downloaded weight file at a path of your choice and set `MODEL_PATH` accordingly in the inference scripts.

The expected project structure is as follows:

```
FG-IAA/
FGAesQ_Inference/
   β”œβ”€β”€utils/
        β”œβ”€β”€ FGAesQ.py               # Model definition
        β”œβ”€β”€ DiffToken.py            # Differential token preprocessing
        β”œβ”€β”€ data_utils.py
        └── clip_vit_base_16_224.pt
   β”œβ”€β”€ inference_series.py         # Series-mode inference
   β”œβ”€β”€ inference_single.py         # Single-image inference
   β”œβ”€β”€ requirements.txt
 README.md
```

### 3. Run Inference

FGAesQ supports two inference modes: **Series Mode** for photo series ranking, and **Single Mode** for individual image scoring.

---

#### πŸ–ΌοΈ Mode 1 β€” Single Image / Folder Scoring

Use `inference_single.py` to score a single image or all images within a folder.

**Configuration** (edit the `main()` function in `inference_single.py`):

```python
MODEL_PATH = "path/to/your/model.pt"   # Path to the pre-trained weights
INPUT_PATH = "path/to/image_or_folder" # Single image file or folder of images
OUTPUT_TXT = "path/to/output.txt"      # Output txt path (folder mode only; set None to auto-generate)
DEVICE     = "cuda"
BATCH_SIZE = 128
```

**Run:**

```bash
python inference_single.py
```

**Output format** (`single_result.txt`):

```
Total: 3
============================================================

  1. photo_A.jpg                                      0.872314
  2. photo_B.jpg                                      0.751203
  3. photo_C.jpg                                      0.634891
```

- **Single image**: the predicted aesthetic score is printed directly to the terminal.
- **Folder**: a ranked list of all images with scores is saved to `OUTPUT_TXT`.

---

#### πŸ“‚ Mode 2 β€” Photo Series Ranking

Use `inference_series.py` to rank images within multiple photo series simultaneously.

The input folder should contain one sub-folder per series, with image files named in the format `{series_id}-{index}.jpg` (e.g., `000009-01.jpg`, `000009-02.jpg`).

```
input_folder/
 000009/
   β”œβ”€β”€ 000009-01.jpg
   β”œβ”€β”€ 000009-02.jpg
   └── 000009-03.jpg
 000010/
   β”œβ”€β”€ 000010-01.jpg
   └── 000010-02.jpg
 ...
```

**Configuration** (edit the `main()` function in `inference_series.py`):

```python
MODEL_PATH    = "path/to/your/model.pt"   # Path to the pre-trained weights
INPUT_FOLDER  = "path/to/series_folder"   # Root folder containing all series sub-folders
OUTPUT_FOLDER = "path/to/series_result"   # Output directory for per-series result txt files
DEVICE        = "cuda:0"
BATCH_SIZE    = 64
MAX_SIZE      = 2048  # Max image resolution (long edge). Use None for no limit.
                      # Recommended: 2048 if many images exceed this resolution.
```

**Run:**

```bash
python inference_series.py
```

**Output format** (one `{series_id}_result.txt` per series in `OUTPUT_FOLDER`):

```
Series: 9
Count: 3
============================================================

Ranking: 000009-02.jpg  000009-01.jpg  000009-03.jpg

Scores:  0.8812  0.7654  0.6231

Order: 000009-02.jpg > 000009-01.jpg > 000009-03.jpg
```

Each output file contains the predicted ranking and aesthetic scores for all images in that series, sorted from best to worst.

---

## Citation

If you find this work useful, please cite our paper!

```bibtex
@article{yang2026fine,
  title={Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks},
  author={Yang, Zhichao and Wang, Jianjie and Zhang, Zhixianhe and Xie, Pangu and Sheng, Xiangfei and Chen, Pengfei and Li, Leida},
  journal={arXiv preprint arXiv:2603.03907},
  year={2026}
}
```