Fiona1019 commited on
Commit
cd159db
·
verified ·
1 Parent(s): 0ef6fa9

Upload PP-DocLayoutV3 OpenVINO IR

Browse files
Files changed (6) hide show
  1. README.md +182 -0
  2. config.json +107 -0
  3. inference.bin +3 -0
  4. inference.xml +0 -0
  5. inference.yml +100 -0
  6. preprocessor_config.json +36 -0
README.md ADDED
@@ -0,0 +1,182 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ pipeline_tag: object-detection
4
+ tags:
5
+ - PaddleOCR
6
+ - PaddleOCR-VL
7
+ - OpenVINO
8
+ - openvino-ir
9
+ - intel
10
+ - AIPC
11
+ - ocr
12
+ - layout
13
+ - layout_detection
14
+ - document-parsing
15
+ language:
16
+ - en
17
+ - zh
18
+ - multilingual
19
+ library_name: openvino
20
+ base_model:
21
+ - PaddlePaddle/PP-DocLayoutV3
22
+ base_model_relation: quantized
23
+ ---
24
+ <div align="center">
25
+
26
+ <h1 align="center">
27
+
28
+ PP-DocLayoutV3 · OpenVINO IR
29
+
30
+ </h1>
31
+
32
+ <p align="center">Layout Analysis Module of PaddleOCR-VL-1.5 — converted to OpenVINO™ IR for local inference on Intel CPU / GPU / NPU</p>
33
+
34
+ [![OpenVINO](https://img.shields.io/badge/Runtime-OpenVINO-1A73E8)](https://github.com/openvinotoolkit/openvino)
35
+ [![repo](https://img.shields.io/github/stars/PaddlePaddle/PaddleOCR?color=ccf)](https://github.com/PaddlePaddle/PaddleOCR)
36
+ [![Base model](https://img.shields.io/badge/Base-PP--DocLayoutV3-orange)](https://modelscope.cn/models/PaddlePaddle/PP-DocLayoutV3)
37
+ [![License](https://img.shields.io/badge/license-Apache_2.0-green)](./LICENSE)
38
+
39
+ **🔥 [Official Website](https://www.paddleocr.com)** |
40
+ **📝 [Technical Report](https://arxiv.org/pdf/2601.21957)**
41
+
42
+ </div>
43
+
44
+ ---
45
+
46
+ ## Introduction · 简介
47
+
48
+ This repository hosts the **OpenVINO™ IR** build of **PP-DocLayoutV3**, the layout-analysis module of
49
+ **PaddleOCR-VL-1.5**. The original PaddlePaddle weights have been converted to OpenVINO Intermediate
50
+ Representation (`inference.xml` + `inference.bin`) so the model runs **fully locally** on **Intel CPU,
51
+ integrated/discrete GPU, and NPU** via the OpenVINO runtime — no cloud service and no PaddlePaddle
52
+ runtime required.
53
+
54
+ > 本仓库提供 **PP-DocLayoutV3 的 OpenVINO™ IR 版本**,它是 **PaddleOCR-VL-1.5** 的版面分析(layout)模块。
55
+ > 模型已从 PaddlePaddle 权重转换为 OpenVINO 中间表示(`inference.xml` + `inference.bin`),可在
56
+ > **Intel CPU / 集显 / 独显 / NPU** 上**完全本地**运行,无需联网、无需安装 PaddlePaddle。
57
+
58
+ **PP-DocLayoutV3 is specifically engineered to handle non-planar document images.** It directly predicts
59
+ multi-point bounding boxes for layout elements (rather than standard two-point boxes) and determines the
60
+ logical reading order for skewed and curved surfaces within a single forward pass, significantly reducing
61
+ cascading errors. It is an essential component of PaddleOCR-VL-1.5, providing the layout analysis that
62
+ drives high-precision parsing of real-world documents.
63
+
64
+ ### Model Architecture
65
+
66
+ <div align="center">
67
+ <img src="https://raw.githubusercontent.com/cuicheng01/PaddleX_doc_images/refs/heads/main/images/paddleocr_vl_1_5/PP-DocLayoutV3.png" width="800"/>
68
+ </div>
69
+
70
+ ---
71
+
72
+ ## What's in this repo · 文件说明
73
+
74
+ | File | Description |
75
+ | --- | --- |
76
+ | `inference.xml` | OpenVINO IR network topology |
77
+ | `inference.bin` | OpenVINO IR weights |
78
+ | `inference.yml` | Preprocessing config (resize 800×800, normalization) + 25-class label list + `draw_threshold` |
79
+ | `config.json` | Model config |
80
+ | `preprocessor_config.json` | Image-processor config |
81
+
82
+ - **Architecture:** DETR-style detector
83
+ - **Input:** `image` `[1, 3, 800, 800]` (BGR, resized to 800×800, no keep-ratio) + `scale_factor` `[1, 2]`
84
+ - **Default score threshold:** `0.5` (`draw_threshold` in `inference.yml`)
85
+ - **Layout classes (25):** abstract, algorithm, aside_text, chart, content, display_formula, doc_title,
86
+ figure_title, footer, footer_image, footnote, formula_number, header, header_image, image,
87
+ inline_formula, number, paragraph_title, reference, reference_content, seal, table, text,
88
+ vertical_text, vision_footnote
89
+
90
+ ---
91
+
92
+ ## Usage · 使用方法
93
+
94
+ ### 1. Recommended — as part of the PaddleOCR-VL OpenVINO pipeline
95
+
96
+ This model is the layout stage of an end-to-end document-parsing pipeline. Pair it with
97
+ [`FionaGu1019/PaddleOCR-VL-1.5-ov`](https://modelscope.cn/models/FionaGu1019/PaddleOCR-VL-1.5-ov)
98
+ (the recognition VLM) to get layout detection → reading order → text/table/formula recognition.
99
+
100
+ ```python
101
+ from modelscope import snapshot_download
102
+
103
+ # Download both stages of the pipeline
104
+ layout_dir = snapshot_download("FionaGu1019/PP-DocLayoutV3-ov")
105
+ vl_dir = snapshot_download("FionaGu1019/PaddleOCR-VL-1.5-ov")
106
+ ```
107
+
108
+ ### 2. Standalone OpenVINO inference
109
+
110
+ ```python
111
+ import cv2
112
+ import numpy as np
113
+ import openvino as ov
114
+
115
+ model_dir = "PP-DocLayoutV3-ov" # local path or snapshot_download(...) result
116
+ core = ov.Core()
117
+ compiled = core.compile_model(f"{model_dir}/inference.xml", "GPU") # "CPU" / "GPU" / "NPU"
118
+
119
+ # Preprocess: resize to 800x800 (no keep-ratio), CHW, float32
120
+ img = cv2.imread("document.jpg")
121
+ h, w = img.shape[:2]
122
+ resized = cv2.resize(img, (800, 800), interpolation=cv2.INTER_LINEAR)
123
+ blob = resized.astype(np.float32).transpose(2, 0, 1)[None] # [1,3,800,800]
124
+ scale_factor = np.array([[800 / h, 800 / w]], dtype=np.float32) # [1,2]
125
+
126
+ results = compiled({"image": blob, "scale_factor": scale_factor})
127
+ # Outputs are DETR detections (class id, score, box); filter by draw_threshold=0.5
128
+ # and map class ids via the label_list in inference.yml.
129
+ ```
130
+
131
+ > Tip: enable on-disk kernel caching with
132
+ > `core.set_property({"CACHE_DIR": ".ov_cache"})` to avoid re-compiling kernels on every run
133
+ > (a large speedup on GPU).
134
+
135
+ ---
136
+
137
+ ## Visualization
138
+
139
+ ### Light Variation
140
+ <div align="center">
141
+ <img src="https://raw.githubusercontent.com/cuicheng01/PaddleX_doc_images/refs/heads/main/images/paddleocr_vl_1_5/layout_lighting.jpg" width="800"/>
142
+ </div>
143
+
144
+ ### Skewing
145
+ <div align="center">
146
+ <img src="https://raw.githubusercontent.com/cuicheng01/PaddleX_doc_images/refs/heads/main/images/paddleocr_vl_1_5/layout_skew.jpg" width="800"/>
147
+ </div>
148
+
149
+ ### Screen-photo
150
+ <div align="center">
151
+ <img src="https://raw.githubusercontent.com/cuicheng01/PaddleX_doc_images/refs/heads/main/images/paddleocr_vl_1_5/layout_screen.jpg" width="800"/>
152
+ </div>
153
+
154
+ ### Curving
155
+ <div align="center">
156
+ <img src="https://raw.githubusercontent.com/cuicheng01/PaddleX_doc_images/refs/heads/main/images/paddleocr_vl_1_5/layout_curv.jpg" width="800"/>
157
+ </div>
158
+
159
+ ---
160
+
161
+ ## Notes · 说明
162
+
163
+ - This is a **format conversion** of the official PaddlePaddle model to OpenVINO IR; the network weights
164
+ and detection behaviour are intended to match the original PP-DocLayoutV3. For the original weights see
165
+ [PaddlePaddle/PP-DocLayoutV3](https://modelscope.cn/models/PaddlePaddle/PP-DocLayoutV3).
166
+ - For best results across Intel hardware, prefer **GPU** when available and fall back to **CPU**.
167
+
168
+ ## Citation
169
+
170
+ If you find PP-DocLayoutV3 helpful, feel free to give the original project a star and citation.
171
+
172
+ ```bibtex
173
+ @misc{cui2026paddleocrvl15multitask09bvlm,
174
+ title={PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing},
175
+ author={Cheng Cui and Ting Sun and Suyin Liang and Tingquan Gao and Zelun Zhang and Jiaxuan Liu and Xueqing Wang and Changda Zhou and Hongen Liu and Manhui Lin and Yue Zhang and Yubo Zhang and Yi Liu and Dianhai Yu and Yanjun Ma},
176
+ year={2026},
177
+ eprint={2601.21957},
178
+ archivePrefix={arXiv},
179
+ primaryClass={cs.CV},
180
+ url={https://arxiv.org/abs/2601.21957},
181
+ }
182
+ ```
config.json ADDED
@@ -0,0 +1,107 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_dropout": 0.0,
3
+ "activation_function": "silu",
4
+ "anchor_image_size": null,
5
+ "architectures": [
6
+ "PPDocLayoutV3ForObjectDetection"
7
+ ],
8
+ "attention_dropout": 0.0,
9
+ "backbone": null,
10
+ "backbone_config": {
11
+ "model_type": "hgnet_v2",
12
+ "arch": "L",
13
+ "return_idx": [0, 1, 2, 3],
14
+ "freeze_stem_only": true,
15
+ "freeze_at": 0,
16
+ "freeze_norm": true,
17
+ "lr_mult_list": [0, 0.05, 0.05, 0.05, 0.05],
18
+ "out_features": ["stage1", "stage2", "stage3", "stage4"]
19
+ },
20
+ "backbone_kwargs": null,
21
+ "batch_norm_eps": 1e-05,
22
+ "box_noise_scale": 1.0,
23
+ "d_model": 256,
24
+ "decoder_activation_function": "relu",
25
+ "decoder_attention_heads": 8,
26
+ "decoder_ffn_dim": 1024,
27
+ "decoder_in_channels": [
28
+ 256,
29
+ 256,
30
+ 256
31
+ ],
32
+ "decoder_layers": 6,
33
+ "decoder_n_points": 4,
34
+ "disable_custom_kernels": true,
35
+ "dropout": 0.0,
36
+ "encode_proj_layers": [
37
+ 2
38
+ ],
39
+ "encoder_activation_function": "gelu",
40
+ "encoder_attention_heads": 8,
41
+ "encoder_ffn_dim": 1024,
42
+ "encoder_hidden_dim": 256,
43
+ "encoder_in_channels": [
44
+ 512,
45
+ 1024,
46
+ 2048
47
+ ],
48
+ "encoder_layers": 1,
49
+ "eos_coefficient": 0.0001,
50
+ "eval_size": null,
51
+ "feature_strides": [
52
+ 8,
53
+ 16,
54
+ 32
55
+ ],
56
+ "hidden_expansion": 1.0,
57
+ "id2label": {
58
+ "0": "abstract",
59
+ "1": "algorithm",
60
+ "2": "aside_text",
61
+ "3": "chart",
62
+ "4": "content",
63
+ "5": "formula",
64
+ "6": "doc_title",
65
+ "7": "figure_title",
66
+ "8": "footer",
67
+ "9": "footer",
68
+ "10": "footnote",
69
+ "11": "formula_number",
70
+ "12": "header",
71
+ "13": "header",
72
+ "14": "image",
73
+ "15": "formula",
74
+ "16": "number",
75
+ "17": "paragraph_title",
76
+ "18": "reference",
77
+ "19": "reference_content",
78
+ "20": "seal",
79
+ "21": "table",
80
+ "22": "text",
81
+ "23": "text",
82
+ "24": "vision_footnote"
83
+ },
84
+ "initializer_range": 0.01,
85
+ "is_encoder_decoder": true,
86
+ "label2id": {},
87
+ "label_noise_ratio": 0.5,
88
+ "layer_norm_eps": 1e-05,
89
+ "learn_initial_query": false,
90
+ "matcher_alpha": 0.25,
91
+ "matcher_bbox_cost": 5.0,
92
+ "matcher_class_cost": 2.0,
93
+ "matcher_gamma": 2.0,
94
+ "matcher_giou_cost": 2.0,
95
+ "model_type": "pp_doclayout_v3",
96
+ "normalize_before": false,
97
+ "num_denoising": 100,
98
+ "num_feature_levels": 3,
99
+ "num_queries": 300,
100
+ "positional_encoding_temperature": 10000,
101
+ "torch_dtype": "float32",
102
+ "use_pretrained_backbone": false,
103
+ "use_timm_backbone": false,
104
+ "global_pointer_head_size": 64,
105
+ "mask_feature_channels": [64, 64],
106
+ "x4_feat_dim": 128
107
+ }
inference.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9310c72cb0be540eeb431e1966710a01ae1425b93dd23c0be697f2f0abe0a389
3
+ size 66958738
inference.xml ADDED
The diff for this file is too large to render. See raw diff
 
inference.yml ADDED
@@ -0,0 +1,100 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ mode: paddle
2
+ draw_threshold: 0.5
3
+ metric: COCO
4
+ use_dynamic_shape: false
5
+ Global:
6
+ model_name: PP-DocLayoutV3
7
+ arch: DETR
8
+ min_subgraph_size: 3
9
+ Preprocess:
10
+ - interp: 2
11
+ keep_ratio: false
12
+ target_size:
13
+ - 800
14
+ - 800
15
+ type: Resize
16
+ - mean:
17
+ - 0.0
18
+ - 0.0
19
+ - 0.0
20
+ norm_type: none
21
+ std:
22
+ - 1.0
23
+ - 1.0
24
+ - 1.0
25
+ type: NormalizeImage
26
+ - type: Permute
27
+ label_list:
28
+ - abstract
29
+ - algorithm
30
+ - aside_text
31
+ - chart
32
+ - content
33
+ - display_formula
34
+ - doc_title
35
+ - figure_title
36
+ - footer
37
+ - footer_image
38
+ - footnote
39
+ - formula_number
40
+ - header
41
+ - header_image
42
+ - image
43
+ - inline_formula
44
+ - number
45
+ - paragraph_title
46
+ - reference
47
+ - reference_content
48
+ - seal
49
+ - table
50
+ - text
51
+ - vertical_text
52
+ - vision_footnote
53
+ Hpi:
54
+ backend_configs:
55
+ paddle_infer:
56
+ trt_dynamic_shapes: &id001
57
+ image:
58
+ - - 1
59
+ - 3
60
+ - 800
61
+ - 800
62
+ - - 1
63
+ - 3
64
+ - 800
65
+ - 800
66
+ - - 8
67
+ - 3
68
+ - 800
69
+ - 800
70
+ scale_factor:
71
+ - - 1
72
+ - 2
73
+ - - 1
74
+ - 2
75
+ - - 8
76
+ - 2
77
+ trt_dynamic_shape_input_data:
78
+ scale_factor:
79
+ - - 2
80
+ - 2
81
+ - - 1
82
+ - 1
83
+ - - 0.67
84
+ - 0.67
85
+ - 0.67
86
+ - 0.67
87
+ - 0.67
88
+ - 0.67
89
+ - 0.67
90
+ - 0.67
91
+ - 0.67
92
+ - 0.67
93
+ - 0.67
94
+ - 0.67
95
+ - 0.67
96
+ - 0.67
97
+ - 0.67
98
+ - 0.67
99
+ tensorrt:
100
+ dynamic_shapes: *id001
preprocessor_config.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_valid_processor_keys": [
3
+ "images",
4
+ "do_resize",
5
+ "size",
6
+ "resample",
7
+ "do_rescale",
8
+ "rescale_factor",
9
+ "do_normalize",
10
+ "image_mean",
11
+ "image_std",
12
+ "return_tensors",
13
+ "data_format",
14
+ "input_data_format"
15
+ ],
16
+ "do_normalize": true,
17
+ "do_rescale": true,
18
+ "do_resize": true,
19
+ "image_mean": [
20
+ 0,
21
+ 0,
22
+ 0
23
+ ],
24
+ "image_processor_type": "PPDocLayoutV3ImageProcessor",
25
+ "image_std": [
26
+ 1,
27
+ 1,
28
+ 1
29
+ ],
30
+ "resample": 3,
31
+ "rescale_factor": 0.00392156862745098,
32
+ "size": {
33
+ "height": 800,
34
+ "width": 800
35
+ }
36
+ }