File size: 13,886 Bytes
2f382c4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
# Dependencies, attribution, and weight downloads

This guide is written for a first-time user. AccessPath is inspired by and
adapted from prior work on amodal completion, segmentation, geometry, and
visual 3D reconstruction. It explains both **what is actually used at runtime**
and **what is only cited or retained as an optional comparison**. Do not
download a model merely because its name occurs in a related-work note.

## 1. What AccessPath implements

AccessPath implements an accessibility-focused workflow inspired by these
research directions, while adding project-level adaptation around external
models:

1. prompt selection and target/obstacle mask proposal orchestration;
2. constrained visible/hidden/amodal mask construction and validation;
3. mask-restricted 2D inpainting candidate generation and selection;
4. depth/geometry diagnostics, quality gates, rendering, and review bundles;
5. an adapter that sends an RGB image plus a three-value mask to an external
   visual-3D backend.

It does **not** claim authorship of SAM 3, Stable Diffusion, Depth Anything V2,
VGGT, Amodal3R, TRELLIS, pix2gestalt, Open-World AMODAL, or Amodal Completion
in the Wild. Where one of these is a runtime backend, its code, license, and
weights remain separate; where it is listed as related work, AccessPath does
not import or execute it.

## 2. Exact dependency status

| Method/model | Status in this repository | Where in AccessPath | What a user must do |
| --- | --- | --- | --- |
| [SAM 3](https://github.com/facebookresearch/sam3) | Used by the automatic mask-proposal stage. | `tools/accessibility_mask_proposals.py` | Install it in its own environment and request access to [`facebook/sam3`](https://huggingface.co/facebook/sam3). |
| [Stable Diffusion inpainting](https://huggingface.co/sd-legacy/stable-diffusion-inpainting) | Used by the default GPU 2D completion stage. | `tools/accessibility_2d_completion.py` | Download an inpainting checkpoint under its own terms. |
| [Depth Anything V2](https://github.com/DepthAnything/Depth-Anything-V2) | Used by the default full-pipeline depth/geometry diagnostic. | `accessibilityamodal/depth.py` | Install the official repository and download the selected checkpoint. |
| [Amodal3R](https://sm0kywu.github.io/Amodal3R/) | Used by the default visual 3D stage through a thin adapter. | `tools/accessibility_3d_completion.py` | Install the official runtime and obtain its model assets. |
| [TRELLIS image large](https://huggingface.co/microsoft/TRELLIS-image-large) | Upstream dependency reported by the Amodal3R model card. | External Amodal3R runtime | Obtain it only following the upstream Amodal3R instructions. |
| [VGGT](https://github.com/facebookresearch/vggt) | **Optional** depth/point-map engine; not selected by the default Slurm launcher. | `accessibilityamodal/depth.py` | Install/download only when running `--depth-engine vggt`. |
| [pix2gestalt](https://github.com/cvlab-columbia/pix2gestalt) | **Not executed.** Kept only as a related-work/comparison URL in a baseline manifest. | `tools/accessibility_fast_2d_baseline.py` | No installation or checkpoint is required for AccessPath. |
| [Open-World AMODAL](https://github.com/saraao/amodal) | **Not executed.** Kept only as a related-work/comparison URL. | `tools/accessibility_fast_2d_baseline.py` | No installation or checkpoint is required for AccessPath. |
| [Amodal Completion in the Wild](https://github.com/Championchess/Amodal-Completion-in-the-Wild) | **Not executed.** Mentioned as a compatible external research backend only. | `accessibilityamodal/pipeline.py` | No installation or checkpoint is required for AccessPath. |

Therefore, no pix2gestalt, Open-World AMODAL, or Amodal-Wild code/weights are
needed for the commands in this repository. Their links are present so readers
can distinguish the accessibility-adapted AccessPath workflow from related
amodal-completion work.

## 3. Before starting

You need Linux, Git, Python, a CUDA-capable GPU, and a Slurm installation for
the all-in-one `pipeline` command. The direct `2d` and `3d` commands can be
run outside Slurm, but still need a compatible CUDA environment.

Clone the anonymous release:

```bash
git clone https://huggingface.co/anonymous-accesspath/AccessPath
cd AccessPath
```

Create a general utility environment. Install the PyTorch build that matches
your CUDA driver using the [official PyTorch selector](https://pytorch.org/get-started/locally/),
then install the Python packages used by the project wrappers:

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
# Install a CUDA-compatible PyTorch build here, following pytorch.org.
python -m pip install -r requirements/runtime.txt
python -m pip install diffusers transformers accelerate safetensors huggingface_hub
```

Do not place downloaded checkpoints inside the Git clone. Keep them in a
separate local `models/` directory and pass their paths through environment
variables. This avoids accidentally committing large or license-restricted
files.

## 4. Download each runtime asset

### 4.1 SAM 3: mask proposals

SAM 3 is used only to propose a visible target mask and an obstacle mask. The
proposal is not ground truth; inspect it before using it for an experiment.

1. Read and accept the access terms on [`facebook/sam3`](https://huggingface.co/facebook/sam3).
2. Follow the official [SAM 3 installation guide](https://github.com/facebookresearch/sam3).
   SAM 3 may require a newer Python/PyTorch/CUDA combination than the remaining
   pipeline, so a dedicated environment is recommended.
3. Authenticate and download the official files:

```bash
hf auth login
hf download facebook/sam3 sam3.pt config.json --local-dir /path/to/models/sam3
```

4. Record the environment and paths for AccessPath:

```bash
export SAM3_PYTHON=/path/to/sam3-environment/bin/python
export SAM3_REPO=/path/to/sam3-source
export SAM3_CHECKPOINT=/path/to/models/sam3/sam3.pt
```

### 4.2 Stable Diffusion: 2D inpainting

The 2D stage calls Diffusers' inpainting pipeline. Obtain an inpainting model
from [`sd-legacy/stable-diffusion-inpainting`](https://huggingface.co/sd-legacy/stable-diffusion-inpainting)
or use another compatible checkpoint only after checking its model card and
license. A typical download is:

```bash
hf download sd-legacy/stable-diffusion-inpainting \
  --local-dir /path/to/models/stable-diffusion-inpainting
export SD_MODEL=/path/to/models/stable-diffusion-inpainting
```

The 2D implementation changes only the reviewed completion envelope and saves
its selected result plus metadata under the requested output directory.

### 4.3 Depth Anything V2: depth and geometry diagnostics

The default full launcher expects the large indoor metric checkpoint:
[`Depth-Anything-V2-Metric-Hypersim-Large`](https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large).
Install source code from the official [Depth Anything V2 repository](https://github.com/DepthAnything/Depth-Anything-V2), then download the exact checkpoint:

```bash
git clone https://github.com/DepthAnything/Depth-Anything-V2 /path/to/Depth-Anything-V2
hf download depth-anything/Depth-Anything-V2-Metric-Hypersim-Large \
  depth_anything_v2_metric_hypersim_vitl.pth \
  --local-dir /path/to/models/depth-anything-v2-metric-hypersim-large

export DEPTH_REPO=/path/to/Depth-Anything-V2
export DEPTH_CHECKPOINT=/path/to/models/depth-anything-v2-metric-hypersim-large/depth_anything_v2_metric_hypersim_vitl.pth
```

Depth output is a diagnostic estimate. It is not automatically calibrated
metric ground truth for the input camera and must not be treated as a
navigation-safety measurement.

### 4.4 Amodal3R and TRELLIS: visual 3D completion

AccessPath does not redistribute Amodal3R source code or weights. The adapter
requires an external Python environment where this import succeeds:

```bash
python -c "from amodal3d.pipelines import Amodal3RImageTo3DPipeline; print('Amodal3R runtime ready')"
```

Use the official [Amodal3R project page](https://sm0kywu.github.io/Amodal3R/),
[model card](https://huggingface.co/Sm0kyWu/Amodal3R), and
[paper](https://arxiv.org/abs/2503.13439) for installation and license terms.
The model card states that its implementation is built on TRELLIS and obtains
pretrained assets from [`microsoft/TRELLIS-image-large`](https://huggingface.co/microsoft/TRELLIS-image-large).
For a local snapshot of the published model card repository:

```bash
hf download Sm0kyWu/Amodal3R --local-dir /path/to/models/Amodal3R
```

Follow the upstream project instructions for any additional TRELLIS assets and
CUDA rasterizer dependencies. Then configure AccessPath:

```bash
export AMODAL3D_PYTHON=/path/to/amodal3r-environment/bin/python
export AMODAL3D_MODEL=/path/to/models/Amodal3R
export AMODAL3D_TORCH_HOME=/path/to/model-cache/torch
```

The 3D result is a learned visual reconstruction. It is not guaranteed to be
metric, watertight, scale-calibrated, or safe for path-planning decisions.

### 4.5 Optional VGGT geometry route

VGGT is implemented as an optional `--depth-engine vggt` route. It is useful
for comparing a geometry/point-map estimate, but it is **not** the default
AccessPath visual-3D output and it is not required for the default full run.

```bash
git clone https://github.com/facebookresearch/vggt /path/to/vggt
python -m pip install -r /path/to/vggt/requirements.txt
hf download facebook/VGGT-1B --local-dir /path/to/models/VGGT-1B
```

Then invoke the diagnostic engine with `--depth-engine vggt`, pass
`--vggt-repo /path/to/vggt`, and use `--vggt-model /path/to/models/VGGT-1B`.
See the [official VGGT repository](https://github.com/facebookresearch/vggt)
and [model page](https://huggingface.co/facebook/VGGT-1B) for current access,
license, and checkpoint conditions.

### 4.6 Project amodal-mask adapter checkpoint

`AMODAL_CHECKPOINT` is an AccessPath-trained mask-adapter checkpoint, not an
upstream public dependency. It is deliberately withheld during anonymous
review. Its release requires confirmation of training-data permissions,
base-model terms, privacy review, and the paper's release policy.

Without this checkpoint, the **all-in-one** `pipeline` command stops at its
preflight check. Readers can still run the direct 2D and visual-3D commands
with their own reviewed masks. This limitation is intentional and explicit;
the repository does not substitute a different checkpoint silently.

## 5. Run the three supported routes

### Route A — direct 2D completion

Prepare an RGB image and three aligned binary masks: visible target, complete
amodal target, and obstacle. All must have the same width and height as the
RGB image.

```bash
python accesspath.py 2d -- \
  --image inputs/scene.jpg \
  --target-visible-mask inputs/target_visible.png \
  --target-amodal-mask inputs/target_amodal.png \
  --obstacle-mask inputs/obstacle.png \
  --category stairs \
  --model "$SD_MODEL" \
  --output-dir outputs/stairs_2d \
  --device cuda
```

Inspect the selected RGB and the JSON manifest under `outputs/stairs_2d`.

### Route B — direct visual 3D completion

Create one three-value mask at the RGB resolution: `255` is background, `188`
is visible target, and `0` is hidden target. Then run:

```bash
AMODAL3D_TORCH_HOME="$AMODAL3D_TORCH_HOME" \
"$AMODAL3D_PYTHON" tools/accessibility_3d_completion.py \
  --image inputs/scene.jpg \
  --mask inputs/amodal_3value.png \
  --model "$AMODAL3D_MODEL" \
  --output-dir outputs/stairs_3d
```

Inspect the generated manifest and image/mesh assets. Do not interpret a
plausible-looking 3D asset as a physical measurement.

### Route C — full Slurm pipeline

After configuring every required variable above, submit a finite one-image job:

```bash
python accesspath.py pipeline -- \
  --image inputs/scene.jpg \
  --category stairs \
  --output-dir outputs/stairs_full
```

The command submits one GPU job, waits by default, saves its results, and exits
when the job terminates. Add `--no-wait` to return after submission. The output
review order is documented in [REPRODUCIBLE_PIPELINE.md](REPRODUCIBLE_PIPELINE.md).

## 6. Citation and related-work links

Use the official paper/model pages below when citing an external dependency.

```bibtex
@article{wu2025amodal3r,
  title={Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images},
  author={Wu, Tianhao and Zheng, Chuanxia and Guan, Frank and Vedaldi, Andrea and Cham, Tat-Jen},
  journal={arXiv preprint arXiv:2503.13439},
  year={2025}
}

@inproceedings{wang2025vggt,
  title={VGGT: Visual Geometry Grounded Transformer},
  author={Wang, Jianyuan and Chen, Minghao and Karaev, Nikita and Vedaldi, Andrea and Rupprecht, Christian and Novotny, David},
  booktitle={CVPR},
  year={2025}
}

@inproceedings{ozguroglu2024pix2gestalt,
  title={pix2gestalt: Amodal Segmentation by Synthesizing Wholes},
  author={Ozguroglu, Ege and others},
  booktitle={CVPR},
  year={2024}
}

@inproceedings{ao2025open,
  title={Open-World Amodal Appearance Completion},
  author={Ao, Jiayang and Jiang, Yanbei and Ke, Qiuhong and Ehinger, Krista A.},
  booktitle={CVPR},
  year={2025}
}

@inproceedings{zhan2024amodal,
  title={Amodal Ground Truth and Completion in the Wild},
  author={Zhan, Guanqi and Zheng, Chuanxia and Xie, Weidi and Zisserman, Andrew},
  booktitle={CVPR},
  year={2024}
}

@article{yang2024depthanythingv2,
  title={Depth Anything V2},
  author={Yang, Lihe and others},
  journal={arXiv:2406.09414},
  year={2024}
}
```

For convenience, the canonical links are: [Amodal3R paper](https://arxiv.org/abs/2503.13439),
[VGGT paper](https://arxiv.org/abs/2503.11651),
[pix2gestalt paper](https://arxiv.org/abs/2401.14398),
[Open-World AMODAL paper](https://arxiv.org/abs/2411.13019),
[Amodal Completion in the Wild paper](https://arxiv.org/abs/2312.17247), and
[Depth Anything V2 paper](https://arxiv.org/abs/2406.09414).