--- license: mit tags: - onnx - image-to-3d - 3d-reconstruction - triplane - qtmesheditor library_name: onnx pipeline_tag: image-to-3d --- # TripoSR — ONNX export **ONNX re-export of [stabilityai/TripoSR](https://huggingface.co/stabilityai/TripoSR)** (Tripo AI + Stability AI — MIT code + MIT weights): single-image → 3D reconstruction via a triplane radiance field. All credit for the original weights goes to Tripo AI and Stability AI. Exported for **[QtMeshEditor](https://github.com/fernandotonon/QtMeshEditor)** (epic #764), powering `qtmesh generate3d`, the Inspector's **AI: Image → 3D** section, and the `generate_mesh_from_image` MCP tool — local inference via ONNX Runtime, with surface extraction done by a native marching-cubes implementation (upstream's `torchmcubes` is torch/GPU-only). > The files QtMeshEditor downloads at runtime live in the shared > [`fernandotonon/QtMeshEditor-models`](https://huggingface.co/fernandotonon/QtMeshEditor-models) > repo under `triposr/`. This repo is the standalone model card + mirror. ## Files | file | role | |---|---| | `triposr_encoder.onnx` | fp32 encoder (best quality, ~1.7 GB): `image [1,3,512,512]` → `scene_codes [1,3,40,64,64]` (triplane) | | `triposr_encoder_int8.onnx` | int8 encoder tier (~430 MB, slight quality loss) | | `triposr_decoder.onnx` | per-point field decoder: `scene_codes` + `points [1,P,3]` → `density [1,P,1]`, `color [1,P,3]` | An fp16 tier is deliberately absent: TripoSR's attention has a hardcoded Cast-to-float32 that the ONNX fp16 converters cannot rewrite. ## Inference contract - Input image: 512×512 RGB in `[0,1]`, subject isolated (upstream composites over gray-128 and re-pads the foreground to a 0.85 ratio; QtMeshEditor uses U²-Net for background removal first). - Query the decoder in **chunks** — generating the full `res³` grid up front OOMs at high resolutions. - Surface = marching cubes on `density − threshold` at iso 0, threshold 25.0, within radius 0.87. The field is inside-positive; emit flipped winding (`v0,v2,v1`) to keep faces outward. - Output is not +Y-up: bake −90°X then +90°Y to stand the model upright facing forward. - The decoder's colour output can also be used as an image-conditioned colour field for texture baking (per-texel queries after a UV unwrap). ## Reproducing `scripts/export-triposr-onnx.py` in the QtMeshEditor repo (one-time, offline; `transformers==4.35.0`, `torchmcubes` stubbed, frozen ViT positional encoding; emits the int8 variant unless `--no-quant`). ## License MIT (same as the upstream code and weights). Credit: Tripo AI + Stability AI.