title: Repair Guy
emoji: π§
colorFrom: purple
colorTo: red
sdk: gradio
sdk_version: 6.17.3
python_version: '3.12'
app_file: app.py
pinned: false
preload_from_hub:
- nvidia/nemotron-colembed-vl-4b-v2
- openbmb/MiniCPM4.1-8B
- openbmb/MiniCPM-V-4_5
- nvidia/llama-nemotron-embed-vl-1b-v2
license: mit
Repair Guy β a hands-busy mechanic's manual assistant
A page viewer over repair manuals driven by short requests β "go to brake bleeding", "next page", "circle the brake hoses". It finds the right page and points at it; it never writes answers (the manual does the talking). It keeps a compact memory of the turn-so-far, used only to resolve references ("circle the other one", "go back to that bolt"). Everything runs inside the Space (no external endpoints).
Obvious nav (client, instant) β next/previous page, page 412, back, next/previous section: pure frontend state over the rendered page images.
Everything else β
/find, one ZeroGPU call (app/pipelines/agent_ask.py): MiniCPM5-1B β the small text "brain" β sees the conversation so far, the manual's table of contents, and the whole text of the page being viewed (parsed page β text, figures/tables as their descriptions), and calls tools in a loop until a page is shown:- search β semantic-search the manual; the best page is shown and its text fed back so the agent can then circle on it;
- go_to_page β jump straight to a known physical page number (e.g. one read off an index/contents line);
- circle β circle something on the current page: MiniCPM-V 4.5 grounds the target visually (β bbox β SVG circle);
- done β nothing to do, or it isn't in the manual.
Streamed as events; the UI shows tool chips and the resulting page/circle, never model prose. The table of contents is the manual's clean PDF-bookmark chapters plus a per-request fuzzy shortlist of fine parse headings (
app/core/sections.py).
Retrieval is fused β the search tool needs both indexes of a manual, so a manual must be indexed both ways:
- Visual β every page is embedded as an image with Nemotron ColEmbed v2 (multi-vector, late interaction). The search tool embeds the query and scores it against every page with MaxSim (batched torch matmuls, pages streamed from disk via numpy memmap) to shortlist candidate pages β the top page is shown. (Late-interaction visual ranking beat a 1B text rerank of the shortlist in the eval, 0.84 vs 0.68 hit@1, so the search tool takes ColEmbed's top page.)
- Parsed β pages are parsed with
Nemotron Parse v1.2,
figures and tables are described by MiniCPM-V, and heading-based section
chunks are embedded with
Llama Nemotron Embed VL 1B v2
(used by the offline answer eval). On the Space the parsed pages supply the
whole-page text the agent reasons over (
app/core/page_context.py).
Indexing (offline only)
All ingestion runs offline on Modal GPUs β scripts/index_modal.py in the
GitHub repo (--method visual|parsed) β and is pushed to the
library dataset,
laid out as <method>/<doc_id>/. The Space syncs it to /data/preindexed
at startup (and via the π button). The parsed method runs as two Modal
functions because Nemotron Parse requires transformers 5.x while the rest of
the stack is on 4.x.
Local UI development (mock mode)
To iterate on the UI without GPUs, model downloads or the HF library sync, run
with MOCK_MODELS=1. Drop any PDF into app/data/mock_pdfs/ (or set
MOCK_PDF_DIR) and the whole find-and-point UX works over real rendered pages
of that PDF β navigation, and every event the agent emits (go-to-section,
circle-on-this-page, searchβshowβcircle, plus the can't-find reply). A keyword
heuristic stands in for the LLM agent; the section structure and bboxes are faked.
cd app
uv pip install gradio pymupdf pillow numpy huggingface_hub # light deps only
MOCK_MODELS=1 python app.py
The π Sync library button just re-scans the folder, so PDFs added while the app
is running show up without a restart. See pipelines/mock_ask.py.
Space setup
- Persistent storage must be enabled (the library sync lives under
/data). The visual index is the big one: roughly 5β12 MB per page of float16 token embeddings (a 1000-page manual is ~6β12 GB); the parsed index is a few MB per manual. - Optional env vars:
LIBRARY_DATASET_ID,COLEMBED_MODEL_ID,MINICPM_MODEL_ID(the VLM "eyes"),MINICPM_AGENT_MODEL_ID(the 1B "brain"),NEMOTRON_EMBED_MODEL_ID,COLEMBED_ATTN(defaults tosdpa; setflash_attention_2if flash-attn is installed).
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Tracing (Langfuse)
Each find turn is traced to Langfuse when keys are
present, otherwise tracing is a complete no-op and the Space runs unchanged.
Set these as Space secrets (or shell env / a .env for local runs):
LANGFUSE_PUBLIC_KEY=pk-lf-...
LANGFUSE_SECRET_KEY=sk-lf-...
LANGFUSE_HOST=https://cloud.langfuse.com # or https://us.cloud.langfuse.com, or self-hosted
(LANGFUSE_BASE_URL is accepted as an alias for LANGFUSE_HOST.) Keys come from
the Langfuse project's Settings β API Keys.
One turn is one trace β an agent observation named agent-find (input = the
mechanic's request, output = the terminal action), with a child per step:
agent-decide(generation) β the 1B brain picking a tool, with the resident model id and input/output token counts.search(retriever) β the ColEmbed visual lookup and its page hits.ground-circle(generation) β MiniCPM-V placing the circle box.
The frontend sends a per-page-load session_id, so a visit's turns group into
one session in the Langfuse UI. The wiring lives in core/tracing.py (a
no-op-safe wrapper that builds and flushes the trace inside the ZeroGPU worker,
attaching children to the root explicitly rather than via context β the turn is
a streamed generator). The Langfuse working skill is vendored under
.claude/skills/langfuse/.