| --- |
| title: Agentic Space Factory |
| sdk: docker |
| hf_oauth: true |
| hf_oauth_expiration_minutes: 480 |
| hf_oauth_scopes: |
| - read-repos |
| - write-repos |
| - manage-repos |
| - gated-repos |
| - inference-api |
| - jobs |
| - read-billing |
| --- |
| |
| # Agentic Space Factory |
|
|
| Agentic Space Factory turns a Hugging Face model card into a **private, testable Gradio Space** using a Hugging Face Job, Pi, a coding model through Hugging Face Inference Providers, and a private Storage Bucket for run state, traces, reports, generated files, and validation artifacts. |
|
|
| It is a Hugging Face-native productization of the agentic ZeroGPU Spaces workflow: |
|
|
| ```text |
| model card → agentic build job → private Space → hardware if available → live validation → traces/report/artifacts |
| ``` |
|
|
| ## What the app does |
|
|
| - Creates private generated Spaces under the signed-in user's namespace. |
| - Launches long-running build and validation work as Hugging Face Jobs. |
| - Uses Pi with a coding model served through Hugging Face Inference Providers. |
| - Injects the configured GitHub gist as Pi's operational instruction source. |
| - Stores each run in a per-user private Storage Bucket. |
| - Attempts ZeroGPU first when requested, then optional fixed-GPU fallback. |
| - Supports manual hardware selection when automatic assignment is unavailable. Automatic fallback avoids high/restricted tiers such as A100/H200; use Space Settings manually if an account or organization has access. |
| - Distinguishes bootability from real inference. |
| - Measures live inference latency when a generation smoke test succeeds. |
| - Uses measured latency to recommend `@spaces.GPU(duration=...)` when the generated Space uses ZeroGPU. |
|
|
| ## Current UI |
|
|
| The product root `/` is a custom Docker/FastAPI dashboard. `/custom` is kept as an alias. |
|
|
| The UI is organized into three main areas: |
|
|
| ## v122 stability note |
|
|
| The v122 line intentionally rolls back the experimental v121 client-side caching/polling optimizations after real UI testing exposed flicker and broken build/run behavior. The app keeps the v120 Agent recovery UI and the v118/v119 core blockage protocol, but returns to the proven polling and run-selection model for production stability. |
|
|
| ### Left column |
|
|
| - **Run storage** — check or create the signed-in user's private Bucket. |
| - **New Build** — starts from `Tongyi-MAI/Z-Image-Turbo` by default, fast pre-scans the model card, then launches a private Space build only after an acceptable scan verdict. |
| - **Billing & compute** — show payment readiness, link to HF Billing, and centralize the compute-cost note for Jobs, Inference Providers, and safe hardware fallbacks. |
|
|
| ### Center workspace |
|
|
| Two tabs share the main detail area: |
|
|
| - **Active Run** — build progress, run actions, Space/Settings/Job/Artifacts links, manual hardware guidance, inference latency, a vector-icon latest-events feed, and curated trace documents. |
| - **Space Test** — validation run details, actionnable Gradio endpoint picker, endpoint params, validation status, latency, point-based progress, and validation actions. |
|
|
| Selecting a build run opens `Active Run`. Selecting a validation run opens `Space Test`. The Runs Explorer is intentionally compact: each row prioritizes the model or Space, status, relative freshness, and a secondary context line while keeping the raw run ID available in the row. |
|
|
| ### Right column |
|
|
| - **Runs Explorer** — navigation-only list of build and validation runs with filters, search, status badges, and inline delete. |
| - **Run Stats** — aggregate metrics across indexed runs: totals, success/pass rates, validation latency median/p95, and issue counts. |
|
|
|
|
|
|
| ### Agent recovery visibility |
|
|
| When a build or runtime blockage appears, the Active Run panel shows an **Agent recovery** card. It summarizes Pi's diagnosis decision, the selected recovery action (`wait_for_logs`, `factory_rebuild_same_code`, `patch_code`, manual hardware, or blocker), key evidence, and the next step. This makes the blockage protocol visible without turning the UI into a raw log viewer. |
|
|
| The Run traces dock also links to the unified `agent_trace.jsonl`, the diagnosis `REPAIR_DECISION.json`, repair artifacts, and blockage details when they exist. |
|
|
| ## Model pre-scan |
|
|
| New builds are gated by a fast model-card pre-scan. The scan is metadata-only and does not download model weights. It checks: |
|
|
| - model-card guidance and runnable usage examples; |
| - Diffusers signals such as `library_name`, `model_index.json`, pipeline class, and `from_pretrained(...)` examples; |
| - safetensors versus unsafe or legacy weight formats; |
| - gated/private access and unsupported task signals; |
| - custom-code indicators such as `auto_map`, custom pipeline files, or `trust_remote_code` hints. |
|
|
| Verdicts are intentionally product-facing: `safe` and `caution` allow launch, `risky` requires explicit acknowledgement, and `unsupported` blocks launch. `Tongyi-MAI/Z-Image-Turbo` is treated as the known-good default example because it has a clear Diffusers model card and works in the app's default test flow. |
|
|
| ## Run storage |
|
|
| Each signed-in user writes to their own private Storage Bucket. The default bucket is: |
|
|
| ```text |
| https://huggingface.co/buckets/<your-username>/space-factory-runs |
| ``` |
|
|
| For local dev or deployment, configure only the bucket name; the namespace is derived from the signed-in OAuth user: |
|
|
| ```bash |
| SPACE_FACTORY_BUCKET_NAME=space-factory-runs |
| SPACE_FACTORY_BUCKET_MOUNT=/output |
| ``` |
|
|
| Typical run layout: |
|
|
| ```text |
| runs/<run_id>/launch.json |
| runs/<run_id>/state.json |
| runs/<run_id>/summary.json |
| runs/<run_id>/events.jsonl |
| runs/<run_id>/report.md |
| runs/<run_id>/generated/ |
| runs/<run_id>/tests/ |
| runs/<run_id>/artifacts/ |
| runs/<run_id>/traces/raw/ |
| runs/<run_id>/traces/raw/agent_trace.jsonl |
| runs/<run_id>/traces/redacted/ |
| runs/<run_id>/traces/redacted/agent_trace.jsonl |
| runs/<run_id>/logs/ |
| ``` |
|
|
| ## Main workflow |
|
|
| ### 1. Build from model card |
|
|
| Paste a Hugging Face model ID or model-card URL. The app launches an HF Job that: |
|
|
| - fetches model metadata and repo information; |
| - writes a strict `GOAL.md` for Pi; |
| - asks Pi to follow the configured GitHub gist before coding; |
| - generates a Gradio app; |
| - creates and uploads a private Space; |
| - attempts ZeroGPU/fixed-GPU assignment when requested; |
| - runs health/API checks and a live generation smoke test when possible; |
| - writes events, report, traces, and structured artifacts to the user's Bucket. |
|
|
| ### 2. Validate Space |
|
|
| Validation is a separate HF Job. Use it after a build succeeds or after manually selecting hardware in Space Settings. |
|
|
| The validation job: |
|
|
| - waits for the target Space; |
| - discovers Gradio endpoints when possible; |
| - calls the selected endpoint; |
| - checks the returned output type; |
| - stores returned artifacts; |
| - measures latency; |
| - recommends a conservative ZeroGPU duration when relevant; |
| - writes validation state, events, report data, and smoke-test artifacts. |
|
|
| ## Result statuses |
|
|
| - `full_inference_success`: a live generation smoke test returned the expected output type. |
| - `manual_hardware_required`: the Space was generated, but hardware must be selected manually before validation. |
| - `full_inference_candidate_health_passed`: the Space boots and contains inference signals, but generation was not fully smoke-tested. |
| - `health_only`: the Space boots, but no real inference path was validated. |
| - `technical_blocker`: concrete blockers such as multi-GPU requirements, missing licenses, custom CUDA, or unclear usage. |
| - `failed`: the build, runtime, or validation job failed. |
|
|
| ## Guarantees |
|
|
| - Generated Spaces are private by default. |
| - Nothing is published automatically. |
| - Runs are persisted in the configured per-user Storage Bucket. |
| - The app distinguishes health checks from full inference validation. |
| - Hardware assignment is best-effort and can fall back to manual user action. |
| - Raw and redacted Pi traces are kept separate. |
|
|
| ### Unified agent trace |
|
|
| Pi can be invoked more than once during a run: the initial build, a blockage diagnosis pass, and an optional structured repair patch. The Factory now writes a canonical run-level agent journal to `traces/raw/agent_trace.jsonl` and `traces/redacted/agent_trace.jsonl` so those phases read as one auditable agentic loop instead of separate conversations. Specialized stdout logs remain available under `logs/`, but the Active Run `Pi RAW` / `Pi redacted` buttons prefer the unified trace when it exists. |
|
|
|
|
| ## Non-goals and limits |
|
|
| The app does not guarantee that every model card becomes a working Space. Some models require multi-GPU execution, Docker, unavailable CUDA extensions, external services, accepted licenses, very large hardware, or manual fixes. |
|
|
| The app does not mirror live spend/quota totals from Hugging Face Billing. Users should consult the HF Billing dashboard for exact usage. |
|
|
| ## Local development |
|
|
| Install dependencies and run tests: |
|
|
| ```bash |
| pip install -r requirements.txt |
| pytest |
| ``` |
|
|
| The Docker Space starts the app with one uvicorn process from the Dockerfile. Do not add an extra `uvicorn.run(...)` launcher inside `app.py`. |
|
|
| ## Documentation |
|
|
| See: |
|
|
| - `docs/ARCHITECTURE.md` |
| - `docs/PUBLIC_PREVIEW.md` |
| - `docs/SECURITY.md` |
| - `docs/NEXT_STEPS.md` |
|
|
|
|
| Default example model: `Tongyi-MAI/Z-Image-Turbo`. It is treated as a known-good Diffusers text-to-image model for the fast pre-scan calibration and public-preview smoke tests. |
|
|
| ### Pi diagnosis and structured recovery |
|
|
| When the generated Space fails live validation, the factory now enters a diagnosis-first blockage protocol instead of asking Pi to patch immediately. The worker collects HF build/runtime logs, Space status and validation errors, writes `repair/INCIDENT_BRIEF.md`, then asks Pi/Qwen Coder to choose one bounded action in `repair/REPAIR_DECISION.json`: wait for delayed logs, inspect again, force a same-code factory rebuild, patch code, request manual hardware, or declare a technical blocker. |
|
|
| A code repair is allowed only when Pi selects `patch_code` and the Factory accepts the decision. Empty/no-reason HF failures are routed through wait/rebuild/blocker paths rather than speculative code edits. Patch repair still writes `REPAIR_BRIEF.md`, `REPAIR_PLAN.md`, and `REPAIR_SUMMARY.md`, uploads once, then revalidates the live Space with the normal API checks. |
|
|
| Repair mode is intentionally strict: it must preserve the original model and real-inference contract, avoid fake/static outputs, keep a cheap health endpoint, and report blockers instead of hiding runtime errors. |
|
|
| ## v123 — Critical Job launch + failure timeline fix |
|
|
| - Fixed a production-critical HF Job startup failure where the large embedded worker script could make Python fail before startup with `argument list too long`. |
| - The worker script is now persisted as `runs/<run_id>/_worker.py` in the mounted run bucket before launch, and the Job receives only `WORKER_SCRIPT_PATH`. |
| - Kept a small `WORKER_SCRIPT_B64` fallback only for compatibility/manual launch paths; normal app launches no longer put the worker source in env/argv. |
| - Added a failure-log fallback so pre-worker Job failures such as `argument list too long` still create an explicit failed timeline event and red failure point. |
|
|
| ## v124 — Timeline and Run Explorer stabilization |
|
|
| - Fixed the compact Overall Progress timeline so dynamically visible repair/failure groups cannot wrap dots onto a second row. |
| - Removed smooth timeline auto-scroll that could cause visible flicker during polling. |
| - Simplified Run Explorer type/status badges to neutral type labels and short status labels such as Success, Failed, Running and Manual. |
| - Added regression checks for timeline single-row behavior and compact Run Explorer labels. |
|
|
| ## v127 — Billing usefulness + UI event/timeline audit |
|
|
| - Reworked the Billing & compute card into an actionable compute checklist: payment readiness, Pi Inference Providers, build Job cost, generated Space hardware and high-tier GPU guardrails. |
| - Kept exact spend/quotas delegated to the Hugging Face Billing dashboard instead of inventing usage totals. |
| - Audited JS event bindings for one-shot/delegated listeners across navigation, run rows and timeline scroll. |
| - Reduced timeline auto-scroll churn: identical timeline renders no longer force keep-visible scrolling after user interaction. |
| - Confirmed run selection keeps the immediate cached/optimistic render path before network hydration. |
|
|
|
|