fffiloni's picture
Upload 4 files
3fc81c8 verified
|
Raw
History Blame
4.47 kB

Public Preview

Agentic Space Factory turns Hugging Face model cards into private, testable Gradio Spaces using Hugging Face Jobs, Pi, Inference Providers, and per-user Storage Buckets.

What it does

  • Creates private Spaces from model cards.
  • Runs the build inside an ephemeral Hugging Face Job.
  • Uses Pi with an Inference Providers coding model.
  • Injects the shared GitHub gist as Pi's operational instruction source.
  • Stores state, events, raw/redacted traces, reports, generated files, validation results, and output artifacts in a private user Bucket.
  • Attempts ZeroGPU first when requested, then optional fixed-GPU fallback.
  • Supports manual hardware selection when automatic hardware assignment is unavailable.
  • Runs a separate Space validation job with endpoint discovery, output-type checks, latency measurement, and generated artifact capture.

Current UI

The dashboard is split into three working areas:

  1. Left column — Run storage setup, model pre-scan, New Build form, Billing & compute readiness.
  2. Center tabsActive Run for build details and Space Test for validation details.
  3. Right column — Run Explorer for navigation and Run Stats for aggregate history.

The Run Explorer is navigation-only. Build run details live in Active Run; validation run details live in Space Test.

Recommended demo path

  1. Sign in with Hugging Face OAuth.
  2. Check or create the private run Bucket.
  3. Review the default Tongyi-MAI/Z-Image-Turbo example or enter another model card, then run the model pre-scan.
  4. Launch only after the scan returns safe/caution or after explicitly acknowledging a risky verdict.
  5. Watch Active Run: progress timeline, latest events, latency when available, Space/Settings/Artifacts links, and trace documents.
  6. If hardware is manual, open Space Settings and select the recommended hardware.
  7. Open Space Test, prefill from the build, inspect available Gradio endpoints, click an endpoint to use it, and launch validation.
  8. Confirm full_inference_success, measured latency, and saved output artifact.

What it does not promise

The app does not guarantee that every model card can become a working Space. Some models require large or unusual hardware, Docker, multi-GPU execution, custom kernels, external services, accepted gated licenses, or manual fixes.

Production-readiness note

The app is suitable for a controlled public-preview or internal production deployment when deployed with OAuth enabled, per-user Buckets, private generated Spaces, and clear billing expectations.

Before broad public launch, keep monitoring:

  • Job cost and user quota expectations;
  • token redaction in traces;
  • failed build/validation rates;
  • hardware assignment failure modes;
  • latency and ZeroGPU duration recommendations.

Additional preview checklist

  • Pre-scan a few representative model cards before launching Jobs.
  • Confirm that risky models require acknowledgement and unsupported models stay blocked.

UI reliability checks

For preview deployments, verify that repeated refreshes do not duplicate click handlers: endpoint chips should update one API field once, run pagination should move one page per click, and run explorer selection/delete should remain independent.

  • Repair is one bounded automated retry with visible artifacts and no unbounded loop.

Recovery behavior

If a generated Space fails with HF build/runtime errors, especially empty logs or no-reason runtime errors, the Factory asks Pi to diagnose the blockage before any code patch is allowed. The system may wait for logs, re-upload the same workspace to force a rebuild, patch code once, request manual hardware, or stop with an auditable technical blocker.

v120 note: Agent recovery visibility

The UI now exposes Pi diagnosis/recovery state through the Active Run Agent recovery card and links to recovery decision/blockage artifacts. This helps users understand when the Factory is waiting for HF logs, rebuilding the same code, patching, requesting hardware, or declaring a technical blocker.

v122 stability note

The v122 line intentionally rolls back the experimental v121 client-side caching/polling optimizations after real UI testing exposed flicker and broken build/run behavior. The app keeps the v120 Agent recovery UI and the v118/v119 core blockage protocol, but returns to the proven polling and run-selection model for production stability.