A local studio for turning an idea into a production: story, pictures, clips, comics, sound and 3D — with every intermediate result visible and editable.
HocusPocus is an experimental, non-commercial fork of Blizaine/Maestro by Blaine Brown (@blizaine). Maestro already did the hard part: a serious local generation stack on Wan2GP. We keep that foundation and add the missing production layer — a world that can be planned, directed, recovered and revised without starting from zero.
The HocusPocus mark is a quill shaping a cube: imagination becoming a buildable world. The UI is English and Spanish.
Gandalf and Tentri, from a real Video 3D production (world compositor, image lip-sync, MiniMax plates on the screens). Not a mockup.
The same clip inside the running studio (gallery, Wizard, Spanish UI):
Install with Pinokio from https://github.com/IAnMove/hocuspocus. NVIDIA GPU required.
Director takes a brief (or a song) and turns it into reviewable shots, prompts and references. Manual mode lets you approve each step. Auto mode runs analyze → plan → images → clips → assemble, and Productions keeps the whole thread so you can resume or retake one shot.
Example — music video. Drop a track. Director reads BPM, sections and energy, writes shots on the downbeats, generates start frames with the same characters, then runs the video model. Open Productions if a chorus clip fails; regenerate that shot only.
Example — short film. Write “two couriers argue in a rainy alley, then one reveals a glowing cube.” Director writes a screenplay, named cast, continuity across cuts. Pick MiniMax H3 for picture + native stereo audio, or Wan / LTX / Hunyuan Video for picture-only pipelines.
Example — trailer without a song. In Story Lab create a Tráiler cinematográfico. You get a 6–12 beat theatrical arc (cold open → unresolved hook), then generate 15–180 seconds as text-to-video or from approved frames.
Story Lab is the production bible: premise, world rules, locations, cast, relationships, beats. Approve fields, then hand the canon to Comics, Director, trailers or videoclips. Export a .storypack when you want to move the project.
Example. Approve a desert city, three characters and a logline. Open Productions → Comic for a 4-page chapter that does not retell the whole plot. Open Short Film → Story with the same canon and the same identity images. The writing model stays the one you picked on the story, not a silent global default.
Walkthrough with screenshots: Story → Comics → Video.
Comics plans pages and panels, drafts a script you can rewrite, generates art per panel, letters with restraint, and exports PDF / CBZ / PNG. Translation can rewrite balloons without touching artwork.
Example — Comic → AI film. Approve the comic first. The film path feeds the canon and every planned scene to the LLM, strips lettering from the panels, and uses each panel as the real first frame of that shot. You spend video time, not a second comic pass.
Studio is the manual bench: pick image, video or audio, write the prompt, add references and LoRAs, generate. Outputs land in the gallery and are reusable as references, editor clips, 3D plates or comic identities.
| Kind | What ships in the box (among others) |
|---|---|
| Video | MiniMax H3 (picture + stereo audio), Wan 2.1 / 2.2, Hunyuan Video, LTX-2.3 |
| Image | Flux 2 Klein, Qwen Image Edit |
| Audio | ACE-Step 1.5 XL (default new songs), MiniMax Music, Kugelaudio / Qwen3 TTS, MMAudio SFX |
Example — MiniMax H3. Prompt a wide night sea and add Audio: surf, wind, a low cello. Use FL2VA when you have an exact first/last frame from Story. Use Ref2VA when you pass up to 9 images, 3 videos and 3 audio clips as identity/mood references (<Picture 1>, <Video 1>, <Audio 1>).
Hoja de estilos / Style sheet imports ostris/minimax_h3_1k by ostris (@ostrisai): 1,000 MiniMax H3 clips with full multimodal captions (look, action, soundscape, music). Each style keeps author, repo, revision and a preview. Search, apply the prompt language to a new H3 job, or keep the clip as a visual reference.
Example. Download the source from the Style sheet tab. Filter “claymation”. Open a style, copy its visual lead-in into Studio H3, keep your own story beats. That library is how HocusPocus learned what a strong H3 brief looks like — credit ostris when you show the results.
3D runs Hunyuan3D in an isolated env: text, one image, or front/left/right/back views → GLB. Retexture GLB paints a new copy; the source file stays untouched.
Character Creator: one photo → H3 360° turntable → pick front/left/back/right → Hunyuan multi-view mesh.
3D Video is a real compositor (not a video model pretending to be a camera): place GLBs, images, Character Kits, lights and world SFX (portals, circles, beams, auras, missiles — they live in the 3D world, occlude, and export with the shot). Animate can rig a static GLB. Face Rig is the 2D cutout path: mouth overlays on a reviewed pose, not a mesh.
Example. Photo of a courier → Character Creator mesh. Drop it in 3D Video with a street plate. Anchor a magic circle to the character, pause mid-walk, nudge the gizmo — the offset follows the live pose. Export MP4. Operator guide: 3D Video compositor, Character Kits.
Video 3D editor (live compositor: stage, Play, world SFX, lip-sync):
Gandalf speaking in that world (image lips on the mesh, not a baked video):
Video 2.5D compositor (layers, plates, SFX, export):
Image lips / Face Rig — click the mouth on the portrait, place overlays, then play them in the stage:
| Close-up | Placement in the 3D stage | Character Creator → Face Rig |
|---|---|---|
![]() |
![]() |
![]() |
The Wizard is an in-app director: “open the concert scene”, “prepare a 3D showcase”, “make a 5-second clip of the cube in the rain”. MCP exposes the same jobs to external agents (image, video, SFX, scenes, receipts). Switching the footer workspace while a Wizard scene is still loading will not stomp the compositor or wipe undo.
Example. In 3D Video, ask the Wizard to open a saved scene by name and select a layer. If you change output folder mid-load, it aborts instead of importing into the wrong world.
Video Editor trims, splits and reorders clips you already like (H3 MP4s, compositor exports, series handoffs). Export is a queued FFmpeg job. Guide: Video Editor.
Edits (experimental): retake a section, outpaint a frame, prompt-driven replace. Multi-clip is for longer prompt-by-prompt sequences with overlapping continuity.
- Output folders are physical save directories (client A vs B, SFW vs NSFW).
- Workspace collections group projects without moving files. The gallery Workspaces tab is the Director thread dashboard for the active folder — guide.
- CivitAI LoRA browser with one-click install, update badges, and auto-written prompting guides from CivitAI / Hugging Face cards.
- Local LLM (Gemma 4 / Qwen GGUF via llama.cpp) or external OpenAI / Anthropic / compatible endpoints. Unloads after idle so VRAM goes back to generation.
- Themes: Golden Hour, Classic, Onyx.
- LAN: optional share on the local network; optional token auth (
LOREFRAME_LAN_AUTH). - NSFW and experimental gates are opt-in.
Operator index: docs/HOWUSEIT.md.
- Create canon in Story Lab (or skip it and start in Studio).
- Approve a character and a location. Optional: import ostris styles so H3 speaks a consistent look.
- Generate a few clips in Director or one-off shots in Studio.
- For exact cameras, build a 3D Video scene from Hunyuan GLBs / kits / world SFX.
- Assemble in Video Editor. Keep sources in the same output folder.
- If anything dies mid-run, open Productions — do not start from zero.
Every step can also start from an existing image, video, audio file or GLB.
| Minimum | Recommended | |
|---|---|---|
| OS | Windows 10/11 or Linux | Windows 11 or Linux |
| GPU | NVIDIA, 6 GB VRAM | RTX 3090 / 4090 / 5090, 24 GB+ |
| RAM | 16 GB | 32 GB+ |
| Disk | 150 GB free | 500 GB free for a full model shelf |
| Python | Installed by Pinokio | — |
| Card | After models are cached |
|---|---|
| 24 GB | comfortable; short clips in a few minutes |
| 12–16 GB | auto-tune offloads; slower |
| 6–8 GB | works with heavy offload; keep clips short |
AMD GPUs and macOS are not supported (CUDA kernels). First launch downloads weights on demand (often 50–100 GB; the full set can pass 300 GB). Hunyuan3D compiles native extensions: on Windows you want CUDA Toolkit and Visual Studio Build Tools.
- Install Pinokio.
- Discover → paste
https://github.com/IAnMove/hocuspocus, or download from this repo. - Install, then Start. The first job on each model fetches its weights.
Pinokio Install and Update share Windows/Linux recipes with separate Python environments and pinned dependencies per engine. Update also rebuilds the UI. SAM (Inpaint) and UniRig are optional menu installs; UniRig currently has a Linux recipe. See runtime profiles and recovery.
Start verifies and repairs the React build before loading the backend. For a missing or incomplete interface, stop Start, use Repair Web UI, then Start again; models are preserved. Startup logs show the app version, commit, OS and React build ID for bug reports. See React recovery and manual commands.
Reset removes the managed environments, vendor checkouts and UI build, including the Hunyuan model cache in app/ckpts/model3d. Use Install/Update to retry a failed setup; Reset is destructive.
To inspect the selected runtime recipes, use the URL shown by Start (replace the example host and port, including when connecting over LAN):
curl -X GET http://127.0.0.1:7860/api/v1/runtime-capabilitiesimport requests
report = requests.get("http://127.0.0.1:7860/api/v1/runtime-capabilities", timeout=30).json()
print(report["engines"])const report = await fetch('/api/v1/runtime-capabilities').then(response => response.json());
console.log(report.engines);HocusPocus is not a single Python script with a web wrapper. It is a small studio stack:
- Pinokio clones this repo, creates isolated environments, and owns Install / Start / Update / Reset. The published line is
main; day-to-day work lands ondevelopment. - FastAPI (
app/) is the studio server: jobs, gallery, provenance, Director pipelines, 3D, comics, styles, MCP. The React UI (ui/) is the control surface. Both languages share the same commands. - Generation still runs on the WanGP lineage Maestro already integrated — Wan, Hunyuan Video, LTX, Flux, audio models — plus an isolated MiniMax H3 ComfyUI runtime (picture and native stereo audio) and an isolated Hunyuan3D env (meshes). Optional SAM and UniRig live in their own envs so their stacks cannot fight the video venv.
- Director is a multi-pass LLM planner (story → shots → per-model polish, including LoRA guides). It writes artifacts you can inspect. Productions serializes that state so a crash is a resume, not a myth.
- 3D Video renders in the browser (Three.js). World SFX are objects in that scene, not 2D overlays. Export is local composition + FFmpeg, not “ask a video model to pretend it dollied”.
- Style sheet stores ostris’s H3 captions with source attribution. Wizard and MCP call the same backend tools the buttons use, so an agent cannot do a secret second pipeline.
- Auto-tune profiles VRAM on first launch. Jobs report through the footer. Nothing important is only in a chat bubble.
If you are automating: POST /api/v1/generate and POST /api/v1/model3d/generate are the same contracts the UI uses. MCP tools are listed by the running server (tools/list). Do not scrape internal Python.
Lineage in one line: Wan2GP → Maestro → HocusPocus, plus Hunyuan3D, MiniMax H3, ostris’s H3 style corpus, Pinokio, and the models credited below.
We would not exist without the people who shipped the layers we stand on. Tag them if you show HocusPocus.
| Who | What | Links |
|---|---|---|
| Blaine Brown | Maestro — the local studio we forked | @blizaine |
| deepbeepmeep | Wan2GP / WanGP — generation pipeline and non-commercial license we inherit | @deepbeepmeep |
| cocktailpeanut | Pinokio and the original Wan2GP launcher | @cocktailpeanut |
| Who | What | Links |
|---|---|---|
| ostris | minimax_h3_1k — 1,000 H3 clips and multimodal captions that power Style sheet. Also AI Toolkit. | @ostrisai · HF |
The dataset card does not name a license. We keep author, repo and revision on every imported style and do not claim those prompts as ours.
| Who | What | Links |
|---|---|---|
| MiniMax | MiniMax H3 (FL2VA / Ref2VA, native stereo audio), MiniMax Music | @MiniMax_AI · HF |
| Comfy-Org | H3 ComfyUI runtime we isolate | PR |
| Tencent Hunyuan | Hunyuan3D 2 / 2.1 (meshes, paint) and Hunyuan Video | @TencentHunyuan |
| Alibaba / Wan-Video | Wan 2.1 / 2.2 | Wan2.1 |
| Lightricks | LTX-Video 2 / 2.3 | LTX-Video |
| Black Forest Labs | Flux | @bfl_ai |
| Qwen / Alibaba | Qwen image + LLMs | Qwen |
| Gemma 4 (default local Director LLM) | Gemma | |
| Meta | SAM (optional Inpaint) | SAM |
| llama.cpp | Local GGUF inference | llama.cpp |
| CivitAI | LoRA browser and community weights | civitai.com |
| MMAudio | Ambient audio | MMAudio |
| seed-vc (Plachta) | Voice conversion, GPL-3.0, cloned at install from maestro-seedvc | seed-vc |
Other vendored pieces keep their own licenses, including BigVGAN (MIT), FlashVSR sparse-sage (Apache-2.0), IndexTTS2, Rhubarb lip-sync, BS-RoFormer / audio-separator, Three.js, UniRig (optional), and Tencent’s Hunyuan checkouts. Read those LICENSE files before you redistribute.
If we missed a name you shipped into this tree, open an issue — we want the list complete.
WanGP Non-Commercial Evaluation License 1.1, inherited through Maestro from Wan2GP. Summary: LICENSE. Full text: app/LICENSE.txt.
TL;DR: free to use and modify for non-commercial purposes. Outputs you generate are yours to use commercially (with attribution). Commercial use of the software (including hosted APIs) needs a separate license from the WanGP licensor. Third-party models keep their own terms.
Bugs and requests: github.com/IAnMove/hocuspocus/issues.







