Skip to content
ARarmature
v0.1.0 released · thirteen experiments closed · identity holds — even on a hosted tier fed only authored references

armature — You block the shot. The model shoots it.

Stage your character in Blender. Render the control sequence. Let the video model paint the life over it.

The problem

A video model gives you what no renderer can

Motion, light, weight, the way a coat swings and a face catches a lamp. That is the expensive half of animation, and a video model produces it directly.

It cannot be told who is on screen

Or where they are standing. Prompt a character and you get a character; prompt again and you get a different one. There is no handle for this person, here, facing that way, in frame 47.

armature is the handle

The previz scene is ground truth. A canonical mesh is staged and animated, and the render becomes a per-frame control sequence the video model works inside. Twelve experiments in, the idea holds at product level — and every claim behind that sentence traces to a numbered experiment in the repo.

How it works

Three steps. Structure comes from geometry you own; life comes from the model.

1 — Stage

A canonical character mesh is posed, animated and framed in headless Blender. You block the shot: where the character stands, how they move, what the camera does.

2 — Render the control sequence

Per-frame control channels come out of the staged scene — depth, normal, mask, edge, and optionally a skeleton drawn from bone transforms that are already known. Rendered from geometry, never estimated from footage.

3 — Generate inside the structure

The control sequence and a reference stack go to a video model, which generates within them. Identity is meant to be a named, versioned thing that rides in the prompt and the references — not an accident of a lucky frame.

Where this is

armature was founded on 2026-08-10. The thesis it exists to test — does a CG-rendered control sequence hold a character through a video model — is now measured at product level across thirteen closed experiments, judged by the Director’s eye, and a negative result remains a full success here. The counters below are dated; README.md in the repo carries the live ones.

CounterAs of 2026-08-13
Experiments14 closed (one more withdrawn un-run on a falsified premise); E13 ran its full arc — dispatch, zero-spend halt, repair, re-arm, four generations, close — inside 2026-08-13, and E14, the LoRA scene-lever bake-off, closed behind it the same date: both style LoRAs bind on the derivative weights, the character holds on technically_color and fails on the photo-real pair, at zero partner credits
RoutesThree, measured: driven (rig-rendered pose → Animate; proven at shot level, parked, licence-clear for its unpark), free (authored start frame → camera tier at the 6.0 / uni_pc baseline, its LoRA scene-lever priced live by E14), composed (authored references into a hosted identity-lock tier — graduated by E13)
SpendFounding arc: 22 probes at 4 credits each; the E08–E12 arc metered 0 at every submission (GPU-hour billing) under per-experiment ceilings; E13’s four generations are the repo’s first partner-credit spend, inside their pre-stated bracket; E14’s two metered 0 partner credits at a ceiling reached exactly
License mapEvery adopted dependency carries a retrieved licence document; UNVERIFIED is treated as NO; third-party-tier routes carry per-route disclosure (ruled 2026-08-12)
Research grounding24 findings; 34 of 34 citations resolved against a retrieval oracle, zero fabricated
The founding thesisMEASURED — identity holds driven (E08), unanchored (E11), and through a hosted human-trained tier fed only authored references (E13); the camera obeys authored control to one pixel (E11); a handed world holds on two seeds (E12); reference grounds steer model-decided worlds (E13)
What exists todayv0.2.1 — the record’s current marked state, and the first that installs: armature_core on PyPI (armature-studio) and npm (armature-studio), plus the routes, the instrument shelf, the licence map, the experiment record E01–E14, and this page

What is already established

Three things here are evidenced. None of them was measured in this repo — they were measured elsewhere and retrieved, and that is exactly why they are the only things stated as fact.

The closest published precedent

Champ (Zhu et al. 2024, arXiv:2403.14781) compared dense 3D-parametric guidance — depth, normal and semantic maps rendered from a posed body model — against a 2D skeleton alone. FVD 192.34 to 170.20, SSIM 0.672 to 0.773, LPIPS 0.296 to 0.235. Dropping the skeleton entirely still beat skeleton-only on FVD, at 184.24. That is armature’s thesis, measured by someone else, on someone else’s subject matter.

The licensing finding

OpenPose — the most widely used pose extractor in the ControlNet ecosystem — is CMU non-commercial, and is banned here. Depth Anything V2 Small is Apache 2.0 while Large is CC-BY-NC: same family, same page layout, different licence. Rendering control from geometry removes that entire tier by construction rather than by substitution.

What nobody has measured

No quantitative curve exists anywhere for control strength against identity drift. No study compares per-character LoRA against zero-shot conditioning on stylized game art. No head-to-head of depth versus pose versus segmentation on one identity metric. Those gaps are why this repo has an experiment arc instead of a feature list.

How this repo works

Three seats, and the separation is the point: the session that designs an experiment does not grade its results, and the session that runs it does not decide their meaning. Every non-trivial change runs as a numbered experiment — spec written before the work, report written after, advisor ruling last. Metrics are diagnostics; the Director’s eye is the judge.

SeatDoesMust not
DirectorSets direction; judges every artifact by eye
AdvisorWrites specs, rules on reports, folds findings into the repoExecute, or grade its own rulings
ExecutorRuns the spec, measures, reports evidenceDecide what results mean, or judge quality