{"ID":23475091,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.19853","arxiv_id":"2609.19853","title":"PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance","abstract":"Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge. On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director's words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: https://github.com/StudioPiLabs/pace-core","short_abstract":"Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the pl...","url_abs":"https://arxiv.org/abs/2609.19853","url_pdf":"https://arxiv.org/pdf/2609.19853v1","authors":"[\"Bing Duan\",\"Qiang Guo\",\"Linpu Li\",\"Zhijian Mao\",\"Min Zhu\",\"Zhirui Ren\",\"Yiwei Yan\",\"Xi Chu\",\"Xiaoding Li\"]","published":"2026-09-17T08:05:08Z","proceeding":"cs.CV","tasks":"[\"cs.CV\",\"cs.AI\"]","methods":"[\"Diffusion Model\"]","has_code":false,"code_links":[{"ID":639791,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-18T01:09:05.407443952Z","DeletedAt":null,"paper_id":23475091,"paper_url":"https://arxiv.org/abs/2609.19853","paper_title":"PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance","repo_url":"https://github.com/StudioPiLabs/pace-core","is_official":false,"mentioned_in_paper":false,"mentioned_in_github":true,"github_stars":0}]}
