Style-Aligned Camera Angle Manipulation Guided by Gaussian/3D Models

Civitai   |   โ˜•๏ธ Ko-fi (support my work)  

Introduction

In AI image generation, it is typical that the camera angle is not correct, even after specifying it in the prompt. The new paradigm of image editing models such as Qwen Image Edit or FLUX.2 have made it easier to change the camera angle of an existing image, however most of the time it is relegated to fixed azimuths and elevations, which doesn't give you full creative control over an image. Additionally, the style of the image can drift with such dramatic camera angle changes. The AnyAngle lora is designed to improve upon this.

Here's how this works: We take the desired image and transform the scene into a Gaussian Splat (we use Tripo Splat, but you could use any existing gaussian splat generator) or into a 3D model (via Trellis2/Pixal3d). We take that generated splat/model, and import it into Blender. After the import, we create a new Blender camera, change the camera angle to any desired angle, and export the image of the new to-be angle (coarse image). After, we plug that image as a reference as well as the original image into Qwen Image 2.1, and with the AnyAngle lora and edit prompt, the camera angle from the coarse image will transfer to the original image, maintaining style coherency.

So, summing it up, the process goes like this:

Original image --> Generate Gaussian Splat/3D model from image --> Import into Blender and adjust camera angle + render image --> Plug into Qwen Image 2.1 and render:

Usage

Follow the general process from above, and use the following edit prompt:

Change the camera angle from <image2> to <image1>.

The general connections should look like this:

Ensure that the Lora strength is set to 1. Additionally, use CFG 3.0 and 20 steps or more for the most optimal results. However, for boarding such as shot planning and general speedy inference, it is okay to use a turbo lora to reduce latency, however note there will be a slight hit to quality.

Training Regimen

We use various real blender renders--all stylistically diverse-- as well as the SOTA video model MiniMax H3 (for digital illustrations) for our dataset images. We gather images of an both an orignal image (anchor) and a frame at a different camera angle (target). We use the target image and generate a Gaussian Splat/3D model out of it, and we then use that data as the "coarse render" for a control image. We do this lots of times until we gather a good dataset, and then we train the lora for few thousand steps.

Essentially with this training ideology, we are able to consistently keep the style aligned as nothing is hallucinated (at least when it comes to the real Blender renders). However, to ensure that other styles such as illustration and sketch are able to be manipulated, we imploy MiniMax H3 image-to-video to do various turn-around "orbit" renders and camera manipulations where everything stays completely stationary and still, so that the alignment stays as consistent as possible.

More Examples

Where it fails

If the generated splat/3d model is not spacially aware or is too overly coarse (with extreme anatomy or facial deformation), there may be issues with spacial arrangement of items or malformed faces if faces are not well defined or the images are of low resolution. Take for example this image here:

The spacial arrangement of the table and ponytail is clearly misplaced, so the spacial arrangement of the output image is off. This issue can be fixed by manipulating the gaussian/3d model to be placed in the correct spot, or by having a more accurate gaussian splat of the scene (whereby using a 3D world generator/workflow). Additionally, as time progresses there will be newer and better splat/3d model generators, so this issue will be progressively solved as time moves on.

Special Thanks

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for lilylilith/QI_2.1_AnyAngle

Adapter
(58)
this model

Space using lilylilith/QI_2.1_AnyAngle 1