FLUX.2 Klein Schematic LoRA

We developed and released six LoRAs for the image editing model FLUX.2 [klein] 9B that turn a photo into a depth map, a normal map, a human pose, or a cut-out mask. Work that normally needs a dedicated recognition model for each task is done as an edit: redrawing the photo as a different image.
All six run locally. They are less accurate than the dedicated models, but some results are ones only an image generation model gives, such as cut-outs that include the parts of an object hidden from view.
Recognition as image editing
Image recognition tasks (computer vision tasks) such as depth estimation and segmentation usually use a dedicated model for each task. Google DeepMind has published research, Vision Banana, that treats these tasks as image editing.
That work led to our starting question: could FLUX.2 [klein] do the same?
Simply reproducing Vision Banana would not have been very interesting, so on top of depth and normal map estimation we added two tasks: pose estimation and amodal segmentation. We chose them because both are common parts of image generation pipelines. Poses, like depth, are a standard ControlNet input, and segmentation masks are often used to mark the area to redraw when inpainting.
Cutting out what can't be seen
Ordinary segmentation cuts out only what is visible. Amodal segmentation also imagines the parts that can't be seen. If a person stands behind a car, ordinary segmentation cuts out the shoulders and head showing above the roof; amodal segmentation also imagines the body hidden by the car.
Imagining and drawing what can't be seen is what image generation models do best. Where SAM 3.1, a dedicated model, cut out only the visible parts of a bench, the LoRA drew the whole bench, including the backrest and seat hidden behind the people.

Six tasks
The six tasks are depth, normal maps, body pose, whole-body pose including hands and face, segmentation, and amodal segmentation. Training them together in one LoRA made the tasks bleed into each other, so each task has its own LoRA.





Weight, and the difficulty of choosing
There are open problems. One is that this is heavy for a vision task: running a 9B image generation model takes more time and compute than a dedicated recognition model.
The other is segmentation. Not only amodal segmentation but ordinary segmentation too turned out to be hard. Depth maps and normal maps are, in a sense, style conversions, so they come easily. Segmentation adds a harder step: reading the prompt and choosing which thing in the image to cut out.
The training data is published together with how it was made.



