Matatabi AI
  • 日本語
  • English

Matatabi AI

Just as cats cannot resist matatabi,

there must be something people cannot help but be drawn to, too.

We are a company that seeks out what that something is and lets it take root in AI.

About Matatabi AI
An AI development company specializing in content creation.
More
What we do
AI development, technical evaluation, and production expertise.
More
Contact
Send us your questions or inquiries.
More
New tabGitHub(opens in a new tab)New tabComfy with ComfyUI(opens in a new tab)
Kura
Kura's screen showing training progress, settings, a loss graph and GPU status
More
FLUX.2 Klein Schematic LoRA
One painting turned into six diagrams — depth, surface normals, poses and cut-outs
More
ComfyUI Panorama Stickers
More
ComfyUI Workflow Image Export
More
Comfy with ComfyUI
The Comfy with ComfyUI home page
More
ComfyUI Video Stabilizer
More
Cross-Image Try-On LoRA
A woman in a green dress next to a man in a black shirt, and the edited result with the man wearing the green dress
More
© 2026 Matatabi AIAboutContactPrivacy Policy

FLUX.2 Klein Schematic LoRA

2026.06.01Works

One painting turned into six diagrams — depth, surface normals, poses and cut-outs

We developed and released six LoRAs for the image editing model FLUX.2 [klein] 9B that turn a photo into a depth map, a normal map, a human pose, or a cut-out mask. Work that normally needs a dedicated recognition model for each task is done as an edit: redrawing the photo as a different image.

All six run locally. They are less accurate than the dedicated models, but some results are ones only an image generation model gives, such as cut-outs that include the parts of an object hidden from view.

nomadoor/flux-2-klein-9B-schematic-loraThe six LoRAsnomadoor/flux-2-klein-9B-schematic-datasetTraining dataFLUX.2 [klein] Schematic LoRAWrite-up — how the data was made, with detailed results

Recognition as image editing

Image recognition tasks (computer vision tasks) such as depth estimation and segmentation usually use a dedicated model for each task. Google DeepMind has published research, Vision Banana, that treats these tasks as image editing.

That work led to our starting question: could FLUX.2 [klein] do the same?

Simply reproducing Vision Banana would not have been very interesting, so on top of depth and normal map estimation we added two tasks: pose estimation and amodal segmentation. We chose them because both are common parts of image generation pipelines. Poses, like depth, are a standard ControlNet input, and segmentation masks are often used to mark the area to redraw when inpainting.

Cutting out what can't be seen

Ordinary segmentation cuts out only what is visible. Amodal segmentation also imagines the parts that can't be seen. If a person stands behind a car, ordinary segmentation cuts out the shoulders and head showing above the roof; amodal segmentation also imagines the body hidden by the car.

Imagining and drawing what can't be seen is what image generation models do best. Where SAM 3.1, a dedicated model, cut out only the visible parts of a bench, the LoRA drew the whole bench, including the backrest and seat hidden behind the people.

Cut-outs of a person, a bench and a locomotive: visible-only masks compared with masks including hidden parts
Amodal segmentation. Middle: SAM 3.1, which cuts out only what is visible. Right: the LoRA

Six tasks

The six tasks are depth, normal maps, body pose, whole-body pose including hands and face, segmentation, and amodal segmentation. Training them together in one LoRA made the tasks bleed into each other, so each task has its own LoRA.

Depth estimated for three photos by Depth Anything V2 and by the LoRA
Depth. Middle: Depth Anything V2. Right: the LoRA
Normal maps estimated for three photos by Lotus-2 and by the LoRA
Normal maps. Middle: Lotus-2. Right: the LoRA
Body poses estimated for three images by DWPose and by the LoRA
Body pose. Middle: DWPose. Right: the LoRA. For the character in the bottom row, DWPose detects nothing
Whole-body poses including hands and face estimated for three images by DWPose and by the LoRA
Whole-body pose including hands and face. Middle: DWPose. Right: the LoRA
Segmentation masks for three photos from SAM 3.1 and from the LoRA
Segmentation. Middle: SAM 3.1. Right: the LoRA

Weight, and the difficulty of choosing

There are open problems. One is that this is heavy for a vision task: running a 9B image generation model takes more time and compute than a dedicated recognition model.

The other is segmentation. Not only amodal segmentation but ordinary segmentation too turned out to be hard. Depth maps and normal maps are, in a sense, style conversions, so they come easily. Segmentation adds a harder step: reading the prompt and choosing which thing in the image to cut out.

The training data is published together with how it was made.