Open protocols from our experiments with image generators: how we measure a room, how we set the product size and which models hold up. Every result and every miss is published.
19 August 2026
Question: how do you set the size
A furniture try-on lives or dies on one number. A 192 cm sofa has to come out 192 cm wide, otherwise someone buys a thing that will not fit. Image generators do not promise that: they draw an object at whatever size looks good in the frame.
We looked for a way to hand a model the size so that it actually keeps it. Below is the protocol and three experiments: what to set the size with, which models read it and whether the result repeats across rooms.
Earlier rounds looked for a way to give the model a size. This one asks how to verify the model kept it. Until September we had nothing but our own confidence: we showed the finished shot to a text model and asked where the object was.
In the 8 September run all twelve checks came back "unverified". The breakdown showed the coordinate format was not the problem.
Experiment 1
A text model cannot see pixels
The reviewer was shown the finished shot and the target frame, and asked for the object bounds as fractions of the image. It matched the frame instead of measuring what was drawn.
What we asked for
What came back
fractions from 0 to 1
right = 1100 on a 1920-wide frame
bounds of a floor lamp
nearly the frame itself, on an object twice as tall
The model does not measure, it invents something plausible. While the target frame is in front of it, it converges on that frame — that is, it confirms itself.
Experiment 2
The instrument: segmentation, not opinion
The object is found by text-prompted segmentation, and the target frame is never shown to it — there is nothing to conform to. Height and centre are computed in code. The wording of the request turned out to decide everything.
Prompt
Result
sofa
found
3-seat sofa
the same box to three decimals
MORABO 3-seat sofa, 232x89x86 cm
not found
sofa pale green tufted with wooden legs
not found
A product code or a chain of adjectives kills the search as reliably as a typo. A short noun phrase works. The instrument was checked against two hand measurements and matched both exactly; every measurement now runs two phrasings, and if they disagree by more than 10% the check refuses to answer.
Experiment 3
What the instrument found
The first thing the instrument showed was that the error was not in the generator but in our own arithmetic. Scale comes from the depth at a reference point, and that point was landing on a foreground object instead of the floor.
Product
Frame was
Frame now
Truth
plant, 110 cm
141 px
162 px
184 px
floor lamp, 140 cm
150 px
255 px
271 px
The "truth" column was obtained independently: in a room measured with a tape, scale was derived from an object of known size and transferred to the placement point by a depth ratio. Two anchors at different distances gave answers one pixel apart, and the depth ratio matched the physical ratio of scales to within 1.1%.
Experiment 4
A correct frame made the result worse
Three new products, one room, the same tap. Exactly one thing differed: in the second arm the scale came from a measurement against an object of known size rather than from the model — that is, the frame was nearly correct.
Product
Frame from the model
Frame from a measurement
bookcase, 237 cm
−5%
+7%
storage unit, 197 cm
+1%
+5%
armchair, 80 cm
+15%
+25%
The generator draws the object 1.10–1.32 times larger than the frame, and the factor does not depend on where the frame came from: it is a property of the product, not of the arithmetic. So today’s decent accuracy rests on two errors cancelling — an undersized frame and an oversized drawing. Remove one and the other shows in full.
Where the method breaks
What this round does not prove
One room, three products, one shot per arm. A direction, not a statistic.
The mask does not hold the size: on the path with a real alpha mask, 29% of the pixels outside the mask changed anyway — the model redraws the whole frame and treats the mask as a hint.
The measurement only understands English product names and refuses to work when two phrasings disagree. We treat refusal as correct behaviour: better silence than a confident wrong number.
The frames from this round were shot in a private home and are not published. The section runs on tables, as every round has since 31 August.
Method
The door leaf as a ruler
To check a size in a photo you need the scale of the room: how many pixels fit into one centimetre. We take it from the interior door, whose leaf is 210 cm almost everywhere and usually visible in full. That measurement is 1–2% off. The older method, floor-to-ceiling distance, was off by 13 to 43%.
The product dimensions from the shop listing then become a rectangle in pixels, drawn straight onto the room photo in red: that is the correct size. The model gets the same photo with a green frame and a request to draw the product inside those lines and remove the frame. A hit is checked by overlaying the red reference on the result.
Three rooms shot on a phone, 1280 × 960. Products are real IKEA and Ozon listings with their own dimensions. A hit means the top, bottom and sides of the object sit inside the frame, visually within 10%. Each model ran each case once, with no cherry-picking. The rooms are a private home, so the raw frames stay unpublished; every measurement and outcome is in the tables below.
Experiment 1
Three ways to say "110 cm"
One product, a plant 110 cm tall, one room, one model — Seedream 4.5. The only thing that changes is how the model learns the size: from the prompt text, from a cut-out pasted into the frame, or from a rectangle drawn on the photo.
How the size was set
Result
Error
Centimetres in the prompt
a tree about 250 cm
+127%
Pasted cut-out photo
147 см
+34%
Frame drawn on the room photo
115 см
+5%
The model does not trust words at all. A pasted cut-out sets the size geometrically, yet the model still redraws the object bigger. A rectangle inside the picture turned out to be the only language that works.
Experiment 2
Eleven models, one frame
With the method settled, the question moves to the models. A bookcase 80 × 237 cm, the same room and frame, the same prompt word for word. We tested everything on Replicate that accepts a product photo alongside a room photo.
Model
What happened
Per image
Seedream 4.5
top, bottom and sides inside the frame
$0,040
Seedream 5 Lite
inside, body slightly narrower
$0,040
Wan 2.7
width off by 60%, leaning
$0,030
Seedream 4
bigger than the frame, shifted right
$0,030
Nano Banana Pro
melted into the wall, no back panel
$0,040
Qwen Image Edit Plus
the green frame stayed in the shot
$0,030
FLUX 2 Klein 9B
a quarter shorter than the frame
5 сек
FLUX 2 Klein 4B
leaning at an angle
4 сек
Runway Gen-4
redrew the whole room
—
Grok Imagine
ignored the frame
—
AnyDoor
drew a different bookcase
$0,007
Two models out of eleven read the frame. The rest fail in different ways, most often in the same direction: the product comes out bigger than asked. Qwen is its own case, drawing the bookcase correctly but leaving the green frame in the picture.
Experiment 3
Does it repeat in other rooms
A single good shot proves nothing. The same two finalists, plus Wan 2.7 as the cheapest option, ran four new cases in two other flats: a 192 × 94 cm sofa, a 270 cm storage unit deliberately wider than the wall, an 80 × 237 cm bookcase and a 110 cm plant. The room from experiment 2 is not repeated here.
Model
How it misses
Hits
Seedream 4.5
missed only on the plant
3 / 4
Seedream 5 Lite
furniture bigger than the frame
1 / 4
Wan 2.7
inflates the width in every case
0 / 4
The losers err in the same direction, which matters more than the score: this is not randomness but how the model reads the rectangle. It treats it as a hint about position, not a boundary. All three missed on the plant: a living crown has no clean edges, and the frame stops working.
Limitations
Where the method breaks
A rug lies in the floor plane, so it gets a perspective footprint instead of a rectangle. The 250 × 250 cm rug landed on that footprint correctly, and then a rolled-up tube appeared beside it that nobody asked for. The lamp frame came out narrow, 109 pixels in a 1280 pixel frame, and the model simply missed it.
What is fair to say about these numbers: the rooms come from one flat, the products from two shops, the hit is judged by eye, and each model ran each case once. This is not statistics, it is a tool-selection protocol. We publish the method and every outcome, misses included, so the reasoning can be rechecked; the illustrated round on a neutral room is in the next section.
Previous round · 56 results · 14 products
Before the frame: four models on 14 products
We compare models on the same empty room and real IKEA products, not polished demos. Wrong objects, visible masks and broken frames are left untouched.
One source room
Fixed inputs
same dimensions and placement zone
up to three product listing images
no cherry-picking or manual retouching
Gemini and Seedream used one pipeline on Aug 9, 2026. GPT is the previously published baseline in the same room.
Short takes from the newest rounds — only the conclusions that survived scrutiny. Full protocols with every frame live in the research repo.
25 August 2026
Native mask vs frame: 19 production cases
We reran 19 real user try-ons through gpt-image-2 with a true inpainting mask and judged them against the production frame contract. Strict score: mask 11/19, frame 10/19 — a tie with opposite failure modes. The frame ignores placement and tilts the object; the mask never misses the spot but overshoots the size and wipes everything inside the zone. Conclusion: no wholesale switch; the mask earns its keep as a placement-miss retry.
25 August 2026
Depth is not optional
Three arms on identical inputs: production, the v2 frame without a depth map, and v2 with one. The frame alone came out worse than production; only the depth-plus-mask arm beat it. Room scale from a single photo has to come from measured depth, not from the model’s taste.
24 August 2026
Nobody holds the frame
Four cases, same frame contract, four generators: Gemini, Seedream 4.5, FLUX.2 Pro and Flex. None held the drawn frame reliably. Gemini and FLUX.2 Pro came closest; Seedream redrew the scene around the product in 40% of runs — the whole photo, not just the object. A drawn rectangle is a hint, never a guarantee: verification after generation stays mandatory.
23 August 2026
Can you just buy this? We checked
We audited every virtual-staging API on the market against our exact need: room photo, product cutout, placement coordinates. MeltFlex, virtualstaging.ai, VSAI, SofaBrain — none accept coordinates; they restyle rooms or stage generic furniture. The geometry has to stay ours: classical paste and warp own the position, generation only harmonises light and shadow.
22 August 2026
ROI generation: touch only the crop
Instead of handing the model the whole photo, we cut a crop around the placement zone, generate inside it and composite back with a feathered seam. Outside the crop the room stays pixel-identical to the original — the failure mode where a generator repaints your walls disappears by construction. Local drift inside the crop remains the open problem.
22 August 2026
Seedream 5.0: assessed, declined
Official ModelArk access to Seedream 5.0 Pro, Lite and 4.5 on our test rooms: no native mask, and its bbox tags do not lock the size — the one thing we need locked. Production stays where it is.