A ceramic figurine shot on a kitchen counter with a phone held 20 cm away came back from image-to-3D with a bulging forehead, a back that looked like melted wax, and a base fused into a lump of countertop. The same figurine, shot forty minutes later on a grey paper sweep at 1.5 metres with the phone’s 3x lens, four angles, window light and a white bounce card, came back clean enough to decimate and drop straight into a scene. Nothing about the model changed. The photographs changed.
This is the single most under-taught part of the workflow. People iterate on prompts for an hour and never reshoot, when reshooting is faster and fixes more. The reconstruction quality you get is roughly a function of how much unambiguous information about form your photos contain, and a bad photograph contains startlingly little.
What the system is doing with your photo
Image-to-3D is not photogrammetry, and confusing the two leads to the wrong shooting habits. Photogrammetry needs 60–80 overlapping frames and solves for camera positions by matching features. Image-to-3D runs a learned prior: it recognises the object, infers a plausible complete form, and conditions the visible surfaces on your pixels. From a single image it is genuinely inventing the back.
Multi-image reconstruction — typically two to four views — does not triangulate in the classical sense either, but the extra views constrain the guess enormously. The practical consequence is that your views must agree. Four photos of an object that was nudged between shots, or lit differently, or shot at wildly different heights, give the system contradictory evidence and you get an averaged, mushy result. Consistency matters more than count.
Focal length: the mistake almost everyone makes
A phone’s default camera is around 24–26 mm equivalent. At the distance people naturally hold a phone from a small object, that lens produces severe perspective distortion — the near face is enlarged, parallel edges converge hard, and the object appears to taper away from camera. The reconstruction faithfully bakes that distortion into the geometry. That is where the bulging forehead came from.
Shoot longer and step back. On a phone, that means the 2x or 3x telephoto, not a digital crop of the wide. On a camera, 50–85 mm on full frame. The object should fill roughly 70–80% of the frame height at that distance, with a comfortable margin.
| Subject size | Focal length (35 mm equiv.) | Camera distance | Aperture | Views to shoot |
|---|---|---|---|---|
| Small, 5–20 cm (figurine, mug, tool) | 70–85 mm | 1.2–1.8 m | f/8 | 4 keepers from 12 frames |
| Tabletop, 20–60 cm (helmet, small chair, plant) | 50–70 mm | 2.0–3.0 m | f/8–f/11 | 4–6 keepers from 16 frames |
| Furniture, 0.6–2 m | 35–50 mm | 4–6 m | f/11 | 6 keepers from 20 frames |
| Vehicle or large object, 2 m+ | 35 mm | 8–12 m | f/11 | 6–8 keepers from 24 frames |
The aperture column matters more than people expect. Shallow depth of field is flattering in product photography and poison here: a blurred back edge gives the reconstruction no usable information about where the form ends. Stop down until the whole object is sharp front to back, put the camera on a tripod, and drop the shutter speed instead of opening up.
Lighting: flat and boring wins
Whatever lighting is in your photo becomes part of the albedo texture. A dramatic rim light and a deep shadow side will be permanently painted onto the model, and no amount of relighting in-engine will remove them. You want the most uninteresting lighting you can produce.
The reliable setup costs nothing: a north-facing window as a large soft key, and a piece of white foam board on the opposite side, 40–60 cm from the object, to lift the shadow side. Aim for a key-to-fill ratio around 2:1 — the shadow side should be clearly darker but still fully readable. Overcast daylight outdoors works equally well.
Things that ruin an albedo: hard direct sun, a bare bulb producing specular hotspots, mixed colour temperature (tungsten lamp plus daylight window gives you an object that is orange on one side and blue on the other), and a coloured wall bouncing a cast onto the subject. Set white balance manually off a grey card rather than trusting auto, which will drift between shots and give you four views that disagree about the object’s colour.
Kill the contact shadow. A dark pool where the object meets the surface reads as geometry and often gets reconstructed as a skirt around the base. Raise the object on a small clear riser, or shoot on a translucent surface with light underneath, or at minimum bounce light into the base.
Background separation
Use an unbroken mid-grey sweep. Grey works for both light and dark objects, and it does not tint the subject the way a coloured backdrop does. A2 or A1 paper is enough for anything hand-sized; curve it up the wall so there is no visible horizon line.
The rule is simple: the background must never be the same value as the object’s edge. A black camera bag on a black cloth has no silhouette, and silhouette is the strongest signal the system has. If your object is dark, go light. If it is white, go mid-grey — not black, which causes flare and lifted blacks around the edge.
Avoid anything with texture or pattern behind the subject. Wood grain, tiles, keyboards and carpet all get read as surface detail. Reflective subjects will pick up the room; if you must shoot chrome or gloss, tent it with white card on three sides so the reflections are at least uniform.
Turntable coverage
Rotate the object, not the camera. Mark the base position on the sweep with tape, put the camera on a tripod, and turn the object in place. This keeps distance, lighting and framing identical across views, which is exactly the consistency the reconstruction needs.
For a four-view set, shoot 0°, 90°, 180° and 270°, all from a camera elevation about 15–25° above the object’s vertical midpoint. That slight downward angle is deliberate — a set shot purely at eye level tells the system nothing about the top surfaces, and you will get an invented, usually flattened, top. If the top is important, add a fifth frame at roughly 50–60° looking down.
Shoot three times as many frames as you need and cull. Bracket a stop either side, check focus at 100% on each, and discard anything with motion, a shifted object or a changed shadow. The four you submit should look like the same object photographed by the same camera four times, because that is what they are.
Do not shoot the 45° in-between angles instead of the cardinal ones. Front, back and both profiles carry the most form information; the diagonals are mostly redundant with them. And never change the object between frames — no opening a lid at 180°, no removing a strap because it was in the way.
Subjects that will not reconstruct
Some materials defeat the process regardless of technique, and it is cheaper to know that in advance than to shoot for an hour.
| Symptom in the result | Usual cause | Fix |
|---|---|---|
| Bulging or tapered form | Wide lens, camera too close | Longer focal length, step back |
| Melted or featureless back | Single view only | Shoot four views, use multi-image |
| Baked shadows in the albedo | Hard directional light | Diffuse key plus bounce fill, 2:1 ratio |
| Skirt or plinth fused to the base | Dark contact shadow | Raise the object, light the base |
| Thin parts missing (handles, wires, antennae) | Feature below the reconstruction’s effective resolution | Model the thin parts separately, or accept and add them by hand |
| Holes, transparent-looking gaps | Glass, acrylic, or polished chrome | Matte spray or dulling powder, then shoot |
| Fuzzy blob where hair or fur was | Fur has no coherent surface | Reconstruct the body, add fur as cards or a shader |
| Mismatched left and right halves | Views disagree — object moved or light changed | Reshoot on a marked turntable position |
Resolution is rarely the limiting factor. 1024–2048 px on the long edge, cropped to about a 10% margin around the object, is plenty. Do not upscale a small photo before submitting; interpolated detail is not detail, and studios like MeshyFlix downsample the input anyway.
Do and don’t
- Do use a tripod and a 2-second timer. Handheld frames drift in scale and angle between views.
- Do shoot RAW if you can, and apply identical corrections to all four frames as a batch.
- Do include a scale reference in one extra frame, so you know what to type into the import scale field later.
- Don’t use portrait mode, beauty modes, HDR stacking, or anything that fakes depth. They invent edges.
- Don’t cut the object out with a rough selection tool. A harsh matte with a halo is worse than a clean grey background.
- Don’t shoot at f/2.8 because it looks nicer. Sharp everywhere beats pretty.
- Don’t mix a phone frame and a camera frame in the same set.
Build the kit and shoot one object today
Assemble the whole setup once and leave it assembled: a roll of grey background paper, a piece of A2 white foam board, a small tripod, a grey card, a clear acrylic riser, and a strip of tape on the floor marking your camera distance. It costs about 37 € and takes fifteen minutes.
Then run a controlled test. Photograph one object badly on purpose — phone wide-angle, held close, overhead room light — and reconstruct it. Photograph the same object properly and reconstruct it again. Put the two meshes side by side in your viewer. Once you have seen that comparison on an object you know well, the shooting rules stop being a checklist you have to remember and become the only way you would think to do it.