Do generative image models truly understand lighting? We benchmark 16 models by asking them to inpaint light probes into real photographs, then measure how accurately the generated probes reproduce the ground-truth illumination in direction, colour, and radiance distribution.
Accurate modelling of illumination is central to realistic image synthesis and scene understanding. Yet, there is little exploration into whether image generative models truly understand lighting in a physically accurate manner. This work proposes a benchmark to assess the lighting understanding and harmonisation capabilities of generative models. Our key insight is that evaluating lighting understanding only requires testing how well models insert novel objects into real photographs whilst maintaining consistent illumination. We use a multi-illumination dataset with images containing simple objects serving as “light probes”, and prompt models to inpaint the same object onto the original image, then compare the generated results against the ground-truth light probes. We estimate the lighting direction, colour and radiance distribution from the inpainted probes, providing a quantitative measure of illumination accuracy and photometric realism.
Click a finding to see the supporting figure with the relevant conclusion highlighted.
Azimuth
Elevation
Given an image containing real light probes, we mask the probes and prompt a generative model to inpaint diffuse grey spheres. The generated probes are detected, cropped, and their lighting parameters estimated through inverse rendering. We compare against the ground truth across three measures: angular error in light direction, ΔEab colour error, and KL divergence of radiance distributions.
Select a scene, then click any model to inspect its generated probe and estimated light direction against the ground truth.
Drag the slider to compare ground-truth (left) and generated (right) images. Below each pair, the probes show ground-truth (orange border) and generated (blue border) light probes with estimated light directions.
If you find our work useful, please consider citing:
This work uses the Multi-Illumination dataset for ground-truth scene images.
This research was supported by NSERC grant RGPIN 2020-04799 and an NSERC PhD scholarship to JG. JVC was supported by Grants PID2024-162555OB-I00, AIA2025-163919-C52 funded by MCIN/AEI/10.13039/501100011033 and by ERDF “A way of making Europe”, the Generalitat de Catalunya CERCA Program, and the 2025 Leonardo Grant for Scientific Research and Cultural Creation from the BBVA Foundation. The BBVA Foundation accepts no responsibility for the opinions, statements and contents included in the project and/or the results thereof, which are entirely the responsibility of the authors. Compute support was provided by Digital Research Alliance of Canada (RRG 5299).