The challenge for ray tracing has never been whether physically
accurate lighting works. The challenge has always been computational cost. Tracing every possible light interaction quickly becomes prohibitively expensive, particularly in dynamic scenes where geometry, lighting and camera positions are constantly changing. The industry's response has been surprisingly simple: stop trying to
calculate everything. Modern ray-traced systems increasingly operate using sparse sampling.
Instead of tracing every possible light path, they trace only a subset of rays and then reconstruct the final image. At first glance, this sounds counterintuitive. How can fewer samples produce a result that appears comparable to one generated from significantly more rays? The answer lies in the structure of light itself. Light transport is not random. Rays within a local region of a scene
exhibit strong geometric and temporal relationships. Surfaces tend to vary smoothly. Reflections change predictably. Neighbouring pixels frequently observe similar lighting conditions. The resulting light field contains patterns and correlations that can be learned. Machine learning models are remarkably effective at identifying those
patterns. Rather than attempting to simulate every possible light interaction, a
model can learn how information propagates through a scene and infer missing samples from the data that has already been observed. In effect, the ray tracer provides sparse but physically meaningful samples, while the neural network reconstructs the detail between them. This is why AI-assisted denoising and reconstruction have become such powerful tools. They are not inventing an image from nothing;
they are exploiting the geometric structure already present within the lighting solution. The same principle extends naturally beyond a single frame. Across a sequence of frames, lighting typically evolves in predictable
ways as cameras move and objects animate. Small changes in viewpoint often produce correspondingly small changes in illumination. This temporal coherence creates another opportunity for machine learning. Frame generation techniques exploit this relationship by learning how
rendered scenes evolve over time, allowing future frames to be predicted from previously observed data. Again, the objective is not to replace the underlying physics, but to maximise the usefulness of every physically derived sample.
WHY THIS MAKES COMMERCIAL SENSE Whenever AI enters a discussion, there is a tendency to imagine data centre scale infrastructure and enormous generative models. Fortunately, graphics presents a very different problem. The neural networks used for denoising, reconstruction and frame
generation are highly specialised. They do not need the scale or complexity of large language models because they are solving a much narrower task. Their job is to reconstruct lighting information, not to model human language or world knowledge. More importantly, their computational behaviour is predictable. Once trained, running inference through a graphics reconstruction
network requires a known quantity of work. The execution time can be characterised and engineered into a graphics pipeline. Ray tracing alone does not always offer that luxury.
September/October 2026 MCV/DEVELOP | 39
Page 1 |
Page 2 |
Page 3 |
Page 4 |
Page 5 |
Page 6 |
Page 7 |
Page 8 |
Page 9 |
Page 10 |
Page 11 |
Page 12 |
Page 13 |
Page 14 |
Page 15 |
Page 16 |
Page 17 |
Page 18 |
Page 19 |
Page 20 |
Page 21 |
Page 22 |
Page 23 |
Page 24 |
Page 25 |
Page 26 |
Page 27 |
Page 28 |
Page 29 |
Page 30 |
Page 31 |
Page 32 |
Page 33 |
Page 34 |
Page 35 |
Page 36 |
Page 37 |
Page 38 |
Page 39 |
Page 40 |
Page 41 |
Page 42 |
Page 43 |
Page 44 |
Page 45 |
Page 46 |
Page 47 |
Page 48 |
Page 49 |
Page 50 |
Page 51 |
Page 52