
On September 25, 2026, I joined OpenCV Live! episode 226 as a guest of Dr. Satya Mallick and Phil Nelson to talk about something we have been quietly doing at Click-Ins for more than a decade: training production vision models almost entirely on rendered data. This article is an extended version of that conversation. It follows the structure of the talk, includes the material I did not have time to cover live, and ends with the audience Q&A.
▶ Watch the full episode: Deploying Synthetic Data for Real-World Inspections – OpenCV Live! 226
A standard sedan is parked on a city street in broad daylight. To the human eye it is undeniably there. Run the same scene through a state-of-the-art object detector - any recent version of YOLO - and the vehicle registers as nothing at all. The cause is a glitchy, psychedelic camouflage wrap that looks as though a corrupted file had been printed on vinyl and stretched over the chassis. We generated that car, the wrap and the damage on it synthetically, and I confirmed the day before the show that YOLO26 still does not see it.
That image opened my talk because it represents the outer limit of what computer vision has to deal with in the physical world. Dealing with that level of intentional obfuscation requires rethinking how vision models are trained in the first place - and that is what Click-Ins is about.
Click-Ins is an AI vehicle-inspection platform for insurers, rental fleets and dealerships with a strong focus on vehicle damage detection. There is no hardware: the user receives a link, follows a guided photo session on an ordinary phone, and the system returns a condition report within moments. Every damage, down to minor scratches, is segmented, mapped to the specific panel, measured in absolute units (centimetres or inches, using photogrammetry rather than AI), positioned on a body-style map and graded for severity. The result is available both as an interactive HTML breakdown and as structured data through our API.
Customers in this space demand an unprecedented level of accuracy in an extremely noisy domain. Bounding boxes are not enough: when a scratch spans two panels, each panel's share of the damage must be measured separately, and adjacent undamaged parts must not be touched. That requires fine instance segmentation. And because no model is ever 100% accurate, we never replace the human inspector - we empower them with a one-click toggle to remove anything the AI got wrong.
Before synthetic data can be the solution, you have to feel the problem. A camera interprets three-dimensional space purely through captured light and contrast. It produces a 2D image with no depth and no inherent 3D interpretation. Mathematically, the specular highlight on a chrome bumper, or the reflection of a nearby tree, can look identical to a dent or a deep scratch in the clear coat. It is like trying to spot a friend in a funhouse mirror.
During the show I showed real images sent through our system: a panel covered in an adjuster's crayon markings that had to be dissected from the actual damage; a car-sharing livery whose texture camouflaged a dent and a scratch at the bottom of the door; reflections through which the real damage had to be read. Dr. Mallick raised an even subtler case - a scratch on a wall, reflected on the car's surface - and I can add a real one from our own experience: inspections usually happen in parking lots full of other damaged cars, and the dents on the grey car parked next to yours end up reflected on your paintwork. Chase every such pattern with real photographs and you have an endless cycle of acquisition and annotation, performed by annotators who cannot know the ground truth unless they are standing next to the vehicle.
The phrase I use for all of this is the domain gap: the physical world is not a controlled laboratory.
Click-Ins was founded in 2014 as an anti-fraud solution for insurers, and later pivoted to vehicle inspections. Our synthetic data has gone through three generations.
Generation 1 was largely artificial. I built it myself with zero knowledge of 3D or simulation. It worked anyway: a simple object-detection model trained on those crude samples generalised well enough to find the major panels of a real vehicle. Generation 2 introduced more realistic scenes but was still rudimentary and still aimed at object detection. When customers made clear that they needed fine segmentation, we raised funding, hired a professional 3D team, and built Generation 3: full instance segmentation with photorealistic rendering, physically based materials and real HDRI lighting. A professional adjuster cannot tell our current renders from photographs.
Two of those terms deserve a definition I did not have time for on air. Physically based materials describe a surface by its real optical properties - roughness, metalness, specularity, clear-coat parameters - so that light interacts with the render the way it interacts with the actual paint or polymer. An HDRI is not a flat JPEG skybox dropped behind a model. It is a 360-degree, floating-point image captured at a real location that emits accurate light data into the 3D scene, so every shadow, reflection and patch of ambient light in the simulation follows the physics of the place where it was captured.
And yet realism itself is not the goal. The goal is a model that learns and performs well on real images.
This leads to the question any ML practitioner will ask: if the synthetic data simulates physics perfectly, why would a model trained on it still fail in deployment? Because excellent synthetic data does not automatically produce an excellent model. When you measure the distribution of synthetic images against the distribution of real ones, there is always a gap - and the essential question is what kind. A data gap means you need to simulate more, or different, data. A model gap means you need to stop and review your algorithms. (I will not go into our model architecture; parts of it are patented. But the model must be designed to ingest and learn from synthetic data.)
I should also be frank: we do not use 100% synthetic data. It is impossible at this point. Most of our data is synthetic, and we use real data for very specific patterns we cannot simulate or reproduce.
To validate these gaps we have partnered for several years with LatticeFlow, a Swiss AI-governance company. Their platform lets us reason about an image as a data point rather than a picture. A neural network does not see a car the way we do; it compresses the pixels into a high-dimensional vector - an embedding. Map those vectors into a latent space and similar images cluster together: all the front views land in one region, for example.
Now overlay the embeddings of real-world test images on the embeddings of synthetic training images. Where the real points sit and the synthetic points do not, you have found an environment the model does not understand yet. In our case, a cluster of real images stood out that was absent from the synthetic set. Phil spotted the common factor before I revealed it: they were all taken indoors, inside garages, workshops and showrooms. Slicing the evaluation set to indoor images only - the platform lets you mark positive and negative examples to sharpen the slice - showed the model's performance dropping by 17%.
The fix was not to render millions of images. It was to render enough images with the right variety: vehicles under harsh fluorescent tube lighting, on concrete floors, surrounded by typical workshop clutter. When we added that corpus, the new synthetic points joined the real-world distribution gracefully, and the model improved by 19.7% on real indoor images.
The more interesting number is that performance on outdoor images also improved, by 6.3%. On air I summarised this as "the model was overfitting on outdoors". The mechanism deserves a fuller explanation, because it is the heart of why synthetic data works. A model trained only on scratches lit by direct sunlight is prone to shortcut learning: it memorises the 2D pixel pattern of a sunbeam hitting a scratch rather than learning what a scratch is structurally. Show it the same physical scratch under fluorescent tubes, and again under a soft diffused showroom spotlight, and you deny the network that shortcut. It is forced to isolate the invariant feature - the three-dimensional deformation itself - from the environmental variables around it. Deploying synthetic data is not about generating infinite random images; it is about strategically controlling variables to guide what the network learns.
Not every gap is environmental. The LatticeFlow platform also surfaces automatic slices, and one of them showed the model degrading to 45% accuracy. The slice was one vehicle: the Citroën C4 Cactus, with its distinctive textured polymer "Airbump" panels. We had already rendered many Cactuses, so the quantity of data was not the problem.
For this we built an internal tool that anyone can reproduce without buying commercial software. We take a frozen DINOv3 backbone - a foundation model whose weights we never fine-tune, used purely for its generalised understanding of the visual world - and break real photographs into a grid of patches, building a bank of real-world patch features. When we analyse a synthetic render, we patch it the same way and use k-nearest neighbours to measure the distance in feature space between each synthetic patch and the closest real patch. Where the distance is too large, the patch is flagged. The result is a realism heat map, and for the Cactus the entire Airbump panel lit up as unreal, even though the render looked perfectly convincing to a human.
A natural objection is: why not just train a discriminator, as in a GAN, to tell real from fake? In fact we use both, because they answer different questions. A trained discriminator asks, "Have I ever seen anything like this in the entire real corpus?" - and it is notorious for latching onto a single tell, a dead pixel or an artifact, and stopping its evaluation there, the classic path to mode collapse. The frozen-feature kNN asks a different question: "Does this specific texture exist anywhere in my database of physical reality, and how far is it from the nearest thing that does?" It trains no weights and exploits no loophole; it is purely a distance. One is a lie detector that buzzes; the other points to the specific bead of sweat on the forehead.
The heat map then goes to the people who can fix it. Our 3D artists saw the problem immediately: the shader rendered the polymer with incorrect specularity, computing light bounces as if the panel were smooth, glossy plastic. The real Airbump is matte, covered in micro-normals - microscopic surface variations that scatter light. The team rebuilt the material from reference images, adjusting roughness, clear-coat parameters and micro-normal detail, re-rendered, and ran the result through the same procedure. The heat maps were gone. Generating a million more images would never have solved this; the underlying physics was wrong, and only domain expertise could correct it at the source. Of course, the next anomaly will pop up shortly. That is the job.
Dr. Mallick asked the obvious question: why maintain a 3D team at all instead of using a generative approach? The answer is pixel-perfect annotation. Our environment has a CAD component that knows the exact geometry of every part, and that is what annotates every image automatically. A generative model reconstructs the image; it does not preserve the original geometry, the positions of panels or the details we are deliberately simulating, so the labels no longer match the pixels. We have been working with diffusion models since well before the current hype - one of the PhDs involved in the development of our algorithm studied under the professor who invented Stable Diffusion - and every version we tried distorted reality enough to break the annotations.
Where generative AI shines is as an augmentation layer on top of physically accurate renders. After a long effort to make sure the models leave dimensions, panel positions and every engineering aspect of the vehicle untouched, we use them to add the ad-hoc chaos of the real world: dirt, water streaks, sensor noise, even the ruler an adjuster leaves in the frame. Two caveats: the top-tier models that deliver decent results are expensive enough that it rarely pays off at scale, and anything you engineer around a model's quirks is obsolete when the next model ships two months later.
A rule of thumb from our experience: where you would need roughly a million annotated real images, you need perhaps a hundred thousand synthetic ones. Synthetic data generalises because it is abstract and controlled - every scenario can be simulated quickly, including the rare Mercedes or the specific car-sharing shuttle with exactly the damage you need, which you may never find in a real dataset. For the edge cases we still annotate real data, with a dedicated, highly experienced annotation team and a QA team as a guardrail so that bad labels do not leak into training.
Dr. Mallick put the principle better than I did: the information is in the variance of the data, not in the number of images. A million copies of the same image have zero variance. Had we rendered a million unrealistic Cactuses, we would have burned GPU time and team time, overfitted to a wrong pattern, and never addressed the real one. Sometimes a few dozen realistic images in different colours and scenes are enough for a model to generalise.
This brings us back to the opening slide, and to what I think is an underrated selling point of synthetic data. What if the real world is designed to fool the detector?
Researchers create adversarial textures by running object detection in reverse. Using gradient descent against the convolutional filters of a target detector, they compute the pixel pattern that maximises its error, then tile it into an endlessly repeatable print. Applied to a vehicle on vinyl, it suppresses the detector's response from almost any angle - an invisibility cloak for AI. The patterns I showed were published by three different research groups under permissive licences. The original paper targeted YOLOv3; on my machine, YOLO26's largest model still registers nothing on a wrapped car. From some viewpoints, where wheels or antennas are prominent, generic detectors do fire - but as a train, a truck, an oven, even a kite. Never a sedan.
The same image through the Click-Ins pipeline yields a clean instance segmentation of every panel, the windows and the localised damage on the bumper. The reason is architectural. Because we began as a counter-fraud solution, we designed the system from day one on the assumption that it would be attacked. And because we own the 3D pipeline and the CAD model, an adversarial texture is simply something to map onto the model and render across thousands of lighting, weather and viewpoint variations. Think of it as a biological immune system: a model that has only ever seen normal paint has no antibodies for this virus. By rendering synthetic adversarial variants, you synthesise the vaccine. The camouflage loses its efficacy because the network learns the structural geometry of the car beneath the pattern instead of relying on surface-level colour contrast. Robustness is fundamentally a data problem. An adversarial wrap is just another slice to render, simulate, label and train.
I keep a folder on my computer called "exotic patterns". Whenever I travel and see an unusual livery, I photograph it and run it through the system.
Something I only touched on live: the core of our system is an ontology of the domain. The visual simulation produces images, but the annotations come from the ontology, and at inference time we reason over the same ontology to suppress hallucinations. It is a comprehensive approach, not just "render millions of images".
The automatic annotation itself is worth understanding mechanically, because it is the real scalability advantage. In traditional machine learning, annotators spend countless error-prone hours drawing boxes and masks over photographs. In a rendering pipeline the labels are deterministic. The engine built the scene, so it has perfect knowledge of every pixel: in the same pass that produces the RGB image, it emits the depth map, the object-ID pass and the part-ID pass. It never has to guess where the bumper ends and the tyre begins, because they are distinct mathematical objects inside the simulation. Pixel-perfect segmentation masks and classifications arrive instantly, with zero human intervention. If you do choose to generate images with generative AI, make sure you have tooling to annotate them automatically — that is not straightforward.
Because the CAD system and renderer are domain-agnostic, the same pipeline that models a dented polymer bumper on a Citroën processes any asset supplied as a 3D file. We have applied it to boats, to two-wheelers with exposed and complex engine components, to aircraft, and to aerial and satellite imagery of infrastructure. The principles are identical.
How do you pick meaningful images and avoid wasting training time? Control the distribution. Because the data is simulated, you have full command over makes, models, colours, conditions and environments - track every parameter and manage the distribution to match the geography you deploy in, not an arbitrary 50/50 split. Analyse embeddings in different latent spaces, watch for new, under-represented clusters, and measure accuracy per slice. This is not a human task at scale; you need tooling, commercial or your own. And because new vehicles reach the market constantly, you are always chasing - our system detects panels on cars not yet in production precisely because of synthetic data. Start with a decent number, measure, and add only what the measurement demands.
I can only train on synthetic data but will be tested on real photos. Any advice? If you can get real data from the same industry, even if not from your customer, use DINOv3 plus kNN to measure the distance between your synthetic and real distributions and establish a baseline. If you have no real data at all, it is trial and error: train, measure, watch for shortcuts (loss collapsing suspiciously fast is a tell), augment, and iterate until you are confident the remaining gap is not a model gap. Dr. Mallick's addition: beg the client for even a hundred representative images including corner cases, because the real world will land you in situations beyond your imagination. Mine: offer to work with anonymised data - in our domain we blur plates, faces, reflections of faces, and reflections of reflections of faces for customers who require it.
Do camera specifications matter for training data? Absolutely. You must simulate sensors: aperture, barrel and radial distortion, atmosphere, specularity, ambience, diffusivity, glossiness. This is where 3D simulation outperforms generative AI - simulated data uses geometry and HDRI so the physics of light and reflection is preserved, while image-generating models have no knowledge of physics and will happily produce reflections that do not match the backplate. Focus the sensor simulation on your actual input (mobile phones, DSLRs or stationary cameras), but cover a wide range within that. Everything matters, which is why this technology takes twelve to fourteen years to build.
How accurate are human adjusters? One European customer measured a pedantic professional with 30–40 years of experience: not above 80% on damage recognition. We give humans a lot of credit.
We are running out of real data, and annotating insane quantities of it is becoming more complicated and more expensive every year. Simulating the physical world is becoming the next standard for training data - and it should be as realistic as the real thing.
Let me leave you with the question I had planned to close with. Rendering engines now compute photon bounces, micro-normals and sensor noise profiles with growing fidelity; the simulation is approaching a one-to-one mapping of physical reality. At what threshold does a model gain a deeper, more robust understanding of the physical world by studying the pristine mathematics of a simulation rather than the noisy, error-ridden data of the world itself? If a simulation offers absolute control of every physical variable together with flawless deterministic annotation, the simulated environment may soon become the definitive reality for training intelligent systems. It redefines what we consider ground truth.
My sincere thanks to Dr. Satya Mallick, CEO of OpenCV, and Phil Nelson, Director of Content and Creative at OpenCV, for the invitation, for hosting the conversation, and for the sharp questions that made it better. Thank you to the OpenCV project and its global community - OpenCV is a 501(c)(3) non-profit that has powered computer vision for two decades, including much of our own work at Click-Ins.
The original recording, Deploying Synthetic Data for Real-World Inspections – OpenCV Live! 226, is available on the OpenCV YouTube channel: https://www.youtube.com/live/QpowZtYO7sE. Adversarial patterns shown in the presentation were published by their respective researchers under permissive licences.
Questions about synthetic data for your own inspection or QA problem? I am always happy to compare notes: dima@click-ins.com.