The surface record.

A painting has a texture no copy of it shares: the weave under the ground, the ridge a loaded brush leaves, the crack that opened in 1904. This lights that texture from twelve positions on the phone screen, measures which way each point of the surface has to be tilted to explain what the camera saw, and commits twenty four patches of the result to a signed hash. Come back later, capture again, and it tells you whether the thing in front of you is the same physical object. A picture hung on a wall and the face of a statue work as readily as something lying on a table, because the lamp and the camera travel together: you turn the screen toward the thing, whichever way that happens to be, and it talks you through the aiming.

What this can and cannot do. It can tell you that two captures came from the same physical surface, and it can prove that a record has not been edited since it was signed. It cannot prove who made the object, when, or that the record was made before a dispute began, since the signature is your own and the clock is your phone's.

It also cannot tell two supports apart when the only thing it can see is the support. Two canvases cut from one bolt carry the same weave, and if a surface is smooth enough that the weave is all there is, both will match. The verdict below reports that case rather than hiding it.

Records on this phone.

stored locally, never uploaded

How you want to light it.

the choice that decides everything else
the swept light is better by a wide margin
a ball bearing, a marble, any glossy ball

Setting up the screen lamp.

only needed for the screen method
changes the spoken guidance, not the physics
260 mm
6.30 px/mm · not yet calibrated

Lay any bank card on the screen and drag the second slider until the amber rectangle is exactly as wide as the card. Every card in the world is 85.60 mm across, so that one gesture converts the screen into millimeters. Without it the lamp positions are guesses and the geometry is wrong on every phone by a different amount.

Why the light should be a separate thing you carry. the swept light

The screen is a poor lamp, and the reason is geometry rather than brightness. It can only shine from where the glass is, so with the camera at the top of the panel every direction it offers lies inside a cone about twenty degrees wide. Recovering a surface normal from a cone that narrow means inverting a badly conditioned matrix, which multiplies sensor noise by a factor of about a hundred and buries the very texture being measured. No amount of care with the screen fixes that.

So put the phone down and pick up a torch. Stand the phone facing the object, use the back camera, which is the sharp one and focuses at any distance, and move a separate light around the object by hand. Now the light can come from thirty degrees above the surface or from almost edge on, from any side, and raking light at eight degrees is exactly what a conservator uses to make a canvas weave or a brush ridge leap out. On the simulator the light set comes back with a condition number of 2.6, against 143 for the screen.

The difficulty is knowing where the light was. Your hand does not report its position. The answer is the one museums have used since the early two thousands, and it is charmingly cheap: put a glossy sphere in the frame beside the object. A sphere reflects the light as a single bright point, and that point sits wherever the sphere's surface happens to face halfway between the camera and the lamp. So the position of the highlight on the ball is a direct readout of the light direction, frame by frame, with no calibration and no assumptions. The technique is called highlight reflectance transformation imaging. A ball bearing or a marble will do.

The sphere is also a ruler. Its diameter is known, so the number of pixels it spans fixes the scale of the whole image, and the capture is cropped so that the object lands at exactly four pixels per millimeter every time. Two sessions taken from different distances then compare directly, with no scale search and nothing to go wrong.

What it costs. You have to stand the phone up, or lean it, or prop it against something, because the camera must not move during the sweep. You have to work in a dim room, as before. You have to hold the light off for a moment at the start so it can take a dark frame. And you have to own a shiny ball.

In return the measurement gets far better. Driven through this interface by a simulated camera, two completely different sweeps of the same object agree to a median correlation of 0.977 with every one of the twenty four sites clearing the bar and every one agreeing on the transform. Move the camera six millimeters between sessions and it still reads 0.981, and the shift comes back as twenty four pixels at four pixels per millimeter, which is six millimeters. Rotate it three degrees and it reads 0.983 and recovers 2.8 degrees. A different painting scores 0.059. The screen lamp, measured the same way, manages 0.83 against 0.03: workable, but with a tenth of the margin.

How a screen becomes a lamp, and what the camera gets back. the screen method

Photometric stereo is old and simple. Light a surface from a known direction, and a point tilted toward the light is brighter than a point tilted away. Do it from several directions and you can solve for the tilt at every pixel. The usual apparatus is a rig of lamps on a hemisphere. A phone already has a large flat emitter and a camera pointing the same way, so laying the phone face down over the object turns it into that rig, with the lamp being a white disc drawn wherever you like on the glass.

Twelve positions, not eight. The camera sits at the top of the screen and the object sits directly under it, so every lamp position is somewhere below the lens. That is a one sided light set, and it is the central difficulty here. The twelve positions are pushed to the far edges of the panel in both axes to spread the directions as widely as the glass allows. Even then the spread across the short axis is a few degrees, which is why the tool does not pretend to recover a true surface normal.

What is actually solved. Rather than invert the light matrix, which on this geometry amplifies sensor noise by about a factor of ten, the tool forms two weighted sums of the exposures, using the horizontal and vertical components of each lamp direction as the weights, and divides by the mean exposure. The result is a shading gradient: large where the surface slopes, signed by which way it slopes, and independent of how dark or light the paint is. It is the same information a normal map carries, without the step that destroys it.

Six frames per position. The signal is a texture perhaps one part in fifty of the exposure, so it sits close to the sensor noise. Averaging six consecutive frames at each lamp position halves the noise, and a band pass that keeps structure between about half a millimeter and two millimeters, while discarding both the lighting falloff above it and the pixel noise below it, does the rest. Those two steps are the difference between a fingerprint and a random number.

Two dark frames. One before the sequence and one after, averaged and subtracted, which removes whatever the room contributes and lets the drift between them be measured. If the room changes during the capture, the meter says so.

Aiming a screen you cannot see, and measuring the distance instead of guessing it.the aiming

The screen is the lamp, so it always faces the object, which means it always faces away from whoever is holding the phone. There is no arrangement of this instrument in which you watch a preview while you aim. That is as true lying flat over a table as it is turned toward a wall. So the aiming happens out loud: the tool measures the frame two or three times a second, says what is wrong with it, sounds a tone that rises as the framing improves, and starts the capture itself once the frame has been good twice running. You never have to look at it.

The distance measures itself. It used to be a slider you set by eye, and that guess propagated into every light direction and so into the whole solve. It does not have to be a guess. A lamp at offset y from the lens, lighting a patch of surface directly under the lens at distance d, delivers

E(y) = d / (y² + d²)^(3/2)

which is the inverse square law times the cosine of incidence. Flash a lamp near the lens, then one at the far end of the panel, and take the ratio. The albedo cancels, because both flashes light the same patch of the same paint, and the ambient has already been subtracted. What is left contains nothing but the distance:

R = (E_near / E_far)^(2/3) d² = (y_far² − R·y_near²) / (R − 1)

On this geometry that ratio runs from about 1.68 at 200 mm to 1.24 at 320 mm, a wide and easily measured swing. The reading shows in the corner while aiming and is what actually gets used, with the slider left as a fallback for when the measurement fails.

Squareness is the same trick sideways. Compare a lamp at the left edge of the panel against one at the right. On a surface facing the lens squarely the two are equal, and any yaw makes the nearer one brighter, which gives the angle. That matters because a phone held at an angle foreshortens the surface, changing the scale across the frame, and scale is the one error the matching cannot absorb. It waits until you are within twelve degrees and says which way to turn.

Steadiness comes from the accelerometer, using only the variance, so it does not depend on knowing which way up the phone is. A hand that has stopped moving reads a constant acceleration whichever axis gravity falls along.

The arithmetic, written out.the maths

A lamp drawn at screen position p sits, in millimeters, at an offset from the camera lens of (p - camera) / pxPerMM, at a distance d from the object. The direction from the object to that lamp is

L = normalize( (px - cx)/s , (py - cy)/s , d )

with s the calibrated pixels per millimeter. The image is mirrored on the way in, which flips the horizontal axis exactly once more than the face down phone already did, so screen x and image x agree and the signs carry over unchanged.

For a Lambertian surface of albedo rho and unit normal n, exposure k reads I_k = rho (n . L_k). Summing the exposures weighted by the light directions and dividing by their mean gives

g = ( sum_k L_k I_k / N ) / ( sum_k I_k / N )

The albedo cancels in that ratio, and so does any per frame exposure change, because both appear identically above and below the line. What is left depends only on the normal. Its horizontal and vertical parts are the two channels that get hashed.

The band pass. Subtracting a wide box blur from a narrow one keeps a specific range of spatial scales:

b = blur(g, 5 px) - blur(g, 20 px)

At a typical working distance one pixel is about a third of a millimeter, so that window holds the weave, the brush ridge and the craquelure, and drops both the smooth falloff of the lamp and the single pixel noise of the sensor.

Sites. The frame is divided into a six by four grid and the highest energy point in each cell becomes a site. Each site stores a 56 by 56 patch of both channels, standardized to zero mean and unit variance and quantized to one signed byte per sample, which is 6,272 bytes a site and 150 kilobytes a record.

Matching. Normalized cross correlation between the stored patch and the new capture. Correlation is invariant to any change in brightness or contrast, which is what makes it survive a different room and a different exposure.

Why the score is read at one place, and not at each site's best place. the false accept problem

The obvious way to score a re-capture is to let every site hunt around its stored position for the offset that correlates best, and report those numbers. That is wrong, and it is wrong in the direction that matters: it accepts objects it should reject.

Searching a window of plus or minus 34 pixels in steps of two, across seven rotations, is about 4,000 chances per site to find a high correlation. Take the best of 4,000 draws from a null distribution and the result is not the null any more. Measured on the simulator: a genuinely different object scored a median of 0.234 when each site was read at one fixed place, and 0.413 with 83 percent of sites above the per site threshold when every site was allowed its own search. The second number would have passed.

The fix is consensus. Sites still search, but the search only casts a vote for one global transform, a shift and a rotation. The transform the most sites agree on wins, and then every site is read at that one transform, whether it likes it or not. A real match has all twenty four sites agreeing, because they are all looking at one rigid object. An impostor's best offsets scatter, because there is nothing there to agree about.

That is why the verdict reports three numbers rather than one: the median correlation, the fraction of sites above the per site threshold, and the fraction of sites that voted for the winning transform. All three have to clear their bar. The third is the one that stops a lucky search, and it is also the one that catches a repeating weave, because a periodic pattern offers many equally good offsets and the votes split between them.

What the hash commits to, and what it does not. the cryptography

Twenty four site patches become twenty four leaves. Each leaf is hashed with a byte that marks it a leaf, the parameters of the capture, the site index and its coordinates, then the patch bytes. Internal nodes are hashed with a different marker byte, so no internal node can ever be confused for a leaf. An odd node at any level is promoted rather than duplicated, which closes the malleability that duplication opens. The root is signed with ECDSA on P-256 by a key this browser generated and keeps.

This is a commitment, not a reproducible fingerprint. The patches come from a physical measurement with noise in it, so a second capture never produces the same bytes and never produces the same root. Anyone who tells you a hash of a photograph identifies an artwork is selling something. What the root does is fix the enrollment beyond later editing: it proves the twenty four patches being compared against today are the ones committed to on the day of recording, and the signature proves they came from this key.

So verification runs in two independent halves, and both are reported. The record is recomputed from its own patches and checked against its stored root, and the signature is checked against the stored public key. That half is exact, and either passes or fails. Then the new capture is correlated against those patches. That half is statistical, and comes with the three numbers above.

Two kinds of export. The full file carries the patches and can be imported on another phone, which is the only way a record survives this one being lost, and the only way somebody else can check the object. The public file carries the root, the signature and the public key, and nothing else. Publishing the patches would hand a forger the target to sculpt toward, so the public file is what you would attach to a certificate and the full file is what you would keep in two places.

Where this fails, measured rather than guessed. limits

Gloss. Varnish, glass and any wet finish reflect the lamp instead of scattering it, and a mirror tells you about the lamp rather than about the surface. Matte and semi gloss work. The capture meter reports clipped pixels, which is what a specular highlight looks like from here.

Focus. Front cameras on most phones are fixed focus with a near limit somewhere around 200 to 250 millimeters, and blur removes exactly the band that carries the fingerprint. The distance slider starts at 260 mm for that reason, and the meter runs a sharpness estimate and refuses to record below a threshold. If your phone can focus closer, the geometry improves and you should use it.

LiDAR will not help. It is the obvious thing to reach for and it is the wrong instrument twice over. A phone LiDAR resolves depth to a few millimeters over a sparse grid of points, while the fingerprint here lives between a third of a millimeter and two millimeters of relief, so it is three orders of magnitude short of seeing anything that matters. And it is not reachable anyway: Safari on iOS still exposes no WebXR in 2026, so no web page on an iPhone or an iPad can read the depth sensor at all. Where depth is reachable, on Android through WebXR, it is coarse for the same reason. The one thing depth would have been useful for is knowing how far away the object is, and a sixteen millimeter ball bearing does that better, on every phone.

A shared support. Two canvases from one bolt share a weave, and a weave is periodic and strong. The consensus test catches most of that, because a periodic pattern splits the vote, but an object whose only texture is its weave is genuinely not distinguishable from its sibling. The verdict flags a periodic surface when it sees one.

Room light. The screen has to out light the room, so this needs the dimmest room available. The meter reports the ambient level and the drift between the two dark frames.

Numbers. The figures above are for the swept light. Those below are the screen lamp, kept because it needs no apparatus at all. Both come from a simulator that renders a synthetic painted surface with sensor noise, sRGB encoding and exposure drift, feeds it to this page as a camera, and drives the real interface. Under those conditions the same object re-captured scores a median correlation of 0.811, and 0.821 when shifted by sixteen pixels, and 0.810 when shifted and rotated four degrees. In all three every one of the twenty four sites agreed on the transform, and the transform came back exact: the sixteen pixel shift was recovered as sixteen pixels and the four degree rotation as 4.1 degrees. Two different objects score 0.024 and 0.028, with one site in twenty five agreeing. A different painting on the same bolt of canvas is separated by the agreement test rather than by the correlation, which is the case the free scoring described above gets wrong.

That is a calibration against a model, not against fifty real paintings. A model gets the physics right and the world wrong: it has no dust, no varnish, no hand tremor and no phone that decides to change its exposure. Until somebody runs the real experiment, treat the numbers as indicative of the method and not as a false accept rate.