When a document-understanding model cannot read a passport, it does not always say so. Researchers at Shenzhen MSU-BIT University and the Indian Institute of Science found that models fine-tuned to pull fields out of identity documents will instead assemble an answer out of identities memorized during training, returning several correlated fields at once. The paper was posted on August 13 and is due at ACM Multimedia in November.

The team fine-tuned three open models, LLaVA-1.5-hf, Xgen-Phi3 and Idefics2, on identity cards, driver’s licenses and passports, then fed them inputs carrying little or no readable evidence. Trained on DocXPand-25k and shown an image with nothing legible in it, LLaVA-1.5-hf returned at least two of three private fields exactly, family name, given name or document number, in 83.5 percent of cases. The other two models leaked less on the same test, 21.5 percent and 18.9 percent.

The failure did not require a document. Given 1,000 unrelated photographs, LLaVA-1.5-hf produced exact field pairs from its training set 95.5 percent of the time, drawing on 32 distinct identities. Cropped faces from the identity documents, matched to the training distribution, produced a lower rate, 82.5 percent, but surfaced 55 identities.

Text opened a second route. The authors generated 1,300 prompts in three styles: ordinary questions, partly filled records that invited the model to complete the blanks, and direct requests for the document number belonging to a named person. Results split sharply by phrasing. Against LLaVA-1.5-hf trained on DocXPand-25k, plain questions worked best, drawing exact field pairs 95.3 percent of the time, against 81.8 percent for the named-person query and 28.4 percent for the fill-in-the-blank format. The same three probes against the same model trained on the cleaner IDNet returned nothing at all.

What distinguishes this from earlier work on memorization is the mechanism. Extraction is meant to be grounded in the image. With the image uninformative, the models fell back on field relations learned in training, so what came out was not one wrong guess but a set of fields that belonged together.

Both datasets are synthetic. DocXPand-25k is built from fictitious templates with artificially generated names, dates and faces, and IDNet is likewise generated rather than collected, so no real person’s details were exposed. The fine-tuning was also heavy, 5,000 samples over 15 epochs on a single A100, conditions that favor memorization.

Existing defenses traded one problem for another. Measured on a single high-risk pair, given name and document number, the untreated model leaked on 64.2 percent of image probes. Gradient ascent drove that to zero but took the model’s score on the extraction task itself, the job it was built for, from 0.852 down to 0.119. SCRUB, the strongest of six unlearning methods tested, kept extraction close to intact and still leaked on 4.9 percent.

The authors’ own method targets pairs of fields rather than single ones, on the argument that the risk lies in fields being recoverable together. They reported 0.1 percent leakage on image probes and zero on prompt probes with extraction intact, and released their benchmark, DocPrivacyBench, with the code.

Models surfacing what they absorbed has been a running theme. Researchers who decoded 315,320 encrypted reasoning blocks returned by OpenAI, Anthropic and Google APIs pulled 62 live API keys out of them. A single email can plant a false memory in an assistant that carries into later sessions. Identity checks are themselves being rebuilt around AI, with vendors pitching palm biometrics to bind agents to verified humans.

Passport and license scanning is already routine in bank onboarding, border control and gig-economy signup, where a blurred photograph is unremarkable.

Sources: arXiv, DocXPand-25k

–
By the Control Plane Editorial Team