Scanning is not digitization

Digitizing a photo archive is not the same as scanning it. I learned that the practical way, with a historical photography archive that my father, the cinematographer Giorgi Gersamia, spent a lifetime building. Part of it had already been digitized and published through the museum’s earlier website. The physical archive is much larger: photographs, negatives, documents and decades of accumulated information.

When we started digitizing the old negatives, the first images appeared quickly. So did the real problem.

Watch: Scanning Is Not Digitization on YouTube

A scanned photograph is a file, not an archive

A physical photograph can be scanned and uploaded online and still remain practically invisible. Nobody can find it unless they already know it exists. To make a digital archive searchable, each file needs a layer of structure around it:

  • Who is in the photograph?
  • Where was it taken?
  • When?
  • What theme does it belong to?
  • Which collection is it part of?

Without that metadata, you have digital files. With it, you have a system people can search, explore and reuse.

What we built first: one record, one public page

The first job was not choosing an AI tool. It was deciding what we actually had, how it should be structured, and how it could become accessible. That led to two layers that work together:

  • A structured record for every photograph, held in Sanity. This is the single source of truth: the image, its description and its metadata live in one place.
  • A public page on the WordPress site, in Georgian and English, generated from that record.
The Sanity record for the photograph View of Kutaisi (1870s), with English and Georgian titles, next to the published bilingual page on photomuseum.ge
One photograph, one record: the Sanity entry for “View of Kutaisi (1870s)” (left) and the public page it produces on photomuseum.ge (right).

AI helps with the slow parts: drafting descriptions and metadata, translating them, and producing captions. A person reviews every record before it is published. The AI drafts; people decide what goes live.

Why structure comes before volume

It is tempting to scan everything first and organise later. In practice, adding structure afterwards means touching every file twice. Setting up the record format first means each new photograph arrives ready to be found, and publishing it is a review step rather than a project.

Digitizing the physical collection is ongoing work, separate from building the system. The system is built to keep pace with it, not to replace it.

The same problem outside archives

This is not only a heritage problem. Developers have years of project photography sitting in folders. Consultancies have case work no one can find. Manufacturers have catalogues that exist only as PDFs. In each case the material exists; the structure around it doesn’t. Three questions are a good start:

  1. What would someone need to know to find this item without asking you?
  2. Where does the single, correct version of each item live?
  3. Who checks an item before it is published, and how long does that take?

This is the kind of system we build as Content & Publishing Systems. You can see the result at photomuseum.ge and read the full Photomuseum case study.

More in this series

Similar Posts