Ermolaev

Original photograph by Vladimir Petrovich Ermolaev: riders in the mountain landscape of Tuva

Original historical photograph by Vladimir Petrovich Ermolaev, reproduced from the digital collection of the National Archives of the Republic of Tuva. The visible print texture is representative of the material sources that informed this project.

Ermolaev is a PyTorch image-to-image model that translates contemporary photographs toward the tonal, textural, and print character learned from paper copies of photographs associated with Vladimir Petrovich Ermolaev’s photographic legacy in Tuva.

Portrait of Vladimir Petrovich Ermolaev

Historical context

Vladimir Petrovich Ermolaev (28 July 1892 – 1 April 1982) was an ethnographer, journalist, researcher of Tuva, photographer, and the first director of the Tuvan State Museum. He first visited Tuva as a photographer in 1913 and later documented its people, daily life, political history, transport, industry, culture, and wartime experience. Collections of his photographic albums are preserved by the National Archives of the Republic of Tuva.

The model is named in recognition of this photographic heritage. A longer English-language historical note is available in ARTICLE.md. Primary reference: National Archives of the Republic of Tuva — 130th anniversary article. The public archival catalogue is available at fond.gosarhivrt.ru.

What the model learned

The historical training domain was created from paper photographic copies, not original negatives or glass plates. Consequently, the learned style reflects the material qualities of prints and later copies: visible paper texture, compressed tonal range, print softness, surface irregularity, fading, and copy-generation artefacts. The model should not be interpreted as a physically exact simulation of Ermolaev’s camera, film stock, negatives, or darkroom process.

Dataset author and model developer: Valeriy Aldyn-oolovich Irgit.

Examples from the CUT training report

These snapshots come from the original CUT training report. Input is a contemporary-domain photograph and Translated output is the generated historical-print-style image. They are qualitative examples, not a benchmark. The report used different samples at different epochs.

Report snapshot Input Translated output
Epoch 18
Epoch 19
Epoch 20

Architecture and training

  • Method: Contrastive Unpaired Translation (CUT), unpaired A→B translation.
  • Generator: six-block ResNet, 32 base filters, instance normalization, RGB→RGB.
  • Training domains: 500 contemporary-domain images and 898 historical-print-domain images.
  • Training preprocessing: resize to 1152 px and random 1024 px crop.
  • Batch size: 1.
  • Objective: LSGAN with PatchNCE (lambda_NCE=0.7, layers 0,4,8,12, 128 patches).
  • Continued-training learning rate: 5e-5, two constant-rate epochs plus one decay epoch.
  • Recommended checkpoint: continued epoch 2. Its logged average losses were G_GAN=0.4369, G=1.5346, and NCE=1.0978. Loss values are training diagnostics and do not establish perceptual quality.

The model is based on Contrastive Learning for Unpaired Image-to-Image Translation by Park et al. The original CUT implementation is distributed under BSD-style terms; its license is included as CUT_LICENSE.txt.

Available weights

  • weights/ermolaev_epoch2_recommended.pth — recommended continued checkpoint.
  • weights/ermolaev_epoch3_latest.pth — final continued checkpoint.
  • weights/ermolaev_base_epoch20.pth — base training checkpoint represented by the epoch-20 report.

Usage

pip install torch numpy Pillow huggingface_hub
git clone https://huggingface.co/tuva/Ermolaev
cd Ermolaev
python inference.py input.jpg output.png

The script uses CUDA when available and otherwise runs on CPU. Use --max-side 1024 if memory is limited. To select another checkpoint:

python inference.py input.jpg output.png --weights weights/ermolaev_epoch3_latest.pth

Intended use

The model is intended for artistic style transfer, education, experimentation, and exploration of historical photographic aesthetics. It may be useful for visual studies and creative projects when its generated nature is clearly disclosed.

Limitations and responsible use

  • The model can change faces, small objects, writing, clothing details, and culturally meaningful attributes.
  • Output is generated imagery and is not historical evidence, restoration, or a faithful reconstruction of an archival scene.
  • The learned domain includes the texture and degradation of paper copies; it does not reproduce the appearance of original negatives or glass plates.
  • Performance varies with resolution, subject, colour, and modern photographic processing.
  • The training collection is not included. Users must follow the archive’s conditions and applicable rights when consulting or reusing archival images.
  • Do not use generated images to falsify historical records or misrepresent people and events.

Data provenance

The historical visual domain was assembled from paper copies connected with the photographic collections of Vladimir Petrovich Ermolaev held by the National Archives of the Republic of Tuva. The archive catalogue is publicly available for study. No claim is made that the public catalogue grants unrestricted reproduction rights. The repository distributes trained weights and selected model outputs, not the historical training collection.

License

See LICENSE.md. The custom license permits use and redistribution of the model with attribution while preserving third-party rights in photographs, archival records, and personal likenesses.

Downloads last month
28
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for tuva/Ermolaev