Translation methodology
Source text
The Russian text is a digitised copy of the printed volumes made by optical character recognition (OCR). OCR introduces errors — broken words, wrong letters, garbled Latin — and sometimes mixes figure captions into the text.
Cleaning
Before translation, each article is separated from its neighbours, printed tables of contents and caption debris are removed, and the author's signature and bibliography are extracted.
Translation
Articles are translated paragraph by paragraph with a large language model instructed to correct obvious OCR errors, expand the encyclopedia's abbreviations, keep Latin and Greek terminology as written, and translate faithfully without adding or omitting content. Personal names are given in their established English spelling or in standard transliteration.
Each translation is checked automatically for missing paragraphs, untranslated Cyrillic text and abnormal length, and samples are reviewed by hand. Machine translation can still contain mistakes.
What is not translated
Bibliographies are not reproduced. Figures are reproduced from the scanned volumes with their original Russian labels.