<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Combining hermeneutic and computer based methods for investigating reliability of historical texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alptug Güney</string-name>
          <email>alptug.gueney@uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristina Vertan</string-name>
          <email>cristina.vertan@uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Walther v. Hahn</string-name>
          <email>vhahn@informatik.uni-hamburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Hamburg</institution>
          ,
          <addr-line>Hamburg</addr-line>
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>25</fpage>
      <lpage>34</lpage>
      <abstract>
        <p>Within the framework of the project HerCoRe1 we are analyzing two historical works from 18th century “History of Rise and Decay of the Otoman Empire” and “Description of Moldavia” (both written by Dimitrie Cantemir) and investigate them with regard to the historiography of its time. We evaluate the usage of sources by the author and also the reliability of his references. We also seek to shed more light on the motivation behind the writing-process of these works by taking into account the political and cultural dynamics of the time and the position of Cantemir within the Ottoman elite. To determine missing or incorrectly translated parts of the work, the German and English translations are also compared with a copy of the Latin manuscripts. This comparative approach serves also to discuss the causes of the (un)conscious mistakes and omissions in the translations. We are performing this study by means of hermeneutic and IT approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>historical documents</kwd>
        <kwd>uncertainty and vagueness annotation</kwd>
        <kwd>hermeneutics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        their request, he produced the two books which are the target of this
proposal:
- Descriptio antiqui et hodierni status Moldaviae, written in Latin, a
history of his country in which he describes not only pure historical
facts but also traditions, the language, as well as the political and
administration system. Local denominations and toponyms, as well
as names are written in Romanian with Latin script as his intention
was to demonstrate the Latin origin of his folk. The transcriptions
are not standardized and one retrieves for the same toponyms,
several name variations. Quotations as known today were very rare,
there is no bibliography. According to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], as there was practically
no consistent previous work about the region, Cantemir himself
was not particularly careful with indicating sources of knowledge.
The work is accompanied by a map, the first detailed cartography
of the region. The names on the map are in Romanian language.
The Latin original was translated for the first time into German,
and only later - at the middle of the XIXth century - into Romanian.
The Latin manuscript seemed to be lost for a long time, so that the
first Romanian translation was following the German one. The
German translation is containing editorial notes of the translator.
- Historia incrementorum atque decrementorum Aulae Othomanicae,
the history of the Ottoman Empire. In contrast to the previous work
about Moldavia, here Cantemir indicates very carefully the sources of
information. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] supposes the existence of previous works, known in
the western countries, behind this decision. This work was written also
on the request of the Academy in Berlin. Cantemir follows the same
principle: text in Latin, while the toponyms and local denominations
are written this time in Ottoman Turkish. Although there were already
some previous works about the Ottoman Empire, the novelty of his
approach is the quotation of Turkish sources. The reliability of these
sources is untrusted sometimes by Cantemir himself. The original
manuscript (or a copy of it) reaches the western world after Cantemir’s
death, carried by his son to London. Here, a first translation into
English is produced: The history of Raise and Decay of the Ottoman
Empire. The translator reinterprets the texts, probably also being confused
by the presence of Turkish information sources, which at that time
were perceived as completely unreliable. The Latin original remains
lost for centuries and is rediscovered only at the end of the XXth
century in the USA. Thus, the German translation is based on the English
one and inherits the same alterations, and presumably adds new ones.
The Romanian translations, in contrast, use the Latin versions. The last
translation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is being used in this research.
Until now there is no systematic study on the reliability of the text sources
in Cantemir’s works, nor the degree of alterations produced by the
translations of the two works.
      </p>
      <p>Given the fact that both works became standard reference for western
authors until the middle of XIXth century, it is expected that their reception
influenced also following historical material. There is no reprint/new
edition of his works in German or English. There are, however, several
reprints of the Romanian versions. Recent Romanian translations of Decriptio
Moldaviae are done after the original Latin manuscript.</p>
      <p>A lot of works were dedicated to the personality of Dimitrie Cantemir and
its perception in different parts of Europe. A study of the reliability and
consistency of the historical facts (as they are described in the latin copies) and
their translations is practically impossible to be done only with traditional
hermeneutic methods. One needs expertise at the same time in Latin, German,
English, Romanian, Turkish, to enumerate just the main languages used in the
two books, which additionally sum up to a quantity of about 1000 pages. Both
German editions are printed in “Fracture” (“black letter”) script, which
nowadays is very difficult to be read.</p>
      <p>Already in the 1920s it was demonstrated (by using only a selections of texts),
that the translations are not respecting the original all the time. E.g.
information sources indicated by Cantemir were omitted, because they seemed too
unreliable to the translator.</p>
      <p>
        In the XXth century researchers claimed that some of the sources, persons and
facts quoted by Cantemir were not existing at all (e.g. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
      </p>
    </sec>
    <sec id="sec-2">
      <title>But given the:</title>
      <p>•
•
•
geographic distribution of material (originals in libraries in USA and
Russia; translations and copies all across Europe; most part of the
quoted sources in Turkey),
the multilingual character of the materials to be investigated (Latin,
German, Romanian, English, Turkish at least) and</p>
    </sec>
    <sec id="sec-3">
      <title>The Quantity of data which has to be processed in parallel,</title>
      <p>no study about the reliability and consistency of the original and the
translations could have been performed until now.</p>
      <p>In the HerCoRe project we propose a mix of hermeneutic and IT-methods
in order to:</p>
      <p>• compare the Latin copies and the English and German translations,
identify translation mistakes or gaps (made by purpose or not),
search after the quoted works and identifiy related Ottoman sources,
analyse Cantemir‘s writing and discourse style,
assess the importance of the work in the Ottoman studies and
compare them with other works contemporaneous to Cantemir or
followup research about the Ottomans,
develop electronic resources which may be of use for follow-up work
about the Ottoman empire and the history of Balkans.</p>
      <sec id="sec-3-1">
        <title>2. Hermeneutic investigation</title>
        <p>The hermeneutic investigation concentrates on the identification of sources
quoted directly or indirectly by Dimitrie Cantemir, as well as the mentioned
places, persons, events and dates.</p>
        <p>The two works are very different with respect to the quotation style. While
in the “Description of Moldavia” the quotation sources are almost missing,
in the “History of rise and decay of Ottoman Empire” the author refers
explicitely to different sources. However, there is no quotation style like in
modern scientific works. Most references are real quotations or the author
indicates the source of quotation through syntactic phrases followed by a
reformulation of the semantic substance of a text section. Especially these
cases are subject to the hermeneutic investigation.</p>
        <p>By now we identified the main works quoted by Cantemir. These works are
available only in paper form and are written in Ottoman Turkish (with
Arabic alphabet) thus only a manual comparison can be performed.
This systematic comparison led to a very unexpected result: we observed
that linguistic expressions of certitude (e.g. “for sure”, “without any
doubt”) are not an unambiguous indicator of the reliability of the quotation.
E.g.: Cantemir is sure that all investigated sources mention 4 sons of Sultan
Bayazid. However, all reliable sources of the time mention that the sultan
had five sons.</p>
        <p>We do not know why these inconsistencies occur. One possible motivation
is the context in which he wrote the two books: in exile in St. Petersburg,
probably with few notes at hand, that he made in Istanbul. Whilst we
cannot find the reason of the inconsistency, the hermeneutic analysis showed
that a pure automatic annotation (searching for quotation marks) will not
help in the case mentioned above, as the semantics of the quotation mark
does not match the degree of reliability of the quoted information. This is
something which cannot even be inferred by automatic methods; at least
not at this stage, where documents in Ottoman Turkish are rarely digitised
and no linguistic tools are available for this historical language variant.</p>
        <p>A second part of the hermeneutic investigation concerns the collection of
persons, places, and domain specific concepts which are mentioned by the
author. An automatic identification is practically impossible, as names are
not standardised (e.g. for the city of Iasi in Moldavia we identified at least
12 writing variants).</p>
        <p>The third part of the hermeneutic investigation is concerned with the
identification of missing paragraphs in the German and English translation.
First result: all paragraphs written in the Latin original with Arab characters
were systematically omitted. This leads to misunderstandings in the two
translations.
3.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Computer-based approach</title>
        <p>
          Digital methods can facilitate analysis on the reliability of translations but
also of the historical facts claimed by the author [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. In order to be effective,
these methods must consider an intrinsic feature of all natural languages: the
ability of producing and understanding vague utterances. The project
HerCoRe aims at modelling and annotating five levels of vague assertions
1. the text uncertainty (uncertain readings, losses, translations,
multilinguality, etc.),
2. the linguistic vagueness (metonymies, vague adjectives,
comparatives, non-intersectives, hedges, homonyms,),
3. the author reliability (genres, time style, contemporary knowledge),
4. the factual uncertainty (range expressions, time expressions, geo
relations), and
5. historical change (named entities, abbreviations, meaning changes)
We develop an annotation formalism which allows for:
- the mark-up of different types of vagueness and its source; the
implementation of a set of inference rules for the combination of such
vague features to calculate an overall result of their reliability;
- the definition of a similarity measurement of the inferred results
obtained for the same queries on different translations. The system
architecture is presented in figure 1. It relies on annotation on 4 levels
(linguistics, lexical markers for vagueness/uncertainty, ontological
and factual/quotation markers).
        </p>
        <p>
          For the detection of linguistic vagueness we follow a multilingual
approach. We collected the above listed indicators in the three languages
involved in the project (Latin, German and Romanian). Based on [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] we
distinguish between:
- Vague quantifiers, e.g.: some, most of, a few, about, etc.
- Modal adverbs, e.g.: probably, possibly, etc.
- Verbs e.g.: to believe, think, prefer, assume etc.
- Lexical quotation markers , e.g. introduced by quotation marks or
verbs with explicit meaning (say, write, mention),
- Inexact measures and cardinals.
- Complex quantifiers
- Non-intersective adjectives
- Implicit syntactic clues: mainly verb moods such as
conditionaloptative for Romanian, conjunctive mood or past perfect/pluperfect
for Latin, all of them indicating a “counterfactive” or non-reality
(doubt, hear-say, possibility, etc.)
        </p>
        <p>The initial collections of linguistic indicators are enriched through synsets
in the corresponding Wordnets.</p>
        <p>The knowledge base backbone is ensured by a fuzzy ontology modelled in
OWL2. We distinguish between fixed concepts and relations (like
geographical elements: river, mountain, island) and notions for which several “contexts
can be defined. E.g. a geographical notion like “Danube” is within one
historical context a border of the administrative notion “Ottoman empire”, and in
another one the border to the so called administrative notion “Roman empire”.
The historical contexts are specified by further fuzzy data properties (e.g.
time, placement).</p>
        <p>4.</p>
      </sec>
      <sec id="sec-3-3">
        <title>A case study</title>
        <p>
          In “The History of the Growth and Decay of the Ottoman Empire (1734)”
Cantemir tells the story of a battle between the Moldavian Prince Ștefan and
the Ottoman Sultan Bayezid I. Cantemir does not give the exact date of the
battle, but one could think that this can be inferred from other details of the
text. Ștefan attacked the Ottoman camp in Rasboeni. According to the account
of Cantemir, in this first confrontation, Ștefan lost the battle and retreated to
his castle in Neamț. Then, after the inspiring speech of his mother in Neamț,
he drew his soldiers together and stroke back the Ottoman army twice in
succession. After the last defeat in Vaslui, Sultan Bayezid fled back to Edirne [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
In the same paragraphs, Cantemir gives in a footnote some information
among many other detailed and important details - about the Moldavian
Prince Ștefan, who fought two times against Bayezid I.:
“He overthrew the renown’d Matthias King of Hungary, and wrested from
him Transilvanian Alps […] His son Bogdan made Moldavia tributary to
the Turks.”
        </p>
        <p>This account of Cantemirs includes several problems, which can mislead
the reader and even a historian who does not have detailed knowledge about
the Ottoman and Romanian (Wallachian and Moldavian) history.</p>
        <p>Known and proven historical facts:
• There were two sultans with the name Bazeyid in the history of
Ottoman Empire Bayezid I and Bayezid II but only the first one had
the additional name “Yildirim” i.e. the Thunderbold. Cantemir
mentions exactly this appellative and not the numbering (I or II) so
we can exclude any typo or damaged spot in the manuscript. We
checked this information in all translations and the Latin facsimile.
• The reign time of Bayezid I is known for sure (according to
diplomatic text sources): 1389 – 1402.
• The frontiers of the empire at that time leaned already towards the
Danube River and the Ottoman Empire was yet neighbor to
Wallachia and Moldavia, thus a military confrontation with both
principalities is historically possible.</p>
        <p>During this time Wallachia was ruled by several princes: Mircea I
(1386-1395), Vlad I (1394-1397) and then again by Mircea I
(1397-1418).</p>
        <p>
          The Ottoman chronicles report on a battle in 1391 between the
Wallachian Prince Mircea and Bayezid I in a place called Arkaş (in
Romanian Rovine). According to the Ottoman historians, Bayezid
won the battle and Mircea recognized the Ottoman sovereignity
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>
          At the time of Bayezid Wallachia was ruled by Mircea I
(13861395), Vlad I (1394-1397) and then again by Mircea I (1397-1418).
In Moldavia there was just one Ruler called Ștefan (Ștefan I
13941399)
There was a Moldavian ruler Ștefan III (1457-1512) who defeated
the Ottomans in 1475 after a loss in Rasboieni, and who defeated
also the Hungarian King Matthias (known as Matthias Corvinus
[1443-1490]). Moreover, this ruler had a son called Bogdan who
made Moldavia tributary to the Ottomans. These facts are
confirmed by Ottoman chronicles [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Class Ruler and hasName ‘Bayezid’ and ‘has
additionalName Yildirim’ and hasBattlesIn some (Class
Principality and belogsTo some (Historical Region and (liesIn
value NorthDanube) -&gt; Moldavia and Wallachia
Class Ruler and was RullerOf exactly Moldavia and had
BattleWith value ‘Bayezid Yildirim’ -&gt; this will show all
rulers which fulfil the criteria.</p>
        <p>At a closer look there is a strong mismatch between Cantemir and all other
established chronicles, but a historian would not know here how to interpret
the text:
• Is it referring to a battle against Moldavia or Wallachia?
• Which Ruler opposed Bayezid Yildirim?
• Where took the battle place?</p>
        <p>An historian using only traditional methods (source inspection, reflection,
re-evaluation of text) will face here a bunch of unsure and contradictory
information, very difficult to resolve. An historian with less background
knowledge about the Romanian history will be tempted to interpret wrongly
the text section, either choosing the wrong rulers or the wrong place.</p>
        <p>The HerCoRe System aims at helping historians in their interpretation, and
suggests different reading paths. From the ontology and additional annotations
the following inferences are possible:
•</p>
        <p>Class Ruler and wasRulerof value ‘OttomanEmpire’ and
has Name Bayezid and has Additional Name Yildirim -&gt;
Bayezid Yildirim 1389-1402</p>
        <p>Continuing this queries to the ontology, or invoking a complex inference
chain the system will propose following solutions:
• The paragraph is about a battle in Wallachia in a place called
Rovine and against a ruler Mircea I. -&gt; contradicts Cantemir text:
(it is not a Moldavian king but a Wallachian). The user may infer
also that the place called ‘Rasboe’ by Cantemir might be the place
known as ‘Rovine’
• The paragraph is about a battle in Moldavia against the Ruler
Ștefan I, and Cantemir mistakes the information about Ștefan I for
Ștefan III
• The paragraph is about the battle in Răsboieni against the
Moldavian prince Ștefan III and the mismatch here is about the Sultan
(Bayezid I or Bayezid II). Thus, the battle place here called by</p>
        <p>Cantemir ‘Rasboe” would be ‘Răsboieni’.</p>
        <p>All this information will have attached a score indicating a degree of truth.
The latter one e.g. will have a lower score when introducing an additional
scoring parameter: the metadata. The metadata will say, that the mentioned
paragraph is within a chapter about the sultan Bayezid I, so it is less probable
that Cantemir mistakes the name of the Sultan.</p>
        <p>HerCoRe System does not aim at proposing a final solution. This decision
is left entirely to the user/researcher, who is the hermeneutic subject and also
can store the inference paths presented above as motivation for his choice.
5.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Conclusions and further work</title>
        <p>In this article we intend to show how hermeneutic and IT methods can be
combined in order to investigate the reliability of historical texts (original
and their translations). We show that a deep analysis can be performed only
by the combination of the two approaches. Current research focus on the
semi-automatic annotation and development of the ontology. Further work
concerns the implementation of the reasoner and the visualisation of results.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Remarks on cooperation</title>
        <p>In our project we need the cooperation between computer scientists,
linguists (Latin, German and Romanian) as well as researchers in turcology.
We cooperate also very close with the editor and translator into Romanian of
the two works. The cooperation already revealed to date interesting aspects:
- The simple availability of raw texts in digital form does not help. E.g. the</p>
        <p>German translation of the History of Ottoman Empire is available from
the German Text Archive. However, the text was digitized for
visualisation purposes. The usage of the underlying text versions leads to an
unsorted mixture of paragraphs written by the Cantemir, his side notes and
the comments of the translator. These parts, marked correspondingly in
the TEI version are melted in the .TXT version. We had to invest
additional work on separating these distinct parts.</p>
        <p>Mark-up in the editions have to be considered, as they can enrich the
text. However, first one has to know the semantics of the mark-up.
The ontological formalisation of the notions mentioned in the text was a
great help for the humanist researchers leading to a better reflexion of the
used notions.</p>
        <p>Many of the computational linguistics approaches had to be revised given
the particularities of the historical text.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Babinger</surname>
          </string-name>
          , Franz,
          <year>1927</year>
          ,
          <string-name>
            <surname>Die</surname>
          </string-name>
          <article-title>Geschichtsschreiber der Osmanen und ihre Werke</article-title>
          . Leipzig
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Dimitrie</given-names>
            <surname>Cantemir</surname>
          </string-name>
          ,
          <article-title>Istoria măririi şi decăderii Curţii othmane, 2 volume, editarea textului latinesc şi aparatul critic Octavian Gordon, Florentina Nicolae, Monica Vasileanu, traducere din limba latină Ioana Costa, cuvânt înainte Eugen Simion, studiu introductiv Ştefan Lemny</article-title>
          , Bucureşti,
          <source>Academia Română-Fundaţia Naţională pentru Ştiinţă şi Artă</source>
          ,
          <year>2015</year>
          .
          <source>ISBN 978-606-555-135-0 (978-606-555-136-7</source>
          ,
          <fpage>978</fpage>
          -
          <lpage>606</lpage>
          -555)
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lemny</surname>
          </string-name>
          , Stefan,
          <year>2010</year>
          , Cantemirestii
          <article-title>-Aventura europeana a unei familii princiare din secolul al XVIII-lea</article-title>
          , Polirom Publishing House.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Parmaksızoğlu</surname>
          </string-name>
          , İsmet (Ed.),
          <string-name>
            <surname>Hoca</surname>
            <given-names>Sadeddin</given-names>
          </string-name>
          ,
          <article-title>Tâc'üt-tevârih, Bd</article-title>
          . III, Ankara,
          <year>1979</year>
          ,
          <fpage>153</fpage>
          -
          <lpage>158</lpage>
          ; Abdülkadir Özcan, „Boğdan“, TDV İslam Ansiklopedisi, Bd. VI,
          <volume>269</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Pinkal</surname>
          </string-name>
          , Manfred, Semantische Vagheit:
          <article-title>Phänomene und Theorien</article-title>
          .
          <source>In: Linguistische Berichte 70</source>
          .
          <year>1980</year>
          . 1-
          <fpage>26</fpage>
          . und 72.
          <year>1981</year>
          . 1-
          <fpage>26</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Pinkal</surname>
          </string-name>
          , Manfred,
          <year>1985</year>
          <article-title>Logik und Lexikon: Die Semantik des Unbestimmten</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Unat</surname>
          </string-name>
          , Faik Reşit (Ed.),
          <string-name>
            <surname>Mehmed</surname>
            <given-names>Neşrî</given-names>
          </string-name>
          ,
          <article-title>Kitâb-ı Cihan-Nümâ, Bd</article-title>
          . I, Ankara,
          <year>1949</year>
          ,
          <volume>327</volume>
          ; Parmaksızoğlu, İsmet (Ed.),
          <string-name>
            <surname>Hoca</surname>
            <given-names>Sadeddin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tâc'</surname>
          </string-name>
          üt-tevârih, Bd. I, Ankara,
          <year>1979</year>
          ,
          <fpage>200</fpage>
          -
          <lpage>201</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Vertan</surname>
          </string-name>
          , Cristina and v. Hahn, Walther,
          <year>2014</year>
          ,
          <article-title>Discovering and Explaining Knowledge in Historical Documents</article-title>
          , In: Kristin Bjnadottir, Stewen Krauwer, Cristina Vertan and Martin Wyne (Eds.),
          <source>Proceedings of the Workshop on “Language Technology for Historical Languages and Newspaper Archives” associated with LREC</source>
          <year>2014</year>
          , Reykjavik Mai 2014, p.
          <fpage>76</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>