Unknown writing systems are deciphered using statistical distribution and bilingual inscription analysis.
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
5 sources for · 0 against
Peer-reviewed literature demonstrates that the decipherment of unknown writing systems involves both statistical distribution analysis and comparative matching against known language corpora and adjacent writing systems.
This paper discusses the possible use of unconventional algorithms on analysis and categorization of the unknown text, including documents written in unknown languages. Scholars have identied about ten famous manuscripts, mostly encrypted or written in the unknown language. The most famous is the Voynich manuscript, an illustrated codex hand-written in an unknown language or writing system. Using carbon-dating methods, the researchers determined its age as the early 15th century (between 1404-1438). Many professional and amateur cryptographers have studied the Voynich manuscript, and none has deciphered its meaning as yet, including American and British code-breakers and cryptologists. While there exist many hypotheses about the meaning and structure of the document, they have yet to be conrmed empirically. In this paper, we discuss two dierent kinds of unconventional approaches for how to handle manuscripts with unidentied writing systems and determine whether its properties are characterized by a natural language, or is only historical fake text.
This paper introduces a novel method for addressing the challenge of deciphering ancient scripts. The approach relies on combinatorial optimisation along with coupled simulated annealing, an advanced technique for non-convex optimisation. Encoding solutions through k-permutations facilitates the representation of null, one-to-many, and many-to-one mappings between signs. In comparison to current state-of-the-art systems evaluated on established benchmarks from literature and three new benchmarks introduced in this study, the proposed system demonstrates superior performance in enhancing cognate identification results.
Keywords: ancient script decipherment, combinatorial optimization, k-permutations, coupled simulated annealing, evaluation benchmarks status released display-pdf yes is-olf no is-manuscript no is-preprint no is-journal-matter no is-scanned no is-retracted no Received 2025 Feb 21; Accepted 2025 Apr 23; Collection date 2025. 1 Introduction Numerous ancient scripts around the world remain undeciphered, with many of them dating back millennia. The challenges in deciphering these scripts stem from factors such as insufficient inscriptions, the absence of known language descendants utilizing these scripts, and uncertainty about whether the symbols truly form a writing system.
In literature, numerous contributions address these subproblems, offering computational methods tailored to each, frequently focusing on a particular script. The sequential tasks typically involve: (a) determining if a set of symbols genuinely constitutes a writing system, followed by (b) devising procedures to segment the symbol stream into individual signs. Subsequently, (c) reducing the set of signs to the minimal collection for the given writing system, thereby forming the alphabet (or syllabary, or sign inventory), and identifying all allographs.
This task proves intricate due to variations introduced by scribe writing styles and the evolution of symbols over time, complicating the identification and management of allographs. In addressing this challenge, Skelton ( 2008 ) and Skelton and Firth ( 2016 ) applied phylogenetic systematics to the realm of writing systems. Their focus was particularly on Linear B, a pre-alphabetic Greek script. Through this method, they scrutinized the evolution of the Linear B script over time, taking into account scribal hands as an additional source of variation. This application showcased the efficacy of phylogenetic analysis in understanding the development of writing systems. Born et al.
1.5 Define signs values and match sign sequences with a known language Every contemporary endeavor to decrypt ancient scripts using computational tools relies on contrasting a missing script or language wordlist with words from a deciphered and known language.
In this system, a recurrent Neural Network (NN) is employed to establish the mapping between lost and known signs, despite the advantage of using contextual information to perform the task, it lacks the adaptability necessary for addressing two practical decipherment challenges. Firstly, paleographers often possess partial knowledge about the mapping of certain signs, and this information needs to be incorporated into the system. Secondly, real inscriptions are frequently broken or damaged, leading to unreadable signs, requiring the incorporation of uncertainty into the system, potentially through the use of wildcards or other special symbols.
We reduced the number of non-cognate words, creating a dataset of 2,214 cognate words, 1,119 unpaired Ugaritic words, and 1,108 Old Hebrew words without corresponding cognates. These words were randomly selected from the dataset proposed in Snyder et al. ( 2010 ). Linear B/Mycenaean Greek - LB/MG . Linear B, a syllabic writing system employed for Mycenaean Greek dating back to approximately 1450 BC. Luo et al. ( 2019 ) curated a dataset by extracting pairs of Linear B and Greek words from a compiled lexicon, eliminating some ambiguous translations and resulting in 919 cognate pairs.
Inside round parentheses, the maximum Accuracy value obtained in our experiments is indicated. † Results for NeuroCipher computed or recomputed by us simulating a real setting and using the code in Luo et al. ( 2019 ). ‡ To enable the system to converge toward meaningful results we had to provide the number of cognates in the dataset, information not available in real settings. Our system exhibits superior accuracy compared to any other work across almost all benchmark datasets, with a substantial margin.
(b) Access to an extensive cognate list is crucial, yet in most real cases, only two word lists are available for matching, without any assurance that cognates from the lost language truly exist in the lexicon of the known language. (c) In natural language processing (NLP), evaluations are typically conducted on well-established test beds and the studies discussed earlier focused on well-known correspondences to demonstrate system effectiveness. On the contrary, testing these systems on real cases involving unknown writing systems and their corresponding languages presents an entirely different set of challenges and uncertain comparanda.
We aim to contribute insights that may finally address longstanding problems unresolved for centuries. Funding Statement The author(s) declare that no financial support was received for the research and/or
4 https://github.com/ftamburin/EditDistanceWild 5 https://github.com/structurely/csa 6 The Tower of Babel, https://starlingdb.org . Data availability statement The datasets and codes presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found at: https://github.com/ftamburin/CSA_OptMatcher . Author contributions FT: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Software, Validation, Writing – original draft, Writing – review & editing.
The Phaistos Disk is an ancient artifact from Crete. At each side of the disk, a series of unknown signs is written along a spiral. Professional archaeologists expect that we will only learn what it is until similar objects are found. A statistical analysis in this article shows what it is not: it is not a one‐dimensional text, since there are relations between the signs in adjacent windings of the spiral. Three patterns of such relations have been identified. A Monte Carlo simulation of one of them has been performed, using a model of the spiral form. It is concluded that the probability of this pattern being coincidental is small, well below the conventional threshold.
Oracle bone inscriptions (OBIs) are ancient Chinese scripts originated in the Shang Dynasty of China, and now less than half of the existing OBIs are well deciphered. To date, interpreting OBIs mainly relies on professional historians using the rules of OBIs evolution, and the remaining part of the oracle’s deciphering work is stuck in a bottleneck period. Here, we systematically analyze the evolution process of oracle characters by using the Siamese network in Few-shot learning (FSL). We first establish a dataset containing Chinese characters which have finished a relatively complete evolutio
Everything we examined (5)
This check searched the claim as stated. It did not run a separate search for evidence against it.