My scholarly development in Corpus Linguistics has not progressed as a series of separate activities, but rather as a continuous process moving from the strengthening of disciplinary foundations toward technological integration, industrial application, and institutional development. This roadmap is based on my academic experience, international collaborations, and future research directions. Each subsequent phase does not replace the preceding one; rather, it builds cumulatively on the achievements of earlier phases to enable broader contributions.
The initial phase was marked by the establishment of conceptual foundations in Corpus Linguistics and language technology. During this period, I gained my initial understanding of the relationship between language and technology under the supervision of Prof. Nam Jee Sun at Hankuk University of Foreign Studies. During this phase, my understanding of Corpus Linguistics was combined with early experience in developing language-processing tools, including translation systems, language pedagogy, and basic corpus-based automatic morphosyntactic annotation (Prihantoro, 2011). This experience shaped my view that the study of language cannot be separated from language computing and digital technology (Berry, 2012).
The second phase expanded these foundations through doctoral study at Lancaster University, particularly through CASS and UCREL. During this period, interactions with leading scholars in Corpus Linguistics, including Prof. Tony McEnery, Prof. Paul Baker, and Dr. Andrew Hardie, strengthened both my methodological and technological expertise. In addition to developing my research capacity, this phase resulted in innovations in automated morphological analysis of natural language texts, including SANTI-morf (Prihantoro, 2022b, 2024b; Silberztein, 2016), which served as an important foundation for the development of more independent corpus technologies.
The third phase is characterised by the integration of research, technological development, and industrial application. Innovations such as parameter files for POS tagging and lemmatisation (Prihantoro, 2025b) have been adopted in international platforms such as Sketch Engine, CQPweb, and LancsBox, thereby extending the impact of my work from academic contexts to global practice. During this phase, I also began serving as a consultant for BRIN and the Language Development and Cultivation Agency. The development of CORTEX represents a significant milestone in this phase, marking a transition towards corpus systems developed within the national academic environment while maintaining a global orientation. Although this phase places strong emphasis on technology development for industrial and institutional applications, scholarly knowledge development remains central, as demonstrated by publications contributing to the advancement of knowledge (Prihantoro, 2024a, 2025a; Prihantoro & Gillings, 2025; Prihantoro & Ishikawa, 2025; Prihantoro et al., 2025).
The fourth phase is directed towards institutional strengthening through the development of an academic ecosystem grounded in Digital Humanities. The primary objective of this phase is to integrate research, education, and the development of interdisciplinary academic programmes connecting the humanities, technology, and data science. One of the main objectives is to develop master's and doctoral programmes in Digital Humanities at Universitas Diponegoro.
This phase will also involve the development of adaptive curricula that respond to technological change and evolving industry needs. Thus, the roadmap will extend beyond software development towards the establishment of a scholarly ecosystem in which Corpus Linguistics serves as a centre of integration for research, education, and innovation in Indonesia.