05 Fakultät Informatik, Elektrotechnik und Informationstechnik

Permanent URI for this collectionhttps://elib.uni-stuttgart.de/handle/11682/6

Browse

Search Results

Now showing 1 - 10 of 17
  • Thumbnail Image
    ItemOpen Access
    Cross-lingual citations in English papers : a large-scale analysis of prevalence, usage, and impact
    (2021) Saier, Tarek; Färber, Michael; Tsereteli, Tornike
    Citation information in scholarly data is an important source of insight into the reception of publications and the scholarly discourse. Outcomes of citation analyses and the applicability of citation-based machine learning approaches heavily depend on the completeness of such data. One particular shortcoming of scholarly data nowadays is that non-English publications are often not included in data sets, or that language metadata is not available. Because of this, citations between publications of differing languages (cross-lingual citations) have only been studied to a very limited degree. In this paper, we present an analysis of cross-lingual citations based on over one million English papers, spanning three scientific disciplines and a time span of three decades. Our investigation covers differences between cited languages and disciplines, trends over time, and the usage characteristics as well as impact of cross-lingual citations. Among our findings are an increasing rate of citations to publications written in Chinese, citations being primarily to local non-English languages, and consistency in citation intent between cross- and monolingual citations. To facilitate further research, we make our collected data and source code publicly available.
  • Thumbnail Image
    ItemOpen Access
    Advances in clinical voice quality analysis with VOXplot
    (2023) Barsties von Latoszek, Ben; Mayer, Jörg; Watts, Christopher R.; Lehnert, Bernhard
    Background: The assessment of voice quality can be evaluated perceptually with standard clinical practice, also including acoustic evaluation of digital voice recordings to validate and further interpret perceptual judgments. The goal of the present study was to determine the strongest acoustic voice quality parameters for perceived hoarseness and breathiness when analyzing the sustained vowel [a:] using a new clinical acoustic tool, the VOXplot software. Methods: A total of 218 voice samples of individuals with and without voice disorders were applied to perceptual and acoustic analyses. Overall, 13 single acoustic parameters were included to determine validity aspects in relation to perceptions of hoarseness and breathiness. Results: Four single acoustic measures could be clearly associated with perceptions of hoarseness or breathiness. For hoarseness, the harmonics-to-noise ratio (HNR) and pitch perturbation quotient with a smoothing factor of five periods (PPQ5), and, for breathiness, the smoothed cepstral peak prominence (CPPS) and the glottal-to-noise excitation ratio (GNE) were shown to be highly valid, with a significant difference being demonstrated for each of the other perceptual voice quality aspects. Conclusions: Two acoustic measures, the HNR and the PPQ5, were both strongly associated with perceptions of hoarseness and were able to discriminate hoarseness from breathiness with good confidence. Two other acoustic measures, the CPPS and the GNE, were both strongly associated with perceptions of breathiness and were able to discriminate breathiness from hoarseness with good confidence.
  • Thumbnail Image
    ItemOpen Access
    Knowledge distribution in German drama : an annotated corpus
    (2024) Andresen, Melanie; Krautter, Benjamin; Pagel, Janis; Reiter, Nils
  • Thumbnail Image
    ItemOpen Access
    Computational sentence‐level metrics of reading speed and its ramifications for sentence comprehension
    (2025) Sun, Kun; Wang, Rong
    The majority of research in computational psycholinguistics on sentence processing has focused on word‐by‐word incremental processing within sentences, rather than holistic sentence‐level representations. This study introduces two novel computational approaches for quantifying sentence‐level processing: sentence surprisal and sentence relevance. Using multilingual large language models (LLMs), we compute sentence surprisal through three methods, chain rule, next sentence prediction, and negative log‐likelihood, and apply a “memory‐aware” approach to calculate sentence‐level semantic relevance based on convolution operations. The sentence‐level metrics developed are tested and compared to validate whether they can predict the reading speed of sentences, and, further, we explore how sentence‐level metrics take effects on human processing and comprehending sentences as a whole across languages. The results show that sentence‐level metrics are highly capable of predicting sentence reading speed. Our results also indicate that these computational sentence‐level metrics are exceptionally effective at predicting and explaining the processing difficulties encountered by readers in processing sentences as a whole across a variety of languages. The proposed sentence‐level metrics offer significant interpretability and achieve high accuracy in predicting human sentence reading speed, as they capture unique aspects of comprehension difficulty beyond word‐level measures. These metrics serve as valuable computational tools for investigating human sentence processing and advancing our understanding of naturalistic reading. Their strong performance and generalization capabilities highlight their potential to drive progress at the intersection of LLMs and cognitive science.
  • Thumbnail Image
    ItemOpen Access
    Detecting protagonists in German plays around 1800 as a classification task
    (2018) Reiter, Nils; Krautter, Benjamin; Pagel, Janis; Willand, Marcus
    In this paper, we aim at identifying protagonists in plays automatically. To this end, we train a classifier using various features and investigate the importance of each feature. A challenging aspect here is that the number of spoken words for a character is a very strong baseline. We can show, however, that a) the stage presence of characters and b) topics used in their speech can help to detect protagonists even above the baseline.
  • Thumbnail Image
    ItemOpen Access
    Between welcome culture and border fence : a dataset on the European refugee crisis in German newspaper reports
    (2023) Blokker, Nico; Blessing, André; Dayanik, Erenay; Kuhn, Jonas; Padó, Sebastian; Lapesa, Gabriella
    Newspaper reports provide a rich source of information on the unfolding of public debates, which can serve as basis for inquiry in political science. Such debates are often triggered by critical events, which attract public attention and incite the reactions of political actors: crisis sparks the debate. However, due to the challenges of reliable annotation and modeling, few large-scale datasets with high-quality annotation are available. This paper introduces DebateNet2.0 , which traces the political discourse on the 2015 European refugee crisis in the German quality newspaper taz . The core units of our annotation are political claims (requests for specific actions to be taken) and the actors who advance them (politicians, parties, etc.). Our contribution is twofold. First, we document and release DebateNet2.0 along with its companion R package, mardyR . Second, we outline and apply a Discourse Network Analysis (DNA) to DebateNet2.0 , comparing two crucial moments of the policy debate on the “refugee crisis”: the migration flux through the Mediterranean in April/May and the one along the Balkan route in September/October. We guide the reader through the methods involved in constructing a discourse network from a newspaper, demonstrating that there is not one single discourse network for the German migration debate, but multiple ones, depending on the research question through the associated choices regarding political actors, policy fields and time spans.
  • Thumbnail Image
    ItemOpen Access
    Towards a resource for multilingual lexicons : an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
    (2026) Han, Lifeng; Mohamed, Najet Hadj; Rassem, Malak; Jones, Gareth J. F.; Smeaton, Alan F.; Nenadic, Goran
    In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with annotations of multi-word expressions (MWEs), named AlphaMWE. The MWEs include verbal MWEs (vMWEs) defined in the PARSEME shared task that have a verb as the head of the studied terms. The annotated vMWEs are also bilingually and multilingually aligned manually. The languages covered include Arabic, Chinese, English, German, Italian, and Polish, of which, the Arabic corpus includes both standard and dialectal variations from Egypt and Tunisia. Our original English corpus is taken from the PARSEME shared task in 2018. We performed machine translation of this source corpus followed by human post-editing and annotation of target MWEs. Strict quality control was applied for error limitation, i.e., each MT output sentence received first manual post-editing and annotation plus a second manual quality rechecking. One of our findings during corpora preparation is that accurate translation of MWEs presents challenges to MT systems, as reflected by the outcomes of human-in-the-loop metric HOPE. To facilitate further MT research, we present a categorisation of the error types encountered by MT systems in performing MWE-related translation. To acquire a broader view of MT issues, we selected four popular state-of-the-art MT systems for comparison, namely Microsoft Bing Translator, GoogleMT, Baidu Fanyi, and DeepL MT. Because of the noise removal, translation post-editing, and MWE annotation by human professionals, we believe the AlphaMWE data set will be an asset for both monolingual and cross-lingual research, such as multi-word term lexicography, MT, and information extraction.
  • Thumbnail Image
    ItemOpen Access
    Resources for Turkish natural language processing : a critical survey
    (2022) Çöltekin, Çağrı; Doğruöz, A. Seza; Çetinoğlu, Özlem
    This paper presents a comprehensive survey of corpora and lexical resources available for Turkish. We review a broad range of resources, focusing on the ones that are publicly available. In addition to providing information about the available linguistic resources, we present a set of recommendations, and identify gaps in the data available for conducting research and building applications in Turkish Linguistics and Natural Language Processing.
  • Thumbnail Image
    ItemOpen Access
    Sense through time : diachronic word sense annotations for word sense induction and lexical semantic change detection
    (2024) Schlechtweg, Dominik; Zamora-Reina, Frank D.; Bravo-Marquez, Felipe; Arefyev, Nikolay
    There has been extensive work on human word sense annotation, i.e., manually labeling word uses in natural texts according to their senses. Such labels were primarily created for the tasks of Word Sense Disambiguation (WSD) and Word Sense Induction (WSI). However, almost all datasets annotated with word senses are synchronic datasets, i.e., contain texts created in a relatively short period of time and often do not provide the creation date of the texts. This ignores possible applications in diachronic-historic settings, where the aim is to induce or disambiguate historical word senses or changes in senses across time. To facilitate investigations into historical WSD and WSI and to establish connections with the task of Lexical Semantic Change Detection (LSCD), there is a crucial need for historical word sense-annotated data. Hence, we created a new reliable diachronic WSD/WSI dataset ‘DWUG DE Sense’. We describe the preparation and annotation and analyze central statistics. We then describe a thorough evaluation of different prediction systems for jointly solving both WSI and LSCD tasks. All our systems are based on a state-of-the-art architecture that combines Word-in-Context models and graph clustering techniques with different hyperparameter settings. Our findings reveal that using the WSI task as optimization criterion yields better results for both tasks even when the LSCD task is the focal point of optimization. This underscores that although both tasks are related, WSI seems to be more general and able to incorporate the LSCD task.