05 Fakultät Informatik, Elektrotechnik und Informationstechnik

Permanent URI for this collectionhttps://elib.uni-stuttgart.de/handle/11682/6

Browse

Search Results

Now showing 1 - 10 of 70
  • Thumbnail Image
    ItemOpen Access
    Supervised semantic proximity noise and disagreement detection
    (2024) Choppa, Tejaswi
    The quality and reliability of annotated data are crucial for the development of Ma­chine Learning models. In this work, we particularly focus on word sense annotation in context (a.k.a. Word-in-Context, WiC). WiC datasets in real-world contexts of­ten exhibit significant disagreement. As a result, information is lost when instances are discarded during the creation of the gold label by adjudicating the annotations through majority or median judgment. Recent advancements have sought to ad­dress this issue by incorporating disagreement data through novel label aggregation methods (Uma et al., 2022). Modeling this disagreement is important because, in a real-world scenario, we often do not have clean data. We need to predict on samples where high disagreement is expected and which are inherently difficult to categorize. Predicting disagreement can help detect or filter highly complex samples. Through this thesis, we aim to build machine learning models that predict human disagreement in annotated text instances. Moreover, we focus on data with noise instances where annotators cannot confidently assign a label or the data does not fit predefined categories. We aim to measure both disagreement and noise, as they both stem from a common source: ambiguity. By modeling these aspects, we aim to design modeling approaches that predict not only the semantic proximity label but also the annotator disagreement, as well as data noisiness.
  • Thumbnail Image
    ItemOpen Access
    Exploring the effects of enriched English language input on language model efficiency
    (2024) Zeller, Tom
    Recent years have seen the advent of large-scale language modeling as exemplified by transformer-based models like GPT or variants of the BERT architecture. These models, which are trained on massive datasets and using compute unattainable by actors that are not of the scale of the biggest tech companies, have shown impressive feats of syntactic and semantic understanding. Naturally, interest has risen in making these models more efficient, in terms of compute as well as data requirements. Research in this area can be seen as primarily motivated by two factors: reducing the barrier for smaller actors like research institutes or end consumers to train and execute state-of-the-art models, as well as reducing the carbon footprint of these models. To achieve this goal, model compression techniques like quantization, pruning or distillation are utilized. This work aims to explore a different, less model-centric and more data-centric approach: Modifying the training and inference data, by enriching it with syntactic and semantic information. To this end, a lexical resource is created which maps English words to a form where individual characters represent values of a range of semantic and syntactic features, providing lexical information that is accessible to all model types that operate on tokens at the sub-word or character-level. Different features and methods of representation are discussed, and their effect on model performance is evaluated by pretraining a small GPT-family model and fine-tuning on downstream tasks of the SuperGLUE benchmark. Given a fixed amount of data and compute, the experiments show a performance advantage for a character-level model trained using the enriched data.
  • Thumbnail Image
    ItemOpen Access
    Comparison of distributional and visual nearest neighbors
    (2025) Naber, Sven
    This thesis investigates how semantic concepts are represented across textual and visual embedding spaces, focusing on the abstract-concrete continuum. Using 5,448 English nouns and their embeddings from both distributional language models (e.g., Word2Vec, GloVe) and vision models (e.g., ViT, DINOv2, CLIP), it compares neighborhood structure via a normalized alignment score (NAS). Results show that alignment is primarily driven by input modality rather than model architecture, with strong local overlap for concrete concepts and more diffuse agreement for abstract ones. Mean aggregation of image embeddings improves visual consistency but cannot fully bridge modality-specific limitations. The findings provide a starting point for further exploration of semantic spaces.
  • Thumbnail Image
    ItemOpen Access
    More reliable retrieval augmented generation for domain-specific question-answering through domain-infused soft prompts
    (2025) Nassar, Zeina
    Transformer-based large language models (LLMs) have revolutionized the field of artificial intelligence, enabling advancements in various applications due to their exceptional reasoning capabilities. However, these models often face challenges such as hallucination, outdated knowledge, and non-transparent reasoning, which limit their reliability in critical tasks like question answering (QA). QA systems, especially in domain-specific contexts like car manuals, require high accuracy and reliability to ensure user safety. Achieving this is complicated by the need for precise retrieval and interpretation of domain-specific information, which can be hindered by unfamiliar keywords or ambiguous questions. Open-domain QA relies on external knowledge repositories and typically follows a retriever-reader framework to locate and process evidence. In contrast, domain-specific QA often lacks sufficient gold-standard datasets, making it essential to explore techniques like retrieval-augmented generation (RAG), which combines retrieved context with the question as input to LLMs. While RAG improves grounding, hallucinations still occur when models rely on pre-trained knowledge over retrieved evidence. This thesis investigates the role of prompting techniques in improving faithfulness and reducing hallucinations in domain-specific QA. By testing domain-specific and domain-agnostic discrete prompts as well as soft prompting methods, this work aims to identify strategies for generating more accurate and grounded responses. The study addresses key research questions on the effectiveness of domain-specific information in prompts, the best way to incorporate such information, and whether dynamic soft-prompts outperform static ones in domain-specific QA scenarios. The findings aim to contribute to building more reliable and factual QA systems.
  • Thumbnail Image
    ItemOpen Access
    Plug-and-play domain adaptation for neural machine translation
    (2023) Kadiķis, Emīls
    Neural machine translation has emerged as a powerful tool, yet its performance heavily relies on training data. In a fast-changing world, dealing with out-of-domain data remains a challenge, prompting the need for adaptable translation systems. While fine-tuning is a proven effective adaptation method, it is not always feasible due to data availability, memory, and computational constraints. This thesis introduces a dynamic plug-and-play method inspired by controllable text generation to enhance machine translation across various domains without fine-tuning. This method, called Plug-and-Play Neural Machine Translation (PPNMT), uses a mono-lingual domain-specific bag-of-words to push the hidden state of the decoder through backrpopogation, making the output more in-domain. The method is tested on two types of domains: formality, gender (where the source language does not make a distinction between these aspects, but the target language does), and fine-grained technical domains (which are more based on topic inherent in the text on both the source and target sides). The method performs reasonably well for adapting the translation to different formality levels and, to a lesser extent, grammatical genders, even with an incredibly simple bag-of-words. However, it struggles with adapting the model to technical domains, and a fine-tuning baseline outperforms the proposed method in anything but very low few-shot settings in all tried domains. Despite that, the method shows some interesting behaviour, adapting to the formality on a level that goes beyond just using formal pronouns.
  • Thumbnail Image
    ItemOpen Access
    Exploring retrieval-augmented language modeling for material prediction of vehicle components
    (2024) Wagner, Frederik
    Jüngste Fortschritte im Bereich natural language processing (NLP), insbesondere bei großen Sprachmodellen (large language models, LLMs) wie ChatGPT, zeigen das Potenzial für ihre Anwendung bei einer Vielzahl von Aufgaben in speziellen Domänen. In der Automobilbranche könnten sie beispielsweise zur Unterstützung bei der Reparatur eines Fahrzeugs eingesetzt werden. Diese Arbeit befasst sich mit dem Problem der Vorhersage geeigneter Materialien für Fahrzeugkomponenten, wie z. B. Bremsscheiben. Es soll ermittelt werden, ob LLMs sowohl auf Allgemein- als auch auf domänenspezifisches Wissen zurückgreifen können, um genaue Vorhersagen über Komponentenmaterialien zu treffen, ohne dass eine umfangreiche Feinabstimmung (fine-tuning) erforderlich ist. Erreicht wird dies durch retrieval-augmented generation (RAG), wobei relevante Informationen aus externen Quellen abgerufen und zur Verbesserung der Modelleingabe des LLMs verwendet werden. In dieser Arbeit werden drei Ansätze verglichen: ein Standard-LLM-Modell, ein einfacher RAG-Ansatz und eine iterative RAG-Methode namens Chain-of-Verification (CoVe). In dieser Arbeit wird auch ein eigenes Annotationstool entwickelt, um eine menschliche Evaluierungsstudie zu erleichtern, da es keinen Goldstandard-Datensatz gibt. Die Ergebnisse zeigen, dass LLMs bei der Materialvorhersage gut abschneiden, und obwohl beide RAG-Ansätze die Vorhersagequalität nicht signifikant verbessern, verschlechtern sie sie auch nicht. Diese Forschungsarbeit kommt zu dem Schluss, dass LLMs mit oder ohne Retrieval-Ergänzung eine vielversprechende Lösung für die Materialvorhersage bei Fahrzeugkomponenten bieten, auch wenn es noch Herausforderungen bei der Bewertung, der Hyperparameter-Optimierung und dem Daten-Retrieval gibt.
  • Thumbnail Image
    ItemOpen Access
    Gender bias in dependency parsing
    (2023) Go, Paul Stanley
    Recent high-profile advances in natural language processing (NLP) have spurred interest into identifying and rectifying socially harmful problems common in NLP systems such as gender bias. Unfortunately, many works which attempt to tackle the issue of gender bias suffer from methodological deficiencies such as the assumption of a binary and immutable concept of gender. We scrutinize one such work which found gender bias in dependency parsing and evaluate if the claims have merit. Our results were inconsistent with the gender bias findings of that paper, and further investigations through error analysis and treebank analysis revealed methodological flaws which artificially introduced differences between their female and male data sets. Mistakes made during preprocessing compromised the outcome; therefore, their results do not prove the existence of gender bias in dependency parsing. Through our findings, we suggest a different methodology for identifying and alleviating syntactic bias that is more inclusive for everyone-no matter their gender.
  • Thumbnail Image
    ItemOpen Access
    Automatic classification of abstractness in English rigid nouns
    (2023) Saponaro, Alberto
    The main difference between (i) Mass-Count Languages (such as English) and (ii)Classifiers Languages (such as Chinese) is that (i) encode the information about nouns’ countability in their grammar and (ii) employ a classification system of classifiers to distinguish between individuals or substance. If the mass-count distinction is a characteristic of mass-count language, the substance-individuals denotation seems to be a concept universally available for all humans. Another concept that appears to be universally accessible and linked to the countability status of English nouns is the notion of abstractness. Then, mass nouns usually refer to an abstract object, and this is confirmed from the distribution of abstractness in the dataset. This thesis’ objective is to provide a model for the classification of rigid nouns (count or mass only) that is capable to generalize on the degree of abstractness. Additionally, it tests if a model trained with the same set of features is capable of rating the abstractness of those nouns. To accomplish these tasks, several sets of features are being identified based on syntactic and semantic properties of nouns that describe the mass-count distinction. The results indicate that the first model M1, a mass-count classifier that predicts the countability class of a rigid noun, provides reliable predictions and can generalize on the degree of abstractness of the targets. The second model M2, an abstractness rate predictor that assigns an abstractness rate from 1 to 5 to a rigid noun, is incapable of providing reliable ratings and cannot generalize on the countability status of the targets. A third model M3, an abstract-concrete (binary) classifier that predicts the abstractness class of a rigid noun, provides reliable predictions and can generalize on the countability status of the targets. Given that those results concerns rigid nouns only, further research can be conducted by examining the abstractness of elastic nouns. However, there is the need of an annotation that rates abstractness of nouns senses.
  • Thumbnail Image
    ItemOpen Access
    Multimodal OCR post-correction on German historical documents
    (2023) Wu, Nianheng
    Optical Character Recognition (OCR) post-correction is essential to digitalizing historical documents, increasing transcription accuracy, and reducing manual effort. Previous works often handle this as a text-to-text translation problem. However, the orthography of many languages, including German, has evolved across centuries, leading to many "irregular" spellings. Thus, a text-only system would face many uncertainties. Therefore, combining image features with text should be meaningful. The rise of large-scale pretrained models has brought new opportunities in this field. In this work, I will: 1) Introduce a dataset that includes historical German documents from 1783 to 1903 based on Deutsches Textarchiv with aligned golden transcription, OCR-ed textline, and their corresponding textline image; 2) Present a multimodal OCR post-correction system that combines CLIP image encoder, a pretrained image feature model, with ByT5, a byte-based language model. According to my experiments, this model outperforms the state-of-the-art text-only model.