Institut für Maschinelle Sprachverarbeitung Universität Stuttgart Pfaffenwaldring 5B D-70569 Stuttgart Master thesis Active Learning Strategies for Deep Learning Based Question Answering Models Kuan-Yu Lin Studiengang: M.Sc. Computational Linguistics Examiners: Prof. Dr. Ngoc-Thang Vu Dr. Antje Schweitzer Supervisors: Maximilian Schmidt Start of the work: 01.08.2023 End of the work: 31.01.2024 Erklärung (Statement of Authorship) Hiermit erkläre ich, dass ich die vorliegende Arbeit selbstständig verfasst habe und dabei keine andere als die angegebene Literatur verwendet habe. Alle Zitate und sinngemäßen Entlehnungen sind als solche unter genauer Angabe der Quelle gekenn- zeichnet. Die eingereichte Arbeit ist weder vollständig noch in wesentlichen Teilen Gegenstand eines anderen Prüfungsverfahrens gewesen. Sie ist weder vollständig noch in Teilen bereits veröffentlicht. Die beigefügte elektronische Version stimmt mit dem Druckexemplar überein.1 (Kuan-Yu Lin) 1Non-binding translation for convenience: This thesis is the result of my own independent work, and any material from work of others which is used either verbatim or indirectly in the text is credited to the author including details about the exact source in the text. This work has not been part of any other previous examination, neither completely nor in parts. It has neither completely nor partially been published before. The submitted electronic version is identical to this print version. Abstract Question Answering (QA) systems enable machines to understand human language, requiring robust training on related datasets. Nonetheless, large, high-quality datasets are only sometimes available due to cost restrictions. Active learning (AL) addresses this challenge by selecting the data with high information value as small subsets for model training, considering computa- tional resources while preserving performance. There are many different ways to detect the information value of the data, which in turn leads to a variety of AL strategies. In this study, we aim to investigate the performance change of the QA system after applying various AL strategies. In addition, we use the BatchBALD strategy, compared with its predecessor, the BALD strat- egy, to inspect the advantages of batch querying in data selection. Eventually, we propose Unique Context Selection (UC) and Unique Embedding Selection Methods (UE) to enhance the sampling effectiveness by ensuring maximal di- versity of context and embedding within querying samples, respectively. Ob- serving the experimental results, we learn that each dataset has its own AL strategy that brings out its best results, and there is no universal optimal AL strategy for QA tasks. BatchBALD maintains the modeling results similar to BALD in the regular setting while significantly reducing computation time, though this feature is not practiced in the low-resource setting. Finally, UC could not enhance the effectiveness of AL since half of the datasets used in this study consisted of more than 65% unique contexts. However, the effect of UE enhancement deviates across datasets and AL strategies, but it can be observed that most of the AL strategies with the best effect of UE enhance- ment can increase by more than 0.5% F1. Compared with context, a feature of datasets is limited to natural language processing tasks; embedding is more generalized and has a good enhancement effect, which is worth studying in depth. 2 Contents 1 Introduction 10 2 Theory 12 2.1 Active Learning (AL) . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2 Question Answering (QA) . . . . . . . . . . . . . . . . . . . . . . . . 13 3 Method 14 3.1 Strategies of Active Learning (AL) . . . . . . . . . . . . . . . . . . . 14 3.1.1 Uncertainty-based Strategies . . . . . . . . . . . . . . . . . . . 14 3.1.2 Diversity-based Strategies . . . . . . . . . . . . . . . . . . . . 16 3.1.3 Hybrid Strategies . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.1.4 MC Dropout . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.2 Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.3 Data Environments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.3.1 Regular Setting . . . . . . . . . . . . . . . . . . . . . . . . . . 20 3.3.2 Low-resource Setting . . . . . . . . . . . . . . . . . . . . . . . 20 3.4 Effective Sample Selection . . . . . . . . . . . . . . . . . . . . . . . . 22 3.4.1 Batch Bayesian Active Learning by Disagreements (Batch- BALD) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 3.4.2 Unique Context Selection Method (UC) . . . . . . . . . . . . 22 3.4.3 Unique Embedding Selection Method (UE) . . . . . . . . . . . 25 4 Results 28 4.1 Experimental Plan . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 4.2 Evaluation Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 3 4.3 Experiment Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 4.3.1 Regular Setting . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4.3.2 Low-resource Setting . . . . . . . . . . . . . . . . . . . . . . . 33 5 Discussion 37 5.1 Baseline performance across various datasets . . . . . . . . . . . . . . 37 5.2 BALD vs. BatchBALD . . . . . . . . . . . . . . . . . . . . . . . . . . 37 5.3 Unique context Selection Method (UC) . . . . . . . . . . . . . . . . . 38 5.4 Unique Embedding Selection Method (UE) . . . . . . . . . . . . . . . 39 6 Conclusion 41 A The results of each experiment with the score of each querying iteration. 49 A.1 Regular Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 A.2 Low-resource Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 B The results of each experiment with the standard deviation. 52 B.1 Baseline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 B.2 Unique Context (UC) . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 B.3 Unique Embedding (UE) . . . . . . . . . . . . . . . . . . . . . . . . . 64 C Comparison 70 4 List of Tables 1 Introduction of datasets used in experiments in this study . Con- text refers to the form or domain of the dataset. Exp. Setting shows the environment settings used for each dataset. OC/RD refers to the dataset as either open-domain or restricted-domain. # unlab. data pool is the unlabeled data pool size. # Test is the size of the testing set. # lab. data (% of D) refers to the number of queried labeled data and its percentage of unlabeled data. . . . . . . . . . . . . . . . . . . 30 2 Equation of F1 and the calculation with example in 4.2. . . . . . . . 31 3 Overall performance of baseline QA models, those enhanced with UC, and those enhanced with UE on SQuAD is presented below. UC dif- ference represents the F1 of UC subtracted from the F1 of the base- line, and similarly, UE difference represents the F1 of UE subtracted from the F1 of the baseline. We marked the highest F1 and the most difference in first place in red, second place in blue, and third in green. 32 4 Overall performance of baselines QA model on BioASQ, DROP, TQA, NewsQA, SearchQA, and NQ. We marked the ranked F1 in first place in red, second place in blue, and third in green. . . . . . . . . . . . . 33 5 Comparison of F1 of UC and the difference between baseline and UC in the low-resource setting on BioASQ, DROP, TQA, NewsQA, SearchQA, and NQ. The difference represents the F1 of UC sub- tracted from the F1 of the baseline. We marked the ranked F1 in first place in red, second place in blue, and third in green. . . . . . . . . . 35 6 Comparison of F1 of UE and the difference between baseline and UE in the low-resource setting on BioASQ, DROP, TQA, NewsQA, SearchQA, and NQ. The difference represents the F1 of UE subtracted from the F1 of the baseline. We marked the ranked F1 in first place in red, second place in blue, and third in green. . . . . . . . . . . . . 36 7 Comparison of BALD and BatchBALD about F1 and querying time. 38 5 8 Comparison of baseline and UC differences in the number of unique contexts in the query samples across multiple datasets. # unique C (% of D) indicates the difference in number and that number as a percentage of the entire dataset. . . . . . . . . . . . . . . . . . . . . . 39 6 List of Figures 1 Performance of AL strategies on each data querying size in the regular setting. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2 Performance of AL strategies on each data querying size in the low- resource setting. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 3 Performance of AL strategies enhanced with UC on each data query- ing size in the regular setting. . . . . . . . . . . . . . . . . . . . . . . 49 4 Performance of AL strategies enhanced with UE on each data query- ing size in the regular setting. . . . . . . . . . . . . . . . . . . . . . . 49 5 Performance of AL strategies enhanced with UC on each data query- ing size in the low-resource setting. . . . . . . . . . . . . . . . . . . . 50 6 Performance of AL strategies enhanced with UE on each data query- ing size in the low-resource setting. . . . . . . . . . . . . . . . . . . . 51 7 Performance with a standard deviation of AL strategies on SQuAD. . 52 8 Performance with a standard deviation of AL strategies on BioASQ . 53 9 Performance with a standard deviation of AL strategies on DROP . . 53 10 Performance with a standard deviation of AL strategies on TextbookQA 54 11 Performance with a standard deviation of AL strategies on NewsQA . 54 12 Performance with a standard deviation of AL strategies on SearchQA 55 13 Performance with a standard deviation of AL strategies on NQ . . . . 56 14 Performance with a standard deviation of AL strategies enhanced with UC on SQuAD. . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 15 Performance with a standard deviation of AL strategies enhanced with UC on BioASQ. . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 16 Performance with a standard deviation of AL strategies enhanced with UC on DROP. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 7 17 Performance with a standard deviation of AL strategies enhanced with UC on TQA. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 18 Performance with a standard deviation of AL strategies enhanced with UC on NewsQA. . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 19 Performance with a standard deviation of AL strategies enhanced with UC on SearchQA. . . . . . . . . . . . . . . . . . . . . . . . . . . 62 20 Performance with a standard deviation of AL strategies enhanced with UC on NQ. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 21 Performance with a standard deviation of AL strategies enhanced with UE on SQuAD. . . . . . . . . . . . . . . . . . . . . . . . . . . . 64 22 Performance with a standard deviation of AL strategies enhanced with UE on BioASQ. . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 23 Performance with a standard deviation of AL strategies enhanced with UE on DROP. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 24 Performance with a standard deviation of AL strategies enhanced with UE on TQA. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67 25 Performance with a standard deviation of AL strategies enhanced with UE on NewsQA. . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 26 Performance with a standard deviation of AL strategies enhanced with UE on SearchQA. . . . . . . . . . . . . . . . . . . . . . . . . . . 69 27 Performance with a standard deviation of AL strategies enhanced with UE on NQ. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 28 Comparing the behavior of AL strategies on SQuAD in the regular setting, with consideration given to the Baseline, UC, and UE. . . . . 71 29 Comparing the behavior of AL strategies on BioASQ in the low- resource setting, with consideration given to the Baseline, UC, and UE. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 8 30 Comparing the behavior of AL strategies on DROP in the low-resource setting, with consideration given to the Baseline, UC, and UE. . . . . 73 31 Comparing the behavior of AL strategies on TQA in the low-resource setting, with consideration given to the Baseline, UC, and UE. . . . . 74 32 Comparing the behavior of AL strategies on NewsQA in the low- resource setting, with consideration given to the Baseline, UC, and UE. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 33 Comparing the behavior of AL strategies on SearchQA in the low- resource setting, with consideration given to the Baseline, UC, and UE. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 34 Comparing the behavior of AL strategies on NQ in the low-resource setting, with consideration given to the Baseline, UC, and UE. . . . . 77 9 1 Introduction Natural language processing (NLP) serves as the field where computers understand human language to complete specific tasks (Jurafsky and Martin, 2000), enhancing human-machine or human-human communication. Developing NLP models relies on labeled data, not to mention question answering (QA) models. Producing annotated data for QA models is a time-consuming annotation process (Marti Roman, 2022b; Rajpurkar et al., 2018; Zhang et al., 2023). The annotators read textual data, analyze it, and determine the exact part of the context that answers the question. The obstacle in QA task development remains in forming large-scale labeled datasets. Accordingly, this study focuses on using annotated data effectively or reducing the required volume of labeled data by supporting the advancement of QA tasks. To effectively use the data in model training, previous studies have proved prac- tical approaches, including self-supervised methods (Devlin et al., 2019), transfer learning (Yu et al., 2018), data augmentation, and active learning (AL) (Zhang et al., 2023). Notable among them is AL, referred to as query learning (Settles, 2009). AL primarily investigates datasets to employ their informational value. Each sample in the dataset offers distinct information contributions. Training the model on a subset with high information value can attain outcomes comparable to that of a model trained on the entire dataset. By selecting samples with high informational values for labeling by oracles, AL achieves outstanding performance with a small amount of annotated samples. AL methods have found widespread application across various tasks, for example, object detection (Casanova et al., 2020; Haussmann et al., 2020), semantic segmen- tation (Saidu and Csató, 2021; Xie et al., 2022), image classification (Xie et al., 2022; Sener and Savarese, 2018), text classification (Zhang et al., 2017; Li et al., 2012), and name entity recognition (Shen et al., 2018; Chen et al., 2015). In addi- tion, there are several empirical analyses (Chen et al., 2015; Naseem et al., 2021) and comprehensive overviews (Settles, 2009; Aggarwal et al., 2014; Kumar and Gupta, 2020) that focus on AL. Zhan et al. (2022) examined performance among various AL methods with a standardized setting for image classification tasks. Marti Roman 10 (2022b) assessed AL approaches in the context of QA. In previous studies, AL has been applied to computer vision and NLP to solve the problem of insufficient labeled data and, at the same time, to maintain perfor- mance. Many studies have probed and formed diverse strategies with the update of time and research. However, not all strategies readily apply to NLP tasks, par- ticularly within the domain of QA. Some comprehensive and impartial assessments of strategies offer a deeper insight. Nevertheless, only a few studies have conducted performance comparisons of some of the AL methods. Contribution to improving the sampling effectiveness is also crucial to the data querying topic. However, the existing strategies are already various and well-developed. Instead of developing new strategies, using additional selection methods to combine the existing strategies is worthwhile. This study aims to bridge this gap by conveying an overview of all prob- able AL methods in QA tasks and exploring the selection methods to improve the sampling effectiveness. Aligned with the motivation and objectives of this study, the research questions are highlighted as follows: 1. How do the baseline strategies perform in the QA tasks in the regular setting? 2. How do the baseline strategies perform in the QA tasks in the low-resource setting across different datasets? 3. Regarding sampling effectiveness, what is the outcome of the following three enhancement approaches? (a) The application of the BatchBALD strategy. (b) The integration of the unique context selection method. (c) The incorporation of the unique embedding selection method. 11 2 Theory In this chapter, we first describe AL methods in Section 2.1. Afterward, the QA systems and the application of AL in QA models are introduced in Section 2.2. 2.1 Active Learning (AL) AL performs impressively with only a few labeled data (Settles, 2009). As long as the small amount of training data is the most informative set, the model can achieve almost the same accuracy as a supervised method (Boreshban et al., 2023). The learner can generate the query in various problem situations. According to the difference between them, there are three main types of AL approaches: membership query synthesis, stream-based selective sampling, and pool-based sampling. (Settles, 2009) Within membership query synthesis, the learner generates a novel instance from the input knowledge domain and requests its label. However, when the learner- generated instances have no semantic meaning or are unrecognizable, it causes anno- tation difficulties for human annotators (Settles, 2009). Solutions like stream-based and pool-based sampling have been introduced to address this challenge. In stream- based sampling, an instance is sampled from a natural input distribution, and sub- sequently, the learner decides to either label and retain the instance or discard it (Settles, 2009). Pool-based sampling selects the more informative samples from the unlabeled dataset and sends them to the oracle (e.g., a human annotator who is also an expert from a particular knowledge domain) for annotation. Next, the model is trained with the selected labeled samples only. In this case, the cost of annotating is reduced, but the performance is not affected. The defect of classical AL is that it cannot handle high-dimensional data, for instance, text, audio, or image (Tong, 2001). Moreover, AL does not extract features, so the high-value samples are evaluated based on the pre-extracted features (Ren et al., 2021). 12 2.2 Question Answering (QA) The task of question answering (QA) involves the identification of answers to pro- vided questions. QA can be a refined form of information retrieval (IR) (Kolomiyets and Moens, 2011). In contrast to traditional IR, which retrieves an entire document to serve as the answer, QA extracts specific information from the document as an answer. The documents containing the answers are often displayed in natural lan- guage. According to the subject that the documents belong to, we can categorize QA systems into open-domain question answering and restricted-domain question answering (Mollá and Vicedo, 2007). Open-domain QA (ODQA) model can address questions across various subjects or knowledge resources, whereas restricted-domain QA (RDQA) models specialize in specific domains, such as the medical domain. Since QA systems search answer in the documents, the QA tasks are often linked to natural language understanding (NLU) (Kwiatkowski et al., 2019). While NLU in QA proves beneficial to users, it displays considerable obstacles in development. As mentioned by Kwiatkowski et al. (2019), the development of QA faces four primary challenges. The first one is determining the approaches and resources for generating questions. The second is formulating methods for annotating and gathering answers. Then, the next is establishing measures to assess and ensure annotation quality. Lastly, select appropriate evaluation criteria. Not only the quality of annotated answers is crucial, but also the efficiency of labeling (Marti Roman, 2022b). Decreasing the need for labeled data could improve the viability of developing QA tasks. The performance of applying AL in QA tasks seems promising. Boreshban et al. (2023) compared several different query strate- gies with the Stanford Question Answering Dataset (SQuAD) dataset and achieved state-of-the-art results using only 40% of the training data. Marti Roman (2022b) estimated the performance of the language model applying Random Sampling (Random), Entropy Sampling (Entropy), and Bayesian Active Learning with Pretrained Language Models (BALM). The results showed that BALM did not outperform Random and Entropy and differed from the finding of Mar- gatina et al. (2022). 13 3 Method This chapter introduces three distinct categories of AL strategies dedicated to query- ing the most informative data points in Section 3.1. Afterward, a detailed description of the datasets applied to the QA model is presented in Section 3.2. In Section 3.3, we illustrate the data environment settings and their corresponding algorithms. Even- tually, Section 3.4, divided into three subsections, outlines the BatchBALD strategy and two sample selection approaches proposed by us. We aim to enhance sample efficiency through the support provided by these methods in Section 3.4. 3.1 Strategies of Active Learning (AL) There are various querying strategies, also referred to as scoring functions, in AL used to estimate the informativeness of a data point in the dataset. According to the characteristics of their sampling strategy, they are classified into three categories: uncertainty-based, diversity-based, and hybrid (Zhan et al., 2022). 3.1.1 Uncertainty-based Strategies Uncertainty-based strategies assess how much uncertainty a model has about clas- sifying a data point (Zhan et al., 2022). Then, extract the one that the model finds most uncertain. There are two kinds of uncertainties in machine learning: aleatoric uncertainty and epistemic uncertainty (Settles, 2009; Nguyen et al., 2019; Senge et al., 2014; Hüllermeier and Waegeman, 2021). Aleatoric uncertainty implies in- herent randomness, which affects the outcome of the data generation process and creates variability. Epistemic uncertainty arises from a lack of knowledge about the best model during the modeling or learning. Margin The probability difference between a sample’s two most probable predictions is mea- sured. The smaller the margin a sample has, the harder it is for the model to classify 14 it. The classifier has less confidence in the labels. Thus, the accurate annotation of these instances helps the most in distinguishing them. These samples provide more information value for understanding the dataset (Netzer et al., 2011; Boreshban et al., 2023). In Equation (1), the first two most likely labeled to the sample x are denoted as ŷ1 and ŷ2, respectively. The selected instances, denoted by x, are added to the training dataset. x∗ Margin = argmin x [p(ŷ1|x)− p(ŷ2|x)] (1) Entropy Margin only considers the two most likely labels while sampling x. Entropy (Shan- non, 1948) is more suitable for querying cases with numerous labels. The sample x has maximal predictive entropy and will be queried. x∗ Entropy = argmax x [− ∑ k p(ŷi = k|x) log p(ŷi = k|x)] (2) Least Confidence (LC) According to the model, LC selects the instance x that has the least confidence to have the labeled as ŷ (Wang and Shang, 2014). x∗ LC = argmax x [−p(ŷ∗|x)] (3) Mean Standard Deviation (MeanSTD) The selected instance x is expected to maximize the average standard deviation of the predicted probabilities across all c classes.(Kampffmeyer et al., 2016; Gal et al., 2017) The model parameters are denoted as θ. x∗ MeanSTD = argmax x 1 c ∑ c √ Eq(θ)[p(y = c|x,θ)2]− Eq(θ)[p(y = c|x,θ)]2 (4) 15 Bayesian Active Learning by Disagreements (BALD) The model parameters are denoted as θ. D = (xi, yi) n i=1 refers to the data that has been observed. p(θ|D) represents the posterior distribution across the param- eters. This strategy aims to leverage Shannon’s Entropy (Shannon, 1948) to re- duce uncertainty regarding the parameters. With data containing high entropy, denoted as H[y|x, D], the model is moderately uncertain about y, while with low Eθ∼p(θ|Dl)[H[y|x,θ]], the individual parameter settings are highly confident. The pos- terior parameter estimates express the greatest disagreement regarding the outcome (Houlsby et al., 2011; Gal et al., 2017). x∗ BALD = argmax x H[y|x, D]− Eθ∼p(θ|Dl)[H[y|x,θ]] (5) 3.1.2 Diversity-based Strategies Diversity-based strategies involve the selection of the most representative sample (Ash et al., 2020). Once the selected set of samples is labeled, it provides a surrogate for the entire dataset. KMeans KMeans is an unsupervised machine learning algorithm that uses clustering to clas- sify data. The data are aggregated in multiple clusters due to the similarity of un- known features. The various clusters highlight the diversity of the data. KMeans in AL is to query the closest to the centroid of each cluster. CoreSet The CoreSet, which is the subset of the dataset that can represent the entire set itself, is the sample selection criterion. It is selected to minimize the maximum distance between any point and its nearest selected point to ensure that theCoreSet exhibits a high degree of diversity and captures the essential characteristics of the entire dataset (Sener and Savarese, 2018). 16 3.1.3 Hybrid Strategies The above two categories have their strength: uncertainty-based score functions are more precise in measuring data informativeness; alternatively, diversity-based strategies possess the data selection in batch size. The hybrid strategies balance and take advantage of both strengths. Batch Active learning by Diverse Gradient Embeddings (BADGE) Deep neural networks are typically optimized utilizing gradient-based techniques. The uncertainty is assessed using the weighted KMeans++ initialization on the gradient embeddings for the parameters of feature representative layer (Ash et al., 2020). 3.1.4 MC Dropout Gal and Ghahramani (2016) used the Monte-Carlo dropout (MC-dropout) as a Bayesian approximation to estimate the uncertainty of the neural network model. The strategies can be applied to the MC-dropout output (Beluch et al., 2018). Zhan et al. (2022) suggested that the MC-dropout version obtained superior accu- racy compared to the original version due to an enhanced divergence in uncertainty scores among samples, attributable to the increased uncertainty introduced by MC- dropout. 3.2 Data QA systems derive machine reading comprehension skills from the QA datasets, which typically contain three key components: context, question, and answer. The following example from SQuAD v1.1 (Rajpurkar et al., 2016) demonstrates the stan- dard instance structure in QA datasets. • Context: The university’s library system is divided between the main li- brary and each of the colleges and schools. The main building is the 14-story 17 Theodore M. Hesburgh Library, completed in 1963, which is the third building to house the main collection of books. The front of the library is adorned with the Word of Life mural designed by artist Millard Sheets. This mural is pop- ularly known as ”Touchdown Jesus” because of its proximity to Notre Dame Stadium and Jesus’ arms appearing to make the signal for a touchdown. • Question: What is the name of the main library at Notre Dame? • Answer: Theodore M. Hesburgh Library The fundamental process of QA systems is searching for the answer span within the provided context to answer the posed questions. Over the years, abundant QA datasets have been suggested. The subsequent sections outline several datasets in detail, aiming to provide a thorough understanding of their characteristics and the challenges they offer in developing and evaluating QA systems. Stanford Question Answering Dataset (SQuAD) SQuAD(Rajpurkar et al., 2016) is the most widely used reading comprehension dataset, and it is a collection of crowdsourcing question-answer pairs obtained from Wikipedia articles. The questions and answers were created together according to Wikipedia passage. BioASQ BioASQ (Tsatsaronis et al., 2015) is designed to improve the development of biomed- ical semantic indexing and question answering. The question-and-answer pairs are created based on related domain snippets (PubMed). Discrete Reasoning Over Paragraphs (DROP) DROP (Dua et al., 2019) contains question-answer pairs based on Wikipedia para- graphs created with the crowdsourcing method. Solving some of the questions in DROP requires reasoning operation, which involves more content comprehension ability than other datasets. 18 TextbookQuestionAnswering (TQA) TQA (Kembhavi et al., 2017) includes the context of textbooks from middle school science lessons. The questions are assembled from Life, Earth, and Physical Science. NewsQA NewsQA(Trischler et al., 2017) is a vast collection of crowdsourcing samples created based on over 10,000 news pieces from CNN. The questions are produced while the accessible sources are the headline and summary of the articles, but the answers are established while having full access to the articles. It can be more challenging than other QA datasets due to the extra reasoning process for the QA model to find the answer. SearchQA In the case of SearchQA(Dunn et al., 2017), the dataset is built using existing question-answer pairs obtained by crawling J! Archive, a Jeopardy TV show. The contexts are augmented with snippets retrieved by the Google search engine, so they are in diverse forms and from distinct domains. Natural Questions (NQ) NQ (Kwiatkowski et al., 2019) consists of real users’ questions from Google.com and the Wikipedia page that may or may not contain an answer. QA model must read the whole page to know if there is an answer. It makes the task more realistic and challenging than other QA datasets. There are four possible answer annotations: both long and short answers, either long or short answers, and no answer with empty annotation. The short answer is more concise than the long answer. 3.3 Data Environments The main objective of AL is to enhance model training when faced with limited data. However, evaluating the performance of AL in conventional data scenarios is 19 equally essential. We propose two different data environments to simulate varying scenarios to achieve this. The first is the regular setting characterized by a larger original quantity of data, as outlined in Section 3.3.1. The second is the low-resource setting, where a smaller original number of data or a subset of the entire dataset is employed, as elaborated upon in Section 3.3.2. 3.3.1 Regular Setting In the regular setting, strategies queried data from the complete SQuAD dataset. As depicted in Algorithm 1, initially, a small labeled dataset r∗ is randomly se- lected from Dpool and added to Dlable to serve as the initial training data for the learner model. Furthermore, instances x∗ are sampled after the learner Θ operate the querying strategy ϕ(·, ·) on Dpool. Finally, oracles label instances x ∗, which are added into Dlabel and are removed from Dpool to prevent duplicate selections. The entire querying process iterates T times. 3.3.2 Low-resource Setting In the low-resource condition, the amount of data is limited, so we do not have the initiated labeled data selected randomly. Hence, we use the knowledge from QA models built on SQuAD through the transfer learning method. Jukić and Šnajder (2023) highlighted the outperformance of the cooperation of AL and parameter- efficient fine-tuning (PEFT). Adding an Adapter resolved the lack of robustness, while the queried training samples for each iteration were too small to train the model. In Algorithm 2, the QA model Θ´ is trained on a large complete QA dataset. Then, sample the instances x∗ after the model Θ´ utilizing the querying strategy ϕ(·, ·) to query from Dpool. Labeled x∗, add it to Dlabel, and remove it from Dpool. Then, train the learner model Θ´ by Dlabel. Later, move on to the query process and iterate it T times. 20 Algorithm 1 AL Methods in the Regular Setting Input: Labeled dataset Dlabel, unlabeled dataset Dpool, initial random acquisition data r∗, queried samples x∗, AL strategy ϕ(·, ·) Ensure: Model Θ 1: Dlabel ← ∅ 2: r∗ ← RANDOM(r, ·) 3: Dlabel ← Dlabel ∪ label(r∗) 4: Dpool ← Dpool \ r∗ 5: Θ← learn the model based on Dlabel 6: for t = 1, 2, 3, ..., T do 7: Dpool ← Dpool \Dlabel 8: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 9: Dlabel ← Dlabel ∪ label(x∗) 10: Dpool ← Dpool \ x∗ 11: Θ← update the model based on Dlabel 12: end for Algorithm 2 AL Method in the Low-resource Setting Input: Model Θ´, labeled dataset Dlabel, unlabeled dataset Dpool, queried samples x∗, AL strategy ϕ(·, ·) Ensure: Model Θ´ 1: for t = 1, 2, 3, ..., T do 2: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 3: Dlabel ← Dlabel ∪ label(x∗) 4: Dpool ← Dpool \ x∗ 5: Θ← update the model based on Dlabel 6: end for 21 3.4 Effective Sample Selection With the support from AL, we can use a set of informative data to train the QA systems. To achieve an improvement in selecting the most effective sample, we intro- duced Batch Bayesian Active Learning by Disagreements (BatchBALD) (Kirsch et al., 2019) strategy as an advanced version of BALD in Section 3.4.1. In addition, we propose two sample selection approaches to boost the queried data with the uniqueness of context and embedding. The unique context selection method (UC) and the unique embedding selection method(UE) are illustrated in Section 3.4.2 and Section 3.4.3, respectively. 3.4.1 Batch Bayesian Active Learning by Disagreements (BatchBALD) BatchBALD (Kirsch et al., 2019) extendsBALD to a batch form to select {x∗ 1, ...x ∗ b}. The sequence of data point x1, ...,xb and y1, ..., yb are presented as x1:b and y1:b in the Equation 6. The strategy evaluates the mutual information between a joint of many samples and the model parameters. Kirsch et al. (2019) demonstrated that Batch- BALD improved data efficiency with high-dimensional image data and decreased the data querying time. {x∗ 1:b}BatchBALD = argmax x1:b H[y1:b|x1:b, D]− Eθ∼p(θ|Dl)[H[y1:b|x1:b,θ]] (6) 3.4.2 Unique Context Selection Method (UC) Context is one of the critical elements in a QA dataset. Some datasets form ques- tions and answers based on a given context, e.g., SQuAD, while others start with a set of question-answer pairs, later augmented by attaching relevant context, e.g., SearchQA. The context is considered an essential feature in the QA task. To enhance the effectiveness of the sample, we ensure that each data selection is highly represen- tative at the context level. Algorithm 3 extends Algorithm 1, with additional steps. Following instance sampling in line 8, we ensure the uniqueness of all contexts from each other and the context of Dlabel. N stands for the number of targeted queried 22 samples. If any context from the samples exists in the context of Dlabel or if there are overlaps in the contexts of queried samples, those instances are excluded from Dpool, and the process in line 8 is repeated to obtain a new set of samples. Algorithm 4, an extension of Algorithm 2, shows the selection process in the low-resource setting, spanning from line 3 to line 8. Algorithm 3 AL Methods with UC in the Regular Setting Input: Labeled dataset Dlabel, unlabeled dataset Dpool, initial random acquisition data r∗, queried samples x∗, AL strategy ϕ(·, ·), number of targeted queried samples N , unique item selector U(·), retrieve the context of data C(·) Ensure: Model Θ 1: Dlabel ← ∅ 2: r∗ ← RANDOM(r, ·) 3: Dlabel ← Dlabel ∪ label(r∗) 4: Dpool ← Dpool \ r∗ 5: Θ← learn the model based on Dlabel 6: for t = 1, 2, 3, ..., T do 7: Dpool ← Dpool \Dlabel 8: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 9: N ← #(x∗) 10: while C(x∗) in C(Dlabel) or #(U(C(x∗))) < N do 11: x∗´← samples from x∗ , whose context exist in C(Dlabel) or overlap with other samples 12: Dpool ← Dpool \ x∗´ 13: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 14: end while 15: Dlabel ← Dlabel ∪ label(x∗) 16: Dpool ← Dpool \ x∗ 17: Θ← update the model based on Dlabel 18: end for 23 Algorithm 4 AL Method with UC in the Low-resource Setting Input: Model Θ´, labeled dataset Dlabel, unlabeled dataset Dpool, queried samples x∗, AL strategy ϕ(·, ·), number of targeted queried samples N , unique item selector U(·), retrieve the context of data C(·) Ensure: Model Θ´ 1: for t = 1, 2, 3, ..., T do 2: x∗ ← argmaxx∈Dpool ϕ(x,Θ´) 3: N ← #(x∗) 4: while C(x∗) in C(Dlabel) or #(U(C(x∗))) < N do 5: x∗´← samples from x∗ , whose context exist in C(Dlabel) or overlap with other samples 6: Dpool ← Dpool \ x∗´ 7: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 8: end while 9: Dlabel ← Dlabel ∪ label(x∗) 10: Dpool ← Dpool \ x∗ 11: Θ´← update the model based on Dlabel 12: end for 24 3.4.3 Unique Embedding Selection Method (UE) Contextual word embeddings offer valuable information for integrating into the NLP tasks (Devlin et al., 2019), including QA tasks (Rajpurkar et al., 2016). To optimize the effectiveness of our queried samples, we apply clustering to the embeddings of each data point inspired by the KMeans algorithm. Subsequently, we ensure that each selected data point originates from a distinct cluster. The process of selecting unique embedding from queried samples in the regular setting is elaborated in the Algorithm 5 from line 9 to line 14, extended from Algorithm 1. In Algorithm 5, queried instances x∗ are required to have unique embeddings both from each other and fromDlabel.N records the number of samples targeted for querying in iteration t. If the count of unique embeddings in the queried samples, #(U(E(x∗))), is less than N , the model continues querying instances until obtaining the set of queried samples with unique embeddings. Furthermore, Algorithm 6 is an extension of Algorithm 2, illustrating the unique embedding selection procedure in the low-resource setting from line 3 to line 8. 25 Algorithm 5 AL Methods with UE in the Regular Setting Input: Labeled dataset Dlabel, unlabeled dataset Dpool, initial random acquisition data r∗, queried samples x∗, AL strategy ϕ(·, ·), number of targeted queried samples N , unique item selector U(·), retrieve the embedding of data E(·) Ensure: Model Θ 1: Dlabel ← ∅ 2: r∗ ← RANDOM(r, ·) 3: Dlabel ← Dlabel ∪ label(r∗) 4: Dpool ← Dpool \ r∗ 5: Θ← learn the model based on Dlabel 6: for t = 1, 2, 3, ..., T do 7: Dpool ← Dpool \Dlabel 8: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 9: N ← #(x∗) 10: while E(x∗) in E(Dlabel) or #(U(E(x∗))) < N do 11: x∗´← samples from x∗ , whose embedding exist in E(Dlabel) or overlap with other samples 12: Dpool ← Dpool \ x∗´ 13: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 14: end while 15: Dlabel ← Dlabel ∪ label(x∗) 16: Dpool ← Dpool \ x∗ 17: Θ← update the model based on Dlabel 18: end for 26 Algorithm 6 AL Method with UE in the Low-resource Setting Input: Model Θ´, labeled dataset Dlabel, unlabeled dataset Dpool, queried samples x∗, AL strategy ϕ(·, ·), number of targeted queried samples N , unique item selector U(·), retrieve the embedding of data E(·) Ensure: Model Θ´ 1: for t = 1, 2, 3, ..., T do 2: x∗ ← argmaxx∈Dpool ϕ(x,Θ´) 3: N ← #(x∗) 4: while E(x∗) in E(Dlabel) or #(U(E(x∗))) < N do 5: x∗´← samples from x∗ , whose embedding exist in E(Dlabel) or overlap with other samples 6: Dpool ← Dpool \ x∗´ 7: x∗ ← argmaxx∈Dpool ϕ(x,Θ) 8: end while 9: Dlabel ← Dlabel ∪ label(x∗) 10: Dpool ← Dpool \ x∗ 11: Θ´← update the model based on Dlabel 12: end for 27 4 Results This chapter illustrates the plan of experiments and the comprehensive results of our study. We provide detailed insights into the plan of experiments we have conducted in Section 4.1, the evaluation matrix for the experiments in Section 4.2, and then explain the result of experiments in Section 4.3. 4.1 Experimental Plan The AL in this study is implemented utilizing DeepAL+ toolkits (Zhan et al., 2022). Each experiment is repeated in 5 trials. Each trial in the regular setting uses the same random seed to split the initial labeled pool and the remaining for the unlabeled pool. The overall score is the average of the performance of 5 trials. AL strategies This study strives to gain a comprehensive understanding of the performance of distinct AL strategies across different data environment settings. In the regular setting, we applied Margin, LC, Entropy, MeanSTD, BALD, KMeans and a greedy version of CoreSet denoted as KCenter. Additionally, we included the MC- dropout version of Margin, Entropy, and LC, labeled as MarginD, EntropyD, and LCD, correspondingly. Due to time constrain, BADGE was not applied in the regular setting. In the low-resource setting, we employed AL strategies, including BADGE. Furthermore, we implemented Random Sampling strategy, denoted as Random, in both data environment settings. Random samples data from the whole dataset, denoted as Dpool, randomly. In this case, all the instances have the same probability of being selected. Model This study used the fine-tuning RoBERTa (Liu et al., 2019) as essential learners. RoBERTa stands for Robustly optimized BERT approach and is an optimized version of another common language model option, BERT (Devlin et al., 2019). 28 Marti Roman (2022b) proved that RoBERTa outperforms BERT in QA tasks, so it is a better fit for this study. This study does not aim to optimize the QA model for AL, so we keep the experimental setting simple and the hyperparameters described below the same for all experiments. 1. Training epoch: 3 2. AdamW parameter optimizer with: a) No warm-up steps b) Learning rate: 3e-5 The diversity-based and hybrid strategies rely on estimating the features ex- tracted from hidden layers. In the use case of RoBERTa as a QA model, the features from the second-to-last hidden layer are more representative than the ones from the last hidden layer (Devlin et al., 2019). Then, our features are extracted from the second-to-last hidden layer of the QA model. Data This experiment runs the QA model on several datasets. Due to the effectiveness of implementation, we accessed the datasets from MRQA (Fisch et al., 2019), a consolidated collection of QA datasets in a unified format. To evaluate each strategy, we used training and testing data. The following subsections illustrate datasets in detail, and Table 1 shows their distribution. In the regular setting, we evaluate the QA model utilizing the SQuAD v1.1 from MRQA (Fisch et al., 2019). The entire training and testing set are used for this purpose. The experiment setup is presented as follows: 1. Initial random acquisition data r∗: 500 2. Queried samples x∗: 500 3. Querying process iteration T : 4 29 Dataset Context Exp. Setting OD/RD # unlab. data pool # lab. data (% of D) # Test SQuAD Wikipedia Regular Open-domain 86,588 2500 (2.9%) 10,507 BioASQ Medical Articles Low-resource Restricted-domain 1,354 200 (14.8%) 150 DROP Wikipedia Low-resource Open-domain 1,353 200 (14.8%) 150 TQA Textbook Low-resource Restricted-domain 1,353 200 (14.8%) 150 NewsQA New Articles Low-resource Restricted-domain 10,000 200 (2%) 4,212 SearchQA Web Snippets Low-resource Open-domain 10,000 200 (2%) 16,980 NQ Wikipedia Low-resource Open-domain 10,000 200 (2%) 12,836 Table 1: Introduction of datasets used in experiments in this study . Context refers to the form or domain of the dataset. Exp. Setting shows the environment settings used for each dataset. OC/RD refers to the dataset as either open-domain or restricted-domain. # unlab. data pool is the unlabeled data pool size. # Test is the size of the testing set. # lab. data (% of D) refers to the number of queried labeled data and its percentage of unlabeled data. In the low-resource setting, our evaluation of QA systems involves multiple datasets. The dataset selected for this study is sourced from MRQA (Fisch et al., 2019). From BioASQ in MRQA, comprising 1,504 samples, 10% (150 samples) are allocated to the testing set, while the remaining 1,354 samples stay in the unla- beled data pool. Similarly, for the 1,503 DROP samples in MRQA, which feature questions with extractive answers, we separate 150 samples for the testing, leaving 1,353 samples for the unlabeled data pool. TQA, also from MRQA (Fisch et al., 2019), comprises 1,503 samples, excluding ’True or False’ questions and those with diagrams. 150 samples are reserved for testing, and the remaining 1,353 samples are utilized for the unlabeled data pool. Within MRQA (Fisch et al., 2019), NewsQA instances lacking answers or exhibiting answers without annotator agreement are ex- cluded. Subsequently, we utilize a subset of 10,000 training examples from NewsQA in MRQA and incorporate the entire NewsQA testing set available in MRQA. Addi- tionally, we leverage a subset of 10,000 SearchQA samples from the training set and employ the complete testing set from MRQA. As outlined by Fisch et al. (2019), MRQA selects 104,071 NQ samples featuring short answers, utilizing long answers as context. In this study, we compile a subset of 10,000 samples and utilize the com- plete testing set of NQ from MRQA. The experimental setup for the low-resource condition is presented as follows: 1. Dataset D′ for pre-training the model Θ′: SQuAD v1.1 2. Queried samples x∗: 50 30 3. Querying process iteration T : 4 4.2 Evaluation Matrix A widely used measurement for evaluating the performance of QA tasks is the F1- score (F1)(Rajpurkar et al., 2016). The F1 compares the average overlapping amount between the predicted and ground truth tokens. In this evaluation, the predicted and ground truth tokens are calculated as bag-of-words, wherein the order is not considered. The primary objective of F1 is to give a high score when the predicted tokens precisely match all the ground truth tokens, excluding any irrelevant ones. Each sample is assigned its F1, and the overall score is then averaged F1 across all samples. Table 2 further clarifies the detailed formulas (Marti Roman, 2022a) for calculating F1 with an example of predicted and labeled answers, elucidated in a list below. Example predicted tokens: [”the”, ”14-story”, ”Theodore”, ”M.”, ”Hesburgh”, ”Library”] Example ground truth/labeled tokens: [”Theodore”, ”M.”, ”Hesburgh”, ”Library”] Measure Equation Example Precision |(Labeled tokens)∩(Predicted tokens)| |(Predicted tokens)| |[”Theodore”, ”M.”, ”Hesburgh”, ”Library”]| |[”the”, ”14-story”, ”Theodore”, ”M.”, ”Hesburgh”, ”Library”]| = 4 6 Recall |(Labeled tokens)∩(Predicted tokens)| |(Labeled tokens)| |[”Theodore”, ”M.”, ”Hesburgh”, ”Library”]| |[”Theodore”, ”M.”, ”Hesburgh”, ”Library”]| = 4 4 F1-score 100 ∗ 2 ∗ precision∗recall precision+recall 100 ∗ 2 ∗ 0.67∗1 0.67+1 = 80.24 Table 2: Equation of F1 and the calculation with example in 4.2. 4.3 Experiment Results This section details the overall F1 score results of the baseline QA systems and the versions using UC and UE enhancements on each dataset. 31 4.3.1 Regular Setting Baseline Unique Context Unique Embedding AL Strategies F1 ± SD F1 ± SD UC difference F1 ± SD UE difference Full 92.0671 - - - - Random 78.4594 ± 0.3153 78.24 ± 0.3327 -0.2194 78.657 ± 0.1486 0.1976 Margin 74.7216 ± 0.8332 73.8666 ± 0.6051 -0.855 74.7874 ± 0.5659 0.0658 LeastConf 79.5868 ± 0.3878 79.8476 ± 0.7351 0.2608 80.008 ± 0.4439 0.4212 Entropy 79.5892 ± 0.2888 79.7804 ± 0.6226 0.1912 79.9546 ± 0.1747 0.3654 MarginD 72.9154 ± 0.5847 73.2052 ± 0.6751 0.2898 73.2292 ± 0.487 0.3138 LeastConfD 79.1824 ± 0.1518 79.5002 ± 0.1095 0.3178 79.3274 ± 0.4435 0.145 EntropyD 79.6106 ± 0.2226 79.676 ± 0.2271 0.0654 79.6988 ± 0.2835 0.0882 MeanSTD 79.8626 ± 0.3015 79.713 ± 0.2973 -0.1496 79.763 ± 0.3767 -0.0996 BALD 79.9088 ± 0.2653 79.5792 ± 0.3494 -0.3296 79.6592 ± 0.4149 -0.2496 KMeans 78.4282 ± 0.386 78.8634 ± 0.2347 0.4352 78.4786 ± 0.4814 0.0504 KCenter 77.7264 ± 0.8187 77.6218 ± 0.4661 -0.1046 77.6922 ± 0.2914 -0.0342 Table 3: Overall performance of baseline QA models, those enhanced with UC, and those enhanced with UE on SQuAD is presented below. UC difference represents the F1 of UC subtracted from the F1 of the baseline, and similarly, UE difference represents the F1 of UE subtracted from the F1 of the baseline. We marked the highest F1 and the most difference in first place in red, second place in blue, and third in green. Figure 1: Performance of AL strategies on each data querying size in the regular setting. 32 4.3.2 Low-resource Setting BioASQ DROP TQA NewQA SearchQA NQ AL Strategies F1 ± SD F1 ± SD F1 ± SD F1 ± SD F1 ± SD F1 ± SD Full 65.8710 57.3988 54.5397 67.6726 71.5673 72.2290 Random 59.3349 ± 0.7915 51.9847 ± 0.75 45.8526 ± 0.8533 60.0351 ± 0.1789 37.0129 ± 0.9111 60.9138 ± 0.2852 Margin 60.1523 ± 0.8392 50.426 ± 0.6348 44.6131 ± 0.5171 59.7398 ± 0.1013 38.0298 ± 0.4868 60.4181 ± 0.291 LeastConf 58.4648 ± 0.7418 51.4066 ± 0.6333 46.7906 ± 0.7323 60.3695 ± 0.1675 37.1042 ± 0.9414 61.2088 ± 0.2247 Entropy 58.356 ± 0.6685 51.4125 ± 0.7225 46.7132 ± 0.7651 60.5174 ± 0.1717 37.1632 ± 0.6403 61.0201 ± 0.2019 MarginD 59.8669 ± 0.7098 50.4725 ± 0.6542 42.0347 ± 0.5204 59.4778 ± 0.1726 30.2362 ± 0.5695 58.2668 ± 0.3324 LeastConfD 58.6037 ± 0.8341 49.5594 ± 0.7156 46.688 ± 0.6052 60.0145 ± 0.1288 35.1549 ± 0.3814 61.6221 ± 0.2266 EntropyD 58.9144 ± 0.6553 49.5365 ± 0.6188 46.4892 ± 0.5551 59.9633 ± 0.1283 35.3589 ± 0.4386 61.526 ± 0.2374 MeanSTD 59.4096 ± 0.7591 52.2442 ± 0.7879 44.3608 ± 0.6697 59.6292 ± 0.1422 31.1528 ± 1.0246 61.188 ± 0.1851 BALD 59.1574 ± 0.5738 52.1718 ± 0.7187 44.0773 ± 0.8951 59.613 ± 0.1329 30.9229 ± 0.7563 60.9946 ± 0.2083 BatchBALD 60.2796 ± 0.7107 50.4278 ± 1.0094 44.6456 ± 0.7648 59.5572 ± 0.1633 37.7102 ± 0.7377 60.2574 ± 0.3779 BADGE 59.1261 ± 0.7021 51.5724 ± 0.6355 43.6183 ± 0.6868 59.6685 ± 0.1419 27.9376 ± 0.6521 60.7058 ± 0.2824 KMeans 59.4739 ± 0.8224 51.4251 ± 0.7257 44.6171 ± 0.7672 59.7566 ± 0.1704 31.3474 ± 1.0092 60.913 ± 0.303 KCenter 59.0172 ± 0.6507 51.4481 ± 0.9468 46.4568 ± 1.4967 59.8206 ± 0.2162 33.2055 ± 0.8957 60.6227 ± 0.2606 Table 4: Overall performance of baselines QA model on BioASQ, DROP, TQA, NewsQA, SearchQA, and NQ. We marked the ranked F1 in first place in red, second place in blue, and third in green. 2Even if the F1 difference of Random is at the highest among others, we would not consider it as the best performance of the UC on the NQ, considering the strategy itself has no computation of data informativeness and that the differences of the other AL strategies are, on average, lower. 33 (a) BioASQ (b) DROP (c) TQA (d) NewsQA (e) SearchQA (f) NQ Figure 2: Performance of AL strategies on each data querying size in the low-resource setting. 34 BioASQ DROP TQA AL Strategies F1 ± SD difference F1 ± SD difference F1 ± SD difference Random 59.3682 ± 0.7723 0.0333 51.0139 ± 0.8944 -0.9708 45.9923 ± 0.6722 0.1397 Margin 60.2679 ± 0.8053 0.1156 50.5443 ± 0.753 0.1183 44.372 ± 0.4674 -0.2411 LeastConf 58.4648 ± 0.7418 0.0 50.9689 ± 0.697 -0.4377 46.9358 ± 0.6812 0.1452 Entropy 58.356 ± 0.6685 0.0 50.746 ± 0.6827 -0.6665 46.822 ± 0.6975 0.1088 MarginD 59.8722 ± 0.7237 0.0053 50.8628 ± 0.7883 0.3903 42.7328 ± 0.7607 0.6981 LeastConfD 58.6037 ± 0.8341 0.0 50.0587 ± 0.7359 0.4993 46.7886 ± 0.6059 0.1006 EntropyD 58.9144 ± 0.6553 0.0 49.95 ± 0.7249 0.4135 46.6581 ± 0.6027 0.1689 MeanSTD 59.3699 ± 0.6688 -0.0397 52.0091 ± 0.8188 -0.2351 44.7365 ± 0.5532 0.3757 BALD 59.0275 ± 0.7702 -0.1299 51.7309 ± 0.8154 -0.4409 44.1821 ± 0.8501 0.1048 BADGE 59.144 ± 0.8025 0.0179 51.2108 ± 0.588 -0.3616 43.8096 ± 0.9063 0.1913 KMeans 59.2321 ± 0.764 -0.2418 50.7334 ± 0.7822 -0.6917 44.3442 ± 0.8737 -0.2729 KCenter 59.0669 ± 0.7335 0.0497 51.0278 ± 1.0738 -0.4203 45.9394 ± 1.2752 -0.5174 NewsQA SearchQA NQ AL Strategies F1 ± SD difference F1 ± SD difference F1 ± SD difference Random 60.0936 ± 0.2108 0.0585 37.0129 ± 0.9111 0.0 62.3823 ± 0.8318 1.46852 Margin 59.7626 ± 0.1263 0.0228 38.0298 ± 0.4868 0.0 60.4472 ± 0.2883 0.0291 LeastConf 60.3257 ± 0.1736 -0.0438 37.1042 ± 0.9414 0.0 61.1959 ± 0.1837 -0.0129 Entropy 60.526 ± 0.1477 0.0086 37.1632 ± 0.6403 0.0 61.0014 ± 0.1587 -0.0187 MarginD 59.444 ± 0.1353 -0.0338 30.2362 ± 0.5695 0.0 58.0904 ± 0.4643 -0.1764 LeastConfD 60.0829 ± 0.136 0.0684 35.1549 ± 0.3814 0.0 61.8388 ± 1.3003 0.2167 EntropyD 59.9883 ± 0.1127 0.025 35.3589 ± 0.4386 0.0 61.5971 ± 0.2455 0.0711 MeanSTD 59.6683 ± 0.1739 0.0391 31.1528 ± 1.0246 0.0 61.2109 ± 0.1906 0.0229 BALD 59.5151 ± 0.1875 -0.0979 30.7217 ± 0.7847 -0.2012 61.0601 ± 0.1713 0.0655 BADGE 59.7062 ± 0.1064 0.0377 29.0398 ± 0.7478 1.1022 60.7176 ± 0.268 0.0118 KMeans 59.7517 ± 0.124 -0.0049 32.4368 ± 0.627 1.0894 60.7913 ± 0.3484 -0.1217 KCenter 59.8777 ± 0.2281 0.0571 33.4291 ± 0.5886 0.2236 60.6441 ± 0.24 0.0214 Table 5: Comparison of F1 of UC and the difference between baseline and UC in the low-resource setting on BioASQ, DROP, TQA, NewsQA, SearchQA, and NQ. The difference represents the F1 of UC subtracted from the F1 of the baseline. We marked the ranked F1 in first place in red, second place in blue, and third in green. 35 BioASQ DROP TQA AL Strategies F1 ± SD difference F1 ± SD difference F1 ± SD difference Random 59.4874 ± 0.8272 0.1525 51.3067 ± 0.7261 -0.678 46.135 ± 1.0098 0.2824 Margin 60.3954 ± 0.5001 0.2431 51.3325 ± 0.7907 0.9065 45.0371 ± 1.2421 0.424 LeastConf 58.9594 ± 0.4445 0.4946 51.3114 ± 0.396 -0.0952 47.4874 ± 0.8092 0.6968 Entropy 58.8017 ± 0.6322 0.4457 51.599 ± 0.6738 0.1865 47.6978 ± 0.9358 0.9846 MarginD 60.029 ± 0.5388 0.1621 51.106 ± 1.0071 0.6335 41.808 ± 0.869 -0.2267 LeastConfD 58.9523 ± 0.573 0.3486 50.6483 ± 0.915 1.0889 46.1153 ± 0.7764 -0.5727 EntropyD 59.3175 ± 0.7086 0.4031 50.6299 ± 0.7377 1.0934 46.1799 ± 0.7191 -0.3093 MeanSTD 59.8422 ± 0.6457 0.4326 51.4074 ± 1.0592 -0.8368 43.7594 ± 0.8274 -0.6014 BALD 59.8231 ± 0.6628 0.6657 51.0219 ± 1.0207 -1.1499 43.179 ± 0.8665 -0.8983 BADGE 59.384 ± 0.6041 0.2579 51.4892 ± 0.6683 -0.0832 44.4 ± 1.1402 0.7817 KMeans 59.593 ± 0.5672 0.1191 51.2643 ± 0.9091 -0.1608 45.396 ± 0.5461 0.7789 KCenter 59.1756 ± 0.5283 0.1584 51.1726 ± 0.6335 -0.2755 45.1059 ± 0.7409 -1.3509 NewsQA SearchQA NQ AL Strategies F1 ± SD difference F1 ± SD difference F1 ± SD difference Random 60.1759 ± 0.152 0.1408 37.2601 ± 0.6218 0.2472 61.1156 ± 0.2533 0.2018 Margin 59.8385 ± 0.0961 0.0987 36.4116 ± 0.5299 -1.6182 59.8205 ± 0.3806 -0.5976 LeastConf 60.3417 ± 0.1099 -0.0278 34.6189 ± 0.5297 -2.4853 60.8625 ± 0.2469 -0.3463 Entropy 60.4755 ± 0.1601 -0.0419 34.3551 ± 0.8176 -2.8081 60.721 ± 0.3648 -0.2991 MarginD 59.463 ± 0.2264 -0.0148 30.572 ± 1.1313 0.3358 59.0336 ± 0.4124 0.7668 LeastConfD 60.2281 ± 0.1283 0.2136 35.4331 ± 0.7234 0.2782 61.7766 ± 0.2034 0.1545 EntropyD 60.1517 ± 0.0992 0.1884 35.0389 ± 0.7196 -0.32 61.7219 ± 0.2663 0.1959 MeanSTD 59.9631 ± 0.207 0.3339 31.7181 ± 0.9403 0.5653 61.049 ± 0.2553 -0.139 BALD 59.8409 ± 0.1971 0.2279 31.4257 ± 1.1808 0.5028 60.9826 ± 0.1906 -0.012 BADGE 59.5198 ± 0.1212 -0.1487 28.8987 ± 0.9768 0.9611 60.6351 ± 0.2562 -0.0707 KMeans 59.6801 ± 0.1403 -0.0765 32.3124 ± 0.885 0.965 60.5317 ± 0.335 -0.3813 KCenter 59.8297 ± 0.1742 0.0091 33.3717 ± 1.2132 0.1662 60.7625 ± 0.4269 0.1398 Table 6: Comparison of F1 of UE and the difference between baseline and UE in the low-resource setting on BioASQ, DROP, TQA, NewsQA, SearchQA, and NQ. The difference represents the F1 of UE subtracted from the F1 of the baseline. We marked the ranked F1 in first place in red, second place in blue, and third in green. 36 5 Discussion This chapter analyzes the experiment results and answers the research questions. 5.1 Baseline performance across various datasets In the regular setting,BALD,MeanSTD, andEntropyD become the top-performing strategies, attaining F1 scores of 79.9088%, 79.8626%, and 79.6104%, respectively. QA systems trained on the entire SQuAD dataset achieved 92.0671% F1. However, models trained on queried samples, including only 2.9% of the entire dataset, still obtained maximally 79.9088% F1, whose difference is only 12.1583% F1. Addition- ally, as described in 1, the score increases approximately 14.5% generally, illustrating the model underwent intensive growth, by AL’s strategic data selection. We observed that different strategies display varying impacts on distinct datasets. Table 4 demonstrated all datasets’ overall result across AL strategies in the low- resource setting.BatchBALD performed the best on BioASQ and achieved 60.2796% F1. MeanSTD outperformed on DROP and obtained 52.2442% F1. LC had the best performance with 46.7906% F1. Entropy extracted the profitable data from NewsQA and obtained 60.5174% F1 on the QAmodel. The best strategy on SearchQA is Margin and scored 38.0298% F1. QA model attained the best F1 on NQ using LCD, scoring 61.6221%. The dropout version performed better than the original version on LC and Entropy on NQ. Moreover, based on Figure 2, querying 200 samples results in a mild growth of approximately 1% to 2% F1 on most datasets, contributing to incremental benefits in QA model training. 5.2 BALD vs. BatchBALD BatchBALD is proposed to be the advanced version of BALD, particularly focus- ing on efficient sampling (Kirsch et al., 2019). As demonstrated in Table 7, Batch- BALD incomparably reduced querying run time compared to BALD when query- ing data from SQuAD. Despite a noteworthy 76.56% F1, which closely aligns with 37 BALD’s 79.9088% F1, BatchBALD excels particularly in SQuAD. However, its impact weakened while querying smaller datasets, indicating a threshold for effec- tive run-time reduction, requiring the unlabeled pool to surpass a certain size. Re- garding F1 performance, QA models implementing BatchBALD outperform those using BALD on BioASQ, TQA, and SearchQA. Notably, in the case of SearchQA, a substantial improvement of 6.7873% F1 marks it as a considerable advancement across the datasets in this study. Datasets AL Strategies F1 ± SD Time (hour) SQuAD BALD 79.9088 ± 0.2653 5.7 BatchBALD 76.56 ± 0.3879 0.45 BioASQ BALD 59.1574 ± 0.5738 0.09 BatchBALD 60.2796 ± 0.7107 0.1 DROP BALD 52.1718 ± 0.7187 0.1 BatchBALD 50.4278 ± 1.0094 0.1 TQA BALD 44.0773 ± 0.8951 0.17 BatchBALD 44.6456 ± 0.7648 0.2 NewsQA BALD 59.613 ± 0.1329 1.42 BatchBALD 59.5572 ± 0.1633 1.63 SearchQA BALD 30.9229 ± 0.7563 2.62 BatchBALD 37.7102 ± 0.7377 2.44 NQ BALD 60.9946 ± 0.2083 1.08 BatchBALD 60.2574 ± 0.3779 0.99 Table 7: Comparison of BALD and BatchBALD about F1 and querying time. 5.3 Unique context Selection Method (UC) UC and UE display divergent enhancement in the querying process across varied datasets. The results of UC in Table 3 surmise a complementarity between LC and UC, gaining the highest 79.847% F1, ranking third place among baselines in Table 4. UC remarkably advances Kmeans strategy on SQuAD, although its F1 score does not achieve a top-three position or exceed baselines. Table 5 evidenced the overall 38 performance of UC across datasets in the low-resource setting. Most F1 differences between UC and baseline among AL strategies is 0 on BioASQ and SearchQA and close to 0 on NewsQA and NQ. In Table 8, no data shares the context in SearchQA, and only a few samples in BioASQ and NQ have a shared context. Therefore, UC has proven to be less effective in them. The querying samples in them can easily have a unique context. UC improved MarginD, LCD, EntropyD (the dropout versions) on DROP; however, none of the final F1 scores surpass those of baselines. In TQA, UC intensively increases F1 when using MarginD, yet the first place in F1 is still LC in both baseline and UC ranking. AL Strategies BioASQ DROP TQA NewsQA SearchQA NQ # unique C. (% of D.) 1341(99%) 259(19%) 361(27%) 6471(65%) 10000(100%) 9517(95%) Random 0 73 69 0 0 0 Margin 1 92 76 6 0 7 LC 0 91 66 4 0 1 Entropy 0 92 61 6 0 1 MarginD 1 88 66 12 0 4 LeastConfD 0 83 72 7 0 3 EntropyD 0 101 68 8 0 3 KMeans 0 87 67 6 0 1 KCenter 1 91 79 4 0 1 MeanSTD 1 127 58 6 0 2 BALD 1 118 60 2 0 1 Badge 0 87 62 2 0 0 Table 8: Comparison of baseline and UC differences in the number of unique contexts in the query samples across multiple datasets. # unique C (% of D) indicates the difference in number and that number as a percentage of the entire dataset. 5.4 Unique Embedding Selection Method (UE) It is noteworthy that UE extensively raised the performance of LC, increasing 0.4212% F1 to achieve 80.008% F1, becoming the best score among all other strate- gies, as indicated in Table 3. UE evidenced the most excellent effectiveness using LC on SQuAD. When employed Margin on BioASQ, UE achieved the best 60.3954% 39 F1 and also outperformed as the first position among baselines on BioASQ, as de- tailed in Table 6 and 4. On TQA, UE obtained 47.6978% F1 utilizing LC, surpassing the highest F1 among baselines. The model using LCD gained the highest F1 in 61.7766% among UE and baseline strategies on NQ. The most substantial improve- ment, the increment of 1.0934% F1, was observed in EntropyD on DROP, although the score does not make the result reach the first place among F1 in UE. Significant improvements of UE, exceeding 0.5% F1, were inspected in BALD on BioASQ, EntropyD on DROP, Entropy on TQA, KMeans on SearchQA, and LCD on NQ. However, UC only achieved the more enormous improvements in MarginD on TQA and KMeans on SearchQA. This investigation exposes that the embedding is a more discriminative feature of data value than the context. 40 6 Conclusion Active learning is invaluable in a scenario with limited data resources. The QA sys- tem, learning from a queried set of data containing only around 15% of SQuAD, achieves impressive F1 over 70%. In the case of BatchBALD, the runtime is sub- stantially reduced by 83% compared to the querying process of BALD in the regular setting, but not significant in the low-resource setting. Selecting valuable data for training is not only beneficial in the regular setting but also in the low-resource set- ting. However, different dataset domains affect AL strategies differently, and there is no universally superior AL strategy for any QA datasets. Unique Context Se- lection Method and Unique Embedding Selection Method are designed to support AL strategies on QA systems. In contrast, the unique context distribution in some datasets limits the visibility of UC’s potential achievements. Nevertheless, the im- provements supported by UE are essential across various datasets. Not all kinds of data include context, making the UC enhancement method hard to generally employ on other tasks, especially those not in NLP. BADGE is compu- tationally expensive in the regular setting, so it is excluded in comparing strategies in the regular setting but included in the low-resource setting. Investigating the outcome of enhancing BatchBALD with UC or UE can be another exciting re- search direction. We compared the performance of QA systems across models, AL strategies, and datasets. A comprehensive survey with no performance comparison between multiple models, for example, BERT and RoBERTa Large, is worthwhile. Some strategies, such as Loss Prediction Loss (LPL), are omitted from this study due to time constraints during implementation. They are also worth further explor- ing in QA systems. A broader comparison of the performance of diverse dataset domains can be insightful to see how each domain’s features interact with different strategies. 41 References Charu C. Aggarwal, Xiangnan Kong, Quanquan Gu, Jiawei Han, and Philip S. Yu. Active Learning: A Survey. In Data Classification. Chapman and Hall/CRC, 2014. ISBN 978-0-429-10263-9. Num Pages: 36. Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. Deep Batch Active Learning by Diverse, Uncertain Gradi- ent Lower Bounds, February 2020. URL http://arxiv.org/abs/1906.03671. arXiv:1906.03671 [cs, stat]. William H. Beluch, Tim Genewein, Andreas Nürnberger, and Jan M. Köhler. The Power of Ensembles for Active Learning in Image Classification. pages 9368–9377, 2018. URL https://openaccess.thecvf.com/content{_}cvpr{_}2018/html/ Beluch{_}The{_}Power{_}of{_}CVPR{_}2018{_}paper.html. Yasaman Boreshban, Seyed Morteza Mirbostani, Gholamreza Ghassem-Sani, Seyed Abolghasem Mirroshandel, and Shahin Amiriparian. Improving question answering performance using knowledge distillation and active learning. Engineer- ing Applications of Artificial Intelligence, 123:106137, August 2023. ISSN 0952- 1976. doi: 10.1016/j.engappai.2023.106137. URL https://www.sciencedirect. com/science/article/pii/S0952197623003214. Arantxa Casanova, Pedro O. Pinheiro, Negar Rostamzadeh, and Christopher J. Pal. Reinforced active learning for image segmentation, February 2020. URL http://arxiv.org/abs/2002.06583. arXiv:2002.06583 [cs]. Yukun Chen, Thomas A. Lasko, Qiaozhu Mei, Joshua C. Denny, and Hua Xu. A study of active learning methods for named entity recognition in clinical text. Journal of Biomedical Informatics, 58:11–18, December 2015. ISSN 1532- 0464. doi: 10.1016/j.jbi.2015.09.010. URL https://www.sciencedirect.com/ science/article/pii/S1532046415002038. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre- 42 http://arxiv.org/abs/1906.03671 https://openaccess.thecvf.com/content{_}cvpr{_}2018/html/Beluch{_}The{_}Power{_}of{_}CVPR{_}2018{_}paper.html https://openaccess.thecvf.com/content{_}cvpr{_}2018/html/Beluch{_}The{_}Power{_}of{_}CVPR{_}2018{_}paper.html https://www.sciencedirect.com/science/article/pii/S0952197623003214 https://www.sciencedirect.com/science/article/pii/S0952197623003214 http://arxiv.org/abs/2002.06583 https://www.sciencedirect.com/science/article/pii/S1532046415002038 https://www.sciencedirect.com/science/article/pii/S1532046415002038 training of Deep Bidirectional Transformers for Language Understanding, May 2019. URL http://arxiv.org/abs/1810.04805. arXiv:1810.04805 [cs]. Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs, 2019. Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Guney, Volkan Cirik, and Kyunghyun Cho. SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine, June 2017. URL http://arxiv.org/abs/1704.05179. arXiv:1704.05179 [cs]. Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. Mrqa 2019 shared task: Evaluating generalization in reading comprehension, 2019. Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Repre- senting model uncertainty in deep learning, 2016. Yarin Gal, Riashat Islam, and Zoubin Ghahramani. Deep Bayesian Active Learning with Image Data. In Proceedings of the 34th International Conference on Machine Learning, pages 1183–1192. PMLR, July 2017. URL https://proceedings.mlr. press/v70/gal17a.html. ISSN: 2640-3498. Elmar Haussmann, Michele Fenzi, Kashyap Chitta, Jan Ivanecky, Hanson Xu, Donna Roy, Akshita Mittel, Nicolas Koumchatzky, Clement Farabet, and Jose M. Al- varez. Scalable Active Learning for Object Detection. In 2020 IEEE Intelligent Vehicles Symposium (IV), pages 1430–1435, October 2020. doi: 10.1109/IV47402. 2020.9304793. ISSN: 2642-7214. Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel. Bayesian Active Learning for Classification and Preference Learning, December 2011. URL http://arxiv.org/abs/1112.5745. arXiv:1112.5745 [cs, stat]. Eyke Hüllermeier and Willem Waegeman. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine Learning, 110:457–506, 2021. 43 http://arxiv.org/abs/1810.04805 http://arxiv.org/abs/1704.05179 https://proceedings.mlr.press/v70/gal17a.html https://proceedings.mlr.press/v70/gal17a.html http://arxiv.org/abs/1112.5745 Josip Jukić and Jan Šnajder. Parameter-efficient language model tuning with active learning in low-resource settings, 2023. Daniel Jurafsky and James H. Martin. Speech and Language Processing: An Intro- duction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Prentice Hall PTR, USA, 1st edition, 2000. ISBN 0130950696. Michael Kampffmeyer, Arnt-Borre Salberg, and Robert Jenssen. Semantic segmenta- tion of small objects and modeling of uncertainty in urban remote sensing images using deep convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2016. Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5376–5384, 2017. doi: 10.1109/CVPR.2017.571. Andreas Kirsch, Joost van Amersfoort, and Yarin Gal. Batchbald: Efficient and di- verse batch acquisition for deep bayesian active learning. CoRR, abs/1906.08158, 2019. URL http://arxiv.org/abs/1906.08158. Oleksandr Kolomiyets and Marie-Francine Moens. A survey on question an- swering technology from an information retrieval perspective. Information Sci- ences, 181(24):5412–5434, December 2011. ISSN 0020-0255. doi: 10.1016/j.ins. 2011.07.047. URL https://www.sciencedirect.com/science/article/pii/ S0020025511003860. Punit Kumar and Atul Gupta. Active Learning Query Strategies for Classification, Regression, and Clustering: A Survey. Journal of Computer Science and Technol- ogy, 35(4):913–945, July 2020. ISSN 1860-4749. doi: 10.1007/s11390-020-9487-4. URL https://doi.org/10.1007/s11390-020-9487-4. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton 44 http://arxiv.org/abs/1906.08158 https://www.sciencedirect.com/science/article/pii/S0020025511003860 https://www.sciencedirect.com/science/article/pii/S0020025511003860 https://doi.org/10.1007/s11390-020-9487-4 Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, An- drew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics, 7:453–466, November 2019. ISSN 2307-387X. doi: 10.1162/tacl a 00276. URL https://direct.mit.edu/tacl/article/43518. Lianghao Li, Xiaoming Jin, Sinno Jialin Pan, and Jian-Tao Sun. Multi-domain ac- tive learning for text classification. In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining, KDD ’12, pages 1086–1094, New York, NY, USA, August 2012. Association for Comput- ing Machinery. ISBN 978-1-4503-1462-6. doi: 10.1145/2339530.2339701. URL https://dl.acm.org/doi/10.1145/2339530.2339701. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A Robustly Optimized BERT Pretraining Approach, July 2019. URL http: //arxiv.org/abs/1907.11692. arXiv:1907.11692 [cs]. Katerina Margatina, Löıc Barrault, and Nikolaos Aletras. On the Importance of Effectively Adapting Pretrained Language Models for Active Learning, March 2022. URL http://arxiv.org/abs/2104.08320. arXiv:2104.08320 [cs]. Salvador Marti Roman. Active learning for extractive question answering, 2022a. Salvador Marti Roman. Active Learning for Extractive Question Answering. PhD thesis, 2022b. URL https://urn.kb.se/resolve?urn=urn:nbn:se:liu: diva-186761. Diego Mollá and José Luis Vicedo. Question Answering in Restricted Domains: An Overview. Computational Linguistics, 33(1):41–61, 03 2007. ISSN 0891-2017. doi: 10.1162/coli.2007.33.1.41. URL https://doi.org/10.1162/coli.2007.33. 1.41. Usman Naseem, Matloob Khushi, Shah Khalid Khan, Kamran Shaukat, and Moham- mad Ali Moni. A Comparative Analysis of Active Learning for Biomedical Text 45 https://direct.mit.edu/tacl/article/43518 https://dl.acm.org/doi/10.1145/2339530.2339701 http://arxiv.org/abs/1907.11692 http://arxiv.org/abs/1907.11692 http://arxiv.org/abs/2104.08320 https://urn.kb.se/resolve?urn=urn:nbn:se:liu:diva-186761 https://urn.kb.se/resolve?urn=urn:nbn:se:liu:diva-186761 https://doi.org/10.1162/coli.2007.33.1.41 https://doi.org/10.1162/coli.2007.33.1.41 Mining. Applied System Innovation, 4(1):23, March 2021. ISSN 2571-5577. doi: 10.3390/asi4010023. URL https://www.mdpi.com/2571-5577/4/1/23. Number: 1 Publisher: Multidisciplinary Digital Publishing Institute. Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and An- drew Y. Ng. Reading Digits in Natural Images with Unsupervised Fea- ture Learning. In NIPS Workshop on Deep Learning and Unsupervised Fea- ture Learning 2011, 2011. URL http://ufldl.stanford.edu/housenumbers/ nips2011{_}housenumbers.pdf. Vu-Linh Nguyen, Sébastien Destercke, and Eyke Hüllermeier. Epistemic uncertainty sampling, 2019. Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ Questions for Machine Comprehension of Text, June 2016. URL https: //arxiv.org/abs/1606.05250v3. Pranav Rajpurkar, Robin Jia, and Percy Liang. Know What You Don’t Know: Unanswerable Questions for SQuAD, June 2018. URL http://arxiv.org/abs/ 1806.03822. arXiv:1806.03822 [cs]. Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B. Gupta, Xiaojiang Chen, and Xin Wang. A Survey of Deep Active Learning. ACM Computing Surveys, 54(9):180:1–180:40, October 2021. ISSN 0360-0300. doi: 10.1145/3472291. URL https://dl.acm.org/doi/10.1145/3472291. Isah Charles Saidu and Lehel Csató. Active Learning with Bayesian UNet for Effi- cient Semantic Image Segmentation. Journal of Imaging, 7(2):37, February 2021. ISSN 2313-433X. doi: 10.3390/jimaging7020037. URL https://www.mdpi.com/ 2313-433X/7/2/37. Ozan Sener and Silvio Savarese. Active Learning for Convolutional Neural Networks: A Core-Set Approach, June 2018. URL http://arxiv.org/abs/1708.00489. arXiv:1708.00489 [cs, stat]. 46 https://www.mdpi.com/2571-5577/4/1/23 http://ufldl.stanford.edu/housenumbers/nips2011{_}housenumbers.pdf http://ufldl.stanford.edu/housenumbers/nips2011{_}housenumbers.pdf https://arxiv.org/abs/1606.05250v3 https://arxiv.org/abs/1606.05250v3 http://arxiv.org/abs/1806.03822 http://arxiv.org/abs/1806.03822 https://dl.acm.org/doi/10.1145/3472291 https://www.mdpi.com/2313-433X/7/2/37 https://www.mdpi.com/2313-433X/7/2/37 http://arxiv.org/abs/1708.00489 Robin Senge, Stefan Bösner, Krzysztof Dembczyński, Jörg Haasenritter, Oliver Hirsch, Norbert Donner-Banzhoff, and Eyke Hüllermeier. Reliable classification: Learning classifiers that distinguish aleatoric and epistemic uncertainty. Informa- tion Sciences, 255:16–29, 2014. Burr Settles. Active Learning Literature Survey. Technical Report, University of Wisconsin-Madison Department of Computer Sciences, 2009. URL https:// minds.wisconsin.edu/handle/1793/60660. Accepted: 2012-03-15T17:23:56Z. C. E. Shannon. A mathematical theory of communication. The Bell System Techni- cal Journal, 27(3):379–423, July 1948. ISSN 0005-8580. doi: 10.1002/j.1538-7305. 1948.tb01338.x. Conference Name: The Bell System Technical Journal. Yanyao Shen, Hyokun Yun, Zachary C. Lipton, Yakov Kronrod, and Animashree Anandkumar. Deep Active Learning for Named Entity Recognition, February 2018. URL http://arxiv.org/abs/1707.05928. arXiv:1707.05928 [cs]. Simon Tong. Active learning: theory and applications. PhD thesis, 2001. Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. NewsQA: A Machine Comprehension Dataset, February 2017. URL http://arxiv.org/abs/1611.09830. arXiv:1611.09830 [cs] version: 3. George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. An overview of the bioasq large- scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 16(1):1–28, 2015. Dan Wang and Yi Shang. A new active labeling method for deep learning. In 2014 International Joint Conference on Neural Networks (IJCNN), pages 112–119, July 2014. doi: 10.1109/IJCNN.2014.6889457. ISSN: 2161-4407. 47 https://minds.wisconsin.edu/handle/1793/60660 https://minds.wisconsin.edu/handle/1793/60660 http://arxiv.org/abs/1707.05928 http://arxiv.org/abs/1611.09830 Yichen Xie, Masayoshi Tomizuka, and Wei Zhan. Towards General and Effi- cient Active Learning, March 2022. URL http://arxiv.org/abs/2112.07963. arXiv:2112.07963 [cs]. Jianfei Yu, Minghui Qiu, Jing Jiang, Jun Huang, Shuangyong Song, Wei Chu, and Haiqing Chen. Modelling Domain Relationships for Transfer Learning on Retrieval-based Question Answering Systems in E-commerce. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, WSDM ’18, pages 682–690, New York, NY, USA, February 2018. Association for Computing Machinery. ISBN 978-1-4503-5581-0. doi: 10.1145/3159652.3159685. URL https://dl.acm.org/doi/10.1145/3159652.3159685. Xueying Zhan, Qingzhong Wang, Kuan-hao Huang, Haoyi Xiong, Dejing Dou, and Antoni B. Chan. A Comparative Survey of Deep Active Learning, July 2022. URL http://arxiv.org/abs/2203.13450. arXiv:2203.13450 [cs]. Ye Zhang, Matthew Lease, and Byron Wallace. Active Discriminative Text Rep- resentation Learning. Proceedings of the AAAI Conference on Artificial Intelli- gence, 31(1), February 2017. ISSN 2374-3468. doi: 10.1609/aaai.v31i1.10962. URL https://ojs.aaai.org/index.php/AAAI/article/view/10962. Number: 1. Zhisong Zhang, Emma Strubell, and Eduard Hovy. A Survey of Active Learning for Natural Language Processing, February 2023. URL http://arxiv.org/abs/ 2210.10109. arXiv:2210.10109 [cs]. 48 http://arxiv.org/abs/2112.07963 https://dl.acm.org/doi/10.1145/3159652.3159685 http://arxiv.org/abs/2203.13450 https://ojs.aaai.org/index.php/AAAI/article/view/10962 http://arxiv.org/abs/2210.10109 http://arxiv.org/abs/2210.10109 A The results of each experiment with the score of each querying iteration. A.1 Regular Setting Figure 3: Performance of AL strategies enhanced with UC on each data querying size in the regular setting. Figure 4: Performance of AL strategies enhanced with UE on each data querying size in the regular setting. A.2 Low-resource Setting 49 (a) BioASQ (b) DROP (c) TQA (d) NewsQA (e) SearchQA (f) NQ Figure 5: Performance of AL strategies enhanced with UC on each data querying size in the low-resource setting. 50 (a) BioASQ (b) DROP (c) TQA (d) NewsQA (e) SearchQA (f) NQ Figure 6: Performance of AL strategies enhanced with UE on each data querying size in the low-resource setting. 51 B The results of each experiment with the stan- dard deviation. B.1 Baseline Figure 7: Performance with a standard deviation of AL strategies on SQuAD. 52 Figure 8: Performance with a standard deviation of AL strategies on BioASQ Figure 9: Performance with a standard deviation of AL strategies on DROP 53 Figure 10: Performance with a standard deviation of AL strategies on TextbookQA Figure 11: Performance with a standard deviation of AL strategies on NewsQA 54 Figure 12: Performance with a standard deviation of AL strategies on SearchQA 55 Figure 13: Performance with a standard deviation of AL strategies on NQ B.2 Unique Context (UC) 56 Figure 14: Performance with a standard deviation of AL strategies enhanced with UC on SQuAD. 57 Figure 15: Performance with a standard deviation of AL strategies enhanced with UC on BioASQ. 58 Figure 16: Performance with a standard deviation of AL strategies enhanced with UC on DROP. 59 Figure 17: Performance with a standard deviation of AL strategies enhanced with UC on TQA. 60 Figure 18: Performance with a standard deviation of AL strategies enhanced with UC on NewsQA. 61 Figure 19: Performance with a standard deviation of AL strategies enhanced with UC on SearchQA. 62 Figure 20: Performance with a standard deviation of AL strategies enhanced with UC on NQ. 63 B.3 Unique Embedding (UE) Figure 21: Performance with a standard deviation of AL strategies enhanced with UE on SQuAD. 64 Figure 22: Performance with a standard deviation of AL strategies enhanced with UE on BioASQ. 65 Figure 23: Performance with a standard deviation of AL strategies enhanced with UE on DROP. 66 Figure 24: Performance with a standard deviation of AL strategies enhanced with UE on TQA. 67 Figure 25: Performance with a standard deviation of AL strategies enhanced with UE on NewsQA. 68 Figure 26: Performance with a standard deviation of AL strategies enhanced with UE on SearchQA. 69 Figure 27: Performance with a standard deviation of AL strategies enhanced with UE on NQ. C Comparison 70 Figure 28: Comparing the behavior of AL strategies on SQuAD in the regular setting, with consideration given to the Baseline, UC, and UE. 71 Figure 29: Comparing the behavior of AL strategies on BioASQ in the low-resource setting, with consideration given to the Baseline, UC, and UE. 72 Figure 30: Comparing the behavior of AL strategies on DROP in the low-resource setting, with consideration given to the Baseline, UC, and UE. 73 Figure 31: Comparing the behavior of AL strategies on TQA in the low-resource setting, with consideration given to the Baseline, UC, and UE. 74 Figure 32: Comparing the behavior of AL strategies on NewsQA in the low-resource setting, with consideration given to the Baseline, UC, and UE. 75 Figure 33: Comparing the behavior of AL strategies on SearchQA in the low-resource setting, with consideration given to the Baseline, UC, and UE. 76 Figure 34: Comparing the behavior of AL strategies on NQ in the low-resource setting, with consideration given to the Baseline, UC, and UE. 77