05 Fakultät Informatik, Elektrotechnik und Informationstechnik

Permanent URI for this collectionhttps://elib.uni-stuttgart.de/handle/11682/6

Browse

Search Results

Now showing 1 - 10 of 48
  • Thumbnail Image
    ItemOpen Access
    Locking-enabled security analysis of cryptographic circuits
    (2024) Upadhyaya, Devanshi; Gay, Maël; Polian, Ilia
    Hardware implementations of cryptographic primitives require protection against physical attacks and supply chain threats. This raises the question of secure composability of different attack countermeasures, i.e., whether protecting a circuit against one threat can make it more vulnerable against a different threat. In this article, we study the consequences of applying logic locking, a popular design-for-trust solution against intellectual property piracy and overproduction, to cryptographic circuits. We show that the ability to unlock the circuit incorrectly gives the adversary new powerful attack options. We introduce LEDFA (locking-enabled differential fault analysis) and demonstrate for several ciphers and families of locking schemes that fault attacks become possible (or consistently easier) for incorrectly unlocked circuits. In several cases, logic locking has made circuit implementations prone to classical algebraic attacks with no fault injection needed altogether. We refer to this “zero-fault” version of LEDFA by the term LEDA, investigate its success factors in-depth and propose a countermeasure to protect the logic-locked implementations against LEDA. We also perform test vector leakage assessment (TVLA) of incorrectly unlocked AES implementations to show the effects of logic locking regarding side-channel leakage. Our results indicate that logic locking is not safe to use in cryptographic circuits, making them less rather than more secure.
  • Thumbnail Image
    ItemOpen Access
    Development of an infrastructure for creating a behavioral model of hardware of measurable parameters in dependency of executed software
    (2021) Schwachhofer, Denis
    System-Level Test (SLT) gains traction not only in the industry but as of recently also in academia. It is used to detect manufacturing defects not caught by previous test steps. The idea behind SLT is to embed the Design Under Test (DUT) in an environment and running software on it that corresponds to its end-user application. But even though it is increasingly used in manufacturing since a decade there are still many open challenges to solve. For example, there is no coverage metric for SLT. Also, tests are not automatically generated but manually composed using existing operating systems and programs. This master thesis introduces the foundation for the AutoGen project, that will tackle the aforementioned challenges in the future. This foundation contains a platform for experiments and a workflow to generate Systems-on-Chip (SoCs). A case study is conducted to show an example on how on-chip sensors can be used in SLT applications to replace missing detailed technology-information. For the case study a “power devil” application has been developed that aims to keep the temperature of the Field Programmable Gate Array (FPGA) it runs on in a target range. The study shows an example on how software and parameters influence the extra-functional behavior of hardware.
  • Thumbnail Image
    ItemOpen Access
    Rigorous compilation for near-term quantum computers
    (2024) Brandhofer, Sebastian; Polian, Ilia (Prof.)
    Quantum computing promises an exponential speedup for computational problems in material sciences, cryptography and drug design that are infeasible to resolve by traditional classical systems. As quantum computing technology matures, larger and more complex quantum states can be prepared on a quantum computer, enabling the resolution of larger problem instances, e.g. breaking larger cryptographic keys or modelling larger molecules accurately for the exploration of novel drugs. Near-term quantum computers, however, are characterized by large error rates, a relatively low number of qubits and a low connectivity between qubits. These characteristics impose strict requirements on the structure of quantum computations that must be incorporated by compilation methods targeting near-term quantum computers in order to ensure compatibility and yield highly accurate results. Rigorous compilation methods have been explored for addressing these requirements as they exactly explore the solution space and thus yield a quantum computation that is optimal with respect to the incorporated requirements. However, previous rigorous compilation methods demonstrate limited applicability and typically focus on one aspect of the imposed requirements, i.e. reducing the duration or the number of swap gates in a quantum computation. In this work, opportunities for improving near-term quantum computations through compilation are explored first. These compilation opportunities are included in rigorous compilation methods to investigate each aspect of the imposed requirements, i.e. the number of qubits, connectivity of qubits, duration and incurred errors. The developed rigorous compilation methods are then evaluated with respect to their ability to enable quantum computations that are otherwise not accessible with near-term quantum technology. Experimental results demonstrate the ability of the developed rigorous compilation methods to extend the computational reach of near-term quantum computers by generating quantum computations with a reduced requirement on the number and connectivity of qubits as well as reducing the duration and incurred errors of performed quantum computations. Furthermore, the developed rigorous compilation methods extend their applicability to quantum circuit partitioning, qubit reuse and the translation between quantum computations generated for distinct quantum technologies. Specifically, a developed rigorous compilation method exploiting the structure of a quantum computation to reuse qubits at runtime yielded a reduction in the required number of qubits of up to 5x and result error by up to 33%. The developed quantum circuit partitioning method optimally distributes a quantum computation to distinct separate partitions, reducing the required number of qubits by 40% and the cost of partitioning by 41% on average. Furthermore, a rigorous compilation method was developed for quantum computers based on neutral atoms that combines swap gate insertions and topology changes to reduce the impact of limited qubit connectivity on the quantum computation duration by up to 58% and on the result fidelity by up to 29%. Finally, the developed quantum circuit adaptation method enables to translate between distinct quantum technologies while considering heterogeneous computational primitives with distinct characteristics to reduce the idle time of qubits by up to 87% and the result fidelity by up to 40%.
  • Thumbnail Image
    ItemOpen Access
    Design for reliability in advanced technologies using machine learning
    (2024) Klemme, Florian; Amrouch, Hussam (Prof. Dr.-Ing.)
    This thesis focuses on the standard cell library, which is one of the core entities in the digital circuit design flow, to demonstrate the challenges and opportunities of advanced technology nodes. The standard cell library serves as a technology interface between the foundry and the circuit designer, enabling automatic mapping of high-level circuit descriptions to the technology of the foundry through the process of logic synthesis. In the past decade, the standard cell library has been continuously adapted to keep up with the demands of shrinking process nodes. This includes, e.g., the integration of more accurate timing models, process variation, or signal integrity for cross-talk and noise in the circuit. This thesis takes this development to the next level and presents approaches to bring machine learning and transistor self-heating into the standard cell library.
  • Thumbnail Image
    ItemOpen Access
    Exploring stochastic computing for edge computing : from architectures to applications
    (2025) Sengupta, Roshwin; Polian, Ilia (Prof. Dr.)
    Der wachsende Bedarf an energieeffizienter Signalverarbeitung und Klassifikation in Edge- und Near-Sensor-Systemen erfordert die Entwicklung kompakter, stromsparender Hardwarelösungen, die unabhängig von der Cloud betrieben werden können. Herkömmliche binäre Implementierungen digitaler Filter und neuronaler Netzwerke sind zwar genau, jedoch häufig ressourcenintensiv und daher weniger geeignet für solche energie- und flächenkritischen Umgebungen. Stochastic Computing (SC) hat sich als vielversprechende Alternative erwiesen, da es durch die Verwendung probabilistischer Bitströme und vereinfachter arithmetischer Einheiten erhebliche Einsparungen bei Fläche und Energie ermöglicht. Diese Arbeit untersucht den Einsatz von SC in verschiedenen Signalverarbeitungs- und neuronalen Netzwerkarchitekturen. Beginnend mit dem Entwurf SC-basierter digitaler Filter, einschließlich Finite- und Infinite-Impulse-Response-Varianten (FIR und IIR), wurde der Einfluss unterschiedlicher stochastischer Zahlengeneratoren (SNGs) und Adderarchitekturen analysiert. Es konnte gezeigt werden, dass SC-Filter in fehlerfreien Szenarien die Fläche um bis zu 49% und den Energieverbrauch um bis zu 64% reduzieren können, bei nur geringem Genauigkeitsverlust gegenüber binären Referenzdesigns. Aufbauend auf diesen Erkenntnissen wurden eine SC-basierte Fast Fourier Transform (SCFFT) sowie eine neuartige SC-basierte Continuous Wavelet Transform (SCWT) für die Analyse nichtstationärer Signale entwickelt. Diese Entwürfe erreichen Energieeinsparungen von 60-80% und bieten somit eine effiziente Alternative zu konventionellen Implementierungen in ultraniedrigleistungsfähigen Systemen. Zur Lösung von Klassifikationsaufgaben in Edge-Systemen wurde SC auch auf Long Short-Term Memory (LSTM)-Netzwerke erweitert. Durch eine Designraum-Analyse von vollständig binären, vollständig stochastischen und hybriden LSTM-Architekturen konnte gezeigt werden, dass vollständig stochastische LSTMs Einsparungen von bis zu 47% bei der Fläche und 86% beim Energieverbrauch erzielen, bei nur minimalem Genauigkeitsverlust. Zudem wurde der Einfluss von Aktivierungsfunktionen wie ReLU und tanh im SC-Kontext untersucht, wobei sich zeigte, dass ihre Auswahl einen wesentlichen Einfluss auf Effizienz und Leistung der Netzwerke hat. Da reale Edge-Anwendungen häufig mit unsicheren Energiebedingungen und störbehafteten Umgebungen konfrontiert sind, wurde in dieser Arbeit auch die Fehlertoleranz SCbasierter Architekturen umfassend analysiert. Durch gezielte Injektion von Bitfehlern in kritischen Komponenten wie SNGs, Addierwerken oder Aktivierungsfunktionen wurde der Einfluss auf Genauigkeit und Robustheit untersucht. Die Experimente zeigten, dass unterschiedliche Designentscheidungen, etwa die Wahl des SNG-Typs oder der Adderstruktur, erheblichen Einfluss auf die Fehlerresilienz haben. Das bedeutet, dass Fehlertoleranz in SC nicht automatisch gegeben ist, sondern durch sorgfältige Architekturentscheidungen explizit gestaltet werden muss. Beispielsweise übertreffen unsere SC-FIR-Filter unter moderaten Fehlerbedingungen sogar binäre Filter mit Triple Modular Redundancy (TMR). Auch bei LSTM-Netzen zeigt sich, dass Konfigurationen mit Sobol-basierten SNGs und tanh-Aktivierung unter Fehlerinjektion besonders robust sind. Eine Erhöhung der Bitstromlänge verbessert zwar die Robustheit, erhöht jedoch auch die Latenz, was die Notwendigkeit eines gezielten Designs unter Abwägung von Fläche, Energie, Genauigkeit und Fehlertoleranz unterstreicht. Basierend auf diesen Erkenntnissen wurde das Wavelet-Assisted Stochastic-Enabled Neural Network (WASENN) für die menschliche Aktivitätserkennung (HAR) vorgestellt. WASENN kombiniert SC-basierte convolutional Neural Netwerk (CNN)- und LSTMSchichten mit einer Wavelet-Vorverarbeitung und ermöglicht eine präzise und energieeffiziente Klassifikation auf ressourcenbegrenzten Geräten. Evaluierungen auf den Datensätzen UCI HAR und WISDM zeigten, dass die Wavelet-Vorverarbeitung sowohl die Klassifikationsgenauigkeit als auch die Hardwarekompaktheit verbessert. Gleichzeitig reduziert der Einsatz von SC den Flächenbedarf um 32% und den Energieverbrauch um 74%, bei nur minimalem Verlust an Klassifikationsgenauigkeit. Abschließend liefert diese Dissertation eine umfassende Untersuchung stochastischen Rechnens als praktikable Entwurfsstrategie für energieeffiziente, fehlertolerante und kompakte Hardwarearchitekturen für Signalverarbeitung und neuronale Netzwerke. Durch Innovationen im Filterentwurf, in der Wavelettransformation, in sequenziellen Netzmodellen sowie in der Systemintegration wird der Weg geebnet für den robusten Einsatz von intelligenter Datenverarbeitung direkt am Sensor in zukünftigen Edge-Anwendungen.
  • Thumbnail Image
    ItemOpen Access
    A GPU-accelerated light-field super-resolution framework based on mixed noise model and weighted regularization
    (2022) Tran, Trung-Hieu; Sun, Kaicong; Simon, Sven
    Light-field (LF) super-resolution (SR) plays an essential role in alleviating the current technology challenge in the acquisition of a 4D LF, which assembles both high-density angular and spatial information. Due to the algorithm complexity and data-intensive property of LF images, LFSR demands a significant computational effort and results in a long CPU processing time. This paper presents a GPU-accelerated computational framework for reconstructing high-resolution (HR) LF images under a mixed Gaussian-Impulse noise condition. The main focus is on developing a high-performance approach considering processing speed and reconstruction quality. From a statistical perspective, we derive a joint ℓ1- ℓ2data fidelity term for penalizing the HR reconstruction error taking into account the mixed noise situation. For regularization, we employ the weighted non-local total variation approach, which allows us to effectively realize LF image prior through a proper weighting scheme. We show that the alternating direction method of the multipliers algorithm (ADMM) can be used to simplify the computation complexity and results in a high-performance parallel computation on the GPU Platform. An extensive experiment is conducted on both synthetic 4D LF dataset and natural image dataset to validate the proposed SR model’s robustness and evaluate the accelerated optimizer’s performance. The experimental results show that our approach achieves better reconstruction quality under severe mixed-noise conditions as compared to the state-of-the-art approaches. In addition, the proposed approach overcomes the limitation of the previous work in handling large-scale SR tasks. While fitting within a single off-the-shelf GPU, the proposed accelerator provides an average speedup of 2.46 ×and 1.57 ×for ×2and ×3SR tasks, respectively. In addition, a speedup of 77×is achieved as compared to CPU execution.
  • Thumbnail Image
    ItemOpen Access
    Modeling and investigating total ionizing dose impact on FeFET
    (2023) Sayed, Munazza; Ni, Kai; Amrouch, Hussam
  • Thumbnail Image
    ItemOpen Access
    Thermal effects on monolithic 3D ferroelectric transistors for deep neural networks performance
    (2024) Kumar, Shubham; Chauhan, Yogesh Singh; Amrouch, Hussam
    Monolithic three‐dimensional (M3D) integration advances integrated circuits by enhancing density and energy efficiency. Ferroelectric thin‐film transistors (Fe‐TFTs) attract attention for neuromorphic computing and back‐end‐of‐the‐line (BEOL) compatibility. However, M3D faces challenges like increased runtime temperatures due to limited heat dissipation, impacting system reliability. This work demonstrates the effect of temperature impact on single‐gate (SG) Fe‐TFT reliability. SG Fe‐TFTs have limitations such as read‐disturbance and small memory windows, constraining their use. To mitigate these, dual‐gate (DG) Fe‐TFTs are modeled using technology computer‐aided design, comparing their performance. Compute‐in‐memory (CIM) architectures with SG and DG Fe‐TFTs are investigated for deep neural networks (DNN) accelerators, revealing heat's detrimental effect on reliability and inference accuracy. DG Fe‐TFTs exhibit about 4.6x higher throughput than SG Fe‐TFTs. Additionally, thermal effects within the simulated M3D architecture are analyzed, noting reduced DNN accuracy to 81.11% and 67.85% for SG and DG Fe‐TFTs, respectively. Furthermore, various cooling methods and their impact on CIM system temperature are demonstrated, offering insights for efficient thermal management strategies.
  • Thumbnail Image
    ItemOpen Access
    Secure cryptographic hardware : assessing logic-locking and fault attack vulnerabilities
    (2025) Upadhyaya, Devanshi; Polian, Ilia (Prof. Dr. rer. nat. habil.)
    The protection of hardware implementations of cryptographic primitives against physical attacks and supply-chain threats remains a critical challenge. This thesis investigates the fault attack vulnerabilities and the secure composability of various countermeasures, with a particular focus on logic-locking - a widely adopted design-for-trust technique aimed at safeguarding against intellectual property piracy and overproduction. One of the primary objectives of this work is to explore whether protecting a circuit against one threat inadvertently makes it more vulnerable to another, particularly when logic locking is applied to cryptographic circuits. Two novel attacks that exploit the presence of logic-locking circuitry are introduced as a major contribution of this thesis. Logic-locking typically serves to protect circuits by allowing them to function only when the correct locking key is provided. However, it is demonstrated that the ability to unlock the circuit incorrectly can provide adversaries with new and effective attack vectors. The first attack, Locking Enabled Differential Fault Analysis (LEDFA), is shown to make incorrectly unlocked circuits more susceptible to fault attacks due to the introduction of new propagation paths by the logic-locking circuitry. Experimental evaluations across various ciphers and logic-locking schemes revealed that fault attacks become either possible or consistently easier in the presence of incorrect unlocking. Moreover, it was found that logic-locking can, in some cases, make circuits vulnerable to classical algebraic attacks without the need for any fault injection, a case referred to as Locking Enabled Differential Analysis (LEDA). This vulnerability results in a significant reduction in the cryptographic strength. The success factors behind LEDA are thoroughly investigated, leading to the proposal of a countermeasure designed to enhance the resilience of logic-locked cryptographic circuits. This countermeasure involves restricting cryptographic key bits from being directly integrated into locking subcircuits, thereby mitigating the vulnerabilities facilitating LEDA. Additionally, a Test Vector Leakage Assessment (TVLA) of incorrectly unlocked AES implementation is discussed, highlighting that logic-locking significantly influences side-channel leakage. These findings raise concerns regarding the use of logic-locking in cryptographic circuits, suggesting that it, in fact, compromises rather than enhances security. The second major contribution of this thesis is the development of a methodology for evaluating the vulnerability of cryptographic circuits to fault injection attacks facilitated by clock manipulation. It is well recognized that state-of-the-art fault attacks typically require either a large number of low-precision fault injections (statistical attacks) or very few injections using sophisticated equipment (algebraic attacks) to breach modern cryptosystems. For instance, a well-known fault attack on AES-128 requires only a single fault injection, provided that the fault effects are confined to a specific 8-bit nibble of the state. This research aimed to optimize the probability of achieving the desired faulty state bit patterns during low-cost clock manipulation, thereby combining the advantages of both statistical and algebraic attacks. For this purpose, a comprehensive methodology is developed, which involves extending formal Boolean satisfiability (SAT) models initially designed for waveform-accurate automatic test pattern generation (ATPG) procedures to fault attacks on cryptographic hardware. A distinguishing feature of this analysis is the presence of fixed-yet-unknown secret cryptographic bits that influence the faulty state bit patterns. A model-counting (MC) approach is utilized to calculate the probability of success across different secret cryptographic bit combinations using a novel Vulnerability Index (VI). This methodology provides a robust framework for assessing the susceptibility of cryptographic circuits to such fault injection attacks. The practical implications of these findings are significant for both cryptographic hardware designers and security analysts. A structured approach is offered for security analysts to evaluate and strengthen cryptographic systems against fault injection attacks, ensuring a comprehensive defense strategy.
  • Thumbnail Image
    ItemOpen Access
    Review on resistive switching devices based on multiferroic BiFeO3
    (2023) Zhao, Xianyue; Menzel, Stephan; Polian, Ilia; Schmidt, Heidemarie; Du, Nan
    This review provides a comprehensive examination of the state-of-the-art research on resistive switching (RS) in BiFeO3 (BFO)-based memristive devices. By exploring possible fabrication techniques for preparing the functional BFO layers in memristive devices, the constructed lattice systems and corresponding crystal types responsible for RS behaviors in BFO-based memristive devices are analyzed. The physical mechanisms underlying RS in BFO-based memristive devices, i.e., ferroelectricity and valence change memory, are thoroughly reviewed, and the impact of various effects such as the doping effect, especially in the BFO layer, is evaluated. Finally, this review provides the applications of BFO devices and discusses the valid criteria for evaluating the energy consumption in RS and potential optimization techniques for memristive devices.