<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Papers on COMPASSLAB</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/</link><description>Recent content in Papers on COMPASSLAB</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 29 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://skku-compasslab.github.io/compasslab/publications/papers/index.xml" rel="self" type="application/rss+xml"/><item><title>Token Filtering: Online Attention Pruning via KV Similarity for Efficient LLM Inference</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260929-neurips-token_filtering/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260929-neurips-token_filtering/</guid><description>Token Filtering: Efficient LLM Inference via Online Attention Pruning Based on KV Similarity Modern LLMs deliver strong performance but suffer from high inference latency and large memory footprints, especially during long-context decoding.</description></item><item><title>HARMONY: Cooperative Memory Scheduling for CXL Memory Systems</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260806-pact-harmony/</link><pubDate>Thu, 06 Aug 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260806-pact-harmony/</guid><description>HARMONY: Cooperative Memory Scheduling for CXL Memory Systems CXL memory expands capacity, but the host and the CMM schedule requests using different information.</description></item><item><title>Clover: Storage-Efficient Page Table Replication in Wafer-Scale GPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260728-cal-clover/</link><pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260728-cal-clover/</guid><description>Abstract Wafer-scale GPUs integrate many GPU modules (GPMs) on a single wafer, but their distributed organization makes the address translation non-uniform: on an L2 TLB miss, a page table walker may fetch a leaf-level PTE from a remote GPM across the on-wafer mesh.</description></item><item><title>NeuroMTA: Programmable Simulation Framework for Multi-Tile NPU Architectures</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260720-cal-neuromta/</link><pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260720-cal-neuromta/</guid><description>Abstract Multi-Tile Accelerators (MTAs) have emerged as a promising paradigm to scale the computational throughput of NPUs for modern deep learning workloads.</description></item><item><title>Reducing Metadata and Page Migration Overheads in CXL Based Secure Tiered Memory</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260720-cal-secure-tieredmem/</link><pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260720-cal-secure-tieredmem/</guid><description>Abstract Securing CXL-based tiered memory introduces substantial overhead. To ensure the confidentiality, integrity, and freshness of data stored in CXL memory, transmitting security metadata over the CXL link consumes link bandwidth and increases memory access latency.</description></item><item><title>Sparsity-Aware Fault Tolerance for Reliable CNN Inference on GPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260712-tpds-spars/</link><pubDate>Sun, 12 Jul 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260712-tpds-spars/</guid><description>Abstract With the increasing deployment of Convolutional Neural Networks (CNNs) in mission-critical and safety-critical applications, ensuring fault tolerance during model inference has become a critical requirement.</description></item><item><title>Don't Get Stuck in Traffic: A Case for Contention-Aware Translation Offloading</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260710-micro-cato/</link><pubDate>Wed, 08 Jul 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260710-micro-cato/</guid><description>Don&amp;rsquo;t Get Stuck in Traffic: A Case for Contention-Aware Translation Offloading CXL-attached memory (CMM) is becoming the standard way to scale memory capacity beyond the limits of DDR channels.</description></item><item><title>Boosting LLC Bandwidth Utilization in GPUs through Adaptive Fine-Grained Data Migration</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-boosting-llc/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-boosting-llc/</guid><description>Abstract Modern server-grade GPUs (e.g., NVIDIA A100) integrate hundreds of cores and tens of memory partitions, providing massive compute capability and memory bandwidth.</description></item><item><title>Braid-ZNS: Leveraging Zone Random Write Area for Efficient In-Storage Compression on ZNS SSDs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-braid-zns/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-braid-zns/</guid><description>Abstract Zoned Namespace (ZNS) SSD is an emerging storage solution that reduces device-level garbage collection in conventional SSDs with a block interface.</description></item><item><title>Lupin: Spatial Resource Stealing with Outlier-First Encoding for Mixed-Precision LLM Acceleration</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-lupin/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-lupin/</guid><description>Abstract The rapid growth of Large Language Models (LLMs) exceeds on-chip memory capacity during inference, causing frequent external memory access and bandwidth limitations that hinder data transfer efficiency and overall performance.</description></item><item><title>Scrooge: Accelerating Attention Inference in LLMs via Early Termination Mechanism</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-scrooge/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/260420-date-scrooge/</guid><description>Abstract Large Language Models (LLMs) have demonstrated remarkable performance in natural language processing and are now widely adopted in diverse applications.</description></item><item><title>DDLM: Demand-Aware Dynamic Link Width Management for Energy-Efficient CXL Memory</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250802-iccd-ddlm/</link><pubDate>Sat, 02 Aug 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250802-iccd-ddlm/</guid><description>Abstract Modern data-centric workloads are driving rapid growth in the demand for high-bandwidth, large-capacity memory. Compute Express Link (CXL) has emerged as a key technology for scalable memory expansion through the Peripheral Component Interconnect Express (PCIe)&amp;rsquo;s high-speed serial interfaces.</description></item><item><title>Leveraging Mobile Processors for ISAR Image Generation and Classification in Radar Platforms</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250802-ieee-isar-image-generation/</link><pubDate>Sat, 02 Aug 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250802-ieee-isar-image-generation/</guid><description>Abstract With the growing threats posed by small airborne platforms such as drones in modern warfare, there is an increasing demand for small target identification technology.</description></item><item><title>LibraPIM: Dynamic Load Rebalancing to Maximize Utilization in PIM-Assisted LLM Inference Systems</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250802-pact-librapim/</link><pubDate>Sat, 02 Aug 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250802-pact-librapim/</guid><description>LibraPIM: Dynamic Load Rebalancing to Maximize Utilization in PIM-Assisted LLM Inference Systems “Balance is the New Speed”
The proliferation of Large Language Models (LLMs) has been a key driver of the AI revolution, and to run them efficiently, modern systems use heterogeneous architectures that combine xPUs (like GPUs/NPUs) with Processing-in-Memory (PIM).</description></item><item><title>Minimizing Read Disturb via Localized Page Allocation for Modern NAND Flash-Based SSDs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250802-iccd-minimizereaddisturb/</link><pubDate>Sat, 02 Aug 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250802-iccd-minimizereaddisturb/</guid><description>Abstract To meet the increasing demand for higher storage density, modern NAND flash-based SSDs employ Quad-Level Cell (QLC) technology, which stores four bits per memory cell.</description></item><item><title>Leveraging Chiplet-Locality for Efficient Memory Mapping in Multi-Chip Module GPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250715-micro-clap/</link><pubDate>Tue, 15 Jul 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250715-micro-clap/</guid><description>Abstract While the multi-chip module (MCM) design allows GPUs to scale compute and memory capabilities through multi-chip integration, it also introduces memory system non-uniformity, particularly when a thread accesses resources in remote chiplets.</description></item><item><title>SoftWalker: Supporting Software Page Table Walk for Irregular GPU Applications</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250715-micro-softwalker/</link><pubDate>Tue, 15 Jul 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250715-micro-softwalker/</guid><description>SoftWalker: Supporting Software Page Table Walk for Irregular GPU Applications Modern GPUs are the powerful engines behind today&amp;rsquo;s most demanding applications, from AI and scientific computing to stunning graphics.</description></item><item><title>Evaluating Performance of Modern GPU with Partitioned Last-Level Caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-partitioned_llcache/</link><pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-partitioned_llcache/</guid><description>Abstract For increased computing capability with the need for high parallelism, data-center GPUs (A100, H100, etc.) make their architecture more efficient by splitting Network-on-Chip (NoC) into two partitions to improve local bandwidth.</description></item><item><title>Outlier Matters: A Statistical Analysis of LLM Tensor Distributions and Quantization Effects</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-outliermatters/</link><pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-outliermatters/</guid><description>Abstract As transformer-based Large Language Models (LLMs) grow, deploying them under resource constraints has become increasingly complex, making quantization a vital technique for efficient inference.</description></item><item><title>Tournament Warp Scheduling: A Dynamic Policy Selector Exploiting Sub-cores in GPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-tournamentwarp/</link><pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-tournamentwarp/</guid><description>Abstract Warp scheduling policies have been extensively studied to improve GPU performance, with each policy optimized for different workload characteristics. However, current GPUs typically support only a fixed or minor parameter-tuned scheduling policy in hardware, limiting performance across diverse applications.</description></item><item><title>Towards Efficient Compression in ZNS SSDs: Analysis of Inherent Attributes and Applicability</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-compression_znsssd/</link><pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250707-itc_cscc-compression_znsssd/</guid><description>Abstract Conventional block interface-based SSDs, performing page-level write operations with out-of-place write policy, lead to issues such as Garbage Collection and Write Amplification.</description></item><item><title>Avalanche: Optimizing Cache Utilization via Matrix Reordering for Sparse Matrix Multiplication Accelerator</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250625-isca-avalanche/</link><pubDate>Wed, 25 Jun 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250625-isca-avalanche/</guid><description>Abstract Sparse Matrix Multiplication (SpMM) is essential in various scientific and engineering applications but poses significant challenges due to irregular memory access patterns.</description></item><item><title>PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic Lookup</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250622-dac-pimpal/</link><pubDate>Sun, 22 Jun 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250622-dac-pimpal/</guid><description>PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic Lookup Following the development of technology, the computing devices becomes much faster while the memory have larger capacity.</description></item><item><title>Evaluating Data Compression Algorithms and Their Impact on NAND Flash-Based SSDs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250401-iceic-datacompression/</link><pubDate>Tue, 01 Apr 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250401-iceic-datacompression/</guid><description>Abstract Modern NAND Flash-based SSDs have become a cornerstone of modern data storage solutions due to their speed and efficiency. With the increasing demand for large-scale datasets to train deep learning models, the need for greater storage capacity is also on the rise.</description></item><item><title>Redefining PIM Architecture with Compact and Power-Efficient Microscaling</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250401-iceic-microscaling/</link><pubDate>Tue, 01 Apr 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250401-iceic-microscaling/</guid><description>Abstract With advances in neural network technology, Processing-In-Memory (PIM) has emerged as a solution to per-formance bottlenecks between processors and memory.</description></item><item><title>Buddy ECC: Making Cache Mostly Clean in CXL-Based Memory Systems for Enhanced Error Correction at Low Cost</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-buddy/</link><pubDate>Mon, 31 Mar 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-buddy/</guid><description>Abstract As Compute Express Link (CXL) emerges as a key memory interconnect, interest in optimization opportunities and challenges has grown. However, due to the different characteristics of the CXL Memory Module (CMM) compared to traditional DRAM-based Dual In-line Memory Modules (DIMMs), existing optimizations may not be effectively applied.</description></item><item><title>Improving Address Translation in Tagless DRAM Cache by Caching PTE Pages</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-tdc/</link><pubDate>Mon, 31 Mar 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-tdc/</guid><description>Abstract This paper proposes a novel caching mechanism for PTE pages to enhance the Tagless DRAM Cache architecture and improve address translation in large in-package DRAM caches.</description></item><item><title>SPB: Towards Low-Latency CXL Memory via Speculative Protocol Bypassing</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-spb/</link><pubDate>Mon, 31 Mar 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-spb/</guid><description>Abstract Compute Express Link (CXL) is an advanced inter- connect standard designed to facilitate high-speed communication between CPUs, accelerators, and memory devices, making it wellsuited for data-intensive applications such as machine learning and real-time analytics.</description></item><item><title>Zebra: Leveraging Diagonal Attention Pattern for Vision Transformer Accelerator</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-zebra/</link><pubDate>Mon, 31 Mar 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250331-date-zebra/</guid><description>Abstract Vision Transformers (ViTs) have achieved remarkable performance in computer vision, but their computational complexity and challenges in optimizing memory bandwidth limit hardware acceleration.</description></item><item><title>Don't Cache, Speculate!: Speculative Address Translation for Flash-based Storage Systems</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/250116-ieee-dontcache/</link><pubDate>Thu, 16 Jan 2025 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/250116-ieee-dontcache/</guid><description>Abstract Address translation using a logical-to-physical (L2P) mapping table is essential for the NAND Flash-based SSDs. Unfortunately, the L2P mapping table size increases as SSD capacity increases.</description></item><item><title>Performance Characterization of CXL Memory Expander: Impact on Read and Write Latencies</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/241102-icce-asia-impact/</link><pubDate>Sat, 02 Nov 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/241102-icce-asia-impact/</guid><description>Abstract As data processing requirements and memory de-mands driven by Machine Learning, Artificial Intelligence, and In-Memory Database technologies continue to grow exponentially, the need for systems with greater memory capacity is increasing.</description></item><item><title>A Case for Speculative Address Translation with Rapid Validation for GPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/241101-micro-avartar/</link><pubDate>Fri, 01 Nov 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/241101-micro-avartar/</guid><description>A Case for Speculative Address Translation with Rapid Validation for GPUs Modern GPUs rely on a unified address space to efficiently share data with CPUs, but this comes at a cost.</description></item><item><title>Distributed Page Table: Harnessing Physical Memory as an Unbounded Hashed Page Table</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/241101-micro-dpt/</link><pubDate>Fri, 01 Nov 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/241101-micro-dpt/</guid><description>Abstract Virtual memory systems rely on the page table,a crucial component that maps virtual addresses to physical addresses (i.e., address translation).</description></item><item><title>Rethinking Page Table Structure for Fast Address Translation in GPUs: A Fixed-Size Hashed Page Table</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/241013-pact-rethinkingpagetable/</link><pubDate>Sun, 13 Oct 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/241013-pact-rethinkingpagetable/</guid><description>Rethinking Page Table Structure for Fast Address Translation in GPUs: A Fixed-Size Hashed Page Table Modern GPUs rely on multi-level Radix Page Tables (RPTs) to translate virtual addresses, but irregular applications incur frequent TLB misses and sequential pointer-chasing walks that severely degrade performance.</description></item><item><title>A Two-level Tracking Mechanism to Mitigate RowHammer Attacks</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/240819-isocc-migirationrowhammer/</link><pubDate>Mon, 19 Aug 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/240819-isocc-migirationrowhammer/</guid><description/></item><item><title>Leveraging Algorithm-based Fault Tolerance for Propagation Error Detection in NPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/240819-isocc-faulttolerance/</link><pubDate>Mon, 19 Aug 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/240819-isocc-faulttolerance/</guid><description/></item><item><title>Analysis of L2P Map cache size and hash function correlation in NAND Flash-based SSDs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/240702-itccscc-l2pmap/</link><pubDate>Tue, 02 Jul 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/240702-itccscc-l2pmap/</guid><description/></item><item><title>Don't Cache, Speculate!: Speculative Address Translation for Flash-based Storage Systems</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/240623-dac-dontcache/</link><pubDate>Sun, 23 Jun 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/240623-dac-dontcache/</guid><description/></item><item><title>SAVector: Vectored Systolic Arrays</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/240525-sa-vector/</link><pubDate>Sat, 25 May 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/240525-sa-vector/</guid><description>Abstract Conventional DNN inference accelerators are designed with a few (up to four) large systolic arrays. As such a scale-up architecture often suffers from low utilization, a scale-out architecture, in which a single accelerator has tens of pods and each pod has a small systolic array, has been proposed.</description></item><item><title>Enhanced Positional SECDED: Achieving Maximal Double-Error Correction in Racetrack Memories</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/240109-secded/</link><pubDate>Tue, 09 Jan 2024 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/240109-secded/</guid><description>Abstract Racetrack memory (RM), a highly dense and high-speed spintronic non-volatile memory (NVM) technology, has the potential to revolutionize data storage.</description></item><item><title>Facto-CNN: Memory-Efficient CNN Training with Low-rank Tensor Factorization and Lossy Tensor Compression</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/231111-facto-cnn/</link><pubDate>Sat, 11 Nov 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/231111-facto-cnn/</guid><description>Abstract Convolutional neural networks (CNNs) are becoming deeper and wider to achieve higher accuracy and lower loss, significantly expanding the computational resources.</description></item><item><title>Conveyor: Towards Asynchronous Dataflow in Systolic Array to Exploit Unstructured Sparsity</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/231106-conveyor-sa/</link><pubDate>Mon, 06 Nov 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/231106-conveyor-sa/</guid><description>Abstract Systolic array (SA) architecture efficiently offers parallel computation using a simple data movement across processing elements. However, their rigid structure and synchronous dataflow limit flexibility in handling sparse computations, resulting in underutilized resources and suboptimal performance.</description></item><item><title>Improving Performance and Energy-efficiency of DNN Accelerators with STT-RAM Buffers</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/231025-stt-ram/</link><pubDate>Wed, 25 Oct 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/231025-stt-ram/</guid><description>Abstract DNN inference on mobile and edge devices is challenging due to high computational and storage demands. To accelerate the inference on these devices, various DNN accelerators have been proposed.</description></item><item><title>SparseFT: Sparsity-aware Fault Tolerance for Reliable CNN Inference on GPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/231021-sparseft/</link><pubDate>Sat, 21 Oct 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/231021-sparseft/</guid><description>Abstract Graphics Processing Units (GPUs), while offering exceptional performance for CNN inference tasks, are susceptible to both transient and permanent hardware faults due to the integration of numerous processing elements and advancements in technology scaling.</description></item><item><title>CRAN: A Computational Redundancy-aware Accelerator for Convolutional Neural Networks</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/230709-dac-cran/</link><pubDate>Sun, 09 Jul 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/230709-dac-cran/</guid><description/></item><item><title>CAESAR: A CNN Accelerator Exploiting Sparsity and Redundancy Pattern</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/230625-caesar/</link><pubDate>Sun, 25 Jun 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/230625-caesar/</guid><description>Abstract Convolutional Neural Networks (CNN) have shown outstanding performance in many computer vision applications. However, CNN Inference on mobile and edge devices is challenging due to high computation demands.</description></item><item><title>In-Cache Processing with Power-of-Two Quantization for Fast CNN Inference on CPUs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/230625-in-cache/</link><pubDate>Sun, 25 Jun 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/230625-in-cache/</guid><description>Abstract Convolutional Neural Networks (CNN) demand high computational capabilities, motivating researchers to leverage Processing-In-Memory (PIM) technology to achieve significant performance improvements.</description></item><item><title>Why Address Translation Matter?: Analyzing Page Access Patterns in NAND Flash-Based SSDs</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/230625-why-address/</link><pubDate>Sun, 25 Jun 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/230625-why-address/</guid><description>Abstract The use of NAND Flash-based SSD is on the rise in various domains such as Enterprise, Data centers, Personal computers, and Automotive.</description></item><item><title>On the Positional Single Error Correction and Double Error Detection in Racetrack Memories</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/230222-on-the-positional/</link><pubDate>Thu, 23 Feb 2023 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/230222-on-the-positional/</guid><description>Abtract In the era of non-volatile memories, the racetrack memory is a promising technology to pack hundreds of bits in a magnetic nanowire.</description></item><item><title>Pinning Page Structure Entries to Last-Level Cache for Fast Address Translation</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/221026-pinning-page/</link><pubDate>Wed, 26 Oct 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/221026-pinning-page/</guid><description>Abstract As the memory footprint of emerging applications continues to increase, the address translation becomes a critical performance bottleneck owing to frequent misses on the Translation Lookaside Buffer (TLB).</description></item><item><title>Virtual PTE Storage: Repurposing Last-level Cache to Accelerate Address Translation for Big Data Workloads</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/221026-virtual-pte/</link><pubDate>Wed, 26 Oct 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/221026-virtual-pte/</guid><description>Abstract Address translation is one of the major performance bottlenecks for emerging big data workloads. Since those workloads have large memory footprints and irregular memory access patterns, they suffer from frequent TLB (Translation Lookaside Buffer) misses and frequently incur expensive page walk.</description></item><item><title>On-the-Fly Lowering Engine: Offloading Data Layout Conversion for Convolutional Neural Networks</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/220720-on-the-fly/</link><pubDate>Wed, 20 Jul 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/220720-on-the-fly/</guid><description>Abstract Many deep learning frameworks utilize GEneral Matrix Multiplication (GEMM)-based convolution to accelerate CNN execution. GEMM-based convolution provides faster convolution yet requires a data conversion process called lowering (i.</description></item><item><title>Constructing Large Buffers with Heterogeneous STT-RAM Cells for DNN Accelerators</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/220310-dac-sttmram/</link><pubDate>Thu, 10 Mar 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/220310-dac-sttmram/</guid><description>Abstract Among the various NVM technologies, phase-change-memory (PCM) has attracted substantial attention as a candidate to replace the DRAM for next-generation memory.</description></item><item><title>Don't open row: rethinking row buffer policy for improving performance of non-volatile memories</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/220310-dont-open-row/</link><pubDate>Thu, 10 Mar 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/220310-dont-open-row/</guid><description>Abstract Among the various NVM technologies, phase-change-memory (PCM) has attracted substantial attention as a candidate to replace the DRAM for next-generation memory.</description></item><item><title>Proactively Invalidating Dead Blocks to Enable Fast Writes in STT-MRAM Caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/220310-proactively/</link><pubDate>Thu, 10 Mar 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/220310-proactively/</guid><description>Abstract Spin-Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising emerging memory technology for on-chip caches. It has a low read access time and low leakage power.</description></item><item><title>Exploiting Data Compression for Adaptive Block Placement in Hybrid Caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/220112-exploiting-data/</link><pubDate>Wed, 12 Jan 2022 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/220112-exploiting-data/</guid><description>Abstract STT-RAM (Spin-Transfer Torque Random Access Memory) appears to be a viable alternative to SRAM-based on-chip caches. Due to its high density and low leakage power, STT-RAM can be used to build massive capacity last-level caches (LLC).</description></item><item><title>A Case Study of Quantizing Convolutional Neural Networks for Fast Disease Diagnosis on Portable Medical Devices</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/211229-a-case-study/</link><pubDate>Wed, 29 Dec 2021 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/211229-a-case-study/</guid><description>Abstract Recently, the amount of attention paid towards convolutional neural networks (CNN) in medical image analysis has rapidly increased since they can analyze and classify images faster and more accurately than human abilities.</description></item><item><title>Exploiting Hardware Events to Reduce Energy Consumption of HPC Systems</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/210831-exploiting-hardware/</link><pubDate>Tue, 31 Aug 2021 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/210831-exploiting-hardware/</guid><description>Abstract This paper proposes a novel mechanism called Event-driven Uncore Frequency Scaler (eUFS) to improve the energy efficiency of the HPC systems.</description></item><item><title>CID: Co-Architecting Instruction Cache and Decompression System for Embedded Systems</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/210701-cid/</link><pubDate>Thu, 01 Jul 2021 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/210701-cid/</guid><description>Abstract Code compression is widely used to reduce the footprint of code memory in cost-sensitive embedded systems. However, despite the small code size, the decompressor and the address translator required to support the code compression incur energy and area overheads.</description></item><item><title>Proactive Dead Block Eviction for Reducing Write Latency in STT-MRAM Caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/210131-proactive-dead/</link><pubDate>Sun, 31 Jan 2021 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/210131-proactive-dead/</guid><description>Abstract Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising emerging memory technology for the on-chip caches. It has low read access time and low leakage power.</description></item><item><title>ADAM: Adaptive Block Placement with Metadata Embedding for Hybrid Caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/201018-adam/</link><pubDate>Sun, 18 Oct 2020 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/201018-adam/</guid><description>Abstract Spin-Transfer Torque Random Access Memory (STT-RAM) is a potential alternative for SRAM-based on-chip caches. STT-RAM offers high density and low leakage power, thereby can be used to build a large capacity last-level caches (LLC).</description></item><item><title>Dynamic Rank Subsetting with Data Compression</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/200429-dynamic-rank/</link><pubDate>Wed, 29 Apr 2020 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/200429-dynamic-rank/</guid><description>Abstract In this paper, we propose Dynamic Rank Subsetting (DRAS) technique that enhances the energy-efficiency and the performance of memory system through the data compression.</description></item><item><title>Dead Block-Aware Adaptive Write Scheme for MLC STT-MRAM Caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/200331-dead-block/</link><pubDate>Tue, 31 Mar 2020 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/200331-dead-block/</guid><description>Abstract In this paper, we propose an efficient adaptive write scheme that improves the performance of write operation in MLC STT-MRAM caches.</description></item><item><title>Analyzing the Performance Overhead of Secure Memory</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/191218-analyzing/</link><pubDate>Wed, 18 Dec 2019 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/191218-analyzing/</guid><description>Abstract 메인 메모리의 데이터를 각종 보안공격으로부터 보호하는 것은 매우 중요하다. 최근 Counter-mode 암호화 기법을 적용하여 memory의 보안성을 향상시킨 secure memory 기술이 제안되었다.</description></item><item><title>Performance Characterization of STAR RNA-seq aligner on Multi-core Processor</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/191218-rna/</link><pubDate>Wed, 18 Dec 2019 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/191218-rna/</guid><description>Abstract 생물체의 유전자에서 발현되는 전사체를 분석하면 내부 유전적 병변 또는 외부 자극 및 변화하는 환경에 대한 세포반응을 알 수 있으며, 이를 통해 유전병 또는 암과 같은 질병을 야기하는 유전자군 발굴이 가능하다.</description></item><item><title>Touché: Towards Ideal and Efficient Cache Compression By Mitigating Tag Area Overheads</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/191012-touche/</link><pubDate>Sat, 12 Oct 2019 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/191012-touche/</guid><description>Abstract Compression is seen as a simple technique to increase the effective cache capacity. Unfortunately, compression techniques either incur tag area overheads or restrict cache block placement to only include neighboring addresses.</description></item><item><title>Split-CNN: Splitting Window-based Operations in Convolutional Neural Networks for Memory System Optimization</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/190404-split-cnn/</link><pubDate>Thu, 04 Apr 2019 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/190404-split-cnn/</guid><description>Abstract We present an interdisciplinary study to tackle the memory bottleneck of training deep convolutional neural networks (CNN). Firstly, we introduce Split Convolutional Neural Network (Split-CNN) that is derived from the automatic transformation of the state-of-the-art CNN models.</description></item><item><title>Interpage-Based Endurance-Enhancing Lower State Encoding for MLC and TLC Flash Memory Storages</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/190315-interpage/</link><pubDate>Fri, 15 Mar 2019 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/190315-interpage/</guid><description>Abstract During the past decade, the endurance of NAND flash memory has severely deteriorated. The maximum number of program and erase cycles has fallen significantly with emerging of multilevel cell (MLC) and triple-level cell (TLC) technology, and scaling down of the cell size.</description></item><item><title>Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/181020-attache/</link><pubDate>Sat, 20 Oct 2018 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/181020-attache/</guid><description>Abstract Memory systems are becoming bandwidth constrained and data compression is seen as a simple technique to increase their effective bandwidth.</description></item><item><title>CramSim: controller and memory simulator</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/171002-cramsim/</link><pubDate>Mon, 02 Oct 2017 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/171002-cramsim/</guid><description>Abstract The explosion of digital data and the high computation demands of data analysis have made the memory system a major contributor to the performance and power consumption of modern computing systems.</description></item><item><title>Partial Row Activation for Low-Power DRAM System</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/170204-partial-row/</link><pubDate>Sat, 04 Feb 2017 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/170204-partial-row/</guid><description>Abstract Owing to increasing demand of faster and larger DRAM system, the DRAM system accounts for a large portion of the total power consumption of computing systems.</description></item><item><title>Designing a Resilient L1 Cache Architecture to Process Variation-Induced Access-Time Failures</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/160101-designing/</link><pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/160101-designing/</guid><description>Abstract Continuous scaling of process technology increases variations in transistors. The process variations cause large fluctuations in the access times of static random-access memory (SRAM) cells.</description></item><item><title>A Low-Cost Mechanism Exploiting Narrow-Width Values for Tolerating Hard Faults in ALU</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/150901-a-low-cost-mechanism/</link><pubDate>Tue, 01 Sep 2015 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/150901-a-low-cost-mechanism/</guid><description>Abstract Digital circuits are expected to increasingly suffer from more hard faults due to technology scaling. Especially, a single hard fault in ALU (Arithmetic Logic Unit) might lead to a total failure in processors or significantly reduce their performance.</description></item><item><title>Ensuring Cache Reliability and Energy Scaling at Near-Threshold Voltage With Macho</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/140716-ensuring-cache/</link><pubDate>Tue, 01 Sep 2015 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/140716-ensuring-cache/</guid><description>Abstract Nanoscale process variations in conventional SRAM cells are known to limit voltage scaling in microprocessor caches. Recently, a number of novel cache architectures have been proposed which substitute faulty words of one cache line with healthy words of others, to tolerate these failures at low voltages.</description></item><item><title>Ternary cache: Three-valued MLC STT-RAM caches</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/141019-ternary-cache/</link><pubDate>Sun, 19 Oct 2014 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/141019-ternary-cache/</guid><description>Abstract Spin-transfer torque random access memory (STT-RAM) has become a promising non-volatile memory technology for cache memories. Recently, 2-bit multi-level cell (MLC) STT-RAM has been proposed to enhance data density, but it suffers from low reliability of its read and write operations.</description></item><item><title>AVICA: An access-time variation insensitive L1 cache architecture</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/130318-avica/</link><pubDate>Mon, 18 Mar 2013 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/130318-avica/</guid><description>Abstract Ever scaling process technology increases variations in transistors. The process variations cause large fluctuations in the access times of SRAM cells.</description></item><item><title>Macho: A failure model-oriented adaptive cache architecture to enable near-threshold voltage scaling</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/130223-macho/</link><pubDate>Sat, 23 Feb 2013 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/130223-macho/</guid><description>Abstract Recent interest in CMOS voltage scaling has produced a class of cache architectures which tolerate parametric SRAM failures at low voltage by substituting faulty words of one cache line with healthy words of another line.</description></item><item><title>Skinflint DRAM system: Minimizing DRAM chip writes for low power</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/130223-skinflint/</link><pubDate>Sat, 23 Feb 2013 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/130223-skinflint/</guid><description>Abstract DRAMs are one of the main players of computer system energy consumption due to their large capacities and frequent accesses.</description></item><item><title>Residue cache: a low-energy low-area L2 cache architecture via compression and partial hits</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/111203-residue-cache/</link><pubDate>Sat, 03 Dec 2011 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/111203-residue-cache/</guid><description>Abstract L2 cache memories are being adopted in the embedded systems for high performance, which, however, increases energy consumption due to their large sizes.</description></item><item><title>TLB index-based tagging for cache energy reduction</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/110801-tlb-index-based/</link><pubDate>Mon, 01 Aug 2011 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/110801-tlb-index-based/</guid><description>Abstract Conventional cache tag matching is based on addresses to identify correct data in caches. However, this tagging scheme is not efficient because tag bits are unnecessarily large.</description></item><item><title>Lizard: Energy-efficient hard fault detection, diagnosis and isolation in the ALU</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/101003-lizard/</link><pubDate>Sun, 03 Oct 2010 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/101003-lizard/</guid><description>Abstract Digital circuits are expected to increasingly suffer from more hard faults due to technology scaling. Especially, a single hard fault in the ALU might lead to a total failure in the embedded systems.</description></item><item><title>TEPS: Transient Error Protection Utilizing Sub-word Parallelism</title><link>https://skku-compasslab.github.io/compasslab/publications/papers/090313-teps/</link><pubDate>Fri, 13 Mar 2009 00:00:00 +0000</pubDate><guid>https://skku-compasslab.github.io/compasslab/publications/papers/090313-teps/</guid><description>Abstract Future microprocessors are expected to observe higher transient error rates in combinational logic due to technology scaling and dense integration.</description></item></channel></rss>