Research Overview
My research broadly explores memory-centric computer systems, with a focus on CXL-based memory expansion,
heterogeneous cooperative computing, near-data acceleration, and datacenter memory efficiency.
The topics below are a curated selection grouped by theme. For the complete list, see
Publications.
Bridging the gap between CXL's architectural promise and its real-world behavior, from first-hand hardware
characterization to redesigning the memory and storage systems built on top of it.
- [MICRO'26] CXL AnySSD: A Composable CXL SSD Using a CXL Type-2 Device and Any SSDs
- [OSDI'26] MAC: Metadata Acceleration for Sustainable Performance in Big-Data Systems with CXL DRAM
- [HCDS'26] Characterizing CXL Memory Management in a Production Cluster Management Host Agent
- [MICRO'25] Re-architecting End-host Networking with CXL: Coherence, Memory, and Offloading
- [CAL'25] X-PPR: Post Package Repair for CXL Memory
- [MICRO'24] Demystifying a CXL Type-2 Device: A Heterogeneous Cooperative Computing Perspective
- [MICRO'23] Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices
Offloading memory-intensive OS and runtime services to specialized devices to reduce CPU overhead and improve efficiency.
- [ATC'25] Para-ksm: Parallelized Memory Deduplication with Data Streaming Accelerator
- [CAL'25] Cooperative Memory Deduplication with Intel Data Streaming Accelerator
- [CAL'25] Hardware-accelerated Kernel-Space Memory Compression Using Intel QAT
- [ATC'23] STYX: Exploiting SmartNIC Capability to Reduce Datacenter Memory Tax
Building specialized hardware and near-memory accelerators for AI workloads that are bottlenecked by data movement rather than compute.
- [MICRO'26] AMG: AMX–GPU Cooperative Acceleration of Billion-Scale ANNS via GEMM Reformulation
- [MICRO'26] C3: Collaborative CPU-Commodity DRAM Computation for Reliable LLM Acceleration
- [DAC'26] PAGE: Processing-Using-DRAM Architecture-Circuits Co-Optimization for Efficient Acceleration of General Matrix-Vector Multiplication
- [ISCA'22] Graphite: Optimizing Graph Neural Networks on CPUs through Cooperative Software-Hardware Techniques