Selected Publications
Please note that papers linked here represent author
preprints. The official, published version must be
obtained from the publisher's website or the published print
copy. This material is presented here to ensure timely
dissemination of scholarly and technical work. Copyright and
all rights therein are retained by authors or by other
copyright holders. All persons copying this information
are expected to adhere to the terms and constraints invoked by
each document's copyright terms. In most cases, these works may
not be reposted without the explicit permission of the copyright
holder. Permission is given to make digital or hard copies of
all or part of this material without fee for personal or
classroom use, provided that the copies are not made or
distributed for profit or commercial advantage, and that copies
bear the appropriate copyright notice and the full bibliographic
citation on the first page. Copyrights for components of
this work owned by others must also be honored. To copy
otherwise, to republish, to post on servers, to redistribute to
lists, etc. requires specific permission and/or a fee. In
particular, permission to reprint/republish this material for
advertising or promotional purposes or for creating new
collective works for resale or redistribution to servers or
lists, or to reuse any copyrighted component of this work in
other works, must be obtained from the copyright owner.
Please note further that any opinions, findings,
conclusions, or recommendations expressed in this material are
those of the authors and do not necessarily reflect the views
of the sponsoring agencies, employers, or publishers.
|
A complete list of my publications can be found on my Google Scholar Page (which will also provide
paper links) or DBLP page (but some workshop papers may not appear there)
Recent Highlights
(processing in memory/PIM, programming model, PIM compilation, simulation)
A. R. Noor, F. A. Siddique, A. Kothari, D. Guo, K. Skadron, and V. Adve. "PIM-AutoDSE: Automated ISA and Design Space Exploration for PIM." In Proceedings of the ACM/IEEE International Symposium on Microarchitecture (MICRO), Oct. 2026, to appear.
(LLM GPU characterization)
Z. Fan and K. Skadron. "Regular-Hardware Characterization of Diffusion vs. Autoregressive Language Model Inference: Compute vs. Memory Bottlenecks." In Proceedings of the IEEE International Symposium on Workload Characterization (IISWC), Sep. 2026, to appear. (pdf)
(assertion-based verification, runtime verification)
Y. Seneviratne, T. Tracy II, K. Nandagopal, D. Chinnasame Rani, M. B. Dwyer B. Bingham, K. Skadron. "Periscope: Extending the Scope of Pre-Silicon Assertion-Based Verification with Runtime Monitor Feedback." In Proceedings of the International Conference on Runtime Verification (RV), Springer, Oct. 2026, to appear.
(regular expression processing)
O. Efimov, T. Tracy II, W. Zheng, K. Skadron and K. Angstadt. "Normal Forms for Perl-Compatible Regular Expressions." Tools Track, International Conference on Implementation and Applications of Automata, Aug. 2026.
(processing in memory/PIM, programming model, simulation)
F. Siddique, D. Guo, H. Abbot, K. Durrer, M. Gholamrezaei, M. Baradaran, E. Ermovick, A. Ahmed, Z. Fan, B. Gul, A. Venkat, and K. Skadron. "Characterizing Digital DRAM PIM through Modeling and Benchmarking." ACM Transactions on Architecture and Code Optimization (TACO). 23(2), 22 pp., June 2026, DOI 10.1145/3815581.
(KV cache serving, processing in memory/PIM)
K. Kiyawat and K. Skadron. "Reimagining LLM Inference Infrastructure with Memory-Centric KV Cache Servers." In the 3rd Workshop on Hot Topics in System Infrastructure (HotInfra), in conjunction with ISCA'26, June 2026. (pdf)
(quantization, compression, processing in memory/PIM)
S. Tajdari, A. Shekar, K. Skadron, and S. Dwarkadas. "Tailoring LLM Weight Compression for PIM Architectures." In the Workshop on ML for Computer Architecture and Systems (MLArchSys), in conjunction with ISCA'26, June 2026. (pdf)
(ECC, processing in memory/PIM)
M. Shires, K. Farran, D. Hansen,* and K. Skadron. "Modeling Error Correction for PIM Architectures in PIMeval-PIMbench." In the Workshop on Open-Source Computer Architecture Research (OSCAR), in conjunction with ISCA'26, June 2026. (pdf)
(regular expression processing, cybersecurity, FPGAs)
S. Obla, T. Tracy II, W. U. Hassan, K. Skadron, and J. Hoe. "RapidScan: High-Throughput Parameterized HLS-based Streaming String Matching Library for FPGAs." In Proceedings of the IEEE International Symposium on Field-Programmable Custom Computing Machines (FCCM), as one of 6 finalists in the FCCM Reconfigurable Computing Challenge, with 4 pp. accompanying paper, May 2026. DOI 10.1109/FCCM68464.2026.00079 - Awarded third place in the competition.
(CRISPR, Levenshtein automata, FPGA)
T. Tracy II, O. Yaish, Y. Orenstein, K. Skadron. "Auto-OFFinder: Accelerating off-target discovery for CRISPR using an FPGA." In Proceedings of the RECOMB Satellite Conference on Hardware Acceleration of Bioinformatics Workloads (RECOMB-Arch), affiliated with RECOMB, May 2026.
(processing in memory/PIM, LLM performance modeling)
K. Kiyawat, Y. Seneviratne, Z. Fan, M. Baradaran, K. Skadron. "HARMONI: Hierarchical Evaluation of Memory-centric Systems for LLM Execution." In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Apr. 2026. DOI 10.1109/ISPASS69572.2026.00016
(processing in memory/PIM, databases/data analytics)
A. Shekar, K. Gaffney, M. Prammer, K. Kiyawat, L. Wu, H. Caminal, Z. Fan, Y. Gao, A. Venkat, J. F. Martínez, J. M. Patel, and K. Skadron. "Membrane: Accelerating Database Analytics with DRAM-Based PIM Filtering and Schema Denormalization." ACM Transactions on Architecture and Code Optimization (TACO). 23(1), 24 pp., Mar. 2026, DOI 10.1145/3786775.
(graph analytics, GPUs)
B. Gul, M. Murach, S. Bekiranov, K. Skadron. "gLeiden: Accelerated Community Detection
Algorithms using Directed and Undirected Graphs on GPUs." Bioinformatics Advances, 6(1), DOI 10.1093/bioadv/vbaf327.
(cache design, eDRAM, BEOL)
F. Waqar, J. Kwak, J. Lee, O. Phadke, M. Shon, M. Gholamrezaei, K. Skadron, and S. Yu. "Optimization and Benchmarking of Monolithically Stackable Gain Cell Memory for Last-Level Cache." IEEE Transactions on Computers, 75(3):760-775, Mar. 2026, DOI: 10.1109/TC.2025.3625490.
(processing in memory/PIM, programming model, databases/data analytics)
J. Castrillon, J. Giceva, Y. Hua, K. Keeton, A. Shekar, K. Skadron, T. Wang, H. Zhang. "Declarative Memory Services." In Proceedings of the Conference on Innovative Data Systems Research (CIDR), Jan. 2026.
(processing in memory/PIM, numerical methods)
E. Tang, T. Zhang, W. Bradford, F. A. Siddique, J. Hoe, K. Skadron, and F. Franchetti. "Hardware-Software Co-Design of Iterative Filter-Update Numerical Methods using Processing-In-Memory." In the International Workshop on Memory System, Management and Optimization (MEMO), in conjunction with SC, Nov. 2025.
(processing in memory/PIM, near data processing/NDP, graph processing, GNN)
A. Ahmed, F. Lin, J. Li, and K. Skadron. "TGN-PNM: A Near-Memory Architecture for Temporal GNN Inference on 3D-Stacked Memory." In Proceedings of the 11th International Sympoium on Memory Systems (MEMSYS), Oct. 2025.
(processing in memory/PIM, graph processing)
M. Baradaran, K. Kiyawat, A. Shekar, A. Mughrabi, and K. Skadron. "TriPIM - Exact Triangle Counting on UPMEM PIM for Graph Analytics." In Proceedings of the 11th International Sympoium on Memory Systems (MEMSYS), Oct. 2025.
(processing in memory/PIM, graph proessing)
E. Cukurtas, K. Ranawella, K. Skadron, and M. R. Stan. "IMPRINT: In-Memory Processing with Indirect Addressing Techniques for GPU-hosted HBM-PIM." In Proceedings of the 11th International Sympoium on Memory Systems (MEMSYS), Oct. 2025.
(processing in memory/PIM, bit-serial, compiler)
D. Guo, M. Gholamrezaei, M. Hofmann, A. Venkat, Z. Zhang, and K. Skadron. "PIMsynth: A Unified Compiler Framework for Bit-Serial Processing-In-Memory Architectures." IEEE Computer Architecture Letters (CAL), 24(2):277-80, July-Dec. 2025. DOI: 10.1109/LCA.2025.3600588.
(processing near memory/PNM/PIM, RAG)
D. Quinn, E. Yucel, M. Prammer, Z. Fan, K. Skadron, J. Patel, M. Alian, J. Martínez. "DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware Codesign." In Proceedings of the ACM/IEEE International Symposium on Computer Architecture (ISCA), June 2025.
(graph processing, multi-FPGA)
O. Jaiyeoba, A. Mughrabi, M. Baradaran, B. Gul, and K. Skadron. "Swift: A Multi-FPGA Framework for Scaling Up Accelerated Graph Analytics." In Proceedings of the IEEE International Conference on Field Programmable Technology (FPT), Dec. 2024. (pdf)
(processing in memory/PIM, simulation, benchmarks)
F. A. Siddique, D. Guo, Z. Fan, M. Gholamrezaei, M. Baradaran, A. Ahmed, H. Abbot, K. Durrer, E. Ermovick, K. Nandagopal, E. Ermovick, K. Kiyawat, B. Gul, A. Mughrabi, A. Venkat, and K. Skadron. "Architectural Modeling and Benchmarking for Digital DRAM
PIM." In Proceedings of the IEEE International Symposium on Workload Characterization (IISWC), Sep. 2024. (pdf)
(dynamic graph processing)
A. Ahmed, F. Siddique, and K. Skadron. "GraphTango: A Hybrid Representation Format for Efficient Streaming Graph Updates and Analysis." International Journal of Parallel Programming, Springer, to appear.
(dynamic graph processing, FPGA)
O. Jaiyeoba and K. Skadron. "Dynamic-ACTS - A Dynamic Graph Analytics Accelerator For HBM-Enabled FPGAs." ACM Transactions on Reconfigurable Technology and Systems, 17(3), 29 pp., Sept. 2024.
(processing in storage/computational storage, accelerators, bioinformatics)
L. Wu, M. Zhou, Weihong Xu, A. Venkat, T. Rosing, and K. Skadron. "Abakus: Accelerating k-mer Counting With Storage Technology." ACM Transactions on Architecture and Code Optimization (TACO), 21(1), 26 pp., Jan. 2024.
(graph processing, FPGA)
O. Jaiyeoba and K. Skadron. "ACTS: A Near-Memory FPGA Graph Processing Framework." Proceedings of the IEEE International Symposium on Field-Programmable Gate Arrays (FPGA), Feb. 2023.
(processing in memory/PIM, sorting)
M. Lenjani, A. Ahmed, and K. Skadron. "Pulley: An Algorithm/Hardware Co-optimization for In-memory Sorting." IEEE Computer Architecture Letters, 21(2):109-112, July-Dec. 2022.
-
(processing in memory/PIM)
L. Wu, R. Sharifi, A. Venkat, and K. Skadron. "DRAM-CAM: General-Purpose Bit-Serial Exact Pattern Matching." IEEE Computer Architecture Letters, 21(2):89-92, July-Dec. 2022.
(automata processing, FPGAs, heterogeneous architectures, accelerators)
K. Angstadt, T. Tracy, J.-B. Jeannin, K. Skadron, and W. Weimer. "Synthesizing Legacy String Code for FPGAs Using Bounded Automata Learning." IEEE Micro special issue on Compiling for Accelerators, 42(5):70-77, Sep.-Oct. 2022.
(microarchitecture, microcode, runtime optimization)
L. Moody, W. Qi, A. Sharifi, L. Berry, J. Rudek, J. Gaur, J. Parkhurst, S. Subramoney, K. Skadron, A. Venkat. "Speculative Code Compaction: Eliminating Dead Code via Speculative Microcode Transformations." In Proceedings of the ACM/IEEE International Symposium on Microarchitecture (MICRO), Oct. 2022.
(processing in memory/PIM)
M. Lenjani and K. Skadron. "Supporting Moderate Data Dependency, Position Dependency, and Divergence in PIM-based Accelerators." IEEE Micro special issue on Processing in Memory, published online Dec. 2021, in print Jan/Feb. 2022, 42(1):108-115. DOI 10.1109/MM.2021.3136189
-
(processing in memory/PIM) M. Lenjani, A. Ahmed, M. R. Stan, and K. Skadron. "Gearbox: A Case for Supporting Accumulation Dispatching and Hybrid Partitioning in PIM-based Accelerators." In Proceedings of the IEEE/ACM International Symposium on Computer Architecture (ISCA), June
2022.
-
(processing in memory/PIM, FPGA) S. Mosanu, M. N. Sakib, T. Tracy III, E. Cukurtas, A. Ahmed, P. Ivanov, S. Khan, K. Skadron, M. Stan. "PiMulator: A Fast and Flexible Processing-in-Memory Emulation Platform." In Proceedings of the ACM/IEEE/EDAA/EDAC Conference on Design, Automation and Test in Europe (DATE), Mar.
2022.
-
(automata processing,
heterogeneous architectures, accelerators) E.
Sadredini, R. Rahimi, M. Imani, and K. Skadron. “Sunder:
Enabling Low-Overhead and Scalable Near Data Pattern
Matching Acceleration.” In Proceedings of the IEEE/ACM
International Symposium on Microarchitecture (MICRO), Oct.
2021.
-
(processing in memory/PIM,
bioinformatics) M. Zhou, L. Wu, M. Li, N.
Moshiri, K. Skadron, and T. Rosing. “Ultra Efficient
Acceleration for De Novo Genome Assembly via Near-Memory
Computing.” In Proceedings of the ACM/IEEE/IFIP International
Conference on Parallel Architectures and Compiler Techniques
(PACT), Sept. 2021
(thermal modeling)
J. Han, R. E. West, K. Skadron, and M. R. Stan. "Thermal Simulation of Processing-in-Memory Devices Using HotSpot 7.0." In Proceedings of the 27th International Workshop on Thermal Investigations of ICs (THERMINIC), Sept. 2021.
-
(processing in memory/PIM,
bioinformatics) L. Wu, R. Sharifi, M.
Lenjani, K. Skadron, and A. Venkat. “Sieve: Scalable In-situ
DRAM-based Accelerator Designs for Massively Parallel k-mer
Matching.” In Proceedings of the ACM/IEEE International
Symposium on Computer Architecture (ISCA), June 2021.
-
(automata processing,
heterogeneous architectures, accelerators, FPGAs, linear
temporal logic, verification) T. Tracy II,
L. Tabajara, M. Vardi, and K. Skadron. “Runtime Verification
on FPGAs with LTLf Specifications.” In Proceedings of the
Formal Methods in Computer-Aided Design (FMCAD), Sept. 2020. (pdf)
-
(automata processing,
FPGAs, heterogeneous architectures, accelerators) R.
Rahimi, E. Sadredini, M. Stan, and K. Skadron. “Grapefruit: An
Open-Source, Full-Stack, and Customizable Automata Processing
on FPGAs.” In Proceedings of the IEEE International Symposium
on Field Customizable Computing Machines (FCCM), May 2020.
Nominated for best paper.
-
(automata processing,
heterogeneous architectures, accelerators) E.
Sadredini, R. Rahimi, M. Lenjani, M. Stan, and K. Skadron.
“FlexAmata: A Universal and Efficient Adaption of Applications
to Spatial Automata Processing Accelerators.” In Proceedings
of the ACM International Symposium on Architectural Support
for Programming Languages and Operating Systems (ASPLOS), Mar.
2020.
-
(automata processing,
heterogeneous architectures, accelerators) R.
Rahimi, E. Sadredini, M. Lenjani, M. Stan, and K. Skadron.
“Impala: Algorithm/Architecture Co-Design for In-Memory
Multi-Stride Pattern Matching.” In Proceedings of the IEEE
International Symposium on High Performance Computer
Architecture (HPCA), Feb. 2020. Nominated for best paper.
-
(processing in memory/PIM) M.
Lenjani, P. Gonzalez, E. Sadredini, S. Li, Y. Xie, A. Akel, S.
Eilert, M. R. Stan, and K. Skadron. “Fulcrum: a Simplified
Control and Access Mechanism toward Flexible and Practical
In-Situ Accelerators.” In Proceedings of the IEEE
International Symposium on High Performance Computer
Architecture (HPCA), Feb. 2020.
- (automata processing,
heterogeneous architectures, accelerators, on-chip
interconnect) E. Sadredini, R. Rahimi, V.
Verma, M. Stan, and K. Skadron. “eAP: A Scalable and Efficient
In-Memory Accelerator for In-Memory Processing.” In Proceedings
of the ACM/IEEE International Symposium on Microarchitecture
(MICRO), Oct. 2019.
-
(memory performance
benchmarking) A. Ahmed, K. Skadron.
“Hopscotch: A Micro-benchmark Suite for Memory
Performance Evaluation.” In Proceedings of the International
Symposium on Memory Systems (MEMSYS), Sep.-Oct. 2019.
-
(dynamic graph processing) O.
Jaiyeoba and K. Skadron. “GraphTinker: A High Performance Data
Structure for Dynamic Graph Processing.” In Proceedings
of the IEEE International Parallel and Distributed
Processing Symposium, May 2019.
- (automata processing,
FPGAs, heterogeneous architectures, accelerators) C.
Bo, V. Dang, T. Xie, J. Wadden, M. Stan, and K. Skadron.
"Automata Processing in Reconfigurable Architectures:
In-the-cloud Deployment, Cross-platform Evaluation, and Fast
Symbol-only Reconfiguration." ACM Transactions on
Reconfigurable Technology and Systems (TRETS), 12(2), May
2019, DOI 10.1145/3314576.
-
(automata processing,
heterogeneous architectures, accelerators, debugging) M.
Casias, K. Angstadt, T. Tracy II, K. Skadron, and W.
Weimer. “Debugging Support for Pattern-Matching
Languages and Accelerators.” In Proceedings of the
ACM International Conference on Architectural Support for
Programming Languages and Operating Systems (ASPLOS),
Apr. 2019.
-
(automata processing,
heterogeneous architectures, accelerators) K.
Angstadt, A. Subramaniyan, E. Sadredini, R. Rahimi, W. Weimer,
K. Skadron, and R. Das. “ASPEN: A Scalable In-SRAM
Architecture for Pushdown Automata.” In Proceedings
of the ACM/IEEE International Symposium on Microarchitecture
(MICRO), Oct. 2018.
-
(automata processing,
heterogeneous architectures, accelerators, benchmarks) J.
Wadden, T. Tracy II, E. Sadredini, L. Wu, C. Bo, J. Du,** Y.
Zhou, M. Wallace,* J. Udall,* M. Stan, and K. Skadron.
“AutomataZoo: A Modern Automata Processing Benchmark Suite.”
In Proceedings of the IEEE International Symposium on
Workload Characterization (IISWC), Oct. 2018.
-
(automata processing,
heterogeneous architectures, accelerators, natural
language processing) E. Sadredini, D. Guo,
C. Bo, R. Rahimi, K. Skadron, and H. Wang. “A Scalable
Solution for Rule-Based Part-of-Speech Tagging on Novel
Hardware Accelerators.” In Proceedings of the ACM
SIGKDD Conference on Knowledge Discovery and Data Mining
(KDD), Applied Data Science track, full paper with
poster presentation, Aug. 2018.
-
(automata processing,
heterogeneous architectures, accelerators, simulation) K.
Angstadt, J. Wadden, V. Dang, T. Xie, D. Kramp, W. Weimer, M.
Stan, and K. Skadron. “MNCaRT: An Open-Source,
Multi-Architecture Automata-Processing Research and Execution
Ecosystem.” IEEE Computer Architecture Letters,
published online Dec. 2017, published in print 17(1):84-87,
Jan.-June 2018. DOI 10.1109/LCA.2017.2780105.
-
(automata processing,
heterogeneous architectures, accelerators) J.
Wadden, K. Angstadt, and K. Skadron. “Characterizing and
Mitigating Output Reporting Bottlenecks in
Spatial-Reconfigurable Automata Processing Architectures.” In
Proceedings of the IEEE International Symposium on High
Performance Computer Architecture (HPCA), Feb. 2018.
-
(automata processing,
heterogeneous architectures, accelerators, bioinformatics)
C. Bo, V. Dang, E. Sadredini, K.
Skadron. “Searching for Potential gRNA Off-Target Sites
for CRISPR/Cas9 using Automata Processing across Different
Platforms.” In Proceedings of the IEEE International Symposium
on High Performance Computer Architecture (HPCA), Feb. 2018.
-
(automata processing,
FPGAs, heterogeneous architectures, accelerators) T.
Xie, V. Dang, J. Wadden, K. Skadron, and M. R. Stan. “REAPR:
Reconfigurable Engine for Automata Processing.” In Proceedings
of the International Conference on Field-Programmable Logic
and Applications (FPL), Sept. 2017. (pdf)
-
(automata processing,
heterogeneous architectures, accelerators, frequent tree
mining, pattern mining) E. Sadredini, K.
Wang, and K. Skadron. “Frequent Subtree Mining on the Automata
Processor: Challenges and Opportunities.” In Proceedings
of the ACM International Conference on Supercomputing (ICS),
June 2017. (pdf)
-
(automata processing,
heterogeneous architectures, spatial architectures,
accelerators, routing) J. Wadden, Samira
Khan, and K. Skadron. “Automata-to-Routing: An Open-Source
Toolchain for Design-Space Exploration of Spatial Automata
Processing Architectures.” In Proceedings of the IEEE
International Symposium on Field-Programmable Custom
Computing Machines (FCCM), Apr. 2017. (pdf)
-
(automata
processing, heterogeneous architecture, accelerators,
association rule mining, frequent itemset mining) K.
Wang, E. Sadredini, and K. Skadron. "Hierarchical
Pattern Mining with the Automata Processor." International
Journal of Parallel Programming, Springer, published
online Jan. 2017. DOI:10.1007/s10766-017-0489-y
(preprint
pdf)
- (automata processing,
heterogeneous architectures, accelerators, entity
resolution) C. Bo, K. Wang, J. Fox, and K.
Skadron. “Entity Resolution Acceleration using Micron’s
Automata Processor. In Proceedings of the 2016 IEEE
International Conference on Big Data (BigData), Dec.
2016. (pdf)
-
(automata processing,
heterogeneous architectures, accelerators, programming
languages, compilers) K. Angstadt, W. Weimer, and K. Skadron. "RAPID Programming
of Pattern-Recognition Processors." In Proceedings of the
ACM International Symposium on Architectural Support for
Programming Languages and Operating Systems (ASPLOS),
Apr. 2016. (pdf
| software)
Highlights from Prior Work
(specialized architectures, encryption)
X. Guo, M. El-Hadedy, S. Mosanu, X. Wei, K. Skadron, and M. R. Stan. "Implementation of Configurable AES Primitive with Agile Design Approach." Integration, the VLSI Journal, Elsevier, 85(C):87-96, July 2022.
-
(fuzzing, security) A.
Ahmed, J. Hiser, A. Nguyen-Tuong, J. W. Davidson, and K.
Skadron. “BigMap: Future-proofing Fuzzers with Efficient Large
Maps.” In Proceedings of the IEEE/IFIP International
Conference on Dependable Systems and Networks (DSN), June
2021.
- (power delivery, IR
drop, thermal) S. Rahimipour, Runjie Zhang, Ke Wang,
K. Skadron, F. Z. Rohkani, and M. Stan. "MTTF Enhancement Power-C4 Bump Placement Optimization." IEEE
Transactions on Very Large Scale Integration Systems (TVLSI),
27(7):1633-39, July 2019, DOI 10.1109/TVLSI.2019.2904048.
-
(specialized architectures, encryption) M. El-Hadedy, H.
Mihajloska, D. Gligoroski, A. Kulkarni, D. Stroobandt and K.
Skadron. "A 16-bit Reconfigurable Encryption Processor for
Pi-Cipher." In Proceedings of the 23rd Reconfigurable
Architectures Workshop (RAW), in conjunction with IPDPS, May
2016. Best paper award! (pdf)
- (specialized
architectures, transpose memory, signal processing) M.
El-Hadedy, X. Guo, M. Margala, M. R. Stan, and K. Skadron.
“Dual-Data Rate Transpose-Memory Architecture Improves the
Performance, Power and Area of Signal-Processing Systems.”
Journal of Signal Processing Systems, Springer,
published online Nov. 2016. DOI 10.1007/s11265-016-1199-1.
-
(automata
processing, heterogeneous architecture, accelerators,
association rule mining) K. Wang, E. Sadredini, and
K. Skadron. "Sequential Pattern Mining with the Micron
Automata Processor." In Proceedings of the ACM
International Conference on Computing Frontiers, May
2016. Best paper award! (pdf)
-
(benchmark suites,
accelerators) G. Juckeland, W. Brantley, S.
Chandrasekaran, B. Chapman, S. Che, M. Colgrove, H. Feng, A.
Grund, R. Henschel, W-M. W. Hwu, H. Li, M. S. Mueller, M.
Perimov, P. Shelepugin, K. Skadron, J. Stratton, A. Titov, K.
Wang, M. van Waveren, B. Whitney, S. Wienke, R. Xu, and K.
Kumaran. "SPEC ACCEL: A Standard Application Suite for
Measuring Hardware Accelerator Performance." In Proceedings of the Fifth International Workshop on Performance
Modeling, Benchmarking and Simulation of High Performance
Computer Systems (PMBS14), in conjunction with SC, Nov.
2014.
-
(heterogeneous architectures,
reconfigurable logic, design space exploration) L.
Wang and K. Skadron. "Lumos+: Rapid, Pre-RTL Design Space
Exploration on Accelerator-Rich Heterogeneous Architectures
with Reconfigurable Logic." In Proceedings of the
IEEE International Conference on Computer Design (ICCD),
Oct. 2016. (pdf)
- (power
delivery, thermal) K. Wang, K. Skadron, and M. R.
Stan. “ Closing the Power Delivery/Heat Removal Cycle for
Heterogeneous Multi-Scale Systems.” In Proceedings of
the 22nd International Workshop on Thermal
Investigations of ICs (THERMINIC), Sept. 2016 (pdf).
-
(power
delivery, transient voltage noise) R. Zhang, K.
Mazumder, B. H. Meyer, K. Wang, K. Skadron, and M. R. Stan.
"Transient Voltage Noise in Charge-Recycled Power
Delivery Networks for Many-Layer 3D-IC."
In Proceedings
of the ACM/IEEE International Symposium on Low Power
Electronics and Design (ISLPED), July 2015. (pdf)
-
(GPGPU,
reliability, redundant multithreading) J.
Wadden, A. Lyashevsky, S. Gurumurthi, V. Sridharan, and K.
Skadron. "Real-World Design and Evaluation of
Compiler-Managed GPU Redundant Multithreading." In Proceedings
of the ACM/IEEE International Symposium on Computer
Architecture, June 2014. (pdf)
-
(power
delivery, transient voltage noise) K.
Wang, R. Zhang, B. H. Meyer, M. R. Stan, and K. Skadron.
"Managing C4 Placement for Transient Voltage Noise
Minimization." In Proceedings of the ACM/IEEE Conference
on Design Automation (DAC), June 2014 (please download
from this
listing).
-
(power
delivery, IR drop) K.
Wang, R. Zhang, B. H. Meyer, K. Skadron, and M. R. Stan.
"Walking Pads: Fast Power-Supply Pad-Placement
Optimization." In Proceedings of the ACM/IEEE Asia and
South Pacific Design Automation Conference (ASP-DAC),
Jan. 2014. Best paper candidate. (preprint
pdf)
-
(heterogeneous
architecture, dark silicon) L. Wang and K. Skadron.
"Implications of the Power Wall: Dim Cores and Reconfigurable
Logic." IEEE Micro special issue on Dark Silicon,
33(5): 40-49, Sept.-Oct. 2013.DOI
10.1109/MM.2013.74.
(preprint
pdf | Lumos
software)
-
(gpgpu,
accelerators, heterogeneous architecture, memory,
performance portability) S.
Che, J. Meng, and K. Skadron. "Dymaxion++: A
Directive-based API to Optimize Data Layout and Memory
Mapping for Heterogeneous Systems." In Proceedings
of the Fourth International Workshop on Accelerators and
Hybrid Exascale Systems, in conjunction with IPDPS,
May 2014. (pdf)
-
(performance
prediction, gpgpu, accelerators, manycore, heterogeneous
architecture, benchmarks) S.
Che and K. Skadron. "BenchFriend: Correlating the
Performance of GPU Benchmarks." Sage International
Journal of High Performance Computing Applications
(IJHPCA), 28(2):236-248, May 2014.(preprint
pdf)
-
(gpgpu,
accelerators, manycore, heterogeneous architecture,
benchmarks) S.
Che. B. M. Beckmann, S. K. Reinhardt, and K. Skadron,
"Pannotia: A Characterization of GPGPU Graph Applications,
In Proceedings of the IEEE International Symposium on
Workload Characterization (IISWC), Sept. 2013. (pdf)
-
(gpgpu,
accelerators, manycore, heterogeneous architecture,
programming models and frameworks) L. Szafaryn, T.
Gamblin, B. R. de Supinski, and K. Skadron. "Trellis:
Portability Across Architectures with a High-level Framework."
Elsevier Journal of Parallel and Distributed Computing,
73(10):1400-13, Oct. 2013. (First published online July,
2013.) DOI 10.1016/j.jpdc.2013.07.001.
(preprint
pdf)
-
(reliability)
L. Szafaryn, B. H.
Meyer, and K. Skadron. "Evaluating the Overheads of Soft
Error Protection Mechanisms in the Context of Multi-bit
Errors at the Scope of a Processor Core." IEEE Micro special
issue on Reliability Aware Design, 33(4):56-65, July-Aug.
2013. DOI 10.1109/MM.2013.68.
(preprint
pdf)
-
(CPU-GPU
load balancing) M.
Boyer, K. Skadron, S. Che, N. Jayasena. "Load Balancing in a
Changing World: Dealing with Heterogeneity and Performance
Variability. In Proceedings of the ACM Conference on
Computing Frontiers, May 2013. (pdf)
-
(thermal)
K. Sankaranarayanan, B. H. Meyer, W. Huang, R. J.
Ribando, H. Haj-Hajiri, M. R. Stan, and K. Skadron.
"Architectural Implications of Spatial Thermal
Filtering." Elsevier Integration, the VLSI Journal,
46(1):44-56, Jan. 2013.DOI 10.1016/j.vlsi.2011.12.002
(preprint
pdf)
-
(floorplanning,
thermal, power delivery) G.
Faust, R. Zhang, K. Skadron, M.R. Stan, and B. Meyer.
"ArchFP: Rapid Prototyping of pre-RTL Floorplans." In Proceedings
of the IFIP/IEEE International Conference on Very Large
Scale Integration (VLSI-SoC), Oct. 2012. (pdf
| ArchFP
software)
-
(SIMD,
cache, branch divergence) J.
Meng, J. W. Sheaffer, and K. Skadron. "Robust SIMD:
Dynamically Adapted SIMD Width and Multi-Threading
Depth." In Proceedings of the IEEE International
Parallel & Distributed Processing Symposium (IPDPS),
May 2012. (pdf)
-
(datacenter
co-scheduling, cache) J.
Mars, L. Tang, R. Hundt, K. Skadron, and M. L. Soffa.
"BubbleUp: Increasing Sensible Co-locations for Improved
Utilization in Modern Warehouse Scale Computers." In
Proceedings of the ACM/IEEE International Symposium on
Microarchitecture (MICRO), Dec. 2011. (pdf)
Also in IEEE Micro "Top Picks from 2011 Architecture
Conferences."
-
(gpgpu,
accelerators, heterogeneous architecture, memory,
performance portability) S.
Che, J. W. Sheaffer, and K. Skadron. "Dymaxion:
Optimizing Memory Access Patterns for Heterogeneous
Systems." In Proceedings of the ACM/IEEE
International Conference for High Performance Computing,
Networking, Storage and Analysis (SC), Nov. 2011. (pdf)
-
(reliability,
fault tolerance, redundant execution) B.
Meyer, B. Calhoun, J. Lach, and K. Skadron.
"Cost-effective Safety and Fault Localization using
Distributed Temporal Redundancy." In Proceedings
of the ACM/IEEE International Conference on Compilers,
Architectures, and Synthesis for Embedded Systems (CASES),
Oct. 2011. (pdf)
-
(thermal,
power, scaling, dark silicon) W.
Huang, K. Rajamani, M. R. Stan, and K. Skadron.
"Scaling with Design Constraints -- Predicting the Future of
Big Chips." IEEE Micro special issue on Big
Chips, 31(4):16-29, July/Aug. 2011, DOI 10.1109/MM.2011.42.
(preprint
pdf)
-
(manycore,
GPU architecture, power) M.
Gebhart, D. R. Johnson, D. Tarjan, S. W. Keckler, W. J.
Dally, E. Lindholm, and K. Skadron. "Energy-efficient
Mechanisms for Managing Thread Context in Throughput
Processors." In Proceedings of the ACM/IEEE
International Symposium on Computer Architecture (ISCA), June
2011. (pdf)
-
(gpgpu,
accelerators, manycore, heterogeneous architecture,
benchmarks) S. Che,
J. W. Sheaffer, M. Boyer, L. G. Szafaryn, L. Wang, and K.
Skadron. "A Characterization of the Rodinia Benchmark
Suite with Comparison to Contemporary CMP Workloads."
In Proceedings of the IEEE International Symposium on
Workload Characterization (IISWC), Dec. 2010. (pdf)
-
(manycore,
dynamic cores) M.
Boyer, D. Tarjan, K. Skadron. "Federation: Boosting
Per-Thread Performance of Throughput-Oriented Manycore
Architectures." ACM Transactions on Architecture
and Code Optimization (TACO), 7(4):1-38, Dec. 2010,
DOI 10.1145/1880043.1880046.
(preprint
pdf)
-
(manycore,
cache, coherence) D.
Tarjan and K. Skadron. "The Sharing Tracker: Using
Ideas from Cache Coherence Hardware to Reduce Off-Chip
Memory Traffic with Non-Coherent Caches." In Proceedings
of theACM/IEEE
International Conference for High Performance Computing,
Networking, Storage and Analysis (SC),
Nov. 2010. (pdf)
-
(SIMD,
cache, branch divergence) J.
Meng, D. Tarjan, and K. Skadron. "Dynamic Warp
Subdivision for Integrated Branch and Memory Divergence
Tolerance." In Proceedings of the 37th ACM/IEEE
International Symposium on Computer Architecture, June
2010. (pdf)
-
(manycore,
cache, coherence) J.
Meng and K. Skadron "Avoiding Cache Thrashing due to Private
Data Placement in Last-Level Cache for Manycore
Scaling." In Proceedings of the IEEE International
Conference on Computer Design (ICCD), pp. 282-88, Oct.
2009. (pdf)
-
(power,
real-time, control theory) T. Horvath and K.
Skadron. "Multi-mode Energy Management for Multi-tier
Server Clusters." In Proceedings of the
ACM/IEEE/IFIP International Conference on Parallel
Architectures and Compilation Techniques (PACT), pp.
270-79, Oct. 2008. (preprint
pdf)
-
(thermal)
W. Huang, K. Sankaranarayanan, K. Skadron, R. J.
Ribando, and M. R. Stan. "Accurate, Pre-RTL
Temperature-Aware Processor Design Using a Parameterized,
Geometric Thermal Model." IEEE Transactions on
Computers, 57(9):1277-88, Sept. 2008, DOI
10.1109/TC.2008.64. (pdf)
-
(gpgpu,
fpga, accelerators, heterogeneous architecture) S.
Che, J. Li, J. W. Sheaffer, K. Skadron, and J. Lach.
"Accelerating Compute Intensive Applications with GPUs and
FPGAs." In Proceedings of the IEEE Symposium on
Application Specific Processors (SASP), pp.
101-07, June 2008. (pdf)
-
(simulation
methodology) H. Cook and K. Skadron. "Predictive
Design Space Exploration Using Genetically Programmed Response
Surfaces." In Proceedings of the ACM/IEEE Conference on
Design Automation (DAC), June 2008. (pdf)
-
(gpgpu)
J.
Nickolls, I. Buck, M. Garland, K. Skadron. "Scalable
Parallel Programming with CUDA." ACM Queue,
6(2):40-53, Mar.-Apr. 2008. DOI
10.1145/1365490.1365500 (pdf)
-
(parameter
variations, multicore, thermal, power, leakage) E.
Humenay, D. Tarjan, and K. Skadron. "Impact of Process
Variations on Multicore Performance Symmetry." In Proceedings
of the 2007 Conference on Design, Automation and Test in
Europe (DATE), pp. 1653-58, Apr. 2007. (pdf)
-
(reliability,
thermal) Z. Lu, W. Huang, M. Stan, K. Skadron, and
J. Lach. "Interconnect Lifetime Prediction for
Reliability-Aware Systems." IEEE Transactions on
VLSI Systems, 15(2):159-72, Feb. 2007. (pdf)
-
(branch
prediction, trace cache, power) M. Co, D. A.B.
Weikle, and K. Skadron. "Evaluating Trace Cache Energy
Efficiency." ACM Transactions on Architecture and Code
Optimization (TACO), 3(4):450-76,
Dec. 2006. (Abstract
| pdf)
-
(graphics
architecture, reliability) J. W. Sheaffer, D. P.
Luebke, and K. Skadron. "The Visual Vulnerability Spectrum:
Characterizing Architectural Vulnerability for Graphics
Hardware." In Proceedings of Eurographics/ACM Graphics
Hardware 2006 (GH), pp. 9-16, Sept. 2006. (pdf)
-
(thermal)
S. W. Chung and K. Skadron. "Using on-Chip Event
Counters for High-Resolution, Real-Time Temperature
Measurements." In Proceedings of the IEEE/ASME Tenth
Intersociety Conference on Thermal and Thermomechanical
Phenomena in Electronic Systems (ITHERM), June 2006. (pdf)
- (thermal) W.
Huang, S. Ghosh, S. Velusamy, K. Sankaranarayanan, K. Skadron,
and M. R. Stan. “HotSpot: A Compact Thermal Modeling
Methodology for Early-Stage VLSI Design.” IEEE
Transactions on Very Large Scale Integration (VLSI) Systems,
14(5):501-513, May 2006. (pdf)
Google
Scholar Classic Paper in in Computer Hardware Design for
2006 (3rd most cited paper).
-
(multi-core
architecture, power, thermal) Y. Li, B. C. Lee, D.
Brooks, Z. Hu, and K. Skadron. "CMP Design Space
Exploration Subject to Physical Constraints." In Proceedings
of the Twelfth IEEE
International Symposium on High Performance Computer
Architecture (HPCA), pp. 15-26, Feb. 2006. (pdf)
-
(branch
prediction) D. Tarjan and K. Skadron. "Merging Path
and Gshare Indexing in Perceptron Branch Prediction." ACM
Transactions on Architecture and Code Optimization,
Sept. 2005, 2(3):280-300. (pdf)
-
(thermal,
security) P. Dadvar
and K. Skadron. "Potential Thermal Security
Risks." In Proceedings of the IEEE Semiconductor
Thermal Measurement, Modeling, and Management Symposium
(Semi-Therm 21), pp. 229-34, Mar. 2005. (pdf)
-
(thermal)
K.
Skadron, K. Sankaranarayanan, S. Velusamy, D. Tarjan, M.R.
Stan, and W. Huang. "Temperature-Aware
Microarchitecture: Modeling and Implementation." ACM
Transactions on Architecture and Code Optimization,
1(1):94-125, Mar. 2004.
(pdf)
-
(leakage
power) Y. Li, D.
Parikh, Y. Zhang, K. Sankaranarayanan, M. R. Stan, and K.
Skadron. "State-Preserving vs. Non-State-Preserving
Leakage Control in Caches." In Proceedings of the
2004 Design, Automation and Test in Europe (DATE)
Conference, pp. 22-27, Feb. 2004. (pdf)
[HotLeakage
software home page]
- (power, real-time)
V. Sharma, A. Thomas, T.
Abdelzaher, Z. Lu, and K. Skadron. "Power-Aware QoS
Management on Web Servers." In Proceedings of the
24th International Real-Time Systems Symposium, pp.
63-72, Dec. 2003. (pdf)
(Best student paper!)
-
(branch
prediction) Z. Lu, J.
Lach, M. Stan, and K. Skadron. "Alloyed Branch
History: Combining Global and Local Branch History for
Robust Performance." International Journal of Parallel
Programming, Kluwer, 31(2):137-77, Apr. 2003. (pdf
| Abstract)
-
(write
buffers) K. Skadron and
D.W. Clark. "Design Issues and Tradeoffs for Write Buffers."
In Proceedings of the Third International Symposium on
High-Performance Computer Architecture, pp. 144-55,
February 1997. (postscript
| pdf
| abstract)
Complete
list of Skadron's publications through 2017;
for a complete list of my publications, please see my Google Scholar page.
|